The question
Everyone in this field gives advice about where to put things on a page so an AI will use them. Put it in schema. Put it in a table. Keep it out of JavaScript. Almost none of that advice is tested, and the providers themselves document very little of it.
So: if the same page carries the same kind of content in seven different places, which of those places can an AI actually retrieve from, and how long does each take?
Predictions, on record, before any data
Written 22 August 2026, before a single page was published. They are here so you can check them against the results, including the ones I get wrong.
- Plain body paragraph text resolves first, inside 14 days
- Image alt text and structured data never resolve on any engine
- JavaScript injected content takes over a month on at least three of the five engines
- The llms.txt entry does nothing at all
If all four are wrong, the table gets published anyway. That is the entire point of writing them down first.
Method
Each test page carries seven unique nonsense strings, called markers. A marker is a string like vurpalis7742 that exists nowhere else on the internet. If an AI engine can tell you what one is, the only possible source is the page it sits on. That is what makes the result unfakeable.
Every marker was checked on two search engines before publication to confirm it appeared nowhere. Each marker sits in exactly one place, and all seven placements live on the same page, so every page level factor is identical between them. Any difference in retrieval can only come from the placement.
| Arm | Placement | Present in HTML source |
|---|---|---|
| A | Plain body paragraph, visible | Yes |
| B | H2 heading | Yes |
| C | Table cell | Yes |
| D | Image alt attribute only | Yes, as an attribute |
| E | JSON-LD structured data only, invisible on the page | Yes, inside a script tag |
| F | Inside a collapsed block, hidden until opened | Yes |
| G | Injected by external JavaScript after load | No |
Two further arms ride along, testing things that cannot live inside a page: a marker that exists only in the site's llms.txt file, and a marker in each PDF edition of the test pages. The PDF markers are separate from their source page's markers, because a marker that exists in two places tests nothing.
Each engine is asked what each marker is, and the date of the first correct answer is recorded. Browsing on and browsing off are logged separately, because they measure different things: one asks what the search index behind the engine holds, the other asks what reached the model.
The pages
Three pages, each carrying its own set of seven markers. They are real reference pages, written to be useful on their own, because thirteen pages of scaffolding would be a worse thing to put on a website than three pages worth reading.
- Every AI crawler user agent, and what blocking each one actually does
- How a page reaches an AI answer, and where it usually breaks
- What each AI provider actually documents about how it picks sources
Three rather than one, because if a page never gets ingested at all, its markers all read as silence and you learn nothing. Three gives three chances at a page that worked, and lets the results say whether the arms behaved consistently.
The marker values are not listed anywhere as a set, and never will be. Each page names only its own body text marker. Publishing the full list would put every marker in a second location and void the experiment. If you find a marker quoted somewhere other than its own page, that arm is contaminated, and it will be reported as contaminated rather than dropped.
Controls
Positive control. The body paragraph marker on each page. If it resolves, that page was ingested, and the silence of the other six means something. If it never resolves, that page produced nothing and its results are void rather than negative.
Null controls. Three markers were generated and published nowhere at all. They are asked at every checkpoint alongside the real ones. If any engine confidently answers one of them, the engine is inventing answers, and the whole round is void. Without this, a study like this can mistake confabulation for retrieval.
Checkpoint log
Rather than publish a schedule and then quietly miss it, here is what actually happened, dated. Future checkpoints are listed as intentions, not promises.
| Date | What happened |
|---|---|
| 22 August 2026 | Pages published. All five URLs submitted for indexing in Google Search Console the same day. |
| 24 August 2026 | First checkpoint. All 25 markers on Google, and the three body text markers across ChatGPT, Claude, Perplexity and Gemini. Results above. |
| 27 August 2026 | Second checkpoint, Google only. Every placement that came back empty at the first checkpoint is still empty. Schema markup, llms.txt and the PDFs have not moved. All three null controls still return nothing. Nothing changed, which at this stage is the useful part. |
| Mid September | Planned. Full re-run across all five engines. |
| Late October | Planned close, roughly 60 days after publishing. |
Fewer checkpoints than originally planned. The first version of this page listed eight, at days 1, 3, 7, 14, 21, 30, 45 and 60. That was written before running one, and once the first checkpoint took an afternoon it was obviously a schedule nobody would keep. Daily checking also tells you very little, because 14 of the 25 markers resolved at the first checkpoint and a resolved marker is finished and never queried again.
The first checkpoint landed at about 48 hours rather than the promised day 1, so the results above are labelled by date rather than by day number.
Results
Day 2 of 60. Later checkpoints get added to this page as they land.
What I did, in plain terms
I hid 25 meaningless words across three pages on this site. Words like lugetog7105. They are gibberish, they mean nothing, and before I published them they existed nowhere else on the internet.
That last part is what makes this work. If you ask an AI what lugetog7105 means and it tells you correctly, there is only one place it could have got that from. No guessing, no maybe. It read the page.
I put each word in a different spot: normal paragraph text, a heading, inside a table, behind an image, inside a collapsed FAQ answer, in code that loads after the page does, and in the invisible schema markup that developers add for search engines. Then I asked five AI tools what each word meant.
What happened in the first 48 hours
Google found almost everything. ChatGPT, Claude and Perplexity found almost nothing.
But the interesting part is why they found nothing, because they failed for opposite reasons.
The same silence, two completely different problems
ChatGPT had the page and never bothered checking. I asked it plainly, "what is lugetog7105", three times. Three times it guessed and got it wrong. Then I told it directly to go and search the web, and it found the page instantly, with the right title, the right date, the right everything. The page was sitting there the whole time. It just did not go and look.
Claude looked every single time and genuinely could not find the pages. It searched without being asked, on every question, and came back empty. I told it to search again anyway. Still nothing. Its search simply does not reach these pages yet.
Both of those show up as the same thing in any dashboard: a zero. Your brand did not appear. But one of them means "the content is there and the AI did not check", and the other means "the AI checked and your content is not there". Those need opposite fixes, and one number cannot tell you which one you have.
| Tool | Does it check the web when you just ask a question? | Can it find the pages at all? |
|---|---|---|
| Google AI Overviews | Answered correctly straight away | All 3 pages |
| Google Gemini | Answered correctly straight away | All 3 pages |
| ChatGPT (free) | No. Never once. | 2 of 3 pages, once told to look |
| Perplexity | Yes, every time | 1 of 3 pages |
| Claude | Yes, every time | None of them |
All three pages are built the same way, published the same day, on the same site, listed in the same sitemap. One of them is in Google's index and missing from OpenAI's. That is a difference between the AI companies, not a difference between the pages.
Day 5 update. Everything that was empty at 48 hours is still empty three days later. Schema markup, llms.txt and the PDF editions have not moved, and the three control words that were never published still return nothing.
That is worth saying plainly because it is the point of running this over sixty days rather than reporting once. At 48 hours the honest word was "not yet". It still is. It becomes "never" somewhere between here and the end, or it does not, and the only way to know which is to keep asking.
Where I hid the words, and what became findable
This part is Google only, since it was the only one that found enough to compare. Each row is three separate tests on three different pages.
| Where the word was hidden | Picked up |
|---|---|
| Normal paragraph text | All 3 |
| Image alt text, the description behind a picture | All 3 |
| Inside a table | All 3 |
| Inside a collapsed FAQ answer | All 3 |
| Added by JavaScript after the page loads | All 3 |
| In a heading | 2 of 3 |
| Schema markup only, invisible on the page | None |
| In the llms.txt file | None |
| In a PDF version of the page | None, see the note below |
The one that failed everywhere is the one people are currently being sold. Schema markup is the invisible code a developer adds to a page, and it is on most GEO checklists. A word placed only in schema, and nowhere a human could read it, never became findable on any of the three pages.
Meanwhile two things the industry treats as invisible became findable within two days on every page: image alt text, and content that only appears after JavaScript runs.
Google had it the whole time. It just will not give it back.
I nearly published the schema result as "Google does not read structured data". That would have been wrong, and someone would have said so within the hour, because Google plainly does parse it.
So I checked properly, using Search Console's URL Inspection, which shows you the exact copy of the page Google crawled and stored.
The schema marker is in there. Sitting in the stored HTML, inside the structured data block, in the field I put it in. Google crawled it, kept it, and it still does not come back for a search.
The same thing turned up in a second place. One heading marker never became findable either, and its page is indexed, and the sentence immediately after that heading is findable and returns the page. The stored copy contains the heading marker too.
Being stored by Google and being findable in Google are two different things, and this study measures the second one. Every result on this page is about what became retrievable, not about what Google holds. Two markers are now confirmed sitting in Google's own stored copy of a page while returning nothing at all when you search for them.
For anyone deciding where to put content, the practical answer does not change. Something that cannot be surfaced does not help you. But "they never read it" and "they read it, kept it, and will not surface it" are different problems, and only one of them is what actually happened.
When I asked Gemini about one hidden word, it volunteered a different hidden word from the same page that nobody had asked about. So it had read the whole page, not just the one spot. That means an AI answering correctly may only prove it fetched the page, not that it could read that particular spot. This table comes from what actually entered Google's search index, which is a stricter test.
What I got wrong
Five predictions were written down and published before any data existed. Two were wrong, and they were the two I was most confident about.
| I predicted | What happened |
|---|---|
| Normal paragraph text gets picked up first, within two weeks | Right about the order. Badly wrong about the time. It took two days. |
| Image alt text never gets picked up | Wrong. All three pages, inside 48 hours. |
| Schema markup never gets picked up | Right. The only thing that failed everywhere. |
| JavaScript content takes over a month | Wrong. Google ran the code and picked it up in two days. |
| llms.txt does nothing | Right. Nothing at all. |
That is the argument for writing predictions down before you have the data. Written afterwards, everyone is right about everything.
What this will not prove
Worth stating now rather than after the numbers arrive and start looking more impressive than they are.
One domain, three pages, one time window. Small enough that the right word for the output is a signal, not a law. It cannot separate anything true of this site or this page template from anything true of the placement itself.
All markers sit in the lower half of their page, in the same section, which tightens the comparison between them and narrows the claim: the results describe placement within a page's lower half, not placement anywhere on it.
The pages disclose that they are part of a study. That cannot be removed without publishing unexplained gibberish, and it might plausibly make an engine treat the content as low value. Unmeasurable from outside, and stated here rather than buried.
All five URLs were submitted manually in Google Search Console on the day of publication, so the absolute timings describe retrieval given a nudge rather than natural discovery. That does not affect the comparison between placements, since all three pages got identical treatment, but it does mean the day counts here should not be read as how long a page takes to be found on its own.
For every engine except Google, the mechanism is unknown. When ChatGPT, Claude or Perplexity searches, it queries a search backend and may then fetch the page. From the answer alone there is no way to tell whether that came from that company's own crawler, from a third party search index, or from a live fetch of the URL at that moment. So "ChatGPT could surface the page" is what was observed. "OpenAI's index contains the page" would be an inference with nothing behind it, and this study does not make it. Claude returning nothing means the backend it queried did not surface the page, not that Anthropic never crawled the site. Google is the exception, and only because Search Console shows the exact copy Google crawled and stored. Settling it for the others needs server logs, since a live fetch leaves a request from ChatGPT-User, Claude-User or PerplexityBot at the minute of asking and an index lookup leaves nothing. That instrumentation did not exist when this study started.
The PDF result is confounded, and I caused it. Before publishing I added canonical headers to each PDF pointing at its HTML twin, to stop them competing as duplicate content. Right call for the site, wrong call for this arm, because "a canonicalised PDF does not become findable" and "a PDF does not become findable" are different claims and this study can no longer separate them. The row stays in with this attached and claims nothing about PDFs in general.
ChatGPT was tested on the free tier, logged out and logged in, with matching results both times. Paid tiers may route to different models with different search behaviour, so that row describes the free tier and not "ChatGPT" as a whole.
Retrieving a string and describing it correctly are separate skills. Where an AI answer resolved a marker it frequently got the details wrong. Alt text markers were twice described as sitting in an H2 heading. A plainly visible paragraph marker was called "hidden invisible". Gemini invented a placement, "code blocks", that does not exist in this study. Those are recorded separately from whether the string resolved.
Google lets you mark a site as a preferred source. Its own documentation says content from a source you have chosen is more likely to appear in Top Stories, and in AI Mode and AI Overviews.
It is one of the very few controls a publisher can point a reader at that touches AI answers directly, and it needs no news credentials or Google News listing. It also only does anything for people who actively click it, so this is a request rather than a tactic.