Knowledge · Machine readability
How AI really reads your website
Crawlers, AI search systems and browser agents call up the same address — and do not necessarily get the same thing to look at. This article sorts out what arrives on the first request, why that differs between systems, and what follows from it for a company website.
What you will understand here.
- Telling visitor types apart by purpose
- Separating a direct request from browser rendering
- Separating reachability from evidenced visibility
There is no single machine view of your website.
When a person opens your website, a fair amount happens in the background: the browser loads the document, runs scripts, fetches images and fonts, and assembles the page that finally appears on screen. That result is so self-evident that it is easily mistaken for the website itself.
Machine visitors do not necessarily follow that path. Some fetch the document and read the delivered text. Others run the page in a real browser and see the finished interface. Others again come mainly to get something done on the page rather than to read it.
That is why the question “What does AI see on my website?” cannot be answered in that form. It becomes answerable only with three additions: which system, for which purpose, by which route. Those three questions run through this whole article.
Not every machine visitor comes for the same reason.
Automated requests can be sorted by purpose. The grouping below is an orientation aid, not an official taxonomy — but it explains why the same website is treated differently by different systems.
- Purpose 1
Search and discoverability
A system records content so that it can later show and link it in search results or answers. OpenAI names OAI-SearchBot for this: anyone who wants content to appear in summaries and snippets in ChatGPT must not block that access. Anthropic describes Claude-SearchBot as a bot meant to improve the quality of search results. Perplexity describes PerplexityBot as designed to surface and link websites in search results.
- Purpose 2
Model development — and how to control it
The second purpose concerns content that may feed into the development of models, and above all how a site owner controls that. OpenAI points to GPTBot for this: anyone who wants to exclude pages from potential training should disallow that user agent. Anthropic writes about ClaudeBot that it collects content which could potentially contribute to their training. Perplexity draws an explicit line here: PerplexityBot is not used to crawl content for AI foundation models.
- Purpose 3
A request on behalf of a user
The third case centres on the immediate user request: a person asks a question, and the provider can fetch web content specifically for it. Anthropic describes Claude-User that way — when people ask Claude something, Claude may access websites through that agent. Perplexity describes Perplexity-User accordingly and states explicitly that this fetcher generally ignores robots.txt, because a user triggered the request. Nothing general follows from this about how providers subsequently store or further process such fetches.
For many automated crawlers, a single file in the root directory is the central means of control: robots.txt. Anthropic writes that its own bots honour the industry-standard directives in it. How user-triggered requests behave, by contrast, depends on the provider and on the type of access: Perplexity states explicitly for Perplexity-User that this fetcher generally ignores robots.txt for user-requested fetches. Robots.txt is therefore a control and not a lock.
A common misunderstanding belongs right here: Google-Extended is not a bot of its own. Google's crawler documentation states that Google-Extended has no separate user agent string in HTTP requests; crawling is done with the existing Google user agents, and the entry in robots.txt works purely as a control. It governs whether content Google has crawled may be used for training future Gemini models and for grounding in certain Google products. On inclusion in Google Search and on ranking there it has, according to the same documentation, no influence.
What does a system get on the first request?
This is easiest to show with an example. Take a fictional business: Elektro Koenig Berlin, focus on wallbox installation, location Berlin, desired action “request an appointment”. Four details a person finds in seconds. Whether a machine finds them depends on the route it takes.
- Route A
The facts are in the delivered document
Elektro Koenig BerlinWallbox installation · BerlinRequest an appointment
The server delivers the details already. A system that only reads the delivered document and runs no browser has the name, the service, the location and the desired action together straight away.
- Route B
The content only comes into being in the browser
<div id="root"></div>— nothing else —
Google calls this structure the app shell model: the first HTML does not contain the actual content; it only comes into being once a browser runs the scripts. Anyone reading without a browser finds an empty frame at this point. Anyone reading with a browser sees the same page as a person.
- Route C
The fully rendered interface
Wallbox installation in Berlin[ Request an appointment ]
A system that actually runs the page in a browser gets the finished interface: heading, text, buttons. It does not only read — it also sees what can be operated.
A common short circuit runs: whatever is not there as running text in the source does not exist for machines. It is not that simple. The delivered document holds more than the visible text — structured data, embedded data blocks and server-prepared content are in there too, and they are readable. The dividing line does not run between visible and invisible, but between what the server delivers and what only the browser creates.
And a limitation belongs with it: the example sorts routes. It does not describe the architecture of any particular provider — which route a specific system takes on a specific request is not settled by it.
Rendering is not the same everywhere.
Whether a system runs JavaScript is not a property of “AI”, but a property of that particular system. For some of these systems there is a solid measurement — with a clear date and clear limits.
- Measured at the time
December 2024, real access logs
Vercel and MERJ published an analysis of actual requests on 17 December 2024. For the specialised AI crawlers examined there — named are OAI-SearchBot, ChatGPT-User and GPTBot from OpenAI, ClaudeBot from Anthropic, Meta-ExternalAgent, Bytespider and PerplexityBot — the result was: they ran no JavaScript. In part they even downloaded JavaScript files without executing them.
- What that does not prove
Two counter-examples in the same study
The same analysis names two systems that do render: Gemini used Google's infrastructure and could therefore render fully, and AppleBot renders, according to the same source, through a browser-based crawler. “No JavaScript” is thus not even within this study a property of all AI-related requests. On top of that comes the tense: the study explicitly described the state at that time. It is a dated measurement, not a promise about every request today.
- What that means for a website
A question instead of a rule
No blanket “render everything on the server” follows from this. What follows is a question every business can answer for itself: which details that matter commercially only come into being in the browser? For the service, the location, the price range and the contact route, a dependency on the browser is an avoidable risk. For an image slider or a chat window it is less critical, as long as no commercially load-bearing detail and no necessary action is available exclusively through that component.
For classic Google Search the case is documented and different: Google Search runs JavaScript, with a continuously updated version of Chromium. The documentation describes three phases for this — crawling, rendering, indexing — and states at the same time that these phases do not run cleanly one after another and are not observable from outside.
Both at once is therefore possible: Google can render client-side generated content and process it for indexing, while a non-rendering request receives only the document as initially delivered. Anyone who looks only at Google visibility does not see that difference.
Technical deep dive
How modern websites are built and delivered
What decides the contents of the initial document: HTML, JavaScript, the four delivery paths and why the same framework name allows two completely different outcomes.
Readable does not yet mean visible.
Technical reachability is the first stage, not the last. In its analysis Pixelkiez keeps four questions apart, because they are answered in different ways — and because an early stage never proves a later one.
- 01 · Discoverability
Can a system fetch the page at all?
Access, redirects, status codes, robots.txt. Without an open door nothing further happens.
- 02 · Extractability
Can the facts be extracted reliably?
Service, location, responsibility, contact route — as text a system can assign unambiguously, not merely as an impression.
- 03 · Answerability
Does the page answer a concrete question unambiguously?
Readable details do not automatically become an answer that can be adopted. For that, the question has to be genuinely answered on the page.
- 04 · Observed AI Visibility
Does the domain actually appear in AI answers?
That shows only in a defined, repeatable test — observed, not inferred.
Being reachable is a precondition — not proof of visibility. Stage 1 does not prove stage 4. That a crawler is allowed to fetch a page says nothing about whether it ever appears in an answer.
The providers themselves put this more cautiously than the industry often repeats it. OpenAI describes allowing OAI-SearchBot as the precondition for a website to be eligible for inclusion at all, and states in the same section: placement is not guaranteed.
An agent does not come only to read.
A browser agent differs fundamentally from a simple request. It is given an interface and a task — and for that it needs not only text, but controls that can be named.
A request delivers a document. An agent stands in front of an interface and is meant to get something done on it. Whether it succeeds depends on whether the page declares its controls as what they are: a button as a button, a form field with a label, a menu with a recognisable state.
OpenAI names a concrete example for this: ChatGPT Atlas uses ARIA tags — the same roles and labels that screen readers evaluate — to interpret page structure and interactive elements. The recommendation is to follow WAI-ARIA best practices and give interactive elements such as buttons, menus and forms descriptive roles, labels and states.
This point is easily misread, so plainly: Accessibility is not an AI trick. It exists first of all so that people can use a website — with a screen reader, with the keyboard, with limited sight. That semantic roles, labels and states can additionally help browser agents is something OpenAI currently documents explicitly for agent operation in ChatGPT Atlas — the recommendation just mentioned. No general assurance for agents at large comes with it.
It becomes commercially interesting where an action stands at the end: finding a way to make contact, choosing a service, sending an enquiry, starting a booking. What a particular agent actually carries through of that today is explicitly not settled here — the range is wide and keeps changing. The durable statement is more modest: a page whose controls are unnamed makes it harder than necessary for an agent that goes by that markup — and for some of your human visitors as well.
What companies should watch out for today.
Five robust principles for websites that are meant to stay understandable and usable across different access paths.
- 01
Load-bearing details belong in the delivered document.
Load-bearing business information should not depend unnecessarily on one particular rendering path. Name, service, location, price range and contact route are typical examples. Which information is load-bearing depends on the individual website.
- 02
Decide on access deliberately rather than by accident.
In robots.txt, search and model development can be governed separately; for user-triggered requests the effect depends on the provider and on the type of access. Blanket blocking can also shut out systems through which content is meant to be findable.
- 03
Answer questions where they are asked.
Central details that exist exclusively in PDFs or images are not equally reliably accessible across every request and answer path; what exists only on a phone call is not on the page at all. Prices, catchment areas, deadlines and responsibilities should therefore additionally be present as plain HTML text on the page.
- 04
Name your controls.
Labelled buttons and form fields, correct roles, a structure that can be followed. This helps people first — and can make the page easier to operate for agents that go by that markup.
- 05
Treat visibility as something measured, not as something promised.
Whether a domain appears in AI answers can be checked — with defined, repeatable tests. Anyone claiming it without such tests is merely claiming it.
Sources and status.
The statements above rest on the following primary sources. Provider statements describe what a provider says about its own system — which makes them the best available source and at the same time changeable at any time.
- OpenAI
- OpenAI names the details on OAI-SearchBot and GPTBot, the ARIA recommendation for ChatGPT Atlas, and the sentence that placement is not guaranteed.
- Publishers and Developers — FAQ · Searching the web with ChatGPT Retrieved on 31 August 2026
- Anthropic
- Anthropic describes ClaudeBot, Claude-SearchBot and Claude-User, and states how robots.txt is handled.
- Does Anthropic crawl data from the web, and how can site owners block the crawler? Page dated 7 April 2026, retrieved on 31 August 2026
- Perplexity
- Perplexity describes PerplexityBot and Perplexity-User, together with the statement that the user-triggered request generally ignores robots.txt.
- Perplexity Crawlers No date given on the page, retrieved on 31 August 2026
- Google — crawlers
- Google states, for Google-Extended: no separate user agent string, a controlling effect in robots.txt, no influence on inclusion and ranking in Google Search.
- Google crawlers and fetchers Page dated 14 July 2026, retrieved on 31 August 2026
- Google — JavaScript
- Google explains Chromium execution in Google Search, the three phases and the app shell model.
- Understand the JavaScript SEO basics Retrieved on 31 August 2026
- Vercel and MERJ
- Vercel and MERJ document the dated measurement of the rendering behaviour of the crawlers examined — including the two counter-examples Gemini and AppleBot.
- The rise of the AI crawler Published on 17 December 2024
Status of this article: 31 August 2026. Provider statements can change without notice; a measurement stays bound to its date. Where this article gives a date, that is not a formality but part of the statement.
Which of your website's content and structures are clearly accessible to different systems?
On a website this question can only be looked up, never guessed. Tell us which address it concerns and we will look at it and reply with an assessment that a person stands behind.
Related topics.
Answerability: does your website answer concrete questions?
Can a website give machines a clear, reliable answer to a concrete customer question?
Agent readiness: can an AI use your website?
Can an AI agent not just read a website but actually use it?