Abstrakte Darstellung von API- und Browser-Messungen, die bei derselben KI-Antwort zu unterschiedlichen Ergebnissen fĂĽhren.

API vs. Browser Scraping: What AI Visibility Tools Actually Measure

When two tools report different levels of “ChatGPT visibility” for the same brand, that does not necessarily mean one of them is wrong. They may simply be measuring different things.

One provider may use browser automation to access a chat interface and extract the response generated there. Another may submit the same question through an AI provider’s documented API. And even then, it makes a difference whether the model searches the live web or answers from its existing model knowledge.

All of these methods can produce a plausible answer. But as measurements, they are not the same.

This is one of the fundamental challenges with GEO and AI visibility tools: there is no fixed “ChatGPT ranking” that a tool simply needs to retrieve. The measurement method influences what ultimately appears in the dashboard.

For this article, I therefore did not pit different GEO tools against one another. I do not have my own comparative data on browser-scraping methods to support such a test. Instead, the focus is on the methodological question: What can be controlled through APIs, what does browser automation capture more accurately, and which method is better suited to robust time-series analysis?

The analysis is based on measurements from Achtung.app. Its AI measurements are collected through the documented APIs of OpenAI, Google, Anthropic, and Perplexity—not through automated chat interfaces.

A deep dataset rather than a broad snapshot

Since March 16, 2026, Achtung.app has stored or archived a total of 105,028 AI runs. Around 772,000 brand mentions have been extracted from them.

At first glance, this may sound like a very broad study. But that would be the wrong interpretation. The responses come from 202 tracked search queries for six brands in the DACH region, run repeatedly over time.

The dataset is therefore deep, not broad. It is particularly well suited to observing changes, technical differences, and fluctuations over time. It is not a sample of 105,000 different topics or companies.

This depth is especially useful when examining measurement methodology: repeatedly testing the same search intents reveals changes in platform behavior that can easily go unnoticed in a one-off query.

The limitations of the dataset are equally important:

  • Four of the six brands are my own or are closely associated with me. They account for 64.2% of all runs. As a result, the dataset is heavily concentrated on providers in the Berlin area and on AI visibility itself.
  • Mistral and Grok were also included until early May. The current set of four—OpenAI, Gemini, Perplexity, and Claude—has been in place since May 5, 2026.
  • OpenAI, Gemini, and Perplexity are queried daily. Claude is queried weekly and therefore carries considerably less weight in the dataset.

The figures below therefore describe the panel observed by Achtung.app. They should not automatically be treated as general platform benchmarks.

This limitation matters: without knowing the dataset, measurement conditions, and weighting, it is difficult to assess what a visibility score actually measures—or whether two points in time are comparable at all.

APIs and browsers answer two different questions

Browser scraping has one obvious advantage: it can more closely reflect what a person sees in the relevant end-user interface.

When a chat web application is accessed through automation, the object being observed is that specific web application. An API-based measurement, by contrast, observes a documented interface provided by the vendor. It therefore does not automatically reproduce the exact behavior of a logged-in ChatGPT, Gemini, or Claude session.

For monitoring, however, proximity to the user interface is not the only criterion. At least as important is which variables can be held constant over weeks and months—and documented afterward.

Documented APIBrowser automation / scraping
Object observedDefined interfaceSpecific web interface
ModelModel identifier used can be loggedBackend version may not always be reliably identifiable
Web searchControllable or verifiable, depending on the providerBehavior of the interface is observed
SamplingSettings are partially controllableUsually no direct control
PersonalizationCan be deliberately excludedDepends on session, login, and interface
Control over measurement conditionsHighAdditional UI and session variables
Proximity to the end-user experienceLowerPotentially higher

I therefore find the question “API or scraping—which is more real?” unhelpful. A better question is: Which object do I need to observe in order to support the claim I want to make?

If the goal is to find out what a specific logged-in interface shows a particular person today, browser-based measurement may be closer to the object of interest. If the goal is to compare changes over several months, as many measurement conditions as possible need to be controlled.

More important than API versus browser: did the model search at all?

A model can answer a question from its existing model knowledge or search the live web before responding. In the first case, the measurement primarily reflects which brands and associations are available to the model. In the second, it reflects which information the system currently finds on the web, selects for the query, and incorporates into its response.

For assessing current visibility, these are two different objects of measurement.

The four providers observed by Achtung.app do not behave identically in this respect:

ProviderWeb search in Achtung.app measurements
OpenAIWeb search is part of the measurement run
Anthropic / ClaudeWeb search is part of the measurement run
PerplexityThe Sonar API used provides web-grounded responses
Google / GeminiWeb search is available, but the model decides whether to use it for the response

OpenAI documents that web search can be made mandatory for an API request. OpenAI: Web Search

Anthropic also provides Claude with server-side web search and source citations. Anthropic: Web Search

Perplexity explicitly describes Sonar as an API for web-grounded responses. Perplexity: Sonar API

Gemini differs in one important respect: Google describes a process in which the model analyzes the query and decides whether a Google Search could improve the response. It searches only when it deems a search necessary. Google: Grounding with Google Search

This may sound like a technical detail, but it is critical for measurement. A response produced without a web search should not be silently combined with a search-grounded response when the intended object of measurement is current visibility on the web.

Achtung.app therefore treats a Gemini run as a visibility measurement only when web search can be technically verified. Responses that are not sufficiently grounded are not included as regular measurements in the analysis.

In this case, a missing measurement is more honest than treating an answer from model knowledge as though Gemini had searched the current web.

The model version belongs in the time series—but not necessarily in every methodology article

The next variable is the model itself.

In a web interface, the backend actually being used can change. Even if the interface displays a product or model name, that does not always guarantee the same technical revision over several months.

With an API, by contrast, the model identifier used can be logged for every run. Where a provider offers specific snapshots, a particular revision can also be fixed. Achtung.app therefore stores the model identifier used for each measurement.

I have deliberately omitted the specific model IDs from this methodological comparison. They are not essential to the question of “API or browser?” and change as the measurement system evolves. In empirical analyses whose findings depend on the model version tested, however, this information belongs in the methodology.

The reason is simple: a model change can affect response style, source selection, length, and brand recommendations. If it is not documented in a time series, a technical change may look like a change in brand visibility.

AI responses remain variable even under consistent conditions

API access does not turn generative responses into database queries.

Achtung.app keeps the controllable measurement conditions as consistent as possible across comparable runs. Even so, the responses are not deterministic. Search results, retrieval, backend routing, and providers’ internal processing can still vary.

The significance of this is illustrated by a previously published analysis of 3,313 control pairs collected between June 1 and July 18, 2026. It compared the overlap between the brands mentioned in two consecutive responses generated from identical wording.

Only pairs with identical prompt text qualified for the analysis. Of 15,945 possible pairs of runs, 3,605 met this condition. After excluding 292 pairs in which neither response mentioned a brand, 3,313 control pairs remained.

All 124 search queries tracked during this period contributed control pairs. The pair rate for individual queries ranged from 12.5% to 33.3%. No query was therefore excluded from the study entirely, although this does not mean that all queries were weighted equally.

PlatformAvg. brand overlapCompletely identicalNo brand in common
ChatGPT69%37%10%
Perplexity48.5%9.2%8.4%
Gemini40.7%9.7%24.3%

The three columns represent different metrics and do not add up to 100%. The first shows the average overlap between the sets of brands mentioned; the other two show the proportion of extreme cases among all response pairs.

Gemini has one additional peculiarity: a large share of the pairs with no brand in common occurred because one of the two responses did not mention any brand at all. These cases often do not reflect two entirely different recommendation lists, but rather a shift between mentioning and not mentioning a brand.

The primary source, including a methodology note, weekly trends, and exclusion criteria, is the analysis “AI Response Stability: ChatGPT, Gemini, and Perplexity”. Further interpretation is available in my blog post “Why Doesn’t ChatGPT Mention My Business? 7 Reasons and What 3,313 Answers Reveal”.

This leads to a fundamental conclusion for measurement methodology: A single AI response is a sample, not a ranking position.

This applies regardless of whether the response was collected through an API or a browser. Browser scraping does not eliminate the variability of generative responses.

Response length also distorts simple visibility scores

Another problem arises when tools compare raw brand mention counts across different providers.

In Achtung.app measurements collected between July 1 and August 8, 2026, average response lengths differed considerably:

ProviderAvg. response length
Gemini3,363 characters
OpenAI2,458 characters
Perplexity1,814 characters
Claude1,462 characters

On average, Gemini’s responses were therefore around 2.3 times as long as Claude’s. A longer text naturally offers more opportunities to mention companies, products, and sources.

Simply adding up absolute mentions therefore conflates visibility with response length. More meaningful metrics include the share of relevant queries in which a brand appears, coverage across multiple providers, or position-based values—each with a clearly defined denominator.

Once again, a score can only be interpreted if its denominator is clear.

APIs provide structured source metadata—but not the user’s personal chat session

For search-grounded responses, the sources disclosed by providers can be stored alongside the response. This makes it possible to examine which external pages were identified as sources for an answer and whether this source trail changes over time.

However, the completeness of the source information returned varied by platform in the dataset. Since August 6, Achtung.app has applied a stricter inclusion rule to Gemini: only responses with verifiable web search are counted as visibility measurements.

This is immediately apparent in the source metadata:

ProviderResponses with source URLs before August 6, 2026Since August 6, 2026
OpenAI99.7%99.7%
Perplexity99.5%100%
Gemini84.9%100%

The 100% figure for Gemini does not reflect an improvement in the platform; it is the result of the revised measurement rule.

Before the change, 15.1% of stored Gemini responses did not contain a source URL. In retrospect, however, this does not reliably establish whether Gemini had actually performed no web search. A missing source citation is an indication, not proof. The technical tracking of search usage only became reliable with the new measurement logic introduced on August 6. Earlier values must therefore be treated as unknown, not as zero.

This distinction is critical for a time series. A change in measurement methodology must not later be mistaken for a change in visibility.

The trade-off of the API approach should be stated just as clearly: Achtung.app deliberately measures without personal chat history, memory, login context, or individual personalization. This is a design choice made by the measurement system, not an inherent property of APIs. Conversation state could technically be supplied, but is intentionally excluded from the time series.

This is a limitation when asking, “What does this particular user see in their account?”

For the question, “How does a brand’s generic visibility change over time?”, however, the absence of personalization is helpful. A time series becomes less meaningful if the model, prompt, user history, and interface all change at once and it is no longer possible to distinguish which factor caused the result.

Browser scraping is not automatically wrong—it simply measures something different

It would be too simplistic to dismiss browser scraping as an inherently flawed method.

Anyone investigating what an end user actually sees in a particular product interface has a valid reason to observe that exact interface. Features, routing, presentation, and personalization may be active there in ways that an API does not reproduce identically.

The problem begins when this observation is turned into a universal metric such as “Your ChatGPT visibility is 42%” without disclosing the measurement conditions.

Browser-based measurement systems must also account for changes to the interface and DOM, session state, bot detection, and the relevant service’s terms of use. API-based systems trade greater proximity to the product interface for documented and more controllable access.

For Achtung.app, I deliberately chose the latter approach and documented the measurement methodology publicly: the actual AI measurements come from the providers’ official APIs, not from automated chat interfaces or third parties.

The reason is methodological rather than moral: if a value in a time series falls, ideally it should be the observed visibility that has declined—not an unnoticed change in how it was measured.

Conclusion: measurement methodology is what turns AI responses into a time series

AI visibility does not have a fixed ranking that can simply be retrieved like a traditional search position. Every measurement depends on which model was queried, whether it searched the web, which session context applied, how many times the query was run, and how the resulting text was analyzed.

APIs do not solve all of these problems. Most importantly, they do not automatically reproduce the personal experience of a logged-in user. But they make key measurement conditions more explicit and easier to document over time.

For long-term monitoring, I consider this essential.

The real question to ask an AI visibility tool is therefore not:

“Does it measure ChatGPT?

Instead, it is:

Which version of which platform is being measured under which conditions—and will the same metric still be comparable with today’s result three months from now?

Only once that question has been answered can we assess what a visibility score actually tells us.

FAQ

Does an API measure the same thing as the ChatGPT or Gemini web interface?

No. An API provides access through a documented interface. In the web interface, session history, personalization, the product interface, and internal model controls may also affect the result. The two methods therefore measure different things.

Is API-based measurement more reliable than browser scraping?

That depends on the objective. Browser scraping may more closely reflect a specific user interface. For long-term time series, however, APIs provide measurement conditions that are easier to control and document.

Why do AI visibility tools produce different results?

The tools may use different models, prompt variations, search capabilities, measurement frequencies, and evaluation methods. Two visibility scores can therefore only be compared when the measurement method and reference metric are known.

Is a single query sufficient to measure AI visibility?

No. Generative AI systems may mention different brands and sources even under comparable conditions. Individual responses are samples and should not be interpreted like fixed search engine rankings.

Does Achtung.app scrape the ChatGPT, Gemini, or Claude interfaces?

No. Achtung.app collects its core AI measurement data through the documented APIs provided by OpenAI, Google, Anthropic, and Perplexity. It does not use automated access to chat interfaces or third-party scraping services for this purpose.