ChatGPT alone is used by more than 900 million people every week. Many increasingly rely on AI to find suitable products, software, or service providers. The answers often appear well-founded and surprisingly definitive.
But what happens when you ask exactly the same question again just a few seconds later?
The answer may be significantly different. Companies that were initially recommended suddenly disappear. Other providers are added. In some cases, the two answers do not even have a single brand in common.
For companies, this means that a single AI answer is not yet a reliable basis for decision-making. Asking twice helps you make a better-informed choice.
The Same Question, Different Providers
Imagine a medium-sized company is looking for new CRM software and asks ChatGPT:
“Which CRM systems are suitable for a medium-sized company in Germany?”
The AI names five providers, describes their strengths, and recommends one solution particularly strongly.
This is a good starting point for further research. However, when the same question is asked again, the list may look different: one provider disappears, two new solutions are added, and suddenly a different system is preferred.
Which answer provides a more accurate picture of the market?
This is exactly the question I examined using data from Achtung.app.
What Achtung.app Measured
Achtung.app submits every tracked question twice. The wording, AI model, and technical settings remain the same. There is only about one second between the two requests.
It then compares which brands are mentioned in both answers.
The analysis covers 3,313 pairs of answers collected between June 1 and July 18, 2026. The average overlap between the brands mentioned was:
- ChatGPT: 69 percent
- Perplexity: 48.5 percent
- Gemini: 40.7 percent
ChatGPT therefore produced the most consistent brand lists. Even there, however, only 37 percent of the answer pairs contained exactly the same brands.
Gemini was particularly striking: in 24.3 percent of the answer pairs, there was not a single brand in common. Almost every fourth repetition therefore resulted in a completely different selection.
The study does not assess which recommendation was correct or better. It shows how significantly the selection of companies mentioned can change.
Why Do AI Answers Vary?
An AI system does not work like a database that retrieves a fixed, stored answer for every question.
The answer is assembled again for each request. Depending on the platform, sources are searched, content is selected, information is weighted, and suitable companies or products are identified.
Even small changes in this process can result in different brands being mentioned. An AI answer is therefore not a complete market overview, but rather a snapshot.
What Does This Mean for Business Decisions?
ChatGPT, Gemini, and Perplexity can speed up research. They help businesses find initial providers, develop selection criteria, and discover previously unknown solutions.
A pragmatic approach consists of three steps:
- Ask the same question at least twice.
- Change the wording slightly.
- Evaluate the providers mentioned using your own criteria.
When a company appears in several answers, that is at least a stronger signal than a single mention. AI can support the initial selection process, but it does not replace an assessment of costs, references, requirements, and contractual terms.
What Does This Mean for a Company’s Own AI Visibility?
Companies themselves are also increasingly testing whether ChatGPT and other AI systems recommend them.
When their brand is not mentioned, they may quickly conclude:
“The AI does not know our company.”
However, a single missing mention is not enough to support this conclusion. The company might appear in the second answer or when the question is worded differently.
Conversely, a single recommendation is not proof of consistently strong visibility. The next request might recommend a competitor instead.
AI visibility is therefore not a fixed ranking position. The decisive question is not whether a company appears once, but:
How frequently, and in what context, is a brand mentioned in response to relevant questions?
This question can only be answered meaningfully through repeated measurements.
Why Occasional Tests Can Be Misleading
Many companies define a set of questions, ask each of them once, and repeat the test a few weeks later.
If their own brand is absent from the first test but appears in the second, this may look like an improvement. In reality, both results could fall within the normal level of variation. Similarly, a later omission does not necessarily indicate a decline.
Without repeated measurements, it is difficult to determine whether visibility has genuinely changed or whether two individual snapshots are simply being compared.
Consistency Is Not the Same as Accuracy
A consistent system can provide the same incorrect answer twice. Two different answers may both be reasonable if several providers are suitable for the question.
Companies should therefore consider two issues separately:
- Is the answer factually correct and comprehensible?
- Does the recommendation remain consistent across repeated requests?
ChatGPT was more consistent than Perplexity and Gemini in the study. However, this does not automatically mean that every ChatGPT answer is more accurate or helpful.
Conclusion
AI recommendations often appear more definitive than they really are. Anyone using ChatGPT, Gemini, or Perplexity for business research should therefore repeat important questions and compare the results. Asking twice helps, but the answers still need to be verified.
A company’s own AI visibility cannot be assessed using a single request either. What matters is how frequently a brand is mentioned across different questions, wordings, and points in time.
This is exactly why Achtung.app submits every tracked question repeatedly. Only multiple answers can turn a snapshot into a recognisable pattern.
The full study on the answer consistency of ChatGPT, Gemini, and Perplexity contains all measured results and their development over time.







