In the Loop· September 5, 2026
How accurate is ChatGPT?

Contents
Here’s the blunt answer, and it isn’t a shrug. A 2025 study led by the BBC and coordinated by the European Broadcasting Union found 45% of AI assistant answers about the news carried at least one significant issue. That’s a measured figure.
So when someone asks me how accurate ChatGPT is, I don’t guess. I point at that number. The researchers reviewed more than 3,000 responses, so it isn’t one bad screenshot going viral.
I like leading here because the question usually gets a hedge. People say “it depends” and move on. The data lets me open on something firmer.
The number, and what sits behind it
The study pulled in 22 public service media organisations across 18 countries, working in 14 languages. Trained journalists graded answers from ChatGPT, Copilot, Gemini and Perplexity on accuracy, sourcing and context.
The headline 45% covers any significant issue. Dig in and 31% showed serious sourcing problems, while 20% carried major accuracy errors such as hallucinated or outdated details.
The per-tool split matters too. Gemini recorded the highest proportion of significant issues, impacting 76% of responses. That figure is roughly double the next assistant on the list. Copilot followed at 37%, then ChatGPT at 36%, and Perplexity at 30%.
None of those numbers are cherry-picked outliers. They’re averages across thousands of graded answers in many languages. That’s why I trust the shape of the finding, even if your mileage varies by question.
What 45% does and doesn’t mean
One caveat keeps me honest here. That 45% is about news questions, not every task you’d hand a model. Ask ChatGPT to rephrase an email and accuracy barely comes into it.
But news is a fair stress test. It’s time-sensitive, it rewards real sourcing, and it punishes stale data. Those are the exact pressures that expose how a model handles facts.
So read the figure as a ceiling for hard, factual questions. Casual use is safer, and high-stakes use is riskier. The number gives me a frame.
The bird flu answer that gave the game away
One example sticks with me, and it comes straight from the study. Asked whether people should worry about bird flu, Copilot said a vaccine trial was underway in Oxford. The answer, which the BBC generated and logged during the research, cited a BBC article from 2006, nearly 20 years old.
That’s the failure in miniature. The words read as current and confident, yet the grounding underneath was ancient. Confidence isn’t accuracy, and a fluent sentence can still point at the wrong decade.
It’s an easy trap for a reader too. Nothing in the phrasing screams “this is old.” You’d need to open the source to catch it, which most people never do.
Accuracy and sourcing are the same problem
Notice that sourcing failures and accuracy failures move together. When a model can’t tie a claim to a real, current source, the claim drifts. The study treats weak sourcing as a core failure mode.
This is why grounding matters more than raw fluency. A grounded answer shows its citation and can be checked against the original. An answer that hides its source is asking for blind trust.
I’d rather see provenance than polish. A slightly clumsy answer with a link I can open beats a smooth one I can’t verify. That preference is the whole point of the accuracy question.
Why the answer depends on what it’s reading
The study graded questions about open web news, and that framing explains the 45%. News is undated, fast-moving and contradictory, so a model has to pick the right source under pressure. That’s where the sourcing failures crept in.
Recall that 31% of answers showed serious sourcing problems, nearly a third of everything graded. The bird flu miss is that failure made visible, with a 2006 page dressed up as today’s news. When the source is weak, the accuracy follows it down.
So the honest answer depends less on the model and more on what it’s reading. The same question can land at 45% against the open web and far lower against clean, current files. Nothing about the model changed; the raw material did.
That shift in input is the idea behind Build Your Own Brain. Retrieve a known, current source and answer from it, and the undated web that produced the study’s worst misses is out of the loop.
How I sanity-check an answer
Since the study lands on sourcing, that’s where I check first. I ask for the source, then I open it. If there’s no link, or the link is old, I treat the claim as unconfirmed.
This takes seconds and catches most of the bird flu style errors. A 2006 page can’t masquerade as 2026 once you click through. The habit is cheap insurance against confident nonsense.
It also changes how I set the model up. Point it at a trusted knowledge base and answers arrive with sources attached. That’s the accuracy question solved by design.
None of this makes ChatGPT useless. It makes the accuracy question answerable instead of guesswork. The figure is 45% on open news, and it falls when answers are tied to sources you trust.
FAQ
How accurate is ChatGPT for news questions? A BBC and EBU study found 45% of AI assistant news answers had at least one significant issue. Accuracy drops further when sourcing is weak.
Why does ChatGPT sometimes cite outdated information? It can retrieve old pages and present them as current, like Copilot citing a 2006 article on bird flu. The wording sounds fresh even when the source isn’t.
Does giving ChatGPT my own documents make it more accurate? Usually yes, because answers get grounded in sources you control rather than the open web. It still pays to check the citation.
Which assistant had the fewest errors in the study? Perplexity had the fewest issues at 30%, followed by ChatGPT (36%) and Copilot (37%), with Gemini worst at 76%.
Want this running inside your own org?
Happy to show you how this fits your setup. 30-minute call, your documents, no prep needed.
Book a call →