By breaking the link between a claim and the thing that supports it, and by charging the cost of the break to the newsroom rather than to the machine that made it. Journalism's product was never the sentence. It is the chain that runs from an assertion back to a document, a recording or a person who can be asked again. Generative systems reproduce the sentence with ease and the chain badly, and the largest study yet run on the question found exactly that shape.
This page is built on one dataset, because for once there is a good one. Everything else here is inference and is marked as such.
Forty-five per cent, and the fault is attribution#
In October 2025 the BBC and the European Broadcasting Union published News Integrity in AI Assistants. Twenty-two public service media organisations across 18 countries and 14 languages put a shared set of 30 news questions, taken from questions audiences had actually asked, to the free consumer versions of ChatGPT, Copilot, Perplexity and Gemini. Responses were generated between 24 May and 10 June 2025. The assistants were anonymised and 271 journalists graded 2,709 responses on accuracy, sourcing, separation of opinion from fact, editorialisation and context.
Forty-five per cent of responses carried at least one significant issue. Including responses with lesser problems takes it to 81 per cent.
The composition is the finding. Sourcing was the largest single cause at 31 per cent, running at more than one and a half times the rate of accuracy at 20 per cent, with insufficient context at 14 per cent, editorialisation at 6 per cent and failure to separate opinion from fact at 6 per cent. The report defines the sourcing category to include information not supported by the source cited, no sources at all, and incorrect or unverifiable sourcing claims.
Read that against the way this problem is usually described. The public story about AI and news is fabrication: the invented quote, the event that did not happen. Fabrication is in there, forming the smaller half. The larger half is an answer that is broadly true and cannot be traced, or is traced to something that does not say it. A reader has no way to tell those apart, and neither does a search ranking.
One assistant carried most of the gap#
The averages hide a spread wide enough to change what the study means.
- Gemini: significant issues in 76 per cent of responses.
- Copilot: 37 per cent. ChatGPT: 36 per cent. Perplexity: 30 per cent.
Almost all of Gemini's distance from the others is one category. Its significant sourcing issues ran at 72 per cent, against 24 per cent for ChatGPT and 15 per cent for both Copilot and Perplexity. Forty-two per cent of Gemini responses provided no direct source at all, meaning no URL a reader could open. On accuracy the four were close together, all between 18 and 22 per cent.
A four-times difference between products built on comparable technology, concentrated in one behaviour, is a design decision rather than a limit of what the technology can do. Whether an answer carries a link is a choice somebody makes about the interface. That is worth holding onto, because the 45 per cent describes a current setting rather than a fixed property of machine-generated news answers.
The reputational cost is charged to the byline#
Of the responses that drew on a participating broadcaster's content, 15 per cent misrepresented it, introducing significant inaccuracies including in direct quotes. A further 6 per cent added editorialisation of the assistant's own that a reader would take as the broadcaster's view. Across all responses containing a direct quote, 12 per cent had significant problems with the accuracy of that quote.
The report pairs this with companion audience research it commissioned from Ipsos, and that pairing is the part editors should read twice. It reports that audiences blame AI providers for these errors and hold the media organisations named in the answer responsible, and that 42 per cent of adults say they would trust an original news source less if an AI news summary of it contained errors. Its own summary of the companion work:
Errors made by third-party Gen AI tools create direct reputational exposure for the sources they cite.
A publisher is therefore exposed to a quality failure in a product it does not build, cannot inspect, is not paid by, and in most cases cannot opt out of being cited in. Nothing in the ordinary economics of publishing has that shape. The closest analogy is a supply chain where a distributor can relabel your goods, sell them badly, and have the complaint arrive at your door.
The audience-side numbers here come from research the report cites rather than from the graded dataset, and I have taken them as the report states them rather than from the Ipsos report itself.
What this dataset was not built to settle#
The authors are unusually direct about the limits, and three of them matter for anyone quoting the headline.
- It is not adversarial. The questions were not chosen to trip the assistants up and difficulty was not controlled. So 45 per cent sits somewhere inside what these products do rather than at either edge of it.
- The versions tested are gone. These were free consumer defaults in late May 2025. All four products have shipped new defaults since. The report's own BBC-to-BBC comparison, on a much smaller sample, found significant issues falling from 51 per cent in the earlier round to 37 per cent, which suggests movement in the right direction and says nothing certain about where the four sit today.
- Per-organisation samples are small. Around 30 core responses per assistant per organisation. The report says explicitly that comparisons between countries or languages should be viewed with caution, and it was not designed to make them.
And the question the study cannot reach at all: whether any of this changed what readers believed. Grading an answer against journalistic criteria is not the same as measuring what a person took away from it. Nobody has run that experiment at this scale.
Journalism automated its verification layer first#
Here is where this connects to the rest of this research, along a line the report itself does not draw.
The tasks a generative system does most readily in a newsroom are checking a claim against a document, summarising a long report, pulling the relevant quote, and producing a serviceable first draft. Those are also, almost exactly, the tasks a junior reporter was given for the first two years. Not because they were valuable output, but because doing them repeatedly is how somebody learns which sources hold and which do not, how a press release differs from a finding, and what a quote sounds like when it has been trimmed to mean something it did not mean.
This research calls that pattern missing rungs: the automated steps and the developmental steps are the same steps, and removing them costs nothing visible for several years. Journalism has the sharpest version of it, because the capability being built by those tasks is the capability the machine is worst at. The 31 per cent sourcing failure rate is a description of what an assistant cannot do. It is also a description of what a second-year reporter spends their time learning to do.
The economics push the other way. Search referral traffic is falling, and the report notes the Financial Times has said it saw a decline of 25 to 30 per cent in readers arriving via search, attributed there to a Guardian report rather than to the FT directly. Newsrooms under that pressure cut the layer that looks least like output. That layer is the checking.
The general form of this appears across professions on the verifier's discount. Journalism's specific version has a twist. Beyond being undervalued inside the building, the verification is now performed for free, badly and in public, by systems that put the newsroom's name on the result.
The number a newsroom does not have#
Every editor now knows the industry figure. Almost none know their own.
The EBU published a toolkit alongside the report setting out the method so that any organisation can run it on itself. That is the more useful half of the publication and it has had a fraction of the attention. The industry number tells a newsroom that assistants get news answers wrong. Running the method internally tells it something actionable: which assistants misrepresent our reporting, on which stories, in which language, and whether that changed after the last model update.
Thirty questions, four assistants, a handful of journalists grading blind. It is a week of work, repeatable quarterly, and it converts a talking point into a measurement.
Four things worth doing this quarter#
- Measure your own misrepresentation rate. Use the EBU toolkit rather than inventing a method, so the result is comparable to something.
- Treat sourcing failures as the primary risk, not fabrication. The monitoring most newsrooms have set up looks for invented facts. Most of the damage is real facts attributed to you that you did not report, and an unsupported claim reads as clean.
- Protect the checking work for juniors explicitly. If a trainee never opens the primary document, the newsroom is buying speed against a capability it will need when the assistant is confidently wrong about something that matters. Name it in the training plan or it will be cut, because it never looks like output.
- Ask who carries the reputational cost in every AI licensing conversation. The evidence says it arrives at the publisher. Contracts that treat attribution accuracy as a courtesy rather than a term are mispricing that.
Key sources
- BBC and European Broadcasting Union (2025). News Integrity in AI Assistants: An international PSM study. 21 October 2025. Authors James Fletcher (BBC) and Dorien Verckist (EBU). Full report, PDF.
- BBC and European Broadcasting Union (2025). News Integrity in AI Assistants Toolkit. The method, published so that other organisations can repeat it.
- Brynjolfsson, E., Chandar, B. and Chen, R. (2026). Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence. Stanford Digital Economy Lab.
Every figure above was read from the report itself rather than from coverage of it. Two exceptions are stated in the text: the audience-perception numbers and the Financial Times traffic figure are quoted as the report presents them, from BBC and Ipsos research and from a Guardian report respectively, neither of which was opened at its own source.
Related SuperSkills research#
On the failure mode, what an AI hallucination is and why AI sounds so confident. On who does the checking and what it costs them, who owns verification, the verifier's discount and the invisible work of oversight. On the training pipeline, missing rungs and entry-level jobs. On the method for reading any profession this way, deskilling risk by profession, and on the neighbouring cases, law, medicine and consulting.
About this research#
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The BBC and EBU report was read in full at the primary source and every percentage on this page is quoted from it. No figure is given anywhere here for junior reporter task composition, newsroom headcount or the rate at which checking work has been cut, because no published dataset was found for any of them, and the argument in the second half is therefore marked as inference rather than measurement.
Cite this
Hirji, R. (2026). How will AI change journalism? The SuperSkills Intelligence Company. Last reviewed 31 August 2026. thesuperskills.com/research/how-will-ai-change-journalism
