Download the Full Study
News publishers are losing the human audience on which their business was built. We’ve found that what most of them are doing about it does not work on the models that matter most.
Two years into the AI search era, publisher strategy has settled into a single choice: block or allow. Publishers have written robots.txt files, signed licensing deals, and filed lawsuits, then waited to see what worked.
Between October 2025 and July 2026, Goodie recorded 31 million AI citations and audited access policies at 105 US and UK news publishers. Of 495,000 citations to 37 major news domains, 34 recorded any citations at all. The results say block-or-allow was the wrong question: what matters is which labs comply with a block in the first place.
What the Goodie Team Measured
Citation Data
Goodie monitors brand visibility across ChatGPT, Gemini, Claude, Perplexity, AI Overviews, AI Mode, Copilot, Grok, DeepSeek, Meta AI, and Amazon Rufus. The 31 million citations in this study were recorded between October 2025 and July 2026. A separate industry-index benchmark, built from neutrally sampled prompts, strips selection bias out of every ranking below.
Access Audit
Goodie classified all 105 US and UK news and editorial publishers by robots.txt posture toward 25 AI user agents, using live fetches in July 2026 alongside direct retrievals gathered during the research.
Agent Readiness
Goodie ran its Agent Readiness Audit against flagship publishers across eight categories, including crawl access, rendering, machine-readable structure, and authentication.
Key findings
- News is a small share of what AI cites. News makes up between 1.3% and 1.9% of all citations in any given month, drifting toward the low end. Publishers are competing over a narrow pool.
- Citations concentrate harder than search rankings ever did. The top five publishers take 66% of news citations. The top ten take 84%. Forbes alone holds 33%, more than the next four combined.
- Traffic rank does not predict citation rank. The New York Times is first in visits and sixth in citations. Fox News is top three in visits and holds 0.12%.
- Blocking works on the models that comply. ChatGPT cited nytimes.com zero times across the sample. Gemini: zero. Claude: once.
- Blocking does nothing to the models that don’t. NYT received 37,642 citations from Grok, 8,007 from AI Overviews, and 4,303 from Perplexity, a company it’s currently suing.
- A licensing deal moves citations inside one AI lab and almost nowhere else. AP blocks every OpenAI crawler and still draws 80% of its citations from ChatGPT, because licensed content travels through the contract, not the crawl.
Blocking, licensing, and compliance turn out to be one mechanism: the AI labs a publisher can block are mostly the ones that pay. The labs that pay nothing are mostly the ones that can’t be blocked.
What the Study Data Reveals
1. AI Barely Cites News At All
Across the 31 million citation sample, news domains earned 495,000 citations, between 1.3% and 1.9% of everything AI cited in a given month. The share drifted toward the low end across the window rather than growing.
That’s a narrow pool relative to how much attention publishers put toward winning it. Every ranking in this study, by publisher, by vertical, by AI engine, describes competition over that same small slice, not over AI citations broadly.

2. Citations Concentrate Harder Than Search Rankings
In Goodie’s neutral benchmark, Forbes holds 33% of all news citations. Reuters takes 10%, Business Insider 9%, CNBC 8%, and the Wall Street Journal 6.1%. The top five take 66%. The top ten take 84%.
This is more concentrated than search rankings ever were. A results page had ten blue links plus endless scrolling, so a page-two publisher could still pick up clicks and impressions. An AI answer has three to eight citation slots and nothing after that.

3. Traffic Rank Doesn’t Predict Citation Rank
The New York Times, first in visits at well over 400 million a month, is sixth in citations. CNN, second in visits, ranks sixteenth. Fox News, top three in visits, holds 0.12%. Axios, at roughly 23 million monthly visits, out-cites publishers twenty times its size.
Being visited and being cited are separate outcomes with separate rules, and most publishers still measure only the first one.

4. Citation Leadership Splits By Vertical, Not By Scale
The gap between visits and citations runs by industry as much as by publisher. Forbes leads ten verticals, including insurance at 86% of news citations in that category. Axios owns cybersecurity at 57%. AI models favor service content, rankings, comparisons, explainers, far more than they favor reporting, which is why a vertical-focused publisher can out-cite a much larger newsroom in its specific lane.

5. Every Model Cites Differently
Treating AI as a single channel is the most expensive assumption in the current playbook: in Goodie’s sample, Grok is the single biggest citer of news at 23% of all earned citations, more than ChatGPT and Claude combined.
AI Overviews follows at 19.5%, Gemini at 13.5%, ChatGPT at 11.9%, AI Mode at 10.6%, DeepSeek at 7.9%, and Claude at 5.1%, the lowest share among the major assistants.
Each engine cites for a different reason. ChatGPT’s citations concentrate around its licensing partners. Gemini and AI Overviews cite more broadly, favoring well-structured content. Grok cites at high volume with no published crawler and no opt-out. Perplexity cites broadly too, including publishers actively suing it.

6. Blocking Works On Models That Comply, But Not the Ones Who Don’t
The New York Times blocks more thoroughly than any other newsroom in the sample. It refuses all 13 major AI user agents and backs the block with a legal notice barring text and data mining.
Against the labs that follow robots.txt, the block holds. Across 31 million citations, ChatGPT cited nytimes.com zero times. Gemini: zero. Claude: once.
Against the labs that don’t, it does nothing. The same Times received 37,642 citations from Grok, which publishes no crawler documentation and no opt-out mechanism; 8,007 from AI Overviews, which can’t be refused without leaving Google Search entirely; and 4,303 from Perplexity, which Cloudflare documented running undeclared crawlers with spoofed browser user agents (Perplexity has disputed this characterization).
Together, Grok, AI Overviews, and DeepSeek account for roughly half the news citations in the sample. None of the three offers a functioning opt-out.
Publishers built their access controls around the labs willing to negotiate, while the labs with no opt-out and no deal cite at will. We call this the access paradox.
There’s a cost on the other side too. Wharton and Rutgers researchers tracked the staggered rollout of robots.txt blocking that began in mid-2023 and found a roughly 7% decline in weekly visits within six weeks of blocking, a decline that showed up in human browsing data rather than bot metrics.

7. Licensing Moves Citations in One Lab
AP blocks every OpenAI agent in its robots.txt. AP also signed an OpenAI licensing deal in July 2023. The result: 80% of AP’s citations come from ChatGPT, the very model whose crawlers it blocks.
That is the whole point. When a publisher licenses its content, that content reaches the model through the deal, not through the crawl. Blocking the crawler does nothing to stop it.
The Wall Street Journal shows the same mechanism from the other direction. Its robots.txt blocks every bot by default, then names specific exceptions it allows through, including GPTBot and OAI-SearchBot, the crawlers tied to its reported $250 million OpenAI licensing deal. The deal’s fingerprint is sitting right there in a public text file.
But a deal only converts to citations where the lab’s pipeline actually reads the licensed feed. Business Insider and Politico both sit under Axel Springer’s OpenAI deal. Business Insider ranks third in Goodie’s benchmark. Politico rounds to zero. Fox News, CNN, People, and USA Today all hold Meta licensing deals, and Meta AI produces under 3% of news citations in the sample.
And a deal isn’t required. Reuters has no OpenAI agreement and ranks second overall, on open access and wire authority alone.

8. Most Publishers Are Hard For Machines to Read
Goodie’s Agent Readiness engine found no Web Bot Auth keys and no OpenAPI discovery on any flagship publisher audited. Those are the two building blocks that let a site verify a bot’s identity rather than trusting a spoofable user-agent string, and expose licensed content programmatically rather than leaving it to be scraped.
The New York Times scored 53% on agent readiness. Forbes 45%. The Wall Street Journal 15%. One top-five paper serves crawlers a 43-character JavaScript shell.
llms.txt adoption across the audit is zero. No major AI lab has confirmed using llms.txt for retrieval, so this is a readiness gap worth tracking rather than an urgent fix on its own.
The industry is litigating over the value of its content while remaining difficult for the machines that would cite it to read.
How Publishers Should Respond
This Week: Rewrite Robots.txt By Function
Set separate rules for training, search, and retrieval, per lab. If you block training, still allow the retrieval agents that carry citations. And know what robots.txt cannot do: it doesn’t affect Grok, DeepSeek, AI Overviews, or Copilot. For those, the options are CDN enforcement, gateways, licensing, or litigation
This Quarter: Check What the Crawler Actually Receives
Most AI crawlers don’t execute JavaScript, so a page that looks complete to a reader can arrive nearly empty to a bot. In Goodie’s audit, the most-cited publishers still scored At Risk on agent readiness, echoing the near-empty shell one top-five paper serves crawlers, described earlier. Schema and clean attribution are table stakes. The failure that actually costs citations is content that only exists after client-side rendering.
This Half: Put a Price on the Machine Lane
Gateway metering and pay-per-crawl are live infrastructure now. No lab has committed to paying through them yet, which is why setting terms early costs little.
This Year: Bring Data to the Licensing Table
Publishers walk into AI licensing talks knowing their traffic numbers but not their citation share by model, the number that actually determines what a deal is worth. Measure it before pricing one, not after.
Throughout: Build the Audience You Own
Every move above manages exposure to platforms publishers don’t control. Subscriptions, newsletters, apps, and events are the share they do control. The publishers weathering this best treat direct audience as the goal and AI visibility as the funnel into it, not the other way around.
Read the Full Study
Every publisher in this dataset is either negotiating an AI licensing deal right now or about to be. The full report is the number to bring to that table: citation share by model, publisher by publisher, not just traffic.
Inside: the six robots.txt postures across the top 50, the complete publisher-by-model citation matrix, a lab-by-lab scorecard on compliance and payment, the full deal map across seven buyers, the litigation record, and the specific tripwires that should change a publisher’s posture the moment they trip.
Full methodology and every source behind every figure included.
AI Citations & News Publishers: FAQs
Only against labs that honor the protocol. In Goodie’s sample, ChatGPT, Gemini, and Claude respected The New York Times’ block almost completely. Grok, AI Overviews, and Perplexity did not, and none of the three currently offers a functioning opt-out.
No. A licensing deal routes content to a lab through the contract, not the crawl. AP blocks every OpenAI crawler and still draws 80% of its citations from ChatGPT, because the licensed feed reaches the model regardless of what robots.txt says.
Grok cites news at the highest rate among engines tracked in this study, at 23% of all earned citations, followed by AI Overviews at 19.5% and Gemini at 13.5%. Claude cites news the least among major assistants, at 5.1%.
Not currently. Adoption across Goodie’s audit of flagship publishers was zero, and no major AI lab has confirmed using llms.txt for retrieval.
No. The New York Times ranks first in visits but sixth in citations. Fox News ranks in the top three for visits but holds just 0.12% of citation share, while Axios, a fraction of its size, out-cites publishers twenty times larger.