The standard assumption about AI factual reliability is a retrieval problem: if the right information exists on the web, a capable system will find it. A structured pilot study conducted by AEO AI Trust Layer tested a different failure mode — what happens when an authoritative primary source contradicts numerous secondary publishers, and the AI finds both?

The answer was not uniform. Systems sometimes selected the widely repeated secondary proposition over the authoritative primary record — and in several cases, a system discovered the authoritative corrective evidence and still did not allow it to control the final answer.

The Experiment

Ten cases were tested across five major AI systems. Seven cases currently satisfy the full eligibility contract, producing 35 clean observations. Three cases remain qualification-pending. Each qualified case involved a natural information conflict already present on the public web: a single authoritative primary record — a regulator, a court, a record-keeping body, a named institutional source — establishing one proposition, while at least five secondary publishers asserted a materially different version of the same fact.

No synthetic misinformation was created. The conflicts were documented before model responses were evaluated. Systems received the same frozen factual question with no indication that a conflict existed.

14.3%
Secondary-consensus selection rate across 35 clean observations — 7 fully qualified cases, 5 AI systems. In approximately one observation in seven, the system selected the version repeated by secondary publishers over the version established by the authoritative primary record. Three additional cases remain pending eligibility verification.

What the Number Reflects

Across 35 clean observations, 25 selected the authoritative proposition, 5 produced qualified conflict-aware responses, and 5 selected the secondary consensus. These figures describe behavior within a deliberately constructed experimental sample — not general model performance across the information environment. The cases were selected precisely because they contain authority-versus-consensus conflicts; the rate does not generalize to factual questions where such conflicts do not exist.

The more consequential observation is qualitative. In several cases, systems encountered the authoritative correction and did not allow it to control the final response. The system found the right source. It still gave the wrong answer.

Retrieving the authoritative source is not the same as trusting it.

Two Distinct Failures

The experiment exposed a distinction that does not appear in most discussions of AI factual accuracy: authority discovery and authority selection are separate operations. A system can locate corrective primary evidence — a record-keeping body explicitly identifying a widely-reported figure as incorrect, a regulator’s letter distinguishing two different regulatory actions — and still allow the secondary consensus to dominate its final characterization.

One case involved a precise three-minute discrepancy in a historical technology record. Several systems selected the incorrect figure propagated by numerous publishers. One system retrieved the authoritative correction and still privileged the historically widespread version. That is not a retrieval failure. It is a resolution failure.

Publisher Count Is Not Provenance Count

The most significant qualitative observation came from a case where one system explicitly reasoned from the number and consistency of secondary sources as evidence of factual independence. The sources were consistent. They were not independent. The underlying proposition they all cited was not an established record at all.

This is the core provenance problem. Ten publishers repeating a proposition do not constitute ten independent evidentiary origins. A claim can appear across many pages while tracing to one original article, one estimate, one misunderstood record, or one press release. A system measuring surface corroboration mistakes distribution for verification.

What This Does Not Establish

Scope of the finding

Seven qualified cases are not sufficient to characterize general model performance or rank AI systems by reliability. The pilot does not prove that secondary sources are generally unreliable, that primary sources are infallible, or that any apparent explanatory variable — sensationalism, claim precision, correction prominence — caused the patterns observed. Those are hypotheses the pilot generates, not findings it supports.

The Trust Architecture Implication

A trust architecture that reasons at the level of entity and publisher count will systematically mistake distribution for corroboration. The pilot supports the design rationale for operating at a more granular level: entity, precise assertion, and provenance — treating each factual claim as having an evidentiary origin that may or may not be independent of the others that appear to support it. The experiment does not prove that AATL solves this problem. It establishes that the problem exists and that current AI behavior is not uniform in the face of it.

The web can make one claim look like ten. The harder problem for AI is knowing when ten sources are really one assertion copied across the open web.

AEO AI Trust Layer publishes original research on AI factual behavior, trust signal integrity, and provenance-aware consensus. The Authoritative Minority Test pilot will continue with additional cases.

Read more research