Haus Research claims roughly one-third of audited Perplexity citations attached to numeric claims do not contain the cited number, implying AI-search and RAG systems need claim-level citation verification.
state: expiredheat: lowuncertainty: highconvergesscott: mediumcitation-verification ai-search rag-reliabilityHaus ResearchPerplexity
What is this?
Haus Research reportedly audited Perplexity’s citations attached to numeric claims and found that roughly one-third of the cited sources did not contain the number they were presented as supporting. The finding concerns citation faithfulness—not merely whether a citation exists or is topically relevant—and points to a need for claim-level verification in AI-search and RAG pipelines. The supplied results corroborate the broader citation-support problem, but they do not provide Haus Research’s original methodology, sample size, audit period, or detailed results, so the precise rate cannot be independently assessed here.
Why it matters to Scott
Haus Research’s reported audit supplies external empirical support for Scott’s existing claim that citations must be verified against exact claim-level evidence, not accepted merely because they exist or are topically relevant. This creates a dated-receipts publishing opportunity around Evidence Packages, Provenance-Coupled Work, and verbatim evidence anchoring, but the missing methodology and sample details limit how strongly the one-third figure can be used.
ip:concept.evidence-packageip:framework.provenance-coupled-workip:source.witness-not-oracle-ebookdev:concept.verbatim-source-evidence-anchoringdev:concept.proof-of-read-citation-gateradar:concept.claim-verificationradar:veruscite-citation-checking
queries asked of Scott's wikis
- claim-level citation verification in RAG
- citation faithfulness versus source relevance
- numeric claim validation in AI search
- RAG evaluation beyond recall and precision
- automated entailment checks for cited claims
- provenance and evidence tracing in knowledge systems
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-09-05T19:27:13Z
The stale review adds no substantive evidence, leaving the reported one-third failure rate as unverified testimony rather than a usable benchmark. With no identified methodology release, replication, or other confirming event expected, this episode has faded; expiration does not disprove the citation-faithfulness concern.
2026-09-03T18:46:47Z
The refreshed discussion remains anecdotal amplification, with an unsupported suggestion of coordinated promotion adding provenance skepticism but no substantive counterevidence. Without methodology, sample details, or an independent audit, the one-third figure remains an attributed lead rather than a usable benchmark.
2026-09-02T18:33:31Z
Refreshed discussion remains anecdotal amplification of the known citation-faithfulness problem and adds no independent audit, methodology, or implementation evidence. The one-third figure should still be treated as an attributed lead rather than a benchmark.
2026-09-02T15:47:34Z
New comments add anecdotal reports of citation mismatches, including a similar complaint about Gemini, but they do not independently validate Haus Research’s rate or methodology. The case remains a useful claim-verification lead rather than a quantified reliability benchmark.
2026-09-02T15:35:57Z
grounded: converges/medium — Haus Research’s reported audit supplies external empirical support for Scott’s existing claim that citations must be verified against exact claim-level evidence
2026-09-02T15:32:42Z
case created — The published audit makes a bounded quantitative reliability claim with direct implications for citation-grounded systems.
Decision trace
- 09-06 05:27expireThe stale review adds no substantive evidence, leaving the reported one-third failure rate as unverified testimony rather than a usable benchmark. With no identified methodology release, replication,
- 09-06 05:27alert_silentThere is no new consequential delta or expected near-term confirmation to justify interrupting Scott. Reopen if audit methods, inspectable examples, or independent replication arrive.
- 09-06 05:27alert_routeThere is no new consequential delta or expected near-term confirmation to justify interrupting Scott. Reopen if audit methods, inspectable examples, or independent replication arrive.
- 09-04 04:46repriceThe refreshed discussion remains anecdotal amplification, with an unsupported suggestion of coordinated promotion adding provenance skepticism but no substantive counterevidence. Without methodology,
- 09-04 04:46alert_silentThe new delta is additional discussion rather than validation, rebuttal, or implementation evidence, so it does not change Scott's decisions and can wait for the next briefing.
- 09-04 04:46alert_routeThe new delta is additional discussion rather than validation, rebuttal, or implementation evidence, so it does not change Scott's decisions and can wait for the next briefing.
- 09-03 04:33repriceRefreshed discussion remains anecdotal amplification of the known citation-faithfulness problem and adds no independent audit, methodology, or implementation evidence. The one-third figure should stil
- 09-03 04:33alert_silentThe new delta consists only of repetitive user anecdotes and general discussion; it neither validates Haus Research’s rate nor changes the practical implications already routed, so it can wait for the
- 09-03 04:33alert_routeThe new delta consists only of repetitive user anecdotes and general discussion; it neither validates Haus Research’s rate nor changes the practical implications already routed, so it can wait for the
- 09-03 04:21sensor_dirtycomment_update
- 09-03 03:22sensor_dirtycomment_update
- 09-03 02:21sensor_dirtycomment_update
- 09-03 01:47repriceNew comments add anecdotal reports of citation mismatches, including a similar complaint about Gemini, but they do not independently validate Haus Research’s rate or methodology. The case remains a us
- 09-03 01:47alert_silentThe new delta is discussion amplification and low-detail user testimony, not a consequential new finding; the prior audit has already been routed, and nothing here warrants interrupting the next brief
- 09-03 01:47alert_routeThe new delta is discussion amplification and low-detail user testimony, not a consequential new finding; the prior audit has already been routed, and nothing here warrants interrupting the next brief
- 09-03 01:43alert_shadowThe published audit is a timely external receipt for claim-level citation verification and could immediately inform Scott’s RAG evaluation tests and writing on evidence packages. The report’s existenc
- 09-03 01:43alert_routeThe published audit is a timely external receipt for claim-level citation verification and could immediately inform Scott’s RAG evaluation tests and writing on evidence packages. The report’s existenc
- 09-03 01:35groundHaus Research’s reported audit supplies external empirical support for Scott’s existing claim that citations must be verified against exact claim-level evidence, not accepted merely because they exist
- 09-03 01:32createThe published audit makes a bounded quantitative reliability claim with direct implications for citation-grounded systems.