2026-10-11 16:38 UTC

GSA launched America.gov as an AI 'front door' over 29,000+ government websites claiming up-to-date answers to any question, and a night-one journalist stress test reports 9/15 correct, 6 incomplete, zero hallucinations β€” whether the deployment sustains accuracy under growing independent testing or gets scaled back decides whether AI front doors become the standard citizen interface to US government information.

state: corroboratedheat: mediumuncertainty: mediumconvergesscott: mediumgovernment-ai-deployment ai-evaluationGSAhannotek
Surfaced 2026-09-30T23:54:08Z β€” We stress-tested America.gov on night one with 15 questions from our own reporting. 9 correct, 6 incomplete, zero hallucinations. β€” The case's meaning shifted from 'launch plus one night-one test' to a contested live deployment: independent cross-platform probing shows both genuine utility and documented trivial-request misbehavior, and a single-outlet report claims the chatbot was politically reprogrammed to stop fact-checking β€” putting political steering, not just technical accuracy, on the live question. Three independent evidence lines (journalist stress test, HN probing, news report of reprogramming) pass corroborated on substance; the prior low heat label under-rated day-one top-decile cross-platform velocity (96.8th percentile, accelerating) and the just-opened national-news political angle, which is exactly the window where high heat matters.

What is this?

GSA β€” the US federal agency that manages government-wide technology and procurement β€” is aggressively deploying AI under the White House's July 2025 AI Action Plan: it launched USAi in August 2025 as a shared platform for agencies to test frontier models, prioritized FedRAMP authorization for conversational AI, and is exploring internal chatbots that can draw on multiple vendors (OpenAI, Anthropic, Google). Third-party coverage widely frames AI chatbots as the 'new digital front door' to government services, with prior deployments (NYC's MyCity, NSF grant chatbot, Air Force NIPRGPT) showing both promise and documented accuracy problems. The supplied snippets do not directly corroborate the specific America.gov launch across 29,000+ sites or the night-one 15-question stress test β€” those rest on the case's own evidence title β€” but a commentary piece argues exactly this failure mode: AI front doors give confident but wrong answers when government content is fragmented, buried in PDFs, or not machine-readable. Whether the deployment sustains accuracy under growing independent testing is therefore the live question, and the snippets establish that GSA's broader AI push (USAi, FedRAMP 20x) is real and ongoing into 2026.

Why it matters to Scott

The world is independently enacting Scott's grounded-front-door arguments: GSA shipped a citizen-facing answer engine and a journalist stress-tested it night one on production reporting questions (his capability-audit posture), and the result's shape β€” zero hallucinations but 40% incomplete β€” is a public instance of his answer-failure-classes split between honest coverage/navigation incompleteness and fabrication. His route-invariant-grounding framework also says 15 single-route correct answers certify little about actual grounding, so the live sustained-accuracy question over fragmented, PDF-buried government content (systemic-wrongness risk) is a citable, ongoing test case for the exact claims his evaluation and witness-not-oracle territory makes.
ip:concept.answer-failure-classesip:framework.route-invariant-groundingip:concept.capability-auditip:concept.systemic-wrongnessip:concept.interface-ladder-for-knowledgeradar:concept.government-airadar:concept.llm-evaluationradar:concept.ai-searchradar:concept.ragradar:austria-govgpt-sovereign-rollout
queries asked of Scott's wikis
  • RAG grounding and citation quality evaluation for production chatbots
  • eval harness design and stress-testing methodology for deployed AI systems
  • machine-readable structured publishing and docs-for-AI-retrieval patterns
  • answer-engine front door versus traditional site search as an interface pattern
  • hallucination measurement in grounded domain corpora at scale
  • government AI adoption procurement and FedRAMP conversational AI

Measured heat

now 0 pts/hpeak 96 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 265h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-30 15:08⭐ origin directly observedWe stress-tested America.gov on night one with 15 questions from our own reporting. 9 correct, 6 incomplete, zero hallucinations.
hannotek on r/artificial
β€”
09-30 19:26first on r/artificial Β· published Β· +4.3hTrump Reprograms Government AI Chatbot to Stop Fact-Checking His Lies. Trump officials seem to have realized their AI chatbot was correcting the president’s biggest lies.
esporx
β€”
09-30 19:34first on hacker news Β· published Β· +4.4hAmerica.gov goes crazy on "play Minecraft"
thoughtfullyso
β€”
10-03 17:32first on r/singularity Β· published Β· +74.4hAmerica.gov gets nerfed: no longer contradicts Trump's false narratives [4:44, France 24 English]
Competitive_Travel16
β€”
09-30 15:08amplified on r/artificialreddit.post.1wu7f9w
hannotek
peak 10 Β· 20 comments Β· 4% of case engagement
09-30 19:26amplified on r/artificialreddit.post.1wueara
esporx
peak 200 Β· 21 comments Β· 31% of case engagement
09-30 19:34amplified on hacker news πŸ‘‘hn.story.49913255
thoughtfullyso
peak 127 Β· 44 comments Β· 43% of case engagement
10-02 15:30amplified on hacker newshn.story.49934631
paimapi
peak 47 Β· 29 comments Β· 19% of case engagement
10-03 17:32amplified on r/singularityreddit.post.1wwsxhp
Competitive_Travel16
peak 13 Β· 4 comments Β· 2% of case engagement
09-30 15:20our radar first saw it Β· +0.2hdiscovery anchor: reddit.post.1wu7f9wβ€”
09-30 23:38reached heat=high Β· +8.5h Β· via ledgerβ€”β€”
pace: p84 vs 1188 stories at the 168h mark (now 265h old) β€” ahead of big-tech-ai-guarantee-exposure (1.0x), behind qwen38-27b-16gb-quant-benchmark (1.0x)

Evidence (5) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐We stress-tested America.gov on night one with 15 questions from our own reporting. 9 correct, 6 incomplete, zero hallucinations.
artificial
hannotek1020
🟠 redditTrump Reprograms Government AI Chatbot to Stop Fact-Checking His Lies. Trump officials seem to have realized their AI chatbot was correcting the president’s biggest lies.
artificial
esporx20021
🟧 hnAmerica.gov goes crazy on "play Minecraft"thoughtfullyso12741
🟧 hnA 20-year-long permanent cookie: America.gov and trackingpaimapi4729
🟠 redditAmerica.gov gets nerfed: no longer contradicts Trump's false narratives [4:44, France 24 English]
singularity
Competitive_Travel16134

Interpretation history

Decision trace