GSA launched America.gov as an AI 'front door' over 29,000+ government websites claiming up-to-date answers to any question, and a night-one journalist stress test reports 9/15 correct, 6 incomplete, zero hallucinations β whether the deployment sustains accuracy under growing independent testing or gets scaled back decides whether AI front doors become the standard citizen interface to US government information.
state: corroboratedheat: mediumuncertainty: mediumconvergesscott: mediumgovernment-ai-deployment ai-evaluationGSAhannotek
Surfaced 2026-09-30T23:54:08Z β We stress-tested America.gov on night one with 15 questions from our own reporting. 9 correct, 6 incomplete, zero hallucinations. β The case's meaning shifted from 'launch plus one night-one test' to a contested live deployment: independent cross-platform probing shows both genuine utility and documented trivial-request misbehavior, and a single-outlet report claims the chatbot was politically reprogrammed to stop fact-checking β putting political steering, not just technical accuracy, on the live question. Three independent evidence lines (journalist stress test, HN probing, news report of reprogramming) pass corroborated on substance; the prior low heat label under-rated day-one top-decile cross-platform velocity (96.8th percentile, accelerating) and the just-opened national-news political angle, which is exactly the window where high heat matters.
What is this?
GSA β the US federal agency that manages government-wide technology and procurement β is aggressively deploying AI under the White House's July 2025 AI Action Plan: it launched USAi in August 2025 as a shared platform for agencies to test frontier models, prioritized FedRAMP authorization for conversational AI, and is exploring internal chatbots that can draw on multiple vendors (OpenAI, Anthropic, Google). Third-party coverage widely frames AI chatbots as the 'new digital front door' to government services, with prior deployments (NYC's MyCity, NSF grant chatbot, Air Force NIPRGPT) showing both promise and documented accuracy problems. The supplied snippets do not directly corroborate the specific America.gov launch across 29,000+ sites or the night-one 15-question stress test β those rest on the case's own evidence title β but a commentary piece argues exactly this failure mode: AI front doors give confident but wrong answers when government content is fragmented, buried in PDFs, or not machine-readable. Whether the deployment sustains accuracy under growing independent testing is therefore the live question, and the snippets establish that GSA's broader AI push (USAi, FedRAMP 20x) is real and ongoing into 2026.
Why it matters to Scott
The world is independently enacting Scott's grounded-front-door arguments: GSA shipped a citizen-facing answer engine and a journalist stress-tested it night one on production reporting questions (his capability-audit posture), and the result's shape β zero hallucinations but 40% incomplete β is a public instance of his answer-failure-classes split between honest coverage/navigation incompleteness and fabrication. His route-invariant-grounding framework also says 15 single-route correct answers certify little about actual grounding, so the live sustained-accuracy question over fragmented, PDF-buried government content (systemic-wrongness risk) is a citable, ongoing test case for the exact claims his evaluation and witness-not-oracle territory makes.
ip:concept.answer-failure-classesip:framework.route-invariant-groundingip:concept.capability-auditip:concept.systemic-wrongnessip:concept.interface-ladder-for-knowledgeradar:concept.government-airadar:concept.llm-evaluationradar:concept.ai-searchradar:concept.ragradar:austria-govgpt-sovereign-rollout
queries asked of Scott's wikis
- RAG grounding and citation quality evaluation for production chatbots
- eval harness design and stress-testing methodology for deployed AI systems
- machine-readable structured publishing and docs-for-AI-retrieval patterns
- answer-engine front door versus traditional site search as an interface pattern
- hallucination measurement in grounded domain corpora at scale
- government AI adoption procurement and FedRAMP conversational AI
Measured heat
now 0 pts/hpeak 96 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 265h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion
How the heat travelled
pace: p84 vs 1188 stories at the 168h mark (now 265h old) β ahead of big-tech-ai-guarantee-exposure (1.0x), behind qwen38-27b-16gb-quant-benchmark (1.0x)
Evidence (5) β β canonical anchor
Interpretation history
2026-10-03T18:54:56Z
France 24's segment upgrades the political-steering claim from uncorroborated single-outlet report to multi-outlet documented behavior change, collapsing one of the case's core open questions and pointing the deployment's trajectory at entrench-with-steering rather than scale-back or accuracy-driven hardening. Heat stays medium against a cold numbers line (1.8 pts/h, p50, plateaued) only because a politically steered government chatbot is re-eruption-prone β the watch item is whether US national outlets follow France 24 within a day or two.
2026-10-03T18:25:45Z
evidence attached: reddit.post.1wwsxhp β France 24 documents a behavior change in America.gov's answers, material context for whether the front door sustains scrutiny or gets altered.
2026-10-02T16:54:31Z
The episode has passed its front-page phase: velocity collapsed ~26x from peak (82.7 β 3.2 pts/h), the politics and Minecraft threads plateaued (+1 pt/day), and the only live addition β a 20-year guest-tracking-cookie finding β opens a third scrutiny axis (privacy, alongside accuracy and political steering) at much thinner amplitude. Meaning shifts from 'breaking cross-platform event' to 'contested live deployment under quiet watch': none of the core live questions (reprogramming corroboration, sustained accuracy, scale-back/harden/entrench) has moved.
2026-10-02T16:26:23Z
evidence attached: hn.story.49934631 β Tracking/privacy controversy on America.gov itself materially contextualises whether the AI front-door deployment sustains scrutiny or gets scaled back.
2026-09-30T23:38:02Z
The case's meaning shifted from 'launch plus one night-one test' to a contested live deployment: independent cross-platform probing shows both genuine utility and documented trivial-request misbehavior, and a single-outlet report claims the chatbot was politically reprogrammed to stop fact-checking β putting political steering, not just technical accuracy, on the live question. Three independent evidence lines (journalist stress test, HN probing, news report of reprogramming) pass corroborated on substance; the prior low heat label under-rated day-one top-decile cross-platform velocity (96.8th percentile, accelerating) and the just-opened national-news political angle, which is exactly the window where high heat matters.
2026-09-30T21:37:57Z
evidence attached: hn.story.49913255 β Documented front-door misbehavior on a trivial request is exactly the independent stress-testing that case's hypothesis hinges on.
2026-09-30T21:37:57Z
evidence attached: reddit.post.1wueara β Political reprogramming of the America.gov chatbot bears directly on whether its accuracy-first deployment survives.
2026-09-30T15:38:16Z
grounded: converges/medium β The world is independently enacting Scott's grounded-front-door arguments: GSA shipped a citizen-facing answer engine and a journalist stress-tested it night on
2026-09-30T15:29:11Z
case created β A brand-new government-wide AI deployment is already under independent accuracy testing, a bounded episode with no open case.
Decision trace
- 10-04 05:54repriceFrance 24's segment upgrades the political-steering claim from uncorroborated single-outlet report to multi-outlet documented behavior change, collapsing one of the case's core open question
- 10-04 05:25attachFrance 24 documents a behavior change in America.gov's answers, material context for whether the front door sustains scrutiny or gets altered.
- 10-04 05:24propose_attachFrance 24 documents a behavior change in America.gov's answers, material context for whether the front door sustains scrutiny or gets altered.
- 10-03 03:32sensor_dirtyvelocity_spike
- 10-03 02:54repriceThe episode has passed its front-page phase: velocity collapsed ~26x from peak (82.7 β 3.2 pts/h), the politics and Minecraft threads plateaued (+1 pt/day), and the only live addition β a 20-year gues
- 10-03 02:26attachTracking/privacy controversy on America.gov itself materially contextualises whether the AI front-door deployment sustains scrutiny or gets scaled back.
- 10-03 02:24propose_attachTracking/privacy controversy on America.gov itself materially contextualises whether the AI front-door deployment sustains scrutiny or gets scaled back.
- 10-02 19:38review_screenjev screen: no material development (noul=0.06)
- 10-02 04:22sensor_dirtycomment_update
- 10-02 00:21sensor_dirtycomment_update
- 10-01 21:22sensor_dirtycomment_update
- 10-01 12:20sensor_dirtycomment_update
- 10-01 10:22sensor_dirtycomment_update
- 10-01 09:54pushWe stress-tested America.gov on night one with 15 questions from our own reporting. 9 correct, 6 incomplete, zero hallucinations. β The case's meaning shifted from 'launch plus one night-one
- 10-01 09:38repriceThe case's meaning shifted from 'launch plus one night-one test' to a contested live deployment: independent cross-platform probing shows both genuine utility and documented trivial-req
- 10-01 09:38alert_heldWe stress-tested America.gov on night one with 15 questions from our own reporting. 9 correct, 6 incomplete, zero hallucinations. β The case's meaning shifted from 'launch plus one night-one
- 10-01 09:38alert_routeWe stress-tested America.gov on night one with 15 questions from our own reporting. 9 correct, 6 incomplete, zero hallucinations. β The case's meaning shifted from 'launch plus one night-one
- 10-01 07:37attachDocumented front-door misbehavior on a trivial request is exactly the independent stress-testing that case's hypothesis hinges on.
- 10-01 07:37attachPolitical reprogramming of the America.gov chatbot bears directly on whether its accuracy-first deployment survives.
- 10-01 07:32propose_attachDocumented front-door misbehavior on a trivial request is exactly the independent stress-testing that case's hypothesis hinges on.
- 10-01 07:31propose_attachPolitical reprogramming of the America.gov chatbot bears directly on whether its accuracy-first deployment survives.
- 10-01 03:28sensor_dirtycomment_update
- 10-01 01:38groundThe world is independently enacting Scott's grounded-front-door arguments: GSA shipped a citizen-facing answer engine and a journalist stress-tested it night one on production reporting questions
- 10-01 01:29createA brand-new government-wide AI deployment is already under independent accuracy testing, a bounded episode with no open case.