Griffin's makers claim it is the first Human Interaction Model to pass a video Turing test โ 44% of judges took it for a human versus roughly 3% for competing systems โ and to rank #1 on NVIDIA's full-duplex video benchmark; identification of the vendor and independent testing resolve whether realtime interactive video crossed perceived-human indistinguishability or this is an inflated demo echo.
state: watchingheat: mediumuncertainty: mediumconvergesscott: highfull-duplex-video-models realtime-interactive-models capability-benchmarksNVIDIA
Surfaced 2026-10-04T13:50:36Z โ Tavus's official announcement page (dated "San Francisco, California October 1st, 2026"): "Today we're introducing Griffin, our first Human โ The r/OpenAI echo matured into a second front-page thread (12โ654 pts, 278 comments), completing cross-community spread โ yet day 4 and ~2,000 aggregate points still yield zero independent testing, NVIDIA statement, evaluator access, or debunk; the claim is now hardening into public consensus ('we're cooked') purely through repetition while the safety-gated research preview structurally blocks verification. That widening spread-vs-verification gap is the case's defining fact, but with momentum cooling to 5.8 pts/h off a 241 peak, no new evidence in kind (the magnitude-valve reading counts the same two Reddit threads; the only other object is Tavus's own page), and a 90th-percentile peer rate that just reflects residual votes on aging posts, the episode's attention value has decayed โ cooled to low, awaiting external triggers (independent eval, access change, NVIDIA comment, GA timeline).
What is this?
Griffin is a full-duplex video-to-video conversational model announced October 1, 2026 by Tavus, a San Francisco AI research lab previously known for its conversational video avatar stack (Phoenix 4.5, Sparrow-2, Raven-1); it generates full video frames in real time and can listen, watch, and speak simultaneously. Tavus's own live blind study found 48% of participants on a one-minute video call took Griffin for a real human, versus a โค2โ3% pass rate for prior systems including Tavus's own (the case's 44% figure is a weaker echo; the vendor's materials and launch coverage say 48%), and it is credited #1 on the Video Full-Duplex Benchmark that NVIDIA built and scored. It is a research preview, not a public release โ Tavus itself says safety review must precede general availability. The Turing-test claim rests on Tavus's own study: the supplied material shows no independent replication, and NVIDIA's documented role is building/scoring the duplex benchmark, not certifying the human-indistinguishability claim.
Why it matters to Scott
Converges with his realtime-interaction and evidence-discipline canon: NVIDIA institutionalizing full-duplex video as a scored benchmark category extends the exact deployment class his Real-Time AI Systems / Fast-Slow Split work and Twilio/Ultravox labs map, while the 44โ48% Turing-test number is a vendor-run study with no independent replication โ sitting precisely where his Evidence Class Ladder and Positioning Ladder say an unverified claim must be priced, and a live test case for the radar's benchmark-integrity lineage (same actor as the prior Sparrow-2 case). If the claim holds, his Chat-Era-Trust / Cryptographic-Trust position gains its strongest realtime-video forcing function ('your eyes are no longer a verifier'); if it collapses under independent testing, it is a dated receipt for the same argument โ and either branch touches what he builds, since realtime duplex video is the next substrate for his voice-agent and talking-head prospecting offers.
ip:concept.evidence-class-ladderip:concept.positioning-ladderip:framework.category-transition-lintip:concept.real-time-ai-systemsip:concept.chat-era-trust-modelip:concept.cryptographic-trustip:concept.synthetic-personasdev:project.twiliodev:technology.ultravoxwork:concept.ai-personalised-outbound-prospectingradar:tavus-sparrow-2-conversational-flowradar:concept.voice-agentsradar:concept.benchmark-integrityradar:concept.model-evaluationradar:realtime-venus-open-av-interactionradar:gemini-38-live-release
queries asked of Scott's wikis
- vendor-coined AI category launch framing
- self-reported benchmark claims vs independent evals
- Turing test marketing capability claims
- realtime full-duplex video voice agent stack
- safety-gated research preview release pattern
- synthetic human indistinguishability trust
Measured heat
now 0 pts/hpeak 241 pts/hcomments 0/hpeers p35momentum: steady2 platformsage 266h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion
How the heat travelled
pace: p97 vs 1188 stories at the 168h mark (now 266h old) โ ahead of openai-gpt61-sol-release (1.0x), behind qwen38-flash-dual-3090-speedup (1.0x)
Evidence (4) โ โญ canonical anchor
Interpretation history
2026-10-11T05:43:12Z
NY Post coverage attached (reddit.post.1x2zfuj) is derivative reporting on Tavus's own study โ 0 Reddit engagement confirms it added no new viral energy. Measured heat shows the episode fully decayed: 0 pts/h, 16.7th percentile, steady momentum at 255h age. The spread-vs-verification gap persists unchanged: still no independent replication, third-party eval, NVIDIA statement on the Turing-test claim, or evaluator access. Core uncertainties and Scott relevance unchanged.
2026-10-11T05:29:22Z
evidence attached: reddit.post.1x2zfuj โ Independent media coverage (NY Post) reporting 48% fool rate in job interviews, corroborating the Griffin video-Turing-test claim.
2026-10-03T23:41:47Z
magnitude valve eligible (multi-platform, top-decile engagement) and never alerted; deterministic escalation to deliver
2026-10-02T13:26:38Z
Spread matured into a front-page-scale cross-community episode (anchor 263โ1114 pts, 320 comments; thin r/OpenAI echo; magnitude valve) but remains pure amplification of Tavus's vendor-run 48% study โ two days of heavy attention have produced no independent test, demo access, or credible contradiction. Held at watching/medium on cooling momentum off the 206 pts/h peak (not high: periphery added one 12-pt derivative, not a dozen); the widening spread-vs-verification gap, not engagement, is what would move this case.
2026-10-02T13:24:53Z
evidence attached: reddit.post.1wvtedh โ Same 'first model to pass the Video Turing Test, ~half thought it human' claim echoing into r/OpenAI โ cross-community spread is exactly the resolution evidence for whether this is a real milestone or an inflated demo echo.
2026-10-01T21:48:44Z
origin walked (opencode/cheap-glm, conf 0.93): anchor reddit.post.1wv7q40 -> echo.blog.20450789aa by Tavus (by Hassaan Raza, Co-founder & CEO; Ioannis Patras, Head of Research; Tavus Research Team)
2026-10-01T21:11:18Z
grounded: converges/high โ Converges with his realtime-interaction and evidence-discipline canon: NVIDIA institutionalizing full-duplex video as a scored benchmark category extends the ex
2026-10-01T21:01:48Z
case created โ High-spread first-of-kind capability claim (263 points, 97% ratio) with no open case, but no vendor artifact visible yet so not worth hourly re-observation.
Decision trace
- 10-11 17:16attention_routeThe editor compared this story and chose to keep watching.
- 10-11 16:43repriceNY Post coverage attached (reddit.post.1x2zfuj) is derivative reporting on Tavus's own study โ 0 Reddit engagement confirms it added no new viral energy. Measured heat shows the episode fully dec
- 10-11 16:31attention_routeNew high-relevance case converging on Scott's realtime-interaction and evidence-discipline canon. Vendor-run Turing-test claim with high spread but no independent verification. Knowing before the
- 10-11 16:29attention_candidateattach
- 10-11 16:29attachIndependent media coverage (NY Post) reporting 48% fool rate in job interviews, corroborating the Griffin video-Turing-test claim.
- 10-11 16:28propose_attachIndependent media coverage (NY Post) reporting 48% fool rate in job interviews, corroborating the Griffin video-Turing-test claim.
- 10-05 16:37review_screenjev screen: no material development (noul=0.07)
- 10-05 00:50pushTavus's official announcement page (dated "San Francisco, California October 1st, 2026"): "Today we're introducing Griffin, our first Human โ The r/OpenAI echo matured into a
- 10-04 20:21sensor_dirtycomment_update
- 10-04 10:41repriceThe r/OpenAI echo matured into a second front-page thread (12โ654 pts, 278 comments), completing cross-community spread โ yet day 4 and ~2,000 aggregate points still yield zero independent testing, NV
- 10-04 10:41alert_heldTavus's official announcement page (dated "San Francisco, California October 1st, 2026"): "Today we're introducing Griffin, our first Human โ The r/OpenAI echo matured into a
- 10-04 10:41alert_routeTavus's official announcement page (dated "San Francisco, California October 1st, 2026"): "Today we're introducing Griffin, our first Human โ The r/OpenAI echo matured into a
- 10-04 07:22sensor_dirtyvelocity_spike
- 10-03 23:21sensor_dirtyvelocity_spike
- 10-03 16:21sensor_dirtyvelocity_spike
- 10-03 10:21sensor_dirtycomment_update
- 10-03 07:21sensor_dirtyvelocity_spike
- 10-03 05:21sensor_dirtycomment_update
- 10-02 23:26repriceSpread matured into a front-page-scale cross-community episode (anchor 263โ1114 pts, 320 comments; thin r/OpenAI echo; magnitude valve) but remains pure amplification of Tavus's vendor-run 48% st
- 10-02 23:24attachSame 'first model to pass the Video Turing Test, ~half thought it human' claim echoing into r/OpenAI โ cross-community spread is exactly the resolution evidence for whether this is a real mi
- 10-02 23:23propose_attachSame 'first model to pass the Video Turing Test, ~half thought it human' claim echoing into r/OpenAI โ cross-community spread is exactly the resolution evidence for whether this is a real mi
- 10-02 23:21sensor_dirtyvelocity_spike
- 10-02 21:21sensor_dirtycomment_update
- 10-02 16:21sensor_dirtyvelocity_spike
- 10-02 11:21sensor_dirtycomment_update
- 10-02 09:22sensor_dirtyvelocity_spike
- 10-02 07:48promote_anchororigin walk conf 0.93
- 10-02 07:11groundConverges with his realtime-interaction and evidence-discipline canon: NVIDIA institutionalizing full-duplex video as a scored benchmark category extends the exact deployment class his Real-Time AI Sy
- 10-02 07:01createHigh-spread first-of-kind capability claim (263 points, 97% ratio) with no open case, but no vendor artifact visible yet so not worth hourly re-observation.