2026-10-11 17:10 UTC

Independent evaluation will determine whether the reported 92GB one-bit quantization of Tencent's Hy3 295B preserves coding quality while outperforming its cloud API on a four-RTX-5090 system.

state: expiredheat: lowuncertainty: highconvergesscott: mediumextreme-quantization moe-models local-inferenceTencent

What is this?

Hy3 is a Tencent Hy Team model with 295B total parameters, a Mixture-of-Experts architecture using 21B active parameters, and a 3.8B multi-token-prediction layer. The supplied snippets establish that it can be served through local inference tooling, but they do not substantiate the reported 92GB one-bit quantization, four-RTX-5090 setup, 2.2× API speedup, or preservation of coding quality. Those claims currently rest on the cited announcement/echo and still require independent evaluation; the web answer’s assertion that evaluations already confirm them is unsupported by the listed results.

Why it matters to Scott

The claimed one-bit model would extend Scott’s hardware-aware local-inference and AI-unit-economics positions by testing whether extreme compression can preserve coding capability while beating API throughput on deployable hardware. It is directly relevant to his self-hosted GPU substrate and evaluation doctrine, but remains a conditional convergence until independent, task-specific benchmarks reproduce the quality and speed claims; the radar tracks adjacent quantization cases, not this Hy3 development itself.
ip:concept.evaluation-driven-developmentip:concept.capability-auditip:concept.ai-unit-economicsdev:project.gamepcdev:concept.hardware-aware-local-inferenceradar:concept.local-inferenceradar:bonsai-extreme-quantizationradar:ninfer-qwen-5090-throughput
queries asked of Scott's wikis
  • one-bit quantization quality thresholds
  • coding-agent quality under model quantization
  • local inference versus API economics
  • multi-GPU local model serving
  • MoE quantization and active-parameter efficiency
  • independent evaluation of local model claims

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditOur 1-bit quant of Hy3 295B runs 2.2x faster than the cloud API with no quality loss
LocalLLaMA
ElmBark3848
🟧 echo.x ⭐Atomic Chat’s original announcement said: “1-bit Hy3 running locally is 2.2x faster than its API at the same quality!” It reported 76.9K tokAtomic Chat——

Interpretation history

Decision trace