Independent evaluation will determine whether the reported 92GB one-bit quantization of Tencent's Hy3 295B preserves coding quality while outperforming its cloud API on a four-RTX-5090 system.
state: expiredheat: lowuncertainty: highconvergesscott: mediumextreme-quantization moe-models local-inferenceTencent
What is this?
Hy3 is a Tencent Hy Team model with 295B total parameters, a Mixture-of-Experts architecture using 21B active parameters, and a 3.8B multi-token-prediction layer. The supplied snippets establish that it can be served through local inference tooling, but they do not substantiate the reported 92GB one-bit quantization, four-RTX-5090 setup, 2.2× API speedup, or preservation of coding quality. Those claims currently rest on the cited announcement/echo and still require independent evaluation; the web answer’s assertion that evaluations already confirm them is unsupported by the listed results.
Why it matters to Scott
The claimed one-bit model would extend Scott’s hardware-aware local-inference and AI-unit-economics positions by testing whether extreme compression can preserve coding capability while beating API throughput on deployable hardware. It is directly relevant to his self-hosted GPU substrate and evaluation doctrine, but remains a conditional convergence until independent, task-specific benchmarks reproduce the quality and speed claims; the radar tracks adjacent quantization cases, not this Hy3 development itself.
ip:concept.evaluation-driven-developmentip:concept.capability-auditip:concept.ai-unit-economicsdev:project.gamepcdev:concept.hardware-aware-local-inferenceradar:concept.local-inferenceradar:bonsai-extreme-quantizationradar:ninfer-qwen-5090-throughput
queries asked of Scott's wikis
- one-bit quantization quality thresholds
- coding-agent quality under model quantization
- local inference versus API economics
- multi-GPU local model serving
- MoE quantization and active-parameter efficiency
- independent evaluation of local model claims
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-07-23T03:21:18Z
The claim has faded without independent benchmarks, reproductions, or consequential adoption; the unchanged discussion only repeats skepticism about the original three-task demonstration. The episode can expire unless fresh evaluation evidence revives it.
2026-07-20T23:23:40Z
The new attachment adds no independent benchmark or implementation evidence; discussion remains repetitive skepticism about extrapolating “no quality loss” from three one-shot tasks. The case still hinges entirely on the originating demonstration and should remain a low-temperature seed pending reproducible evaluation.
2026-07-20T22:21:29Z
The expanded discussion adds methodological skepticism rather than corroboration: commenters are converging on the inadequacy of three one-shot tasks for claiming no quality loss. The core speed and quality claim remains unsupported by independent benchmarks.
2026-07-20T15:35:18Z
grounded: converges/medium — The claimed one-bit model would extend Scott’s hardware-aware local-inference and AI-unit-economics positions by testing whether extreme compression can preserv
2026-07-20T15:33:33Z
origin walked (codex/luna, conf 0.88): anchor reddit.post.1v1nc7y -> echo.x.c0e43b0dd5 by Atomic Chat
2026-07-20T15:30:47Z
case created — The claim would materially expand local access to a very large MoE model, but the three-task demonstration does not yet substantiate no quality loss.
Decision trace
- 07-23 13:21expireThe claim has faded without independent benchmarks, reproductions, or consequential adoption; the unchanged discussion only repeats skepticism about the original three-task demonstration. The episode
- 07-21 09:23repriceThe new attachment adds no independent benchmark or implementation evidence; discussion remains repetitive skepticism about extrapolating “no quality loss” from three one-shot tasks. The case still hi
- 07-21 09:20mark_dirtycomment_update
- 07-21 08:21repriceThe expanded discussion adds methodological skepticism rather than corroboration: commenters are converging on the inadequacy of three one-shot tasks for claiming no quality loss. The core speed and q
- 07-21 01:35groundThe claimed one-bit model would extend Scott’s hardware-aware local-inference and AI-unit-economics positions by testing whether extreme compression can preserve coding capability while beating API th
- 07-21 01:33promote_anchororigin walk conf 0.88
- 07-21 01:30createThe claim would materially expand local access to a very large MoE model, but the three-task demonstration does not yet substantiate no quality loss.