2026-10-11 17:12 UTC

Tencent will publicly confirm or release Hunyuan Hy4 as a flagship expert-level model with integrated tool-use capabilities following its reported gray testing.

state: resolvedheat: lowuncertainty: lowconvergesscott: mediumfrontier-models tool-useTencentHunyuan
Surfaced 2026-08-28T06:32:48Z — priced heat=high at reprice: Hy4 has moved from a screenshot-derived gray-test rumor to a first-party public preview-weight release, making independent inspection and harness evaluation possible now. The release establishes the model event but not Tencent’s expert-level comparisons, integrated tool use, licensing practicality, or feasible serving profile.

What is this?

Tencent’s Hunyuan team is reportedly gray-testing Hy4, a larger-parameter flagship model that has appeared for some users in the Yuanbao app labeled as an “expert-level model” and positioned above Hy3 and DeepSeek. Secondary reports say it can use tools to solve problems, while Tencent’s earnings materials reportedly confirm that Hy4 is in training and intended to improve performance and multimodal capabilities. Tencent has not officially released Hy4; the supplied snippets conflict on timing (“soon” versus later in 2026) and do not substantiate the claimed 770B-A49B open-weight preview or 1M-token context.

Why it matters to Scott

A Tencent flagship explicitly designed for tool use would converge with Scott’s view that useful agent capability combines a model with an execution surface, and would create a concrete candidate for his provider-side, trace-backed agent evaluations. The supplied evidence does not establish code-first orchestration, reliable tool performance, open weights, or practical local inference, so the connection remains prospective rather than a validation of those stronger positions.
ip:concept.model-plus-harness-benchmark-unitip:concept.agent-hands-and-eyesdev:concept.agentic-tool-loopdev:project.remote-execdev:concept.trace-backed-agent-comparisonradar:concept.frontier-modelsradar:concept.tool-useradar:concept.agent-evaluation
queries asked of Scott's wikis
  • frontier models with native tool use
  • agent harnesses for model-integrated tool calling
  • open-weight frontier model strategy and sovereignty
  • large MoE local inference economics
  • evaluating expert-level models beyond benchmark claims
  • multimodal models as agent foundations

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (4) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditTencent begins testing its new flagship model Hunyuan Hy4
LocalLLaMA
Nunki0819225
🟧 echo.other ⭐Tencent’s Q2 2026 earnings presentation states: “Training a larger parameter model, Hy4, to be released later this year.” The accompanying eTencent——
🟠 redditTencent/Hy4-preview 770B-A49B weight dropped
LocalLLaMA
Beamsters547135
🟠 redditIntroducing to tencent’s Hy4 preview Open weights: 770B MoE, 49B active, 1M context
singularity
Snoo2683715310

Interpretation history

Decision trace