2026-10-11 17:10 UTC

Independent evaluations will determine whether Z.ai’s released GLM-5.3 delivers frontier-level coding performance and materially stronger practical cybersecurity capabilities.

state: resolvedheat: lowuncertainty: lowconvergesscott: mediumopen-models coding-agents agentic-securityZ.ai

What is this?

GLM-5.3 is an open-weight model from PRC-based Z.ai, formerly Zhipu AI, released on August 14, 2026, with weights published roughly two weeks later; Z.ai attributes its gains over GLM-5.2 to scaled post-training. Independent coding evidence in the case supports frontier-adjacent performance, but does not establish universal frontier parity, production economics, or reliable long-session behavior; results for the distinct Flash variant cannot automatically be transferred to the full model. Z.ai reports major cybersecurity gains, including 84.5% on CyberGym, while CAISI has now published an independent benchmark assessment; however, the supplied CAISI snippet does not clearly map its tabulated scores to models, so it does not by itself establish comparative cyber superiority.

Why it matters to Scott

Independent benchmark lines and Mouse’s task-level, token-accounted run converge with Scott’s Model-Plus-Harness Benchmark Unit and trace-backed comparison doctrine: GLM-5.3 is a credible, economical test candidate, not a validated replacement based on weight-level scores alone. Its open weights and claimed coding-to-vulnerability transfer also make it actionable for Scott’s local inference and security-review work, but unresolved long-session reliability and thin cyber evidence limit the immediate impact.
ip:concept.model-plus-harness-benchmark-unitip:concept.evaluation-driven-developmentdev:concept.trace-backed-agent-comparisonip:concept.ai-unit-economicsip:source.security-reviewer-method-ebookdev:concept.hardware-aware-local-inferenceradar:frontierharness-17x-cost-variationradar:coding-agent-self-report-failure-blindnessradar:openai-gpt-56-cyber-modelradar:peng-agent-vulnerability-research-resultsradar:concept.open-models
queries asked of Scott's wikis
  • model-plus-harness benchmark unit independent evaluation
  • coding-agent cost per accepted outcome
  • long-horizon agent reliability and false completion
  • open-weight local inference quantization tradeoffs
  • coding capability transmutation into vulnerability discovery
  • cyber-agent evaluation refusal behavior and reproducibility

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

08-13 14:00⭐ origin echo-reconstructedThe Z.ai post announces GLM-5.3 and says, “Scaling post-training is all we did for GLM-5.3.” It attributes the gains to post-training, repor
Z.ai on blog (echo) · attributed from hn.story.49294997
—
08-14 05:19first on hacker news · published · +15.3hGLM-5.3: Frontier Coding with Emergent Cyber Capabilities
pella
—
08-14 05:23first on r/LocalLLaMA · published · +15.4hGLM 5.3 Released
jmorant555
—
08-14 06:08first on r/singularity · published · +16.1hGLM 5.3 released: Frontier Coding with Emergent Cyber Capabilities
1a1b
—
08-26 17:07first on r/artificial · published · +315.1hThank you for participating in the Stealth Ox Alpha testing period.
paulrich_nb
—
08-26 19:45first on r/OpenAI · published · +317.8hCurrently GLM 5.3 Flash matches with Sol 5.6 (Max) in Agentic Index (Artificial Analysis)
PilgrimofHaqq2
—
08-14 05:19amplified on hacker news 👑hn.story.49294997
pella
peak 1166 · 574 comments · 13% of case engagement
08-14 05:23amplified on r/LocalLLaMAreddit.post.1vny9zs
jmorant555
peak 1653 · 362 comments · 8% of case engagement
08-14 06:08amplified on r/singularityreddit.post.1vnz30c
1a1b
peak 460 · 77 comments · 2% of case engagement
08-14 07:45amplified on r/LocalLLaMAreddit.post.1vo0r4w
LegacyRemaster
peak 118 · 13 comments · 1% of case engagement
08-14 11:55amplified on r/singularityreddit.post.1vo56qy
1a1b
peak 656 · 106 comments · 3% of case engagement
08-15 06:15amplified on hacker newshn.story.49308147
tosh
peak 1 · 0 comments · 0% of case engagement
85 more amplifiers in ainews.case_chain
08-14 05:21our radar first saw it · +15.3hdiscovery anchor: hn.story.49294997—
08-26 11:29reached heat=high · +309.5h · via ledger——

Evidence (92) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnGLM-5.3: Frontier Coding with Emergent Cyber Capabilitiespella1166574
🟧 echo.blog ⭐The Z.ai post announces GLM-5.3 and says, “Scaling post-training is all we did for GLM-5.3.” It attributes the gains to post-training, reporZ.ai——
🟠 redditGLM 5.3 released: Frontier Coding with Emergent Cyber Capabilities
singularity
1a1b46077
🟠 redditGLM 5.3 Released
LocalLLaMA
jmorant5551648362
🟠 redditGLM 5.3 weights. It might offer the best capacity-to-size ratio.
LocalLLaMA
LegacyRemaster11813
🟠 redditGLM 5.3 finds 2436 unpatched open source vulnerabilities likely missed by Mythos (Project Glasswing)
singularity
1a1b656106
🟧 hnGLM-5.3: How Chinese labs keep stride with the frontiertosh10
🟧 hnPreparing GLM-5.3 for Open Release: A Responsible Path to Cyber Defensepretext20
🟧 hnOpen source, audited by GLM-5.3pama20
🟧 hnGLM-5.3 – The Official Desktop Auditor for Z.AI's Cyber-Enginenioj30
🟧 hnWhat We Learned Moving Our Agent Loops from Anthropic to GLMdennispi186
🟠 redditGLM5.3 Artificial Analysis Benchmarks
LocalLLaMA
anderspitman26366
🟧 hnGLM-5.3 Artificial Analysis Benchmarksapitman12950
🟠 redditGLM-5.3 achieves 60 on the Artificial Analysis Intelligence Index, on par with Kimi K3 and up 7 points from GLM-5.2. Once the weights are released it will be tied as the leading open weights model
singularity
Facelessjoe28442
🟠 redditGLM-5.3 (max) Intelligence, Performance & Price Analysis
singularity
yogthos863
🟧 hnGLM 5.3 Available in OpenRouterherlon21410
🟧 hnGLM-5.3 achieves 60 on Artificial Analysistosh10
🟠 redditGLM-5.3 is out on AA, and I'm fed up with their Intelligence/cost plot
LocalLLaMA
crusaderky1855
🟠 redditGLM 5.3 SlopCodeBench Results
LocalLLaMA
corruptbytes1811
🟠 redditI fingerprinted Ox Alpha: same tokenizer as GLM-5.3 (+75 token offset), z.ai's exact error strings, near-identical temp-0 outputs
singularity
FlunkyGraphics22555
🟧 hnGemini 3.7 Flash, Grok 4.6, GLM-5.3 and DeepSeek V4 Pro joined the frontierstared10
🟧 hnI spent $266 and four AI models to own my tablet. GLM-5.3 finished it in a daydr_pardee69427
🟠 redditAmazon kept shutting down my tablet, so I spent $266 on four AI models to own it
singularity
yogthos53946
🟠 redditFirst serious confirmation. Ox Alpha is GLM-5.3-Flash
LocalLLaMA
MrWidmoreHK474157
🟠 redditConfirmed: Z.AI Made Ox Alpha Stealth Model That Rivals DeepSeek
LocalLLaMA
MrWidmoreHK17119
🟠 redditzai-org/GLM-5.3-Flash · Hugging Face
LocalLLaMA
coder543524131
🟠 redditGLM-5.3-Flash: Frontier Intelligence, Flash Cost
LocalLLaMA
BriguePalhaco1262440
🟧 hnGLM-5.3-FlashPhilpax1075538
🟧 hnTell HN: GLM 5.3 "Flash" appears on DeepSWE with a score of 63%theanonymousone10
🟧 hnzai-org/GLM-5.3-FlashPhilpax30
🟠 redditGLM-5.3-Flash: Frontier Intelligence, Flash Cost
singularity
dydynam19832
🟠 redditGLM 5.3 Flash (Ox Alpha) benchmark comparisons
singularity
elemental-mind16165
🟧 hnGLM-5.3-Flash Intelligence, Performance and Price Analysistheanonymousone13353
🟠 redditThank you for participating in the Stealth Ox Alpha testing period.
artificial
paulrich_nb10
🟠 redditCurrently GLM 5.3 Flash matches with Sol 5.6 (Max) in Agentic Index (Artificial Analysis)
OpenAI
PilgrimofHaqq216346
🟠 redditGLM-5.3 weights will be released tomorrow
LocalLLaMA
serige34440
🟠 redditGLM-5.3 Flash Unsloth GGUF now available
LocalLLaMA
ElementNumber610820
🟠 redditGLM-5.3-Flash (FP8) on 4 x RTX6000 Pro
LocalLLaMA
AutonomousHangOver1219
🟠 redditGLM-5.3-Flash @ DGX Station GB300: ~206 tok/s (single stream), 1M context
LocalLLaMA
funding__secured3834
🟠 reddit[Megathread] GLM-5.3-Flash - former ox-alpha
LocalLLaMA
No_Afternoon_4260267185
🟧 hnShow HN: Warp – Run the 313B GLM-5.3-Flash on a MacBook with 8GB RAMmarcobambini20
🟠 redditZ.ai Served GLM-5.3-Flash Entirely on Chinese AI Chips
singularity
yogthos120
🟠 redditI tried telling GLM 5.3 Flash to continue a bug fix and it went crazy
artificial
ProgrammingGuy_10
🟠 redditglm 5.3 flash on 4 sparks
LocalLLaMA
No_Afternoon_426003
🟠 redditAnyone running GLM-5.3 Flash on 2x RTX PRO 6000 96GB?
LocalLLaMA
No-Paper-557115
🟠 redditzai-org/GLM-5.3 · Hugging Face
LocalLLaMA
jacek2023624140
🟧 hnGLM-5.3 is now open-weightjeudesprits802278
🟧 hnzai-Org/GLM-5.3Philpax30
🟠 redditds4 branch with GLM 5.3 Flash support
LocalLLaMA
lakySK10032
🟠 redditGLM-5.3-Flash Benchmarks on TensorSharp and llama.cpp
LocalLLaMA
fuzhongkai200
🟧 hnGLM5.3 unsloth GGUFs are now upwalrus0121
🟠 redditGLM-5.3 on HF Viewer
LocalLLaMA
Course_Latter503
🟧 hnInteractive Model View zai-org/GLM-5.3ResearchAtPlay40
🟠 redditTerminal Bench 4.0 just dropped, GLM-5.3 is at the same level as Fable 5, accounting for margin of error
LocalLLaMA
SorosAhaverom544110
🟧 hnWhat GLM-5.3 Flash running on Chinese hardware meansymolodtsov10
🟧 hnTinker: GLM 5.3 Fine-Tuningtosh10
🟧 hnBy Opening a Model, a Chinese A.I. Lab May Test the World's Cybersecuritybookofjoe11
🟧 hnGLM-5.3-Flash-GGUFwalrus0181
🟠 redditGLM 5.3 weights are now public
singularity
badumtsssst53375
🟠 redditWhen you say, because I can. Limits of X870e
LocalLLaMA
RedAdo20204569
🟠 redditGLM-5.3-Flash is 100% a step change in agential capability, but I'm not sure it's /reliable/ enough to trust at scale... the long tail of agent work is NASTY when it strikes
LocalLLaMA
me_myself_ai1045
🟧 hnWhat GLM-5.3 Flash running on Chinese hardware meansbirdculture10
🟧 hnShow HN: Our GLM-5.3 Flash Switchless recipe is now out for 4x DGX Sparksalexellisuk21
🟧 hnShow HN: 2x-4x cheaper GLM 5.3 for coding and researchHiteshjain11873
🟧 hnI found two Shopify plugin vulns with GLM-5.3 Flasharm3231
🟧 hnShow HN: Abliterated GLM-5.3 API (84.5% CyberGym, FP8)thomadev021
🟠 redditHere is your chance to take over the world: glm 5.3 abliterated
LocalLLaMA
Terminator857062
🟧 hnAbliterated model large v2: GLM 5.3 84.5% CyberGymgafferongames10
🟠 redditGave a try to Exllamav3 and it's great!
LocalLLaMA
Leflakk2224
🟠 redditGLM 5.3 flash is annoying
LocalLLaMA
AppealSame4367016
🟧 hnGLM-5.3-Flash at 1000 tok/s on RTX PRO 6000davedx10
🟠 redditGLM 5.3 Flash makes a black hole Minecraft mod running locally on 4x RTX PRO 6000 WS
LocalLLaMA
Top-Eye-810427756
🟧 hnGLM-5.3 Uncensoredhkalbasi120
🟠 redditGLM5.3 Flash over DSV4 Flash?
LocalLLaMA
rm-rf-rm5463
🟧 hnI hate benchmarks: Moving a production workflow from GPT to GLM-5.3 Flashahmedelsama21
🟧 hnFast weights and sparse attention in GLM-5.3-Flashsmaddrellmander30
🟠 redditLLM regression in reading comprehension?
LocalLLaMA
GodComplecs1128
🟧 hnA $37 GLM 5.3 red team: the Alloy-modeled auth layer held, but two bugs outsidesvcrunch10
🟠 redditTerminal Bench v4 scores
LocalLLaMA
Ok_Warning214615783
🟠 redditDeepSeek V4.1 Flash at 41 tok/s on 8× A40 on TensorSharp
LocalLLaMA
fuzhongkai04
🟧 hnUsing Modal for GLM Flash on Cadenya's Agent Runtimebtables10
🟠 redditDeepSeek V4.1 Flash on 8× A40: ~40 tok/s Q2_K and ~32 tok/s Q4_K_M with TensorSharp
LocalLLaMA
fuzhongkai30
🟧 hnIs GLM-5.3-Flash Mythos-Level at Cyber?mamiglia30
🟧 hnMouse on frontier harness with glm-5.3-flash
Retrieved article excerpt

Open article · Retrieved 2026-09-15T00:22:15.962622+00:00

[← Blog](https://mouse.dev/blog)

Dark Light

# Mouse on GLM-5.3-Flash

23 of 30 FrontierHarness tasks on Z.ai's Flash model, next to the published GLM-5.3 harness runs, for $6.72 in tokens.

September 14, 2026 · Pete · 4 min read

- Benchmarks
- Harness
- GLM

Community results shared this week ran GLM-5.3 and GLM-5.3-Flash from Z.ai through five coding agent harnesses on the 30 [FrontierHarness](https://frontierharness.org) tasks. Z.ai's own harness, ZCode, came out on top, and was the only one run on Flash. We ran Mouse on GLM-5.3-Flash on the same tasks. It passed 23 of 30, at $0.29 per pass.

Pass rate

77%

23 of 30 tasks

Cost per pass

$0.29

$6.72 for all 30

Cache hit rate

97%

of input tokens

Median time

4:11

per pass

MouseZCodeCodexClaude CodePiOpenCodeDSH CreatorDSH Standard / DSHDSH PTCDSH MinimalOh My PiKimi CodeExo HarnessHermes
$0.2$0.5$1$2$5$10$2050%60%70%80%




Mouse + Kimi K383.3% · $2.79 · RuntaMouse + GLM-5.3-Flash76.7% · $0.29 · RuntaZCode + GLM-5.382.2% · $1.98Claude Code + GLM-5.380.0% · $2.58Codex + GLM-5.366.7% · $1.19Pi + GLM-5.360.0% · $2.10OpenCode + GLM-5.346.7% · $2.28ZCode + GLM-5.3-Flash75.6% · $0.15DSH + DeepSeek V4.1 Flash73.3% · $0.21Codex66.7% · $3.47DSH Creator63.3% · ≥$3.29Claude Code63.3% · ≥$18.34Pi60.0% · $2.43DSH PTC60.0% · ≥$4.58DSH Standard60.0% · ≥$3.46Oh My Pi56.7% · $4.75Kimi Code56.7% · $3.65DSH Minimal56.7% · ≥$4.72Exo Harness53.3% · $1.04OpenCode50.0% · ≥$3.24Hermes50.0% · ≥$2.90

Total recorded cost / passes (log scale)
Pass rate
Names without a model are the published Kimi K3 baselines; the orange line is their cost frontier. ≥ marks incomplete baseline cost coverage.
GLM and DeepSeek rows are third-party local runs at list prices on observed cache. Mouse rows are Runta runs under FrontierHarness's scripts.

FrontierHarness tasks, pass rate against total recorded cost divided by passes. Different models, providers, and runtimes: a side-by-side, not a controlled ranking.

 

## The comparison

The published FrontierHarness board fixes the model at Kimi K3. The GLM results are a separate, community-run extension: the same 30 tasks and scoring, with GLM-5.3 on a Zhipu Max plan instead. Cost is observed tokens multiplied by Z.ai's API list price, not what the plan bills. ZCode was run three times per model, so its rows are averages.

| Harness | Model | Pass rate | Passed | $ / pass | Runs |
| --- | --- | --- | --- | --- | --- |
| ZCode | GLM-5.3 | 82% | 24.67/30 | $1.98 | 26 / 22 / 26 |
| Claude Code | GLM-5.3 | 80% | 24/30 | $2.58 | 1 |
| Mouse | GLM-5.3-Flash | 77% | 23/30 | $0.29 | 1 |
| ZCode | GLM-5.3-Flash | 76% | 22.67/30 | $0.15 | 24 / 21 / 23 |
| Codex | GLM-5.3 | 67% | 20/30 | $1.19 | 1 |
| Pi | GLM-5.3 | 60% | 18/30 | $2.10 | 1 |
| OpenCode | GLM-5.3 | 47% | 14/30 | $2.28 | 1 |

On Flash, Mouse's single run (23) sits within ZCode's three runs (24, 21, and 23). ZCode is cheaper per pass, $0.16 against Mouse's $0.29. On the full GLM-5.3 model, ZCode and Claude Code pass 24 to 25 tasks and cost $2 to $2.60 per pass, so Mouse on Flash comes within one or two tasks of them at a seventh to a ninth of the cost per pass.

## Against Kimi K3

The same harness scored 25 of 30 on Kimi K3 at $2.79 per pass. Flash lost two tasks that K3 passed, both DeepSWE: `expr-try-catch-errors` and `fastapi-deprecation-response-headers`. It did not pass anything K3 failed. The other five misses were shared: `meriyah` on DeepSWE, and `dna-insert`, `polyglot-c-py`, `extract-elf`, and `gcode-to-text` on Terminal-Bench.

The cost difference is mostly the model's price. Flash lists at $0.15 per million input tokens, $0.03 cached, and $0.50 output. Cache hits covered 96.7% of input tokens, so the median passing task cost a little over a cent. The DeepSWE tasks were most of the bill: the nine of them cost $6.10 of the $6.72.

## Terminal-Bench

| Task | GLM-5.3-Flash | Steps | Cost | Time | Kimi K3 |
| --- | --- | --- | --- | --- | --- |
| regex-log | Pass | 26 | $0.05 | 8:56 | Pass |
| openssl-selfsigned-cert | Pass | 26 | $0.01 | 3:28 | Pass |
| polyglot-c-py | Fail | 17 | $0.04 | 9:42 | Fail |
| sqlite-db-truncate | Pass | 12 | $0.01 | 3:46 | Pass |
| git-leak-recovery | Pass | 15 | $0.01 | 3:20 | Pass |
| log-summary-date-ranges | Pass | 18 | $0.01 | 3:12 | Pass |
| constraints-scheduling | Pass | 13 | $0.01 | 3:56 | Pass |
| gcode-to-text | Fail | 57 | $0.13 | 16:26 | Fail |
| dna-insert | Fail | 16 | $0.02 | 6:48 | Fail |
| largest-eigenval | Pass | 15 | $0.04 | 16:48 | Pass |
| merge-diff-arc-agi-task | Pass | 23 | $0.02 | 5:26 | Pass |
| vulnerable-secret | Pass | 16 | $0.01 | 3:16 | Pass |
| extract-elf | Fail | 21 | $0.03 | 7:02 | Fail |
| build-cython-ext | Pass | 98 | $0.17 | 12:16 | Pass |
| kv-store-grpc | Pass | 21 | $0.01 | 3:20 | Pass |
| chess-best-move | Pass | 17 | $0.01 | 4:11 | Pass |
| db-wal-recovery | Pass | 14 | $0.01 | 4:27 | Pass |
| code-from-image | Pass | 8 | $0.00 | 2:13 | Pass |
| modernize-scientific-stack | Pass | 12 | $0.01 | 2:43 | Pass |
| multi-source-data-merger | Pass | 12 | $0.01 | 3:44 | Pass |
| sanitize-git-repo | Pass | 15 | $0.01 | 3:03 | Pass |

## DeepSWE

| Task | GLM-5.3-Flash | Steps | Cost | Time | Kimi K3 |
| --- | --- | --- | --- | --- | --- |
| anko-typed-variable-bindings | Pass | 114 | $0.38 | 20:55 | Pass |
| arktype-json-schema-refs-dependencies | Pass | 211 | $0.82 | 48:27 | Pass |
| fastapi-deprecation-response-headers | Fail | 161 | $0.54 | 42:52 | Pass |
| httpx-multipart-response-parsing | Pass | 96 | $0.34 | 31:36 | Pass |
| expr-try-catch-errors | Fail | 163 | $0.76 | 42:22 | Pass |
| python-statemachine-state-data-scoping | Pass | 250 | $1.55 | 77:48 | Pass |
| katex-multicolumn-array-spans | Pass | 127 | $0.53 | 33:43 | Pass |
| scc-bounded-memory-spilling | Pass | 100 | $0.48 | 29:46 | Pass |
| meriyah-explicit-resource-declarations | Fail | 154 | $0.71 | 39:01 | Fail |

## How the run was made

- FrontierHarness's own trial driver, unmodified, at their eval commit `e837a70`, on Runta with a fresh restore for every task and 4 vCPU / 8 GiB per task. Terminal-Bench through Harbor 0.22.0, DeepSWE through Pier 0.3.1.
- Harness commit `315e2b8` on OpenCode 1.18.27, the same checkpoint as the Kimi K3 run, with only the model changed: `glm-5p3-flash` served by Fireworks.
- Cost is each trial's observed tokens at Z.ai's GLM-5.3-Flash list price, which Fireworks matches. Three attempts on two tasks hit Ubuntu mirror errors while setting up the container, before the model ran; they were re-run and are not counted as failures.
- One run, 2026-09-14. The other rows were run locally by a third party, on a different provider, runtime, and harness versions (Claude Code 2.1.237, OpenCode 1.18.19, Codex 0.148.0, Pi 0.84.2), so this is a side-by-side, not a controlled ranking.

---

GLM results for ZCode, Claude Code, Codex, Pi, and OpenCode are from a community comparison shared this week, building on work by LotusDecoder. Mouse is built on [OpenCode](https://github.com/anomalyco/opencode) and is not affiliated with it or with Z.ai.
Aeroi10
🟠 redditA very unexpected analogy
LocalLLaMA
-dysangel-9312
🟧 hnGLM 5.3 is live on Mistralbiehl40
🟧 hnGLM-5.3-FlashX: Delivering inference speeds of 200 tokens/stheanonymousone51
🟧 hnGLM 5.3 Hosted by Mistralabc4230
🟠 redditJev-style structured decisions with DiffusionGemma GGUF in TensorSharp — benchmarks vs. LocalJev
LocalLLaMA
fuzhongkai01
🟠 redditI don’t believe in z.ai benchmarks
LocalLLaMA
Thin_Pollution8843028
🟠 redditTensorSharp: a local Jev-compatible API, extended to image analysis with DiffusionGemma GGUF
LocalLLaMA
fuzhongkai03
🟠 redditTensorSharp Jev requests can now combine documents, images, video, and audio
LocalLLaMA
fuzhongkai23

Interpretation history

Decision trace