2026-10-11 16:37 UTC

Z.ai claims its GLM-5.3 Infra Agent, guided by localized correctness and performance feedback, helped bring GLM-5.3-Flash serving on Chinese-made accelerators to production in under two weeks with roughly threefold throughput gains and NVIDIA-comparable per-token costs, demonstrating a practical route to agent-assisted inference engineering.

state: watchingheat: highuncertainty: highconvergesscott: mediumagent-harnesses coding-agents ai-infrastructure inference-economicsZ.ai
Surfaced 2026-09-23T18:10:54Z — Z.ai reports that engineers and a GLM-5.3-powered Infra Agent used 'dense feedback' to bring GLM-5.3-Flash from initial adaptation to produc — The Jetson Thor report supplies a separate example of agent-proposed optimizations gated by numerical and behavioral tests, not corroboration of Z.ai's deployment or economics. Despite the cross-platform spread reading, the GLM coverage remains repeated circulation of one account; the new, lightly attended attachment concerns a different implementation rather than an expanding GLM episode.

What is this?

Z.ai says it served GLM-5.3-Flash on a large-scale cluster of Chinese-made accelerators using a dedicated inference engine built on SGLang, with a GLM-5.3-powered infrastructure agent assisting engineers in development and optimization. Its documentation claims a threefold improvement in end-to-end serving performance over its initial baseline on the same hardware, and efficiency and per-token costs comparable to mainstream NVIDIA GPUs. The supplied reporting notes that the chip vendor, NVIDIA comparator, workload and cost methodology were not disclosed, so these remain company claims rather than a controlled, independently verified comparison. The snippets support agent-assisted optimization but do not establish the hypothesis's under-two-week delivery timeline or localized correctness and performance feedback mechanism.

Why it matters to Scott

Z.ai’s reported use of a coding agent to build and optimize production inference infrastructure extends the agent-assisted shipping argument in Scott’s “Your AI Can Code. You Just Don't Know How to Drive It.” into inference-engine engineering, offering a qualified publishing comparison rather than merely another coding demo; the supplied radar hits track related optimization stories, not this development. However, the grounding does not establish the localized verification loop or two-week timeline, and the undisclosed benchmark and cost methodology prevent treating this as validation of Scott’s test-first discipline or a reason to change his CUDA-based serving stack.
ip:source.your-ai-can-code-you-just-don-t-know-how-to-drive-it-ebookradar:codex-autoresearch-gpu-kernel-speedupradar:rooflang-inference-architecture-searchradar:concept.coding-agentsradar:concept.inference-optimizationradar:concept.inference-economics
queries asked of Scott's wikis
  • coding agent harnesses localized correctness performance feedback
  • agent-assisted systems engineering human oversight production validation
  • automated optimization benchmark baselines measurable speedups
  • SGLang inference serving accelerator portability
  • inference economics hardware sovereignty software optimization

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 602h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-16 14:00⭐ origin echo-reconstructedZ.ai reports that engineers and a GLM-5.3-powered Infra Agent used 'dense feedback' to bring GLM-5.3-Flash from initial adaptation to produc
Z.ai on blog (echo) · attributed from hn.story.49737922
—
09-17 08:27first on hacker news · published · +18.4hGLM Built Its Own Inference Infrastructure
whiteros_e
—
09-17 11:11first on r/singularity · published · +21.2hToward Recursive Self-Improvement: How GLM Built Its Own Inference Infrastructure
likeastar20
—
09-17 16:08first on r/LocalLLaMA · published · +26.1hshots fired at dario from glm
Elux91
—
09-17 08:27amplified on hacker news 👑hn.story.49737922
whiteros_e
peak 411 · 260 comments · 79% of case engagement
09-17 11:11amplified on r/singularityreddit.post.1wir16w
likeastar20
peak 91 · 2 comments · 6% of case engagement
09-17 16:08amplified on r/LocalLLaMAreddit.post.1wiy8ga
Elux91
peak 180 · 40 comments · 14% of case engagement
09-22 17:26amplified on hacker newshn.story.49804932
hhuytho
peak 2 · 0 comments · 0% of case engagement
09-17 09:20our radar first saw it · +19.3hdiscovery anchor: hn.story.49737922—
09-23 18:10reached heat=high · +172.2h · via queue+ledger——
pace: p90 vs 1032 stories at the 336h mark (now 602h old) — ahead of cerebras-qwen38-27b-inference-speed (1.0x), behind siri-private-model-substitution (1.0x)

Evidence (5) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnGLM Built Its Own Inference Infrastructure
Retrieved article excerpt

Open article · Retrieved 2026-09-17T09:21:47.959943+00:00

2026-09-17 · Research

# Toward Recursive Self-Improvement: How GLM Built Its Own Inference Infrastructure

[- Try it at Z.ai](https://z.ai)[- Call it at Z.ai](https://z.ai/model-api)[- Z.ai Coding Plan](https://z.ai/subscribe?utm_source=blog&utm_medium=content&utm_campaign=glm5_launch)

As we develop GLM, the model sometimes exhibits capabilities that surprise us, and even unsettle us.

In October 2025, we began researching how to strengthen its cybersecurity capabilities. Our reasoning at the time was straightforward: cybersecurity is a natural extension of coding. A model that can understand complex code should also be able to understand its vulnerabilities. We did not anticipate what would follow. In less than a year, our security partners used GLM to discover thousands of vulnerabilities in real-world codebases. The model began to reshape cybersecurity, while also introducing dangers that had not existed before. To enable responsible use of this capability, we had to design a trusted access program.

The most recent moment that shook us came from a more fundamental shift: GLM is increasingly helping build AI itself. We watched the model complete an infrastructure task that would previously have taken a team of experienced infrastructure engineers weeks. When we realized that this work would directly change how the next generation of models is trained, we became even more convinced: our successors are the AI systems we are creating ourselves.

Frankly, before GLM-4.7, our internal use of GLM for coding involved a certain amount of obligation. It was, after all, our own creation. At that point, product-market fit for coding had yet to arrive. Today, GLM-5.3 has become an indispensable daily coding partner for everyone on the team, and it is moving steadily toward replacing us. If this trend continues, given enough compute and enough time, its endpoint is a system that can design and train its own successor entirely autonomously. This is known as Recursive Self-Improvement, or RSI.

We are not there yet, but early forms of it are already emerging. This article documents one such early example.

Taking a model from its first successful run on new hardware to a high-performance inference service that can reliably handle production traffic is a major systems engineering undertaking. The launch of GLM-5.3-Flash followed the same path. We built a complete production-grade inference service from scratch on a cluster of more than 100,000 Chinese-made AI accelerators. All production inference for GLM-5.3-Flash runs on this system.

This was not easy. No one had previously deployed a cluster of Chinese-made accelerators at this scale. We faced relatively limited chip memory capacity and bandwidth, while also needing to support a new model architecture, a 1M-token context window, and multimodal requests. The ecosystem was immature, kernel support was incomplete, and much of what should have been documented had to be guessed. In the end, the work was completed. It was not done by a team of infrastructure engineers alone. Much of the work was carried out by an Infra Agent powered by GLM-5.3.

What happened next is already familiar. GLM-5.3-Flash was tested through real-world usage on OpenCode and OpenRouter under the anonymous model name Ox-Alpha. Within a week of launch, it became the most-used model on both platforms, processing more than 62 trillion tokens in six days.

We implemented a series of aggressive memory optimizations, including custom approaches that traded compute for bandwidth and communication for device memory. The resulting stack combined several key techniques: intra-node tensor parallelism for linear attention and the LM Head, ReplaySSM, W8A8 quantization, mixed-precision cache quantization using INT8/FP8/BF16, and Layer Split. On top of this, we introduced an Encode-Prefill-Decode (EPD) disaggregated architecture. Together, these optimizations improved end-to-end serving performance by roughly 3×. Both hardware utilization efficiency and per-token cost reached levels comparable to mainstream NVIDIA GPUs.

With the Infra Agent's feedback loop running throughout the optimization process, GLM-5.3-Flash went from initial model adaptation to production readiness in less than two weeks, ultimately tripling end-to-end throughput relative to the initial baseline. Figure 1 shows the performance trajectory of GLM-5.3-Flash from its first successful run to production launch.

> **Figure 1: The end-to-end throughput evolution of GLM-5.3-Flash.**

Throughout this process, we gradually realized that the Infra Agent's engineering effectiveness depended not only on the model's code generation and reasoning capabilities, but even more on whether the system could continuously provide useful feedback that could be traced to specific causes. A codebase provides only static context. Numerical discrepancies, performance regressions, and missed optimization targets in an inference system often arise from dynamic interactions across multiple layers, including kernel implementations, parallelism strategies, communication behavior, memory management, and serving orchestration.

Even if an agent can understand the entire codebase, feedback such as "numerical accuracy test failed," "TTFT increased by 30%," or "output throughput dropped by 20%" after a change still leaves it struggling to determine which layer is responsible, why its current hypothesis is wrong, and what it should test next. End-to-end metrics can tell an agent that results got worse, but they cannot explain why.

Therefore, alongside improving the agent's ability to write and modify code, we need to solve a more fundamental systems problem: **How do we turn sparse end-to-end results into fine-grained, attributable engineering feedback that directly guides the next action?** This is also the key to building an effective feedback loop for the Infra Agent.

# **From End-to-End Metrics to Attributable Feedback**

Traditional inference system optimization has no shortage of tests, logs, profiling tools, or microbenchmarks. But these are usually scattered across different tools and engineering stages. Experienced engineers use the results of a load test to decide what to observe next, progressively inspecting kernel outputs, execution timelines, communication events, or thread states, and connecting information from different tools.

For an agent, unless these observation and validation methods are organized into directly accessible, repeatable workflows, the feedback it can actually use remains sparse. It may know that throughput has missed the target, yet still be unable to determine:

- Is a particular kernel taking too long, or is the compute device sitting idle?
- Is KV Transfer itself too slow, or is higher-level scheduling failing to advance transfers promptly?
- For which input shapes does an optimization work, and under what conditions does it regress?

A single end-to-end metric cannot answer these questions. We therefore incorporated correctness tests, runtime logs, execution traces, runtime events, microbenchmarks, and end-to-end metrics into the agent's iteration workflow, breaking the full system optimization process into steps that can be observed and validated locally. Kernel-level comparisons verify numerical correctness. Microbenchmarks measure local performance under specific input conditions. Execution traces and runtime events reveal the timing relationships among computation, waiting, and communication. The agent can choose the appropriate validation method for its current hypothesis, rather than waiting for a full service deployment and end-to-end load test after every change.

We call this approach "dense feedback." Here, "dense" does not mean feeding the agent as many logs and metrics as possible. Instead, it emphasizes three characteristics.

**First, feedback must be sufficiently local.** Wherever possible, it should be tied to specific engine launch parameters, code changes, kernels, input conditions, threads, execution intervals, or code paths, helping the agent narrow the scope of the problem. For example, instead of reporting that "model accuracy dropped after a fusion optimization," identifying the output differences for a specific request before and after the change gives the agent a better basis for constructing a minimal reproduction and analyzing the cause.

**Second, feedback must be inexpensive and timely to obtain.** Whenever the agent proposes a hypothesis, makes a change, or constructs a set of controlled experiments, it should have an appropriate way to validate it. Questions that can be answered by a kernel test or a local microbenchmark should not require a full service deployment and end-to-end load test every time. Shorter validation cycles help the agent correct course promptly and spend less effort on unproductive hypotheses.

**Third, feedback must support objective verification.** Whether a change is correct and whether performance has improved should be determined by reference implementations, test results, and comparable experimental metrics. Runtime signals can help the agent identify possible causes, but correlations between observations alone cannot establish a root cause. Controlled experiments are still needed to verify whether a change to a specific path produces the expected effect.

Together, these three characteristics determine whether feedback is actionable. Correctness feedback answers, "Is the computation correct?" System behavior feedback identifies, "Where is the time going?" Performance feedback determines, "Which approach works better, and under what conditions?" Validation methods need not follow a fixed sequence. They should match the current hypothesis so that every experiment answers a specific question.

Local validation and end-to-end testing serve different roles in this process. The former eliminates incorrect or ineffective changes early and identifies candidates worth pursuing. The latter confirms whether local gains translate into real serving improvements and whether a proposed change introduces new regressions under actual workloads.

Based on these principles, the GLM-5.3-Flash launch established an optimization loop involving engineers, the Infra Agent, and the experimental environment. Engineers defined objectives and system boundaries. The agent handled analysis, hypotheses, and code changes. The experimental environment provided layered, timely, and verifiable feedback. Together, they transformed a diagnostic process previously connected by engineers' experience into an engineering workflow the agent could execute continuously.

> **Figure 2: The Infra Agent optimization loop built around dense feedback.**

This progression included both performance optimizations that directly increased throughput and bug fixes that did not immediately improve throughput but were essential to launching the system correctly and reliably. The following three cases illustrate how dense feedback helped the agent ensure that the system "computes correctly," explain "why it is not running fast enough," and explore "how to make it run faster."

# **Correctness Feedback: Helping the Agent Determine Whether the Model Computes Correctly**

Inference performance optimization must be grounded in numerical correctness. For the agent, validation begins with establishing exactly which computations the inference engine performs. High-level parallelism strategies change how kernel inputs are partitioned, which execution paths are taken, and how results are combined. Testing a kernel's output only under unpartitioned conditions is not enough to cover its behavior in an actual deployment.

To address this, we established a mapping from the inference engine's parallelism strategies to kernel implementations, converting system-level deployment configurations into kernel-level tasks the agent could verify individually. This mapping helped the agent identify which kernels a para
whiteros_e411251
🟧 echo.blog ⭐Z.ai reports that engineers and a GLM-5.3-powered Infra Agent used 'dense feedback' to bring GLM-5.3-Flash from initial adaptation to producZ.ai——
🟠 redditToward Recursive Self-Improvement: How GLM Built Its Own Inference Infrastructure
singularity
likeastar20912
🟠 redditshots fired at dario from glm
LocalLLaMA
Elux9118040
🟧 hnShow HN: Optimized runtimes for three VLAs on Jetson Thorhhuytho20

Interpretation history

Decision trace