2026-10-11 16:37 UTC

MLC Community claims its released XGrammar-2 guarantees structurally valid complex agent outputs with near-zero serving overhead and integrations across major inference engines, potentially making constrained tool calling a reusable serving primitive.

state: seedheat: lowuncertainty: mediumconvergesscott: mediumstructured-generation tool-calling agent-harnesses inference-servingMLC CommunityxAIDatabricksDeepSeekSGLangvLLMNVIDIA

What is this?

MLC Community released XGrammar-2, an open-source structured-generation engine for agent applications that constrains decoding to produce outputs conforming to specified schemas and tool-calling formats. Its Structural Tag protocol represents reasoning channels, tool calls, OpenAI Harmony-style responses, and custom structures, while caching, JIT compilation, batching, and speculative-decoding support are claimed to provide near-zero serving overhead and major compilation-speed gains. The project reports 100% structural schema accuracy—not guaranteed semantic correctness—and integrations with SGLang, vLLM, TensorRT-LLM, and MLC-LLM; the performance and accuracy figures in the supplied material largely come from the project’s own blog and paper.

Why it matters to Scott

XGrammar-2 advances Scott’s existing structural-reliability boundary by moving schema and tool-call conformance into constrained decoding, potentially replacing part of Ask’s malformed-output parsing and repair layer while retaining downstream semantic validation. It could affect harness design, but the claimed accuracy and near-zero overhead are primarily self-reported, and the named serving integrations do not directly include Scott’s current LiteLLM/Ollama path.
dev:concept.multi-format-tool-call-parsingdev:concept.validation-gated-llm-extractiondev:project.askip:concept.verification-loopsradar:vllm-silent-tool-parser-failuresradar:qwen25-quantization-task-divergenceradar:concept.tool-callingradar:concept.agent-verificationradar:concept.llm-serving
queries asked of Scott's wikis
  • constrained decoding versus validation and retry loops in agent harnesses
  • reliable structured outputs and tool-call protocol design
  • inference-layer primitives for model-independent agent runtimes
  • schema enforcement versus semantic correctness in tool calling
  • local model serving with vLLM SGLang or TensorRT-LLM
  • large tool catalogs and dynamic grammar caching

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 3866h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

05-03 14:00⭐ origin echo-reconstructedIntroduces XGrammar-2 with Structural Tag, constrained generation for complex agent formats, claimed efficiency gains up to 80x, and integra
MLC Community on blog (echo) · attributed from hn.story.49796733
—
09-22 04:15first on hacker news · published · +3398.3hXGrammar-2: Fast, Customizable Structured Generation for Tool Calling and Agents
matt_d
—
09-22 04:15amplified on hacker news 👑hn.story.49796733
matt_d
peak 2 · 0 comments · 98% of case engagement
09-22 04:20our radar first saw it · +3398.3hdiscovery anchor: hn.story.49796733—

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnXGrammar-2: Fast, Customizable Structured Generation for Tool Calling and Agents
Retrieved article excerpt

Open article · Retrieved 2026-09-22T04:21:46.499323+00:00

[MLC](https://blog.mlc.ai/)

- [Home](https://blog.mlc.ai/)

# XGrammar-2: Fast and Customizable Structured Generation for Tool Calling and Agents

May 4, 2026
• 
MLC Community

> **TL;DR.** XGrammar-2 is a major upgrade of XGrammar built for agent applications. It introduces **Structural Tag**, a composable JSON protocol that uniformly expresses OpenAI harmony format, tool calling, reasoning channels, and any custom output structure, exposed directly through serving engines’ API. Multiple **efficiency optimizations**, such as cross-grammar caching, repetition-state compression, and batching and speculative decoding support, ensure fast processing and minimal overhead even for huge structures. XGrammar-2 has been adopted by **xAI, Databricks, DeepSeek, and other leading AI companies** in their products. **SGLang, vLLM, TensorRT-LLM, and MLC-LLM** integrate it for strict tool calling and expose customization through API.

Over the past year, agent applications, from Claude Code to OpenClaw, have grown rapidly in complexity. These systems define sophisticated *harnesses* that LLMs must interact with by producing specific output structures, such as tool calls and structured JSON. As these structures grow more complex, they pose greater challenges for LLMs to follow reliably.

More than a year ago, we released [XGrammar](https://github.com/mlc-ai/xgrammar/), which uses constrained decoding to guarantee 100% structural correctness with near-zero overhead. Since then, many organizations and open-source projects have adopted XGrammar, with active community discussion and contributions. While XGrammar already handles JSON and other common structures efficiently, emerging agent applications demand far more complex structures, raising new challenges in both flexibility and efficiency.

To address these challenges, we are excited to introduce **XGrammar-2**: a major upgrade purpose-built for agent applications. It lets you easily express complex structures for agents, delivers high performance even for very large grammars, offers native cross-platform APIs, and remains fully backward compatible. In this post, we first recap XGrammar and then walk through the key features of XGrammar-2.

Figure 1: XGrammar-2 achieves 100% schema accuracy and delivers higher end-to-end accuracy on tool-calling tasks.

Figure 2: XGrammar-2 delivers up to 80x efficiency gain compared to XGrammar, and achieves near-zero overhead in LLM serving scenarios.

## A Recap of XGrammar

XGrammar uses constrained decoding to ensure LLM outputs conform 100% to a given structure. At each decoding step, constrained decoding produces a mask that blocks invalid tokens according to the structure. During sampling, invalid tokens are assigned zero probability, so only valid tokens will be generated. XGrammar’s key insight is precomputing an efficient token mask cache at compilation time, which substantially reduces mask generation time and achieves near-zero overhead during generation.

Figure 3: Constrained Decoding: Generating Output from a JSON Schema

XGrammar is best used to enforce format constraints, not to change the semantics of an LLM’s response. It helps downstream programs avoid fatal failures from malformed outputs, while keeping the impact on the model’s accuracy minimal. In our experiments, XGrammar ensured 100% valid tool-calling formats and, in many cases, improved tool-calling accuracy by eliminating format-related failures.

## Structural Tag: Abstraction for All Tool Calling and Complex Structures

Figure 4: Workflow of the Structural Tag

Agent applications are pushing LLMs to follow increasingly complex formats. One representative example is the [**OpenAI Harmony Format**](https://developers.openai.com/cookbook/articles/openai-harmony), which splits output into multiple channels, including reasoning, tool calling, and final response, each with its own format. Each open-source model also define their own tool calling formats. Supporting all of these requires significant effort from serving engines and downstream applications, and may still fail to match the official specification.

XGrammar-2 introduces **Structural Tag**, a JSON-based DSL that provides a unified, lightweight, and extensible way to describe the diverse structures agents need, from OpenAI Harmony format to open-source model tool calling protocols and many other custom formats.

For example, a DeepSeek V4 output with reasoning and a tool call looks like this:

```
Let me check the weather in Beijing.</think>
I'll look that up for you.
<|DSML|tool_calls>
<|DSML|invoke name="get_weather">
<|DSML|parameter name="city" string="true">Beijing</|DSML|parameter>
</|DSML|invoke>
</|DSML|tool_calls>
```

There are two distinct parts here. The first part is free-form reasoning that continues until the `</think>` token. The second part is either more free text or a structured tool call, triggered when the model emits the `<|DSML|tool_calls>` marker. The corresponding Structural Tag captures this two-part structure directly:

```
{
  "type": "structural_tag",
  "format": {
    "type": "sequence",
    "elements": [
      /* Reasoning Part */
      { "type": "tag", "begin": "", "content": { "type": "any_text" }, "end": "</think>" },
      /* Output & Tool Calling Part */
      {
        "type": "triggered_tags",
        "triggers": ["<|DSML|tool_calls>"],
        "tags": [
          {
            /* DeepSeek Tool Calling Format */
            "type": "tag",
            "begin": "<|DSML|tool_calls>\n",
            "content": {
              "type": "tag",
              "begin": "<|DSML|invoke name=\"get_weather\">\n",
              "content": {
                  "type": "json_schema", "json_schema": {...}, "style": "deepseek_xml"
              },
              "end": "</|DSML|invoke>\n"
            },
            "end": "</|DSML|tool_calls>\n"
          }
        ],
        "excludes": ["<think>", "</think>"]
      }
    ]
  }
}
```

This example comprises five Structural Tag types, each with a clear role:

1. **Sequence** chains two parts together: the reasoning section followed by the tool-call section.
2. **Tag** matches a `begin` marker, some constrained content, and an `end` marker. Here the reasoning tag has an empty `begin` because the chat template already appends to the prompt.
3. **AnyText** matches arbitrary text until the enclosing tag’s `end` marker, which is exactly what we need for free-form reasoning content.
4. **TriggeredTags** lets the model produce free text by default, but once it emits a trigger string, the output must follow the corresponding structured tag. The `excludes` field prevents `<think>` and `</think>` from appearing in the final output section.
5. **JSONSchema** constrains the tool’s arguments to a given schema. The `style="deepseek_xml"` option tells XGrammar to expect arguments in DeepSeek’s XML parameter format rather than raw JSON.

The key idea behind Structural Tag is that these types are **composable**. JSON Schema, regex, literal strings, and token IDs are all first-class atomic types within the language. By nesting and combining them, you can describe arbitrarily complex output structures, from a simple JSON response to a multi-part reasoning-plus-tool-call format like the one above.

XGrammar ships with **built-in Structural Tags** for common models such as DeepSeek V4, Qwen 3.6, GPT-OSS, and more. The structural tag is already integrated into SGLang, vLLM, TensorRT-LLM, and other serving engines, providing strict tool calling and reasoning support out of the box.

Structural Tag is also exposed as an **OpenAI-compatible response format** by serving engines, so you can customize your own output structure for your agent application:

```
# Assume the client is connected to a hosted SGLang, vLLM, or TensorRT-LLM server.
response = client.chat.completions.create(
    model="deepseek-ai/DeepSeek-V4",
    messages=[...],
    extra_body={
        "response_format": {
            "type": "structural_tag",
            "format": {
                "type": "tag",
                "begin": "<answer>",
                "content": {
                  "type": "json_schema",
                  "json_schema": {
                    "type": "object",
                    "properties": {
                      "status": { "type": "string" },
                      "message": { "type": "string" }
                    }
                  }
                },
                "end": "</answer>",
            }
        }
    }
)
```

For example, we built a multimodal video agent with Molmo-2 model that detects objects in videos and renders annotated outputs. By specifying the desired format with structural tags, including each object’s time range, location, and name, we obtain precise model outputs that can be mapped directly onto the video in downstream processing.

Figure 5: Multimodal Agent Powered by Structural Tags for Precise Output

## Scaling to Complex Structures with Minimal Overhead

The structures agents rely on are also growing larger. Tool calling in particular can involve dozens or even hundreds of tools per session. This puts enormous pressure on structured generation: both grammar preprocessing and mask generation incur substantial overhead, slowing down the entire request. XGrammar-2 introduces a series of optimizations to tackle complex structures and ensure that even very large grammars incur only minimal overhead.

### Cross-grammar Cache

Different grammar structures and different parts of the same grammar often share many common sub-structures. For example, different JSON schemas all share the same string field described by `{"type": "string"}`. These repeated structures can be fully reused during preprocessing. XGrammar-2 implements an automaton-based hierarchical hashing algorithm that automatically finds shared parts within and across grammars, maximizing reuse of grammar preprocessing. In JSON Schema compilation for 50 tools, our experiments show that nearly 50% of structures end up reused.

### Repetition State Compression

Repetition is common in many structures. For example, an array of up to 1M items described by `{"type": "array", "maxItems": 1000000}` contains a large repeated component. If handled naively, preprocessing such a grammar requires `O(repetition_count)` time. XGrammar-2 compresses this to `O(1)` by introducing a new grammar primitive, **repetition**, whose size stays constant regardless of how many repetitions are allowed. We also designed specialized parsing and token mask cache algorithms so that this new primitive is just as easy to handle as any other grammar construct. For complex JSON schema structures, our experiment shows repetition compression reduces the compression from 534 ms to 5.37 ms, a 100x time reduction.

### Serving and Speculative Decoding Support

XGrammar-2 also supports batching and speculative decoding, both key features in modern serving systems. For batching, it provides [**batch APIs**](https://xgrammar.mlc.ai/docs/api/python/grammar_matcher.html#xgrammar.BatchGrammarMatcher) that flexibly combine and process multiple grammar states in one pass on the C++ side, avoiding Python-side loops and reducing batching overhead.

For speculative decoding, XGrammar-2 provides [`traverse_draft_tree`](https://xgrammar.mlc.ai/docs/api/python/grammar_matcher.html#xgrammar.GrammarMatcher.traverse_draft_tree) to traverse a draft tree once and generate masks for all nodes. For finer-grained control, grammar states can also be forked and rolled back to walk through the tree manually.

Figure 6: Overlapping Pattern for Constrained Decoding and Speculative Decoding

This also enables constrained decoding to overlap with speculative decoding. While the target model verifies the draft tree on the GPU, XGrammar walks the same tree on the CPU and generates masks in parallel, reducing overhead further. We collaborated with serving engine teams to integrate this pattern into speculative decod
matt_d20
🟧 echo.blog ⭐Introduces XGrammar-2 with Structural Tag, constrained generation for complex agent formats, claimed efficiency gains up to 80x, and integraMLC Community——

Interpretation history

Decision trace