2026-10-11 16:38 UTC

Makora claims its automation-assisted optimization of Qwen3.5-397B-A17B-FP8 on Ironwood TPUs delivers up to 5Γ— stock vLLM-TPU performance and exceeds B200 performance in its high-interactivity regime, potentially making TPUs more competitive for interactive open-model serving.

state: corroboratedheat: lowuncertainty: mediumnovelscott: mediuminference-economics inference-optimization tpu automated-tuningMakoraGoogleNVIDIAvLLM

What is this?

Makora, an inference-optimization vendor, published a retrospective claiming its 'Swarm' automation tooling spent roughly a month (May–June 2026) optimizing Qwen3.5-397B-A17B-FP8 serving on Google's Ironwood (TPU v7), reaching up to 5Γ— stock vLLM-TPU performance and beating an Nvidia B200 in high-interactivity decode β€” via bespoke Pallas MoE-GEMV kernels that use the TPU's 64MB on-chip VMEM as ring-buffered expert-weight DMA buffers, arguing grouped GEMM wastes work at low concurrency (~5% expert overlap across 3 concurrent requests). Google's own engineering write-up independently reports ~3.1Γ— decode / ~4.7Γ— prefill gains for the same model on Ironwood, integrated natively into vLLM and SGLang, and a cost analysis of the same model puts Ironwood 19–50% ahead of B200/B300 in tokens-per-dollar at interactive rates. A second vendor, Inferact, separately claims its Kimi K3 TPU megakernel hits 709 tok/s low-concurrency decode vs 450 on GB200 β€” turning one vendor claim into a two-company convergent pattern of aggressive custom-kernel TPU serving beating GPUs on interactive decode. Caveats nothing supplied resolves: all headline numbers are self-reported and unreproduced, SemiAnalysis earlier dismissed vLLM-TPU benchmarks as irrelevant because the stock stack was unoptimized (so the '5Γ— stock' baseline is a weak starting point), and whether Makora's gains derive from its automation tooling versus expert hand-tuning is not established.

Why it matters to Scott

A second vendor's parallel claim plus Google's own vLLM/SGLang-integrated Ironwood write-up turn 'TPU competitive with NVIDIA on interactive decode of large open models' from one blog into a convergent pattern, which bears on the open-model hardware-portability and lock-in argument his canon carries (TPU as a credible second source strengthens the design-for-swapping prescription) and on the interactive-decode economics his real-time AI work depends on. It stays a watch item rather than a build/publish trigger: every headline number is self-reported and unreproduced, he runs CUDA/local GPU rather than TPU, and the concretely checkable test is whether Makora's kernel work actually lands upstream in vLLM.
ip:concept.vendor-lock-inip:concept.model-perishabilitydev:concept.hardware-aware-local-inferenceip:concept.real-time-ai-systemsradar:concept.inference-economicsradar:concept.vllmradar:qwen38-kaggle-tpu-servingradar:vllm-hardware-agnostic-modelsradar:vllm-tenstorrent-pluginradar:codex-autoresearch-gpu-kernel-speedup
queries asked of Scott's wikis
  • real-time AI response window latency thresholds interactivity
  • inference economics tokens per dollar hardware tradeoffs
  • coding agent token velocity decode latency experience
  • open-source inference stack vLLM contribution upstreaming
  • TPU vs GPU serving portability hardware lock-in open models
  • automated performance engineering kernel tuning tooling

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 626h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-15 14:00⭐ origin echo-reconstructedMakora reports up to 5Γ— faster Qwen inference than the stock vLLM-TPU fork after a month of optimization, approaching B200 performance in th
Pawel Niegowski, Sofiia Serdiuk, Blazej Tez, Emmanuel Rassou on blog (echo) Β· attributed from hn.story.49741816
β€”
09-17 14:59first on hacker news Β· published Β· +49.0hMaximizing TPU interactivity with automated full stack performance optimization
tripplyons
β€”
09-17 14:59amplified on hacker newshn.story.49741816
tripplyons
peak 2 Β· 0 comments Β· 25% of case engagement
09-24 06:32amplified on hacker news πŸ‘‘hn.story.49826989
xutingl
peak 5 Β· 1 comments Β· 76% of case engagement
09-17 15:20our radar first saw it Β· +49.4hdiscovery anchor: hn.story.49741816β€”
pace: p45 vs 1032 stories at the 336h mark (now 626h old) β€” ahead of agenticos-self-hosted-governance (1.1x), behind anthropic-ci-test-selection-redesign (0.9x)

Evidence (3) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnMaximizing TPU interactivity with automated full stack performance optimization
Retrieved article excerpt

Open article Β· Retrieved 2026-09-17T16:22:57.198193+00:00

Products

[Models](https://www.makora.com/#models)

[Pricing](https://www.makora.com/pricing)

[Blog](https://www.makora.com/blog)

[Contact us](https://www.makora.com/contact)

[Try for free](https://app.makora.com)

Products

[Models](https://www.makora.com/#models)

[Pricing](https://www.makora.com/pricing)

[Blog](https://www.makora.com/blog)

[Contact us](https://www.makora.com/contact)

[Try for free](https://app.makora.com)

Our products

[MakoraInference](https://www.makora.com/model-performance)

[MakoraGenerate](https://www.makora.com/mako-generate)

RESOURCES

[Docs](http://docs.makora.com/)

CASE STUDIES

[Code Translation](https://www.makora.com/use-cases/code-translation)

[Performance Optimization](https://www.makora.com/use-cases/performance-optimization)

COMPANY

[About](https://www.makora.com/about-us)

[Careers](https://jobs.mako.dev)

[Contact Us](https://www.makora.com/contact)

[Blog](https://www.makora.com/blog)

[Pricing](https://www.makora.com/pricing)

[Try for free](https://app.makora.com)

Our products

[MakoraInference](https://www.makora.com/model-performance)

[MakoraGenerate](https://www.makora.com/mako-generate)

RESOURCES

[Docs](http://docs.makora.com/)

CASE STUDIES

[Code Translation](https://www.makora.com/use-cases/code-translation)

[Performance Optimization](https://www.makora.com/use-cases/performance-optimization)

COMPANY

[About](https://www.makora.com/about-us)

[Careers](https://jobs.mako.dev)

[Contact Us](https://www.makora.com/contact)

[Blog](https://www.makora.com/blog)

[Pricing](https://www.makora.com/pricing)

[Contact us](https://www.makora.com/contact)

[Try for free](https://app.makora.com)

# Maximizing TPU interactivity with automated full stack performance optimization software

# Maximizing TPU interactivity with automated full stack performance optimization software

# Maximizing TPU interactivity with automated full stack performance optimization software

# Maximizing TPU interactivity with automated full stack performance optimization software

Makora closed the gap between B200 and Ironwood TPUs in just a few weeks

Makora closed the gap between B200 and Ironwood TPUs in just a few weeks

Written by

Pawel Niegowski

Pawel Niegowski

Sofiia Serdiuk

Sofiia Serdiuk

Blazej Tez

Blazej Tez

Emmanuel Rassou

Emmanuel Rassou

Published on

Sep 16, 2026

Sep 16, 2026

We spent a month (during May/June 2026) optimizing [Qwen-3.5-397B-A17B-FP8](https://huggingface.co/Qwen/Qwen3.5-397B-A17B-FP8) inference using Ironwood TPUs. Our challenge was to investigate whether closing the gap with NVIDIA B200 GPUs was possible for this relatively new model architecture.

We were interested in the following performance optimizations:

- maximizing throughput/TPU with a 10 tok/s/user interactivity floor,
- maximizing throughput/TPU with a 115 tok/s/user interactivity floor,
- maximizing tok/s/user at any concurrency

Importantly, we were not investigating **speculative decoding or multi-token prediction .** Neither were we performing any post-training methods that could impact quality or accuracy. We were primarily focused on speeding up the raw autoregressive token performance.

In a single month, without prior experience with the TPU ecosystem, we significantly improved performance by up to 5x compared to the stock vllm-tpu fork, approaching B200 performance on the high-throughput regime, and outperforming B200 on the high interactivity regime! This confirms that software is a real gap when it comes to performance engineering, and that Makora's automation tools are portable across hardware.

> Higher is more total throughput, more to the right is more tok/s/user, c is the number of concurrent requests.

In other words:

- light grey is what you get if you run `vllm serve` 🐌
- dark grey is what you get if your performance engineer tunes vLLM hyperparameters for you 🐌🏁
- green and purple curves are what you get with Makora 🦈

At the time this post was written, the optimization changes were in the process of being upstreamed to the [vLLM core](https://github.com/vllm-project/vllm) and [vLLM TPU backend](https://github.com/vllm-project/tpu-inference) repositories. Below we will describe the high interactivity optimizations we came up with, then we will outline some of performance automation tools used to achieve our results.

## Interactivity is a strange beast...

Until recently, LLM inference was focused on throughput β€” it was assumed extreme concurrency is all you need, and serving LLMs is only cost-effective at a massive scale. The rise of coding agents changed that dynamic, as reaching 50, 100 or 150 tok/s/user became very valuable for interactive, synchronous work with a human operator.

As LLMs such as Qwen are autoregressive, the number of concurrent requests puts a hard cap on possible parallelism. You cannot reliably generate the n+1-th token without first running the n-th token through the entire model. As such, maximizing low-concurrency interactivity becomes a game of hiding as much memory access latency as possible. Bandwidth and compute capacity are fully utilized only for a fraction of a step's runtime, and the rest of the time is spent starved for data to process.

Qwen-3.5-397B-A17B is a massive model and, due to memory constraints, it must be deployed on 4 TPUs / 8 chiplets. 8-way tensor parallel sharding slices the weights and intermediate buffers into even smaller parts and introduces further waits on collective operations.

## ...and so are TPUs

GPU kernels are designed around hundreds of compute units, each processing several *warps* (Nvidia) or *wavefronts* (AMD) at the same time, with multiple level of schedulers arbitrating which instructions should run at any given time. A block of warps may share a tiny, **kilobytes-sized** memory buffer called shared memory as a scratchpad. A single GPU kernel runs in thousands of parallel, resource-constrained invocations, with each warp delivering a small slice of the result. Even decades later, the legacy of the first [programmable shader](https://en.wikipedia.org/wiki/Shader) processors is visible in GPU design.

A Google TPU doesn't do any of that. Instead, in just a few cycles, a Tensor Core VPU processes up to **128x8** elements each instruction, growing up to **256x256** for MXU operations, with an efficient memory loading pipeline keeping the arithmetic units fed. Where a GPU runs thousands of tiny parallel programs, a TPU has *one* kernel invocation running at a time, with no interruptions.

GPUs have multiple levels of cache β€” L2, L1, sometimes L0. TPU designers dropped the concept of device-managed caches entirely β€” instead, they provided a huge fast-access bank called VMEM. In the Ironwood architecture, each of the two chiplets has **64MB of VMEM** that the programmer can manually utilize for caching, scratch space or intermediate result storage.

If you've ever written a GPU kernel, by this point you should get excited β€” what can we do with all this space and processing power? How many algorithms become straightforward if you don't have to divide work into tiny slices to fit a GPU warp? Let's find out.

## Grouped GEMM is overrated

Mixture of Experts inference is a well-trodden problem and preexisting, high-performance grouped GEMM implementations should, in general, be faster than bespoke custom kernels.

**When decoding one request, or even 3-4 requests at a time, this is completely wrong.**

Grouped GEMM relies on a straightforward observation. If we have many experts and many tokens, we should rearrange the post-attention latents so that each expert's incoming tokens are contiguous in memory. Thus, we can launch many instances of the same kernel, targeting different experts and process each expert's work in parallel.

Now, for a specific example, Qwen-3.5 uses one shared expert and ten routed experts, picked by a routing network for each token out of a set of 512.

So now imagine you have a single request to handle, a single token to decode, and a TPU. Each group has, by definition, a size of *one* as experts cannot be chosen twice. The input is shared for the up projection, and yet in a grouped GEMM kernel it would be duplicated ten times. Moreover, a grouped GEMM kernel is still a GEMM kernel, so it will tile in squares or rectangles for weight reuse... which doesn't happen at all, as in this scenario each expert call is a GEMV and **no weight is used twice!**

You'd think this gets better with a few requests in flight, but we measured the overlap between experts chosen with three concurrent requests β€” only 5% of the chosen experts were duplicated! As such, **at low concurrency grouped GEMM doesn't justify its overhead**.

## GEMV is all you need

So let's write a custom Pallas kernel that will work with these constraints. If you've ever worked with Triton, you should find Pallas easy to read.

A seasoned performance engineer will notice this problem boils down to rotating through the expert weights as fast as possible, and everything else will be hidden in the load latency. In GPU programming, we'd reach for a lot of warps and in-warp double buffering.

On an Ironwood TPU, fortunately, we have a lot of VMEM to spare, and for a few experts at a time, we can cram a full 1/8th TP shard of each into VMEM at once. But we need to make up for just having one mega-warp. So how many experts can we load in parallel, into our VMEM "ring buffer"?

```
# Prime the ring: fire NBUF_ whole-expert weight DMAs up front, so up to
# NBUF_ expert loads are in flight across the HBM engines at once.
for j in range(min(NBUF_, TOP_K_)):
    pltpu.make_async_copy(rhs_ref.at[pl.ds(ids_ref[j], 1)],
                          w_bufs_ref.at[pl.ds(j, 1)],
                          sem_ref.at[j]).start()
    # ... (+ start the matching weight-scale DMA) ...

for i in range(TOP_K_):
    buf = i % NBUF_
    pltpu.make_async_copy(rhs_ref.at[pl.ds(ids_ref[i], 1)],
                          w_bufs_ref.at[pl.ds(buf, 1)],
                          sem_ref.at[buf]).wait()
    # ... (wait on the matching weight-scale DMA) ...
    w_fp8 = w_bufs_ref[buf]
    s = s_bufs_ref[buf]
    w_dequant = (w_fp8.astype(jnp.float32).reshape(K_BLOCKS, QB, N) * s) \
        .reshape(K, N).astype(DTYPE_LHS)
    out = jnp.matmul(lhs_ref[...], w_dequant, preferred_element_type=jnp.float32)
    out_bf16 = out.astype(DTYPE_OUT)   # matmul output rounded to bf16 first
    if FUSE_SILU:                      # combined gate+up β†’ SwiGLU in-kernel
        I = N // 2
        gate = out_bf16[:, :I].astype(jnp.float32)
        up = out_bf16[:, I:].astype(jnp.float32)
        act = (gate * jax.nn.sigmoid(gate) * up).astype(DTYPE_OUT)
        o_scratch_ref[pl.ds(i * M_PAD, M_PAD), :] = act
    else:
        o_scratch_ref[pl.ds(i * M_PAD, M_PAD), :] = out_bf16
    nxt = i + NBUF_                    # prefetch the expert NBUF iters ahead
    if nxt < TOP_K_:                   # into the buffer slot just consumed
        pltpu.make_async_copy(rhs_ref.at[pl.ds(ids_ref[nxt], 1)],
                              w_bufs_ref.at[pl.ds(buf, 1)], sem_ref.at[buf]).start()
```

```
# Prime the ring: fire NBUF_ whole-expert weight DMAs up front, so up to
# NBUF_ expert loads are in flight across the HBM engines at once.
for j in range(min(NBUF_, TOP_K_)):
    pltpu.make_async_copy(rhs_ref.at[pl.ds(ids_ref[j], 1)],
                          w_bufs_ref.at[pl.ds(j, 1)],
                          sem_ref.at[j]).start()
    # ... (+ start the matching weight-scale DMA) ...

for i in range(TOP_K_):
    buf = i % NBUF_
    pltpu.make_async_copy(rhs_ref.at[pl.ds(ids_ref[i], 1)],
                          w_bufs_ref.at[pl.ds(buf, 1)],
                          sem_ref.at[buf]).wait()
    # ... (wait on the matching weight-scale DMA) ...
    w_fp8 = w_bufs_ref[buf]
    s = s_bufs_ref[buf]
    w_dequant = (w_fp8.astype(jnp.float32).reshape(K_BLOCKS, QB, N) * s) \
        .reshape(K, N).astype(DTYPE_LHS)
    out = jnp.matmul(lhs_ref[...], w
tripplyons20
🟧 echo.blog ⭐Makora reports up to 5Γ— faster Qwen inference than the stock vLLM-TPU fork after a month of optimization, approaching B200 performance in thPawel Niegowski, Sofiia Serdiuk, Blazej Tez, Emmanuel Rassouβ€”β€”
🟧 hnInferact's Kimi K3 Megakernel hits 700 tok/s on TPU, Beating GPUxutingl51

Interpretation history

Decision trace