2026-10-11 16:38 UTC

Tenstorrent claims its released vLLM TT Plugin serves supported text and multimodal models on its accelerators through the existing OpenAI-compatible API without modifying vLLM core, potentially enabling backend migration without rewriting clients.

state: watchingheat: lowuncertainty: mediumconvergesscott: mediumllm-serving tenstorrent vllm ai-infrastructureTenstorrentvLLM

What is this?

Tenstorrent’s vLLM TT Plugin is an available backend plugin for serving supported text and multimodal models on Tenstorrent accelerators; the supplied GitHub repository and official documentation identify it as using vLLM’s standard plugin mechanism. The vLLM blog says installation alongside vLLM automatically registers the hardware when TT-Metal’s `ttnn` is importable, while preserving the OpenAI-compatible API, request format, and client code. These snippets support the proposed client-portability benefit, but do not explicitly verify that no vLLM core modifications are required or independently demonstrate migration reliability or performance. Model-support lists differ between the blog and repository, so exact model and hardware coverage needs checking; Moreh’s separate serving product is not evidence of this plugin’s performance.

Why it matters to Scott

Tenstorrent’s released backend extends Scott’s Composable Bespoke position—swappable model dependencies behind stable interfaces—to accelerator-backed serving, offering a concrete portability evaluation for his shared OpenAI-compatible LiteLLM gateway rather than just another anti-lock-in example. The radar tracks vLLM and related portability developments but not this plugin in the supplied hits; no Scott deployment on Tenstorrent, end-to-end gateway/tool-call compatibility, migration reliability, or economic advantage is established.
ip:concept.composable-bespokeip:concept.model-perishabilitydev:technology.litellmradar:concept.vllmradar:concept.llm-servingradar:concept.llm-apisradar:lemonade-local-ai-runtime
queries asked of Scott's wikis
  • OpenAI-compatible endpoints backend portability vendor lock-in
  • vLLM serving infrastructure inference backend migration
  • local inference accelerator economics hardware independence
  • agent harness model-provider abstraction tool-calling compatibility
  • out-of-tree plugins upstream forks maintenance costs

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 842h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-06 14:00⭐ origin echo-reconstructedIntroduces an out-of-tree vLLM plugin with automatic Tenstorrent platform registration, supported text and multimodal architectures, phase-s
Tenstorrent Team on blog (echo) · attributed from hn.story.49655653
—
09-11 09:35first on hacker news · published · +115.6hServing LLMs on Tenstorrent Hardware: Inside the vLLM TT Plugin
ankitg12
—
09-11 09:35amplified on hacker news 👑hn.story.49655653
ankitg12
peak 2 · 0 comments · 98% of case engagement
09-11 10:21our radar first saw it · +116.4hdiscovery anchor: hn.story.49655653—
pace: p11 vs 519 stories at the 720h mark (now 842h old) — behind addom-local-coding-harness (0.5x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnServing LLMs on Tenstorrent Hardware: Inside the vLLM TT Plugin
Retrieved article excerpt

Open article · Retrieved 2026-09-11T10:22:56.611150+00:00

Serving LLMs on Tenstorrent Hardware: Inside the vLLM TT Plugin September 7, 2026 16 min read Tenstorrent Team # hardware # ecosystem Table of Contents Today we are introducing vLLM TT Plugin , which brings Tenstorrent accelerators to vLLM through the standard out-of-tree platform plugin mechanism. Install it alongside vLLM and, whenever ttnn from TT-Metal is importable, Tenstorrent hardware is discovered and registered as a vLLM platform automatically. The serving surface does not change: the same OpenAI-compatible API, the same request format, the same client code. The more interesting part of this project is not that the backend exists. It is that a Tenstorrent device does not look much like a GPU, and vLLM's plugin interfaces turned out to be general enough that we could express those differences - a phase-constrained scheduler, a different data-parallel topology, a sampling path that partly lives on device - entirely outside vLLM core. Supported models The plugin registers Tenstorrent-backed architectures under a TT -prefixed convention, so a checkpoint is picked up by the architecture it declares rather than by name: Model family Architectures Llama 3.1 / 3.2 / 3.3 TTLlamaForCausalLM Llama 3.2 Vision TTMllamaForConditionalGeneration Qwen 2.5 / Qwen 3 TTQwen2ForCausalLM , TTQwen3ForCausalLM Qwen 3.5 / Qwen 3.6 TTQwen3_5ForConditionalGeneration Qwen 2.5-VL / Qwen 3-VL TTQwen2_5_VLForConditionalGeneration , TTQwen3VLForConditionalGeneration Mistral / Mistral 3 TTMistralForCausalLM , TTMistral3ForConditionalGeneration Gemma 3 TTGemma3ForConditionalGeneration Gemma 4 TTGemma4ForCausalLM , TTGemma4ForConditionalGeneration , TTGemma4UnifiedForConditionalGeneration DeepSeek V3 TTDeepseekV3ForCausalLM GPT-OSS 20B / 120B TTGptOssForCausalLM The classes behind those names ship in TT-Metal , alongside the runtime itself: each is a vLLM-facing generator wrapped around a hand-written TTNN implementation of the model. The plugin carries no model code - it registers the names, and tt-metal provides what they resolve to. Because the match is on architecture, one entry can cover several releases - TTQwen3_5ForConditionalGeneration is what serves Qwen/Qwen3.6-27B , for instance. Multimodal coverage is worth calling out, since new backends often stay text-only for a long time: Llama 3.2 Vision, Qwen-VL, Qwen 3.6, Mistral 3, and Gemma 3 all serve through the plugin today. Models do not have to be built into the plugin. Pointing EXTRA_MODELS_DIR at a directory of bundle folders, each holding a vllm_metadata.json and an adapter class, registers architectures at startup under the TT<HFArch> convention. A distribution tool can ship a ready-to-serve model without a source edit, and TT_VLLM_BUILTIN_MODELS=0 narrows the registry to only what was supplied. That covers what you can serve today. The rest of this post is how it works: the design choices a mesh architecture forces on a GPU-shaped serving stack, and what we learned making them. Why a Tenstorrent backend looks different A Tenstorrent system is a mesh of cores and chips connected by an on-fabric network . A single card such as n150 or n300 is already a small mesh; a QuietBox is a larger one; a Galaxy is 32 Wormhole chips wired into a topology the runtime configures directly ( FABRIC_1D , FABRIC_2D , FABRIC_1D_RING ). Programs are compiled and traced against a mesh shape, and the fabric moves data between chips as part of the compiled program rather than as a collective call issued by the host. The models served through this plugin are hand-written TTNN implementations for TT mesh, from a two-chip n300 up to a 32-chip Galaxy. Within that system they run the same parallelization playbook one would use on GPUs - tensor parallelism across chips, data parallelism across submeshes - but expressed in TTNN and compiled into the mesh program rather than configured as runtime ranks. That hand-tuning is what delivers better tokens/$; we will not quote numbers here, current figures live on tenstorrent.com and GitHub . Diagram comparing host-issued collectives on GPUs with a compiled Tenstorrent mesh program Figure 1: Where cross-chip parallelism lives. In a GPU-shaped stack the host issues collectives on every layer and parallelism is a runtime choice expressed as tensor-parallel and pipeline-parallel ranks. On Tenstorrent, the mesh is compiled and traced as one program and the fabric moves data between chips inside it, so the host submits and reads once per step. That compilation model - one traced program for the whole mesh - drives nearly everything downstream: There are no tensor-parallel or pipeline-parallel ranks to configure. A 70B model on Galaxy is not "TP=32 processes"; it is one program compiled for a 32-chip mesh. MESH_DEVICE=TG replaces --tensor-parallel-size , and the plugin rejects -tp / -pp outright rather than pretending to honor them. The parallelism that best fits the (model, mesh) combination is implemented in the model code. The unit of work is a whole traced step. Device execution is dominated by replaying a captured trace for a fixed batch shape, which makes homogeneous, shape-stable batches dramatically cheaper than heterogeneous ones. Sampling can happen on device. Because the mesh program can carry sampling to the end, the token can often come back already chosen, and the host never sees the logits. Each of those is in tension with an assumption somewhere in a GPU-shaped inference stack. The sections that follow are how we resolved them. Plugging in, not forking vLLM's hardware plugin mechanism was introduced in May 2025 with vllm-ascend and vllm-spyre among its first users, and the pluggable-scheduler work that came out of the Spyre effort is what makes our approach viable at all. We depend on it heavily. The plugin registers two entry points: Entry point group Name Target vllm.platform_plugins tt vllm_tt_plugin.entrypoints:platform_plugin vllm.general_plugins tt_model_registry vllm_tt_plugin.entrypoints:register platform_plugin() returns TTPlatform only when ttnn is importable , so installing the package into an ordinary CUDA environment cannot accidentally select the Tenstorrent platform. From there, everything flows through a single handoff. TTPlatform.check_and_update_config() validates the configuration, registers model architectures, and swaps in Tenstorrent-owned runtime classes through vLLM's existing extension points: vLLM config field TT implementation parallel_config.worker_cls vllm_tt_plugin.worker.TTWorker scheduler_config.scheduler_cls vllm_tt_plugin.scheduler.TTScheduler or vllm_tt_plugin.lane_scheduler.TTLaneCoordinator Device-specific options ride on vLLM's generic additional-config namespace rather than through new CLI flags: --additional-config.tt.sample_on_device_mode all --additional-config.tt.fabric_config FABRIC_1D_RING Nothing Tenstorrent-specific lives in vLLM core. That is the property that decides whether a backend stays usable: support tracks vLLM's release cadence rather than ours, and nobody ends up stranded on a fork three months behind upstream. We currently validate against a pinned vLLM release and are widening that window as the plugin's API surface settles. Phase-based scheduling: prefill-only or decode-only steps Upstream vLLM's V1 scheduler is token-budget based, and deliberately so. A request has computed tokens and target tokens; each step hands out more token work subject to budgets. Prefill and decode are not separate modes, which is exactly what allows chunked prefill and mixed-progress batches to fall out naturally. The Tenstorrent path is more constrained. Every scheduling step resolves to one of three outcomes: prefill-only decode-only empty There are no mixed prefill+decode batches. Chunked prefill is supported within that constraint: a prompt that exceeds the per-step token budget is split across multiple prefill steps, and decode-only steps are interleaved between the chunks, so in-flight requests keep advancing while a long prefill is in flight. Prefill work is still admitted first by default, so then the decode steps run with bigger, more efficient batches; if no prefill can be admitted but decode requests are running, the step is decode-only, so progress continues and KV pressure can relax. Timeline comparing upstream token-budget steps with Tenstorrent phase-homogeneous steps Figure 2: The same long prompt under both scheduling models. Upstream spreads it across four chunked steps and mixes decode work for other requests into those same steps. On Tenstorrent a step is still all-prefill or all-decode: the prompt runs as prefill-only chunks with decode-only steps interleaved between them, so every step keeps a stable, traceable shape while in-flight requests keep advancing. This is the design choice most likely to raise an eyebrow, so it is worth being precise about what it costs and what it does not. What it buys. Traced execution rewards batch-shape stability: a step that is uniformly prefill or uniformly decode replays a trace captured for exactly that shape, while a step mixing the two would need a shape the trace was never captured for. The phase separation itself is not a Tenstorrent eccentricity: the largest GPU deployments make the same choice deliberately, running prefill and decode on entirely separate instances - disaggregated serving . The Tenstorrent scheduler applies the same split at step granularity within one engine rather than at instance granularity across a fleet. What it does not cost. Continuous batching still holds in the broad sense. Requests arrive into waiting , may be parked in skipped_waiting while structured-output grammar compiles, are admitted while other requests remain active, can be preempted back, and complete independently. The restriction is within a device step, not across the request lifecycle. What it does cost. The interleave granularity is a whole step. Upstream mixes a prefill chunk and ongoing decode into the same step; the Tenstorrent scheduler alternates, so a decode request still waits out each prefill chunk between its own steps, and each mode switch drains the async decode overlap pipeline described below. Both are scheduling-policy costs, not fundamental limits: nothing in the hardware or in vLLM prevents capturing a mixed-shape step in the future versions. Single-process lane data parallelism on Galaxy This is the part with no analogue elsewhere in vLLM, and the piece we are most interested in feedback on. Some Tenstorrent models - Llama 3.3 70B via TT_LLAMA_TEXT_VER=llama3_70b_galaxy , Qwen3-32B via TT_QWEN3_TEXT_VER=qwen3_32b_galaxy , and GPT-OSS - are served by single-execute generators: one program spanning the entire Galaxy mesh, executed once per step. There is no submesh to give a second engine process. Standard multi-process data parallelism, which assigns each rank its own devices, simply has nothing to partition. But: these models are single- weights and single- execute , yet they keep four independent data-parallel KV caches , each on its own DP submesh. So there is nothing to partition at the process level, and four things to schedule independently. Our initial implementation was to give each DP rank its own process, just as vLLM normally does. However, given that the ranks must negotiate the prefill vs. decode step type, and there is actually only one mesh submit/readout, we needed to modify vLLM core quite a bit - far beyond the scope of the hardware plugin mechanism. The per-rank schedulers did run in parallel, but the extra inter-process scatter/gather on every step cost more than that parallelism won back. The better answer is to put the parallelism inside one engine process: TTLaneCoordinator owns one independent TTScheduler per lane . Each lane has its own waiting and running queues, its own admission decisions, its own KV cache manager, and its own lane-local block ID space. New requests are assigned to the least-loaded lane and stay bound to it. Because the device executes all lanes together, the 
ankitg1220
🟧 echo.blog ⭐Introduces an out-of-tree vLLM plugin with automatic Tenstorrent platform registration, supported text and multimodal architectures, phase-sTenstorrent Team——

Interpretation history

Decision trace