2026-10-11 18:04 UTC

A-Rahim claims the released Kaggle TPU Lab can serve unquantized Qwen3.8-27B with its full 262K context at roughly 130 tokens per second through an OpenAI-compatible endpoint on free Kaggle TPU capacity, potentially making capable long-context inference available at near-zero compute cost.

state: expiredheat: lowuncertainty: highconvergesscott: mediumqwen local-inference inference-economicsA-RahimKaggleAlibaba Qwen

What is this?

A-Rahim claims Kaggle TPU Lab serves an unquantized Qwen3.8-27B model at roughly 130 tokens per second, with its native 262,144-token context, through an OpenAI-compatible endpoint using free Kaggle TPU capacity. A Kaggle model page supports the model’s native 262K context and shows a serving configuration, while Google material confirms Kaggle offers limited free TPU experimentation. The supplied snippets do not independently verify the reported speed, unquantized precision, endpoint compatibility, or whether free-tier limits make sustained near-zero-cost serving practical.

Why it matters to Scott

If independently verified, this extends Scott’s existing cost-tiered, OpenAI-compatible routing architecture with a potentially useful free long-context backend and provides a concrete benchmark against his hardware-aware local inference stack. The practical value remains uncertain because throughput, precision, API compatibility, availability, and Kaggle usage limits are not independently established.
dev:technology.litellmdev:concept.cost-tiered-llm-routingdev:concept.hardware-aware-local-inferenceip:concept.ai-unit-economicsradar:concept.inference-economicsradar:concept.local-inferenceradar:hetzner-free-slm-inference-experimentradar:concept.long-context-inference
queries asked of Scott's wikis
  • free-tier inference economics and hidden constraints
  • TPU serving for local or sovereign inference
  • OpenAI-compatible endpoint as model portability layer
  • long-context inference versus RAG economics
  • unquantized models versus quantization tradeoffs
  • commodity compute access and AI capability diffusion

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐Qwen3.8-27B at ~130 tok/s with full 262k context, on Kaggle's free TPU. OpenAI-compatible endpoint.
LocalLLaMA
A-Rahim6119
🟠 redditQwen 3.8 27B Vs. Qwen 3.6 27B on oMLX
LocalLLaMA
DerTomsn3731

Interpretation history

Decision trace