2026-10-11 17:13 UTC

Sunny Narrator’s author claims a staged Gemma-and-Qwen pipeline can process book-scale literary translation locally at practical throughput on two obsolete Tesla P40 GPUs, making heterogeneous model pipelines a cost-effective option for large creative workloads.

state: expiredheat: lowuncertainty: highconvergesscott: mediumlocal-inference inference-economics multi-model-pipelinesSunny Narratorneowisard

What is this?

Sunny Narrator, also identified as neowisard, reports a local literary-translation workflow that stages Gemma 4 26B-A4B and Qwen 3.6 35B-A3B across two older 24GB Tesla P40 GPUs, claiming roughly 40 tok/s and 50–70 tok/s respectively with MTP speculative decoding. The supplied snippets support the broader feasibility of mixed-model pipelines, local translation, and useful Gemma throughput, but they do not independently verify this exact configuration, book-scale output quality, end-to-end throughput, or cost savings. The Reddit/Habr chronology and attribution are present only in the evidence titles and supplied summary, so those details remain thinly corroborated.

Why it matters to Scott

The claimed working pipeline independently converges with Scott’s hardware-aware local inference, task-aware model routing, and staged long-form content systems, adding a concrete retired-GPU configuration and throughput claim that could affect his own local model-serving choices. Its practical significance depends on independent verification of translation quality, end-to-end throughput, and cost—not just reported token rates.
dev:concept.hardware-aware-local-inferencedev:concept.task-aware-model-routingdev:concept.multi-pass-content-generationdev:project.gamepcip:concept.usable-mass-over-unusable-powerradar:gemma-translator-local-validationradar:dumpstercluster-retired-gpu-inferenceradar:v100-skinny-nvfp4-speculative-decodingradar:concept.inference-economics
queries asked of Scott's wikis
  • local inference economics versus cloud APIs
  • heterogeneous model pipelines for creative workflows
  • obsolete GPU hardware for open-model inference
  • multi-stage LLM translation and quality control
  • speculative decoding and memory-bandwidth constraints
  • book-scale context, chunking, and consistency workflows

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditRunning a 2-model literary book-translation pipeline on 2x Tesla P40: gemma-4-26B-A4B at ~40 tok/s + Qwen3.6-35B-A3B at 50-70 tok/s with MTP spec decode — full llama-server flags inside
LocalLLaMA
neowisard43
🟧 echo.blog ⭐The author’s Habr follow-up predates Reddit and reports the same project, hardware, model pairing, workflow, and speed claims: “gemma4-26B-ANick Kutuzov (@neowisard)——

Interpretation history

Decision trace