2026-10-11 16:37 UTC

llamAmpere’s creator claims the released Ampere-focused llama.cpp fork sustains over 90 tokens per second through 100K tokens of context in its recommended coding configuration, potentially making long-context local agents more responsive on RTX 3090-class hardware.

state: watchingheat: lowuncertainty: highconvergesscott: mediumlocal-inference llama-cpp inference-economicsJakeATXBrief-Tap-6616

What is this?

The case describes llamAmpere as a released, Ampere-focused fork of llama.cpp aimed at local LLM users with RTX 3090 or other 30xx GPUs. It attributes to the creator a claim of more than 90 tokens per second through 100K tokens of context in a recommended coding configuration. The web search returned no results, so the release, performance claim, configuration, and creator’s identity are not independently established; JakeATX and Brief-Tap-6616 are listed without clear roles. Improved responsiveness for long-context local agents remains a proposed implication rather than a demonstrated outcome.

Why it matters to Scott

The Ampere-specific optimization approach converges with Scott’s Hardware-aware local inference practice and offers a concrete benchmark candidate for gamepc’s local serving and Ask’s optional local-agent path. However, the creator’s unverified throughput claim does not establish compatibility with those deployments or improved coding-agent performance; the radar tracks related 3090 long-context work, but the supplied hits do not show this fork already covered.
dev:concept.hardware-aware-local-inferencedev:project.gamepcdev:project.askradar:concept.llama-cppradar:concept.long-context-inferenceradar:burrito-core-gpt-oss-harnessradar:picchio-llama-cpp-bottleneck-diagnostics
queries asked of Scott's wikis
  • local coding agents long-context latency requirements
  • llama.cpp RTX 3090 inference stack
  • local inference economics hardware optimization
  • agent context budgets memory retrieval tradeoffs
  • inference benchmarks throughput context length reproducibility

Measured heat

now 0 pts/hpeak 10 pts/hcomments 0/hpeers p25momentum: steady2 platformsage 638h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-15 02:22 (minted)⭐ origin echo-reconstructedThe creator links this Ampere-focused llama.cpp fork and claims its recommended configuration supports “90+ TPS” through 100K tokens, with c
JakeATX on github (echo) · attributed from reddit.post.1wgmhon · published time unknown
—
09-15 01:32first on r/LocalLLaMA · published · lag ?If you have a 3090, or other 30xx for local LLMs, I have something for you
Brief-Tap-6616
—
09-15 01:32amplified on r/LocalLLaMA 👑reddit.post.1wgmhon
Brief-Tap-6616
peak 114 · 81 comments · 56% of case engagement
09-18 16:17amplified on r/LocalLLaMAreddit.post.1wjuohn
Brief-Tap-6616
peak 10 · 7 comments · 5% of case engagement
09-28 18:35amplified on r/LocalLLaMAreddit.post.1wsmtd4
Brief-Tap-6616
peak 26 · 35 comments · 18% of case engagement
10-09 21:39amplified on r/LocalLLaMAreddit.post.1x1xt1x
Brief-Tap-6616
peak 38 · 35 comments · 21% of case engagement
09-15 02:20our radar first saw it · lag ?discovery anchor: reddit.post.1wgmhon—
pace: p77 vs 1032 stories at the 336h mark (now 638h old) — ahead of jenny-local-coding-harness (1.0x), behind california-data-center-legislation (1.0x)

Evidence (5) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditIf you have a 3090, or other 30xx for local LLMs, I have something for you
LocalLLaMA
Brief-Tap-661611481
🟧 echo.github ⭐The creator links this Ampere-focused llama.cpp fork and claims its recommended configuration supports “90+ TPS” through 100K tokens, with cJakeATX——
🟠 redditLlamAmpere v0.3.1 updated with support for EXL + ternary bonsai, with custom kernels for faster MTP
LocalLLaMA
Brief-Tap-6616107
🟠 reddit95+ TPS through 100K generated for qwen3.8 27b, 262K ctx, on a single 3090
LocalLLaMA
Brief-Tap-66162635
🟠 redditQwen3.8 27b with 200K ctx + MTP on 12GB Ampere Cards
LocalLLaMA
Brief-Tap-66163835

Interpretation history

Decision trace