The case describes llamAmpere as a released, Ampere-focused fork of llama.cpp aimed at local LLM users with RTX 3090 or other 30xx GPUs. It attributes to the creator a claim of more than 90 tokens per second through 100K tokens of context in a recommended coding configuration. The web search returned no results, so the release, performance claim, configuration, and creator’s identity are not independently established; JakeATX and Brief-Tap-6616 are listed without clear roles. Improved responsiveness for long-context local agents remains a proposed implication rather than a demonstrated outcome.
The Ampere-specific optimization approach converges with Scott’s Hardware-aware local inference practice and offers a concrete benchmark candidate for gamepc’s local serving and Ask’s optional local-agent path. However, the creator’s unverified throughput claim does not establish compatibility with those deployments or improved coding-agent performance; the radar tracks related 3090 long-context work, but the supplied hits do not show this fork already covered.
dev:concept.hardware-aware-local-inferencedev:project.gamepcdev:project.askradar:concept.llama-cppradar:concept.long-context-inferenceradar:burrito-core-gpt-oss-harnessradar:picchio-llama-cpp-bottleneck-diagnostics
queries asked of Scott's wikis
- local coding agents long-context latency requirements
- llama.cpp RTX 3090 inference stack
- local inference economics hardware optimization
- agent context budgets memory retrieval tradeoffs
- inference benchmarks throughput context length reproducibility
now 0 pts/hpeak 10 pts/hcomments 0/hpeers p25momentum: steady2 platformsage 638h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion
2026-10-10T02:49:54Z
New release (v0.5?) targets 12GB Ampere cards with 200K context and improved KVaRN; one independent user reports 24 TPS at 200K on a power-limited A2000 12GB. This extends the fork's hardware coverage but does not resolve the core uncertainties: the 90+ TPS-at-100K claim on 3090 still has only one reproduction bundled with quality complaints, no independent benchmark validates coding-agent-grade output at long context, and the fork remains unauditable (no clean upstream PR). The case remains a concrete benchmark candidate for Scott's local serving work, not a verified advance.
2026-10-10T01:44:11Z
evidence attached: reddit.post.1x1xt1x — llamAmpere creator's new release targeting 12GB Ampere cards with 200K context and improved KVaRN directly updates the watching case.
2026-09-29T15:14:21Z
The velocity spike was just the tail of v0.4 voting (v0.4 thread 14→24 pts) plus low-substance questions (quant tiers, Swift conversion); it decayed to 0 pts/h within a day and added no implementations, benchmarks, or new communities. The case's meaning is unchanged: an actively iterating fork that remains a single-reproduction, quality-caveated benchmark candidate rather than a verified long-context advance.
2026-09-28T22:06:48Z
v0.4 confirms llamAmpere as a persisting, iterating 3090 runtime — the maintainer relays positive results across 30xx cards, an independent curator (tomByrer's 3090-AI list) ranks it top-3, and the headline extends to 262K context on a single 3090 — but the 90+ TPS-at-100K claim still has only one independent reproduction, bundled with quality complaints, and a comparative tester still prefers beelllama overall. It remains a concrete benchmark candidate for Scott's local serving work, not an attention-worthy advance.
2026-09-28T21:36:33Z
evidence attached: reddit.post.1wsmtd4 — Maintainer's v0.4 update plus reported results shared back by other 30xx-card users directly advances the fork's throughput and adoption hypothesis.
2026-09-19T15:26:29Z
The latest comments add quoted MTP gains and capacity skepticism, not an identifiable independent benchmark resolving the speed–quality tradeoff. With no demonstrated expansion beyond the existing Reddit discussion, this remains a local benchmark candidate rather than an attention-demanding runtime advance.
2026-09-18T16:41:50Z
The creator now announces v0.3.1 with EXL and ternary Bonsai support plus custom kernels for faster MTP, making this an actively evolving runtime rather than a one-off speed claim. The update does not establish whether these additions preserve quality or sustain the claimed throughput at long context; the creator’s summary of user feedback is not independent corroboration.
2026-09-18T16:22:51Z
evidence attached: reddit.post.1wjuohn — This runtime update materially bears on whether LlamAmpere can make long-context local inference faster on constrained GPUs.
2026-09-15T14:35:39Z
A firsthand tester reports 95 TPS but criticizes the supplied model quantization and turbo3 value-cache quality, shifting this from a pure speed claim to a speed-versus-quality tradeoff worth testing. The report does not establish sustained throughput at 100K context or useful coding-agent performance.
2026-09-15T02:24:57Z
grounded: converges/medium — The Ampere-specific optimization approach converges with Scott’s Hardware-aware local inference practice and offers a concrete benchmark candidate for gamepc’s
2026-09-15T02:22:07Z
case created — A downloadable architecture-specific fork supplies a concrete performance claim distinct from the existing dual-3090 Qwen Flash episode, although the visible excerpt leaves benchmark conditions incomplete.