llama-cpp
band: hotmomentum: stable
score: 1.0
Episodes (50)
Trajectory notes
- 2026-09-10T15:55:50Z: gfx906-llama-cpp-throughput-gains closed (faded) — This is an adjacent example of Scott’s hardware-aware local inference practice, but his documented gamepc serving stack is WSL2/CUDA; the hits establish neither use of gfx906 hardware nor a build decision these maintainer-repor
- 2026-09-09T03:25:27Z: llama-cpp-avx2-iq-prompt-speedup closed (faded) — The radar already tracks this exact PR and validation question on `radar:llama-cpp-avx2-iq-batch-speedup`. If merged and independently validated, the optimization could affect Scott’s hardware-aware local-inference policy and CP
- 2026-09-07T05:26:14Z: gguf-quant-filename-mismatch closed (faded) — The audit independently supports Scott’s provenance-coupled-work position and could justify tensor-level validation in his active local-model stack, since misleading quantization labels may affect memory planning and reproducible pa
- 2026-09-03T20:30:10Z: llama-cpp-numa-weight-mirroring closed (faded) — The proposed NUMA mirroring concretely converges with Scott’s “Hardware-aware local inference” position by making memory placement and pressure explicit runtime policy, and it could materially change dual-socket CPU inference eco
- 2026-09-01T09:30:59Z: llama-cpp-kimi-k3-support closed (faded) — Known via dev:concept.hardware-aware-local-inference and dev:project.gamepc: Scott already treats runtime support, hardware placement, memory pressure, precision, and measured output fidelity as separate requirements for practical loca
- 2026-08-31T17:34:38Z: llama-cpp-lazy-tensor-loading closed (faded) — The radar already tracks essentially the same sparse-MoE memory-reduction thesis in `radar:hotpin-lossless-moe-streaming`, with adjacent llama.cpp expert-streaming work also under observation. This implementation could affect Scott
- 2026-08-31T06:28:56Z: llama-cpp-adaptive-mtp closed (faded) — The radar already tracks substantially the same open validation question in `radar:adaptive-speculative-decoding-300-gpu`, alongside llama.cpp MTP memory cases. Results on end-to-end coding-agent throughput would still bear directly on Sc
- 2026-08-30T15:32:19Z: consumer-gpu-p2p-kernel-fork closed (faded) — The radar already tracks this exact unresolved development in `radar:nvidia-consumer-gpu-p2p-modules`. It bears directly on Scott’s hardware-aware local-inference practice and self-hosted NVIDIA GPU substrate because reliable consum
- 2026-08-30T13:31:00Z: llama-cpp-dense-cpu-ffn-offload closed (faded) — The proposed control directly implements Scott’s hardware-aware local-inference position: accelerator placement and memory pressure should be explicit runtime policy, evaluated at the VRAM/throughput operating point. It could ext
- 2026-08-29T21:30:55Z: llama-cpp-hot-expert-offload closed (superseded) — The radar already tracks this exact development in `radar:llama-cpp-hot-expert-gpu-cache`. It bears directly on Scott’s hardware-aware placement policy and memory-constrained `gamepc` inference stack, but the claimed gain remai