GitHub identifies aikitoria’s project as a fork of tinygrad’s NVIDIA Linux open GPU kernel modules, modified to add peer-to-peer support. A technical report says that patching this driver and bypassing vLLM’s P2P capability check enabled direct GPU-to-GPU transfers on dual RTX 3090s, with a reported 10–30% throughput improvement for Qwen 3.5 35B when combined with a tuned MoE kernel. The supplied snippets do not establish broad independent replication, reliability across other consumer-GPU configurations, or any llama.cpp performance result.
The radar already tracks this exact unresolved development in `radar:nvidia-consumer-gpu-p2p-modules`. It bears directly on Scott’s hardware-aware local-inference practice and self-hosted NVIDIA GPU substrate because reliable consumer-GPU P2P could improve multi-GPU throughput and economics, but the supplied evidence does not yet establish broad reliability or a llama.cpp benefit.
dev:concept.hardware-aware-local-inferencedev:project.gamepcdev:technology.cudaip:concept.ai-unit-economicsradar:nvidia-consumer-gpu-p2p-modulesradar:concept.local-inferenceradar:concept.gpu-infrastructureradar:concept.distributed-inference
queries asked of Scott's wikis
- consumer multi-GPU local inference economics
- peer-to-peer GPU transfers for local models
- NVIDIA software restrictions and hardware sovereignty
- llama.cpp multi-GPU bottlenecks and benchmarks
- open GPU drivers for local AI infrastructure
- dual consumer GPU inference builds
2026-08-30T15:32:19Z
Repeated checks have produced no independent replication, supported-configuration expansion, or workload benchmark, and the originating discussion is dormant. Retire the passive watch until a fresh hardware test or inference result creates a new episode.
2026-08-28T14:38:51Z
The discussion has gone dormant without independent replication, broader hardware support, or workload benchmarks, shifting the case from active validation watch to passive monitoring. The original single-system result remains plausible but materially unverified.
2026-08-26T13:38:55Z
The refreshed discussion remains deployment interest and validation questions rather than independent evidence. Repeated comment updates are now low-value amplification; the case still depends on replicated hardware tests and workload-level inference benchmarks.
2026-08-25T12:36:35Z
The new discussion adds deployment interest but no independent replication, configuration evidence, or inference benchmark. Repeated comment refreshes are now amplification rather than validation; the case still depends on topology-specific testing and measured workload gains.
2026-08-25T10:43:16Z
The refreshed discussion adds no independent replication, supported configuration, or workload benchmark; it remains repetitive amplification around an unvalidated single-system result. The case still hinges on topology-specific testing and measurable inference gains.
2026-08-25T04:27:47Z
The refreshed comments remain requests for topology details and workload benchmarks rather than new validation. The fork still has only one reported dual-RTX-3090 success and no independent reliability or inference-performance result.
2026-08-24T22:33:58Z
The refreshed discussion only identifies the right validation questions—real inference throughput, ReBAR requirements, and PCIe-topology dependence—without supplying a benchmark or independent replication. The case remains a narrow implementation awaiting substantive testing.
2026-08-24T19:58:42Z
No independent test, additional hardware configuration, or inference benchmark has appeared; the minor engagement change adds no substance. The case remains a narrow but relevant implementation awaiting replication.
2026-08-24T19:36:59Z
grounded: known/medium — The radar already tracks this exact unresolved development in `radar:nvidia-consumer-gpu-p2p-modules`. It bears directly on Scott’s hardware-aware local-inferen
2026-08-24T19:34:06Z
case created — The usable kernel artifact and one successful hardware report warrant tracking despite narrow compatibility and minimal validation.