2026-10-11 18:01 UTC

Independent testing will determine whether aikitoria’s open NVIDIA kernel fork reliably enables peer-to-peer transfers on supported dual-consumer-GPU systems and materially improves local inference performance.

state: expiredheat: lowuncertainty: highknownscott: mediumlocal-inference gpu-infrastructure llama-cppaikitoriatinygradNVIDIA

What is this?

GitHub identifies aikitoria’s project as a fork of tinygrad’s NVIDIA Linux open GPU kernel modules, modified to add peer-to-peer support. A technical report says that patching this driver and bypassing vLLM’s P2P capability check enabled direct GPU-to-GPU transfers on dual RTX 3090s, with a reported 10–30% throughput improvement for Qwen 3.5 35B when combined with a tuned MoE kernel. The supplied snippets do not establish broad independent replication, reliability across other consumer-GPU configurations, or any llama.cpp performance result.

Why it matters to Scott

The radar already tracks this exact unresolved development in `radar:nvidia-consumer-gpu-p2p-modules`. It bears directly on Scott’s hardware-aware local-inference practice and self-hosted NVIDIA GPU substrate because reliable consumer-GPU P2P could improve multi-GPU throughput and economics, but the supplied evidence does not yet establish broad reliability or a llama.cpp benefit.
dev:concept.hardware-aware-local-inferencedev:project.gamepcdev:technology.cudaip:concept.ai-unit-economicsradar:nvidia-consumer-gpu-p2p-modulesradar:concept.local-inferenceradar:concept.gpu-infrastructureradar:concept.distributed-inference
queries asked of Scott's wikis
  • consumer multi-GPU local inference economics
  • peer-to-peer GPU transfers for local models
  • NVIDIA software restrictions and hardware sovereignty
  • llama.cpp multi-GPU bottlenecks and benchmarks
  • open GPU drivers for local AI infrastructure
  • dual consumer GPU inference builds

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditYou CAN get P2P working on your dual GPU Setups
LocalLLaMA
luedtek1019
🟧 echo.github ⭐A fork of the open NVIDIA GPU kernel modules adds broader peer-to-peer support and reportedly works on a dual-RTX-3090 system.aikitoria——

Interpretation history

Decision trace