Infermeld's maintainer (do_u_think_im_spooky) claims the released experimental v0.1.0 Linux kit lets a single GGUF run jointly across an AMD (Vulkan) and NVIDIA (CUDA) GPU under llama.cpp, and independent reproduction by mixed-vendor-GPU owners would establish cross-vendor consumer-GPU inference as a practical local setup.
state: seedheat: lowuncertainty: mediumknownscott: lowlocal-inference llama-cpp heterogeneous-gpusdo_u_think_im_spookyInfermeld
What is this?
Per the case, Infermeld is an experimental v0.1.0 Linux kit by maintainer do_u_think_im_spooky claiming that a single GGUF model can run jointly across one AMD GPU (Vulkan backend) and one NVIDIA GPU (CUDA backend) under llama.cpp, with independent reproduction by mixed-vendor owners as the bar for calling cross-vendor consumer inference practical. The supplied search results do not surface Infermeld or its maintainer at all, so the project's specific claims rest entirely on the first-party case evidence. What the results do establish is the surrounding landscape: llama.cpp ships CUDA and Vulkan as distinct backends with Vulkan as the usual cross-vendor path, and mixed-vendor single-process llama.cpp is not unprecedented in the wild โ daimonionnn's multi-gpu-llm-toolkit already documents AMD (ROCm/HIP or Vulkan) plus NVIDIA (CUDA) in one llama-server process without RPC, on workstation-class R9700 + Blackwell cards, with documented ROCm memory bugs and a Blackwell CUDA build pitfall. Benchmark write-ups (d-central, knightli) stress that results never transfer across backends without pinned commit, driver, and flags โ which defines exactly what a credible independent reproduction of the Infermeld claim would have to control.
Why it matters to Scott
Scott's own wikis already carry this exact position: dev:technology.cuda and dev:concept.hardware-aware-local-inference hold the llama.cpp CUDA+Vulkan tensor-split across mixed vendors (including the Vulkan-throughput-penalty caveat), with dev:project.gamepc and its Ollama endpoint as the rig it would run on โ so a hobbyist kit re-packaging the claim adds nothing new on either side, and the grounding shows daimonionnn's multi-gpu-llm-toolkit already documented single-process AMD+NVIDIA llama.cpp. It matters only as lineage under the radar's llama-cpp/local-inference episodes, not as a position shift or a hook he doesn't already have.
dev:technology.cudadev:concept.hardware-aware-local-inferencedev:project.gamepcdev:technology.ollamaradar:concept.llama-cppradar:concept.local-inferenceradar:concept.vulkanradar:concept.cudaradar:janus-vulkan-gguf-runtimeradar:llama-halo-hybrid-strix-haloradar:concept.reproducibility
queries asked of Scott's wikis
- mixed vendor GPU VRAM pooling local LLM inference
- llama.cpp tensor split heterogeneous GPUs CUDA Vulkan
- Vulkan throughput penalty versus CUDA local models
- local inference rig hardware notes upgrades
- CUDA vendor lock-in open weight inference ecosystem
- reproducible benchmark protocol pin build driver backend
Measured heat
now 0 pts/hpeak 5 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 158h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion
How the heat travelled
pace: p52 vs 1247 stories at the 96h mark (now 158h old) โ ahead of android-editable-graph-agent (1.1x), behind ai-vuln-reports-oss-disclosure (0.9x)
Evidence (2) โ โญ canonical anchor
Interpretation history
2026-10-05T04:47:14Z
origin walked (opencode/cheap-glm, conf 0.93): anchor reddit.post.1wxyuxx -> echo.github.b0847d066f by 5p00kyy (GitHub) = u/do_u_think_im_spooky (Reddit), the project maintainer
2026-10-05T04:32:42Z
grounded: known/low โ Scott's own wikis already carry this exact position: dev:technology.cuda and dev:concept.hardware-aware-local-inference hold the llama.cpp CUDA+Vulkan tensor-sp
2026-10-05T04:24:56Z
case created โ First-party maintainer release of a working artifact with a crisp reproduction-resolvable claim (one GGUF spanning AMD+NVIDIA) that no open case covers โ distinct from the AMD-only dual-GPU, ZLUDA-on-AMD, and APU+eGPU cases.
Decision trace
- 10-08 10:38review_screenjev screen: no material development (noul=0.12)
- 10-06 04:21sensor_dirtycomment_update
- 10-05 17:20sensor_dirtycomment_update
- 10-05 15:47promote_anchororigin walk conf 0.93
- 10-05 15:32groundScott's own wikis already carry this exact position: dev:technology.cuda and dev:concept.hardware-aware-local-inference hold the llama.cpp CUDA+Vulkan tensor-split across mixed vendors (including
- 10-05 15:24createFirst-party maintainer release of a working artifact with a crisp reproduction-resolvable claim (one GGUF spanning AMD+NVIDIA) that no open case covers โ distinct from the AMD-only dual-GPU, ZLUDA-on-