AMD-aligned Lemonade's 2026.40 release candidate removes the OpenMOSS ROCm backend as ~40x slower than Vulkan and fixes APU model streaming by sizing against the GTT pool, signaling Vulkan as the practical backend for consumer AMD local inference.
state: resolvedheat: lowuncertainty: lownovelscott: lowlocal-inference amd vulkanAMD
What is this?
Lemonade is an AMD-aligned open-source local LLM server (lemonade-server.ai, lemonade-sdk on GitHub) that exposes multiple inference backends โ including llama.cpp-based Vulkan and ROCm paths, and more recently an experimental vLLM ROCm bundle โ for running models on AMD GPUs and APUs. Its 2026.40 release candidate, per WindowsForum's coverage of the breaking-changes list, removes the OpenMOSS backend's 'rocm'/'rocm_bin' variants on Windows and Linux, and a GitHub issue (#3377) documents the Strix Halo APU problem where Lemonade only sees the small VRAM carve-out and rejects models that would fit in the larger GTT (host-visible) pool โ the sizing bug the RC reportedly fixes. Community evidence supports the direction: HN and r/LocalLLaMA threads report Vulkan consistently outperforming ROCm on consumer AMD hardware (with ROCm regressions blamed), and a Strix Halo hands-on measured Vulkan (mmap) at ~46 t/s vs ROCm at ~37.8 t/s with worse memory behavior. Note: the snippets corroborate the removal and the APU/GTT fix, but the specific '~40x slower' figure appears only in the case's own evidence title, not independently in the supplied material.
Why it matters to Scott
The story lands squarely in Scott's local-inference territory โ backend selection and memory-pool sizing as explicit runtime policy (dev:concept.hardware-aware-local-inference) โ and the GTT-vs-VRAM-carve-out fix is a concrete datapoint on unified-memory sizing, the same terrain as the radar's adaptive KV-cache streaming episode. But it bears on nothing he holds or builds: his serving stack is NVIDIA/CUDA (gamepc, Ollama), the wiki shows no AMD position for Vulkan-vs-ROCm to converge with or challenge, and a better backend for consumer AMD APUs wouldn't change what he runs or argues. This is his kind of topic, not news for him.
dev:concept.hardware-aware-local-inferenceradar:adaptive-kv-cache-streaming
queries asked of Scott's wikis
- local inference backend benchmarks Vulkan vs ROCm vs CUDA โ does he have a position on the practical stack for consumer AMD?
- agent memory / wiki systems running on local models โ self-hosted inference requirements, hardware he runs
- open-weights model sovereignty and avoiding Nvidia/CUDA lock-in โ vendor strategy notes
- llama.cpp, GGUF, model streaming and memory pool sizing โ projects touching inference runtime internals
- APU / unified memory inference (Strix Halo-class hardware) โ any prior writing on local RAG or knowledge systems hardware
- backend abstraction layers in LLM tooling he builds โ which runtimes he targets or wraps
Measured heat
no measured readings yet โ the hourly heat pass fills this in
How the heat travelled
Evidence (3) โ โญ canonical anchor
| source | object | author | score | comments |
| ๐ reddit | Lemonade fixes AMD APU model streaming, drops OpenMOSS ROCm as ~40x slower than Vulkan artificial Retrieved article excerptOpen article ยท Retrieved 2026-09-23T21:46:49.712313+00:00 # Lemonade Fixes AMD APU Model Streaming, Drops OpenMOSS ROCm As ~40x Slower Than Vulkan
Written by [Michael Larabel](https://www.michaellarabel.com/) in [AI](https://www.phoronix.com/linux/AI) on 23 September 2026 at 03:41 PM EDT. [Add A Comment](https://www.phoronix.com/forums/node/1659737)
AI
The AMD-aligned Lemonade open-source project for serving as a local AI server released 2026.39.1 today as well as issuing a release candidate of 2026.40 as their next release. Besides adapting to a new versioning scheme, the Lemonade SDK updates today bring a few notable changes.
The Lemonade 2026.40 release candidate now supports streaming models on AMD integrated GPUs (APUs) by sizing against the APU GTT pool rather than the fixed vRAM carve-out. This came about as some streaming models like DeepSeek-V4-Flash-IQ2XXS-DS4 were failing to run on AMD Ryzen AI Max (Strix Halo). While there is sufficient RAM for doing so, Lemonade had been checking against the smaller carve-out window rather than the addressable GTT. Thus that's being fixed to allow for better streaming model support on AMD iGPUs/APUs.
Another interesting change is the OpenMOSS back-end for ROCm being removed on Windows and Linux. OpenMOSS is an AI research initiative and model ecosystem out of China. The OpenMOSS ROCm back-ends are disabled since they are running around 40x slower than using the Vulkan back-end on the same hardware. It appears there are at least some CPU fall-backs happening with the AMD ROCm code paths for OpenMOSS and what's leading to the big disparity.
OpenMOSS ROCm benchmark
[This merge request](https://github.com/lemonade-sdk/lemonade/pull/3615) took the step of removing OpenMOSS ROCm support until it can be properly fixed and shared the atrocious performance numbers for OpenMOSS ROCm.
Lemonade 2026.40 RC is also adding a new launch agent for JetBrains' Junie. Plus other enhancements as outlined on [GitHub](https://github.com/lemonade-sdk/lemonade/releases/tag/candidate-v2026.40.0).
[Lemonade 2026.39.1](https://github.com/lemonade-sdk/lemonade/releases/tag/v2026.39.1) released today brought configurable vRAM auto-eviction and other smaller changes. It's the first stable release under the new versioning system in relying on the year and week number.
[Add A Comment](https://www.phoronix.com/forums/node/1659737) | Fcking_Chuck | 12 | 1 |
| ๐ง echo.github โญ | Release notes (candidate-v2026.40.0, 23 Sep 2026 18:13 UTC): "Streaming models now run on AMD integrated GPUs by sizing against the APU GTT | Lemonade SDK project (AMD-aligned; release published via github-actions, PR by jeremyfowers) | โ | โ |
| ๐ reddit | vulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4 (โฆ ยท ggml-org/llama.cpp@70c4e15 LocalLLaMA | nickm_27 | 42 | 12 |
Interpretation history
2026-09-25T20:36:15Z
The velocity flag was 4.5x a sub-1-pt/h baseline and decayed to literal zero within a day โ vote-noise on aging posts, not re-ignition. The episode has played out completely: removal confirmed, the Vulkan-practical thesis independently corroborated (llama.cpp coopmat kernels, Strix Halo benchmarks), and no open thread remains that would change the assessment โ closing as proved out rather than letting it fade unresloved.
2026-09-24T15:27:02Z
The Vulkan-practical thesis is no longer just Lemonade's own removal rationale: upstream llama.cpp independently landed int8 coopmat1 Vulkan kernels for RDNA3/RDNA4 with user-measured gains on a 7900XTX, joining the Strix Halo community benchmarks as a second independent line โ the direction is now corroborated, though the episode stays niche and attention-flat.
2026-09-24T15:25:48Z
evidence attached: reddit.post.1wp1vex โ Upstream llama.cpp int8 coopmat gains on 7900XTX independently reinforce the Vulkan-as-practical-consumer-AMD-backend signal.
2026-09-23T23:14:28Z
origin walked (opencode/cheap-glm, conf 0.9): anchor reddit.post.1woh146 -> echo.github.5cafa5f9a7 by Lemonade SDK project (AMD-aligned; release published via github-actions, PR by jeremyfowers)
2026-09-23T22:26:11Z
grounded: novel/low โ The story lands squarely in Scott's local-inference territory โ backend selection and memory-pool sizing as explicit runtime policy (dev:concept.hardware-aware-
2026-09-23T22:20:35Z
case created โ Concrete first-party stack-direction change with measured numbers; a bounded backend-consolidation episode for local AMD inference.
Decision trace
- 09-26 06:36resolveThe velocity flag was 4.5x a sub-1-pt/h baseline and decayed to literal zero within a day โ vote-noise on aging posts, not re-ignition. The episode has played out completely: removal confirmed, the Vu
- 09-25 13:21sensor_dirtycomment_update
- 09-25 06:21sensor_dirtyvelocity_spike
- 09-25 01:27repriceThe Vulkan-practical thesis is no longer just Lemonade's own removal rationale: upstream llama.cpp independently landed int8 coopmat1 Vulkan kernels for RDNA3/RDNA4 with user-measured gains on a
- 09-25 01:25attachUpstream llama.cpp int8 coopmat gains on 7900XTX independently reinforce the Vulkan-as-practical-consumer-AMD-backend signal.
- 09-25 01:24propose_attachUpstream llama.cpp int8 coopmat gains on 7900XTX independently reinforce the Vulkan-as-practical-consumer-AMD-backend signal.
- 09-24 09:14promote_anchororigin walk conf 0.9
- 09-24 08:26groundThe story lands squarely in Scott's local-inference territory โ backend selection and memory-pool sizing as explicit runtime policy (dev:concept.hardware-aware-local-inference) โ and the GTT-vs-V
- 09-24 08:20createConcrete first-party stack-direction change with measured numbers; a bounded backend-consolidation episode for local AMD inference.