2026-10-11 17:11 UTC

AMD-aligned Lemonade's 2026.40 release candidate removes the OpenMOSS ROCm backend as ~40x slower than Vulkan and fixes APU model streaming by sizing against the GTT pool, signaling Vulkan as the practical backend for consumer AMD local inference.

state: resolvedheat: lowuncertainty: lownovelscott: lowlocal-inference amd vulkanAMD

What is this?

Lemonade is an AMD-aligned open-source local LLM server (lemonade-server.ai, lemonade-sdk on GitHub) that exposes multiple inference backends โ€” including llama.cpp-based Vulkan and ROCm paths, and more recently an experimental vLLM ROCm bundle โ€” for running models on AMD GPUs and APUs. Its 2026.40 release candidate, per WindowsForum's coverage of the breaking-changes list, removes the OpenMOSS backend's 'rocm'/'rocm_bin' variants on Windows and Linux, and a GitHub issue (#3377) documents the Strix Halo APU problem where Lemonade only sees the small VRAM carve-out and rejects models that would fit in the larger GTT (host-visible) pool โ€” the sizing bug the RC reportedly fixes. Community evidence supports the direction: HN and r/LocalLLaMA threads report Vulkan consistently outperforming ROCm on consumer AMD hardware (with ROCm regressions blamed), and a Strix Halo hands-on measured Vulkan (mmap) at ~46 t/s vs ROCm at ~37.8 t/s with worse memory behavior. Note: the snippets corroborate the removal and the APU/GTT fix, but the specific '~40x slower' figure appears only in the case's own evidence title, not independently in the supplied material.

Why it matters to Scott

The story lands squarely in Scott's local-inference territory โ€” backend selection and memory-pool sizing as explicit runtime policy (dev:concept.hardware-aware-local-inference) โ€” and the GTT-vs-VRAM-carve-out fix is a concrete datapoint on unified-memory sizing, the same terrain as the radar's adaptive KV-cache streaming episode. But it bears on nothing he holds or builds: his serving stack is NVIDIA/CUDA (gamepc, Ollama), the wiki shows no AMD position for Vulkan-vs-ROCm to converge with or challenge, and a better backend for consumer AMD APUs wouldn't change what he runs or argues. This is his kind of topic, not news for him.
dev:concept.hardware-aware-local-inferenceradar:adaptive-kv-cache-streaming
queries asked of Scott's wikis
  • local inference backend benchmarks Vulkan vs ROCm vs CUDA โ€” does he have a position on the practical stack for consumer AMD?
  • agent memory / wiki systems running on local models โ€” self-hosted inference requirements, hardware he runs
  • open-weights model sovereignty and avoiding Nvidia/CUDA lock-in โ€” vendor strategy notes
  • llama.cpp, GGUF, model streaming and memory pool sizing โ€” projects touching inference runtime internals
  • APU / unified memory inference (Strix Halo-class hardware) โ€” any prior writing on local RAG or knowledge systems hardware
  • backend abstraction layers in LLM tooling he builds โ€” which runtimes he targets or wraps

Measured heat

no measured readings yet โ€” the hourly heat pass fills this in

How the heat travelled

09-22 14:00โญ origin echo-reconstructedRelease notes (candidate-v2026.40.0, 23 Sep 2026 18:13 UTC): "Streaming models now run on AMD integrated GPUs by sizing against the APU GTT
Lemonade SDK project (AMD-aligned; release published via github-actions, PR by jeremyfowers) on github (echo) ยท attributed from reddit.post.1woh146
โ€”
09-23 20:18first on r/artificial ยท published ยท +30.3hLemonade fixes AMD APU model streaming, drops OpenMOSS ROCm as ~40x slower than Vulkan
Fcking_Chuck
โ€”
09-24 13:37first on r/LocalLLaMA ยท published ยท +47.6hvulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4 (โ€ฆ ยท ggml-org/llama.cpp@70c4e15
nickm_27
โ€”
09-23 20:18amplified on r/artificialreddit.post.1woh146
Fcking_Chuck
peak 13 ยท 1 comments ยท 20% of case engagement
09-24 13:37amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1wp1vex
nickm_27
peak 43 ยท 12 comments ยท 80% of case engagement
09-23 21:21our radar first saw it ยท +31.4hdiscovery anchor: reddit.post.1woh146โ€”

Evidence (3) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  redditLemonade fixes AMD APU model streaming, drops OpenMOSS ROCm as ~40x slower than Vulkan
artificial
Retrieved article excerpt

Open article ยท Retrieved 2026-09-23T21:46:49.712313+00:00

# Lemonade Fixes AMD APU Model Streaming, Drops OpenMOSS ROCm As ~40x Slower Than Vulkan

Written by [Michael Larabel](https://www.michaellarabel.com/) in [AI](https://www.phoronix.com/linux/AI) on 23 September 2026 at 03:41 PM EDT. [Add A Comment](https://www.phoronix.com/forums/node/1659737)

AI

The AMD-aligned Lemonade open-source project for serving as a local AI server released 2026.39.1 today as well as issuing a release candidate of 2026.40 as their next release. Besides adapting to a new versioning scheme, the Lemonade SDK updates today bring a few notable changes.
  
   
The Lemonade 2026.40 release candidate now supports streaming models on AMD integrated GPUs (APUs) by sizing against the APU GTT pool rather than the fixed vRAM carve-out. This came about as some streaming models like DeepSeek-V4-Flash-IQ2XXS-DS4 were failing to run on AMD Ryzen AI Max (Strix Halo). While there is sufficient RAM for doing so, Lemonade had been checking against the smaller carve-out window rather than the addressable GTT. Thus that's being fixed to allow for better streaming model support on AMD iGPUs/APUs.
  
   
Another interesting change is the OpenMOSS back-end for ROCm being removed on Windows and Linux. OpenMOSS is an AI research initiative and model ecosystem out of China. The OpenMOSS ROCm back-ends are disabled since they are running around 40x slower than using the Vulkan back-end on the same hardware. It appears there are at least some CPU fall-backs happening with the AMD ROCm code paths for OpenMOSS and what's leading to the big disparity.
  

OpenMOSS ROCm benchmark

  
[This merge request](https://github.com/lemonade-sdk/lemonade/pull/3615) took the step of removing OpenMOSS ROCm support until it can be properly fixed and shared the atrocious performance numbers for OpenMOSS ROCm.
  
   
Lemonade 2026.40 RC is also adding a new launch agent for JetBrains' Junie. Plus other enhancements as outlined on [GitHub](https://github.com/lemonade-sdk/lemonade/releases/tag/candidate-v2026.40.0).
  
   
[Lemonade 2026.39.1](https://github.com/lemonade-sdk/lemonade/releases/tag/v2026.39.1) released today brought configurable vRAM auto-eviction and other smaller changes. It's the first stable release under the new versioning system in relying on the year and week number.

[Add A Comment](https://www.phoronix.com/forums/node/1659737)
Fcking_Chuck121
๐ŸŸง echo.github โญRelease notes (candidate-v2026.40.0, 23 Sep 2026 18:13 UTC): "Streaming models now run on AMD integrated GPUs by sizing against the APU GTT Lemonade SDK project (AMD-aligned; release published via github-actions, PR by jeremyfowers)โ€”โ€”
๐ŸŸ  redditvulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4 (โ€ฆ ยท ggml-org/llama.cpp@70c4e15
LocalLLaMA
nickm_274212

Interpretation history

Decision trace