2026-10-11 16:37 UTC

Artificial Analysis claims its open-source AA-AgentPerf-Local β€” replaying 8 recorded agent trajectories (~168 turns, ~56K-token growing contexts) across DGX Spark, RTX 5090, Ryzen AI Halo, and MacBook Pro M5 Pro with published configs and a maintained leaderboard β€” becomes the reference benchmark shaping local-model and hardware choices for agent work; broad citation, user-submitted results, and expansion to the promised hardware/framework coverage resolve it.

state: watchingheat: lowuncertainty: mediumconvergesscott: highagent-evaluation local-inference inference-economicsArtificial Analysis

What is this?

Artificial Analysis is an independent AI measurement shop that publishes live model- and hardware-benchmark leaderboards. In June 2026 it launched AA-AgentPerf, covered as the first inference benchmark built for agentic workloads: it replays real coding-agent trajectories (up to 200 turns, 100K+ token contexts) against a system under test and reports how many concurrent agents the machine can serve at production latency targets, headlined by an Agents-per-Megawatt metric, with production optimizations (KV cache reuse, speculative decoding, prefill/decode disaggregation) allowed and a live leaderboard open to vendor submissions. The case concerns AA-AgentPerf-Local β€” per the case's evidence titles, an open-source sibling tool for laptops and workstations replaying ~8 recorded trajectories (~168 turns, ~56K-token growing contexts) on DGX Spark, RTX 5090, Ryzen AI Halo and MacBook Pro M5 Pro with published configs β€” but the supplied web snippets only cover the datacenter AA-AgentPerf plus a busy adjacent local-hardware benchmarking scene (LMSYS spreadsheets, TokenMark's 211-config tracker, head-to-head buying guides); they do not directly confirm the -Local variant's specifics, its submission mechanics, or any adoption of it.

Why it matters to Scott

Artificial Analysis has independently productized the core of Scott's replay methodology β€” racing competing systems against a frozen corpus of recorded agent trajectories β€” as an open-source reference benchmark aimed exactly at his local-inference hardware territory (RTX/Mac/Strix-Halo-class boxes like his gamepc and MLX stacks), and the growing-context trajectory design makes prefix/KV-cache reuse the hidden discriminator his Prefix-Caching Economics names. This is a dated-receipts convergence with a tool he'd plausibly use for his own hardware and backend comparisons; two watches: the supplied coverage doesn't yet confirm the -Local variant's specifics or any adoption (that's the resolution condition), and his wrong-unit critique applies if serving-throughput numbers get read as agent capability.
ip:concept.progressive-evaluation-ladderip:concept.counterfactual-design-replaydev:concept.trace-backed-agent-comparisondev:concept.hardware-aware-local-inferenceip:concept.prefix-caching-economicsip:concept.benchmarking-the-wrong-unitradar:concept.agent-benchmarksradar:concept.local-inferenceradar:concept.inference-economicsradar:concept.kv-cache
queries asked of Scott's wikis
  • running coding agents on local models β€” local inference setup
  • agent eval methodology β€” replaying real trajectories vs synthetic prompts
  • local LLM hardware choice β€” DGX Spark / Strix Halo / Mac / RTX
  • agent memory β€” growing context, KV cache, long sessions
  • inference economics β€” local vs API cost per agent
  • open-source eval or benchmark tooling he builds or maintains

Measured heat

now 0 pts/hpeak 12 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 314h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-28 14:00⭐ origin echo-reconstructed"Announcing AA-AgentPerf-Local, our open-source inference testing tool for local AI models - test how fast agentic AI can run on your own la
Artificial Analysis on blog (echo) Β· attributed from reddit.post.1wtzpcr
β€”
09-30 08:41first on r/LocalLLaMA Β· published Β· +42.7hAA-AgentPerf-Local: Benchmarking local AI agents on laptops and workstations
nuclearbananana
β€”
09-30 09:57first on hacker news Β· published Β· +44.0hAA-AgentPerf-Local: Benchmarking local AI agents on laptops and workstations
theanonymousone
β€”
09-30 08:41amplified on r/LocalLLaMAreddit.post.1wtzpcr
nuclearbananana
peak 9 Β· 8 comments Β· 35% of case engagement
09-30 09:57amplified on hacker newshn.story.49906656
theanonymousone
peak 1 Β· 0 comments Β· 4% of case engagement
09-30 21:09amplified on r/LocalLLaMA πŸ‘‘reddit.post.1wugwtw
dh7net
peak 4 Β· 22 comments Β· 54% of case engagement
10-06 14:25amplified on hacker newshn.story.49979017
alexellisuk
peak 1 Β· 1 comments Β· 7% of case engagement
09-30 09:20our radar first saw it Β· +43.3hdiscovery anchor: reddit.post.1wtzpcrβ€”
pace: p57 vs 1188 stories at the 168h mark (now 314h old) β€” ahead of devin-code-scans (1.0x), behind futureos-context-compaction-recall (1.0x)

Evidence (5) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditAA-AgentPerf-Local: Benchmarking local AI agents on laptops and workstations
LocalLLaMA
Retrieved article excerpt

Open article Β· Retrieved 2026-09-30T09:24:18.525378+00:00

[Artificial Analysis](https://artificialanalysis.ai/)

K

[All articles](https://artificialanalysis.ai/articles)

September 29, 2026

# AA-AgentPerf-Local: Benchmarking local AI agents on laptops and workstations

**Announcing AA-AgentPerf-Local, our open-source inference testing tool for local AI models - test how fast agentic AI can run on your own laptop or workstation, and browse our list of serving configurations to plan your next agent setup**

**Key points:**

➀ We’re open sourcing AA-AgentPerf-Local, which replays real agent trajectories on laptop & workstation hardware to test inference performance

➀ We’re releasing initial results for NVIDIA DGX Spark, NVIDIA GeForce RTX 5090, AMD Ryzen AI Halo, and MacBook Pro M5 Pro

➀ The tool and leaderboard will soon expand to cover more hardware, frameworks and models, and will stay updated over time as new releases launch

We already benchmark inference performance on mobile phones and datacenter-scale hardware, and are expanding our coverage to include laptops and workstations. Alongside hosting a leaderboard displaying results from popular model and hardware combinations, we are open-sourcing all code and data required to run AA-AgentPerf-Local to ensure that individuals and companies are able to run their own trials, informing their local AI serving decisions.

AA-AgentPerf-Local replays real agent sessions. Our default workload is 8 recorded agentic tasks, spanning 168 model turns. Each request carries the full conversation so far, as a real agent's would, so context grows to ~56K tokens. Every turn generates exactly its recorded number of tokens, so every system does identical work. Tool execution is skipped by default to isolate inference speed, but when benchmarking your own system, you can also replay the real recorded tool delays or run tool calls live on your CPU.

At launch, we are focusing on the performance a single agent can achieve when able to utilize the entire system; we plan to expand this coverage over time to cover multi-agent systems and scenarios where an agent must run alongside other regular processes.

The initial set of hardware covered on our official page is: NVIDIA DGX Spark (128 GB), AMD Ryzen AI Halo (128 GB), MacBook Pro M5 Pro (64 GB), and NVIDIA GeForce RTX 5090 (32 GB). These have been selected to cover a range of platforms (CUDA, ROCm, Vulkan, Metal), memory capacities, and bandwidths. We will be expanding the featured hardware to include x86 (and other) laptops, AI-focused graphics cards such as the RTX PRO 6000 Blackwell, and more, to give consumers a well-rounded view of performance across different hardware types.

The models featured at launch are: Qwen3.5-9B, Qwen3.8-27B, Qwen3.6-35B-A3B, and Ling 3.0 Flash (124B / 5B active). These initial models span a range of memory requirements and dense/MoE architectures, and are each benchmarked at 4-bit quantizations to reflect realistic serving conditions. The core set of models we feature will shift over time as new open weights models are released. Beyond the featured models, AA-AgentPerf-Local is able to benchmark performance of any OpenAI-compatible inference server the user runs, meaning that any model and config can be tested locally.

Choice of serving configuration is important, with the runtime, quantization and speculative decoding changing results substantially. Where an official off-the-shelf config was available for a system and model pair, we used the published config. We developed our own configs for all other cases. Every config uses speculative decoding (MTP, DFlash or DSpark), and all 14 are published in the [repo](https://github.com/ArtificialAnalysis/aa-agentperf-local) and on the [configs page](https://artificialanalysis.ai/hardware-inference-stack/laptops-workstations/configs) on our website.

**Initial results:**

➀ **Completion time mapped most closely to each model’s active parameter count:** Qwen3.6-35B-A3B (3B active) was the fastest model on every system, e.g. 2.5-3.3x faster than the dense Qwen3.8-27B. However, active parameters are not the whole story, with Ling 3.0 Flash (124B total, 5B active) still finishing behind Qwen3.5-9B (nearly double the active parameter count) on all hardware that can support it.

➀ **The GeForce RTX 5090 was the fastest system for every model that fits in its 32 GB,** achieving completion times >3.5x faster than the other systems. Single-user decoding is heavily influenced by memory bandwidth, and the GeForce RTX 5090 has 1,792 GB/s against 256-307 GB/s for the unified-memory systems.

➀ **The DGX Spark and Ryzen AI Halo are overall similar systems,** with the same amount of unified memory, comparable memory bandwidth, and the same launch MSRP of $4,000. On our default trajectory set, the DGX Spark was 1.4-1.7x faster on three of four models, with the Ryzen AI Halo tying it on Qwen3.5-9B. The gap is far larger than their 7% bandwidth difference - the Spark has greater low-precision compute than the Ryzen AI Halo, enabling it to prefill faster, and its more mature CUDA software likely plays a role in enabling better MoE and speculative decoding performance.

➀ **The MacBook Pro (M5 Pro, 64 GB, 20-core GPU) is the only laptop tested so far,** and exhibited competitive results, finishing within 2–9% of the Ryzen AI Halo on Qwen3.6-35B-A3B and Qwen3.8-27B (though 21% slower on Qwen3.5-9B). It has the most memory bandwidth of the three unified-memory systems (307 GB/s) and its current price of $3,700 is the lowest of the systems tested so far. Its results are likely held back by software maturity and compute available for prefill.

➀ **Despite the agentic trajectories serving 73-93% of prompt tokens from the KV cache,** simply reading each turn’s new input (during prefill) used up substantial proportions of the end-to-end completion time, e.g. 22-41% for Qwen3.8-27B. This was especially impactful on systems with low compute FLOP/s relative to their memory bandwidth, such as the Ryzen AI Halo and potentially the MacBook Pro (MacBook FLOP/s are unpublished).

➀ **Speculative decoding was implemented on all of the most successful configs so far,** e.g. raising Qwen3.8-27B decode speeds ~30-120% above the bandwidth-constraint roofline.

Most hardware prices are significantly inflated vs. their launch MSRP. At current market prices, each of the initial four hardware types increase in price as they decrease in end-to-end completion time, with the MacBook Pro (M5 Pro, 64 GB) the current cheapest and the GeForce RTX 5090 system by far the most expensive. Our page defaults to launch MSRP but allows custom price entry. We’re showing below our best estimate of current market prices.

Decode speeds for Qwen3.8-27B were broadly similar on the three unified-memory devices, which have similar memory bandwidths at 256-307 GB/s. The GeForce RTX 5090 exhibited >5x faster decode speeds than these, owing mostly to its >5x faster memory, at 1,792 GB/s bandwidth.

Prefill speeds for Qwen3.8-27B followed a similar pattern to decode, but the DGX Spark’s disproportionately high compute/bandwidth ratio showed up in faster context reading speeds. This is important in our agentic trajectories as, even with most of each prompt served from cache (~93% on llama.cpp), each session still has ~196K new input tokens against ~31K output tokens. When a tool returns a large file, the agent can pause up to 24 seconds on the Halo and 39 seconds on the MacBook before replying, against about 2 seconds on the GeForce RTX 5090.

AA-AgentPerf-Local and the Laptops & Workstations page will be frequently updated as relevant new hardware and software becomes available. We are planning to implement a range of additional features to better help users run local AI optimally:

➀ A wider range of popular hardware and models

➀ A more comprehensive repository of configs for local AI deployment

➀ More custom inference frameworks, such as Inco Splash, which is currently being tested

➀ User-submitted result leaderboards

➀ Leaderboards featuring live CPU tool-calling

➀ Multi-agent scenarios

Try AA-AgentPerf-Local on your own hardware: <https://github.com/ArtificialAnalysis/aa-agentperf-local>

View the initial results here: <https://artificialanalysis.ai/hardware-inference-stack/laptops-workstations> and every serving configuration here: <https://artificialanalysis.ai/hardware-inference-stack/laptops-workstations/configs>

#### Read the latest

[### GPT-6.1 Sol replaces GPT-6 Sol after just 7 days, with near-Astra intelligence

GPT-6.1 Sol replaces GPT-6 Sol after just 7 days. It scores 1 point below GPT-6 Astra in the Intelligence Index at less than one quarter of the Cost per Task

September 29, 2026](https://artificialanalysis.ai/articles/gpt-6-1-sol-replaces-gpt-6-sol-after-just-7-days-with-near-astra-intelligence)[### Announcing the Artificial Analysis Cyber Index Alliance

The Artificial Analysis Cyber Index Alliance brings together industry partners to set a new standard for evaluating how AI models perform on enterprise cyber defense tasks. The Alliance launches alongside the Artificial Analysis Cyber Index, which combines three partner-contributed and open benchmarks to evaluate how well agents find and fix vulnerabilities.

September 28, 2026](https://artificialanalysis.ai/articles/artificial-analysis-cyber-index)[### Claude Sonnet 5.5 reaches #2 on the Artificial Analysis Intelligence Index

Anthropic's new Sonnet model scores 56, just 2 points behind Opus 5.5 (max), but at the highest Output Tokens per Task we have measured

September 28, 2026](https://artificialanalysis.ai/articles/claude-sonnet-5-5)
nuclearbananana98
🟧 echo.blog ⭐"Announcing AA-AgentPerf-Local, our open-source inference testing tool for local AI models - test how fast agentic AI can run on your own laArtificial Analysisβ€”β€”
🟧 hnAA-AgentPerf-Local: Benchmarking local AI agents on laptops and workstationstheanonymousone10
🟠 redditI need help to benchmark harness/model/hardware combination.
LocalLLaMA
dh7net422
🟧 hnShow HN: RigMark benchmarks local AI the way coding agents use italexellisuk11

Interpretation history

Decision trace