2026-10-11 17:11 UTC

Artificial Analysis claims its Pocket-Scale Inference benchmark provides useful comparative measurements of local LLM performance on smartphones, giving builders a practical basis for selecting on-device models and hardware.

state: expiredheat: lowuncertainty: highconvergesscott: mediumlocal-inference edge-ai inference-economicsArtificial Analysis

What is this?

Artificial Analysis launched “Benchmarking pocket-scale inference,” comparing small language models and mobile phones using quantized builds that fit within 8 GB of memory and are served with llama.cpp. Artificial Analysis administers intelligence evaluations selected for mobile use, while Liquid AI develops and runs inference tests on real devices; Artificial Analysis says it independently validated Liquid AI’s measurement process. Reported measures include wall-clock time for processing a 1,024-token prompt and generating 256 tokens, although the supplied snippets do not establish how predictive the benchmark is of developers’ applications beyond its chosen evaluations.

Why it matters to Scott

Artificial Analysis’s hardware-and-model comparisons converge with Scott’s vendor-neutral capability-audit and hardware-aware local-inference approach, potentially adding a practical selection input for quantized deployments. It is not yet high relevance because the supplied evidence does not show that its mobile evaluations predict Scott’s representative workloads or measure full application-level utility.
ip:concept.capability-auditip:concept.model-perishabilitydev:concept.hardware-aware-local-inferenceradar:intelligence-per-watt-local-ai-metricradar:concept.mobile-inferenceradar:concept.local-inferenceradar:concept.model-evaluation
queries asked of Scott's wikis
  • on-device model selection criteria
  • local inference hardware economics
  • edge AI latency memory and battery tradeoffs
  • quantization effects on model capability
  • benchmark validity for real-world agent workloads
  • local-first AI product architecture

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnBenchmarking Pocket-Scale Inferencesys425908418
🟧 echo.blog ⭐Published a dedicated benchmark of pocket-scale local inference on mobile phones.Artificial Analysis——

Interpretation history

Decision trace