MiniCPM5-2B is an OpenBMB model release: an Artificial Analysis article’s related-news section explicitly lists “OpenBMB releases MiniCPM5-2B,” dated September 7, 2026. The case reports a score of 15 on Artificial Analysis Intelligence Index v4.2 and leadership among open-weight models with at most 4B parameters, but the supplied snippets do not substantiate either claim; the tiny-model listing instead names other leaders, without establishing a comparable date or index version. OpenBMB’s repository snippet describes the separate MiniCPM5-1B as intended for on-device and resource-constrained use, so the 2B release’s actual local-inference performance and release artifacts remain unverified here.
As supplied, this is another instance of Scott’s already-held Model Perishability position: model progress warrants replaceable backends and re-evaluation, with his gamepc/Ollama stack providing a concrete testing destination. The radar hits do not establish prior coverage of this release, but its claimed benchmark lead, local performance and serving compatibility remain unverified, so it does not yet justify changing his model choices or arguments.
ip:concept.model-perishabilitydev:project.gamepcdev:technology.ollamaradar:concept.open-modelsradar:concept.local-inferenceradar:concept.small-models
queries asked of Scott's wikis
- small local models capability thresholds task routing
- on-device inference memory budgets quantization latency
- open-weight models sovereignty deployment control
- benchmark scores versus coding agent reliability
- local models RAG extraction agent memory maintenance
2026-10-04T18:28:09Z
The first genuinely independent hands-on test lands against the case's operative claim: parepeg's LFM2.5 2.6B comparison finds the peer faster, lower-RAM despite more parameters, and better in English, while MiniCPM5 runs RAM-heavy with Chinese leakage — so the 'improving quality for resource-constrained local inference' hypothesis is contradicted in practice, and the never-verified 15/AAII-v4.2 leadership claim has no live conversation left to verify it. The episode ends; what small-model attention remains belongs to LFM2.5 2.6B, not this release.
2026-10-04T18:25:58Z
evidence attached: reddit.post.1wxlw2c — Hands-on comparison finds MiniCPM5-2B stronger at agentic tasks but heavier in RAM and prone to Chinese leakage — material practical context for the case's resource-constrained-value claim.
2026-09-24T03:17:06Z
The newest hands-on demo post turns out to be the original release poster amplifying the same browser-agent demo already on file — low engagement, truncated body, no measured results — so it is repetition, not corroboration; the substantive evidence set is unchanged. The case is cooling (current rate ~0 pts/h against a ~26 pt/h peak), with the benchmark-leadership claim still unverified and the only firsthand capability report a failed tool call.
2026-09-22T15:22:43Z
evidence attached: reddit.post.1wnb30v — A hands-on demo shows MiniCPM5-2B powering a browser-local coding agent, adding practical agent-workflow evidence to the model release.
2026-09-16T14:36:09Z
The new customer-service comparison adds a task-evaluation lead, but its truncated evidence does not identify which model succeeded, and it comes from the original release poster rather than independent validation. Questions about tool-schema parity and clarification policy further limit its value; it does not establish a practical upgrade or overturn the earlier browser-agent failure report.
2026-09-16T14:22:54Z
evidence attached: reddit.post.1why0um — The customer-service task comparison provides practical, though anecdotal, evidence about MiniCPM5-2B's agentic task-completion capability relative to another small model.
2026-09-15T04:25:37Z
A commenter now reports a failed tool call when attempting a simple hello-world app, adding a limited firsthand failure report to the browser-agent implementation claim. This weakens the case for immediately useful local coding capability, but does not distinguish model limitations from integration problems or contradict the claimed benchmark score.
2026-09-14T17:50:20Z
A separate poster now reports a browser-only coding agent using Pi, MiniCPM5-2B and WebGPU, giving the release a more specific implementation lead than benchmark claims or intended uses. However, the supplied evidence is only a title, not an inspected demo or measured result, so it does not yet establish practical coding capability or benchmark leadership.
2026-09-14T17:24:41Z
evidence attached: reddit.post.1wg97cp — The public WebGPU coding-agent demo is concrete deployment evidence for MiniCPM5-2B beyond benchmark claims.
2026-09-10T23:42:54Z
No substantive evidence has arrived since the DSpark pointer: it remains a concrete verification lead, not an established throughput improvement. The underlying release is known, but benchmark leadership and practical value for Scott’s local stack remain unverified.
2026-09-08T22:46:13Z
A newly surfaced comment links a purported MiniCPM5-2B-DSpark variant and claims higher throughput, adding a concrete deployment-artifact lead rather than just prospective use cases. The linked model card has not been inspected, and neither the speed benefit nor benchmark leadership is independently established.
2026-09-08T16:43:59Z
The refreshed discussion remains prospective use cases and unsupported comparisons, leaving MiniCPM5-2B an evaluation lead rather than a demonstrated local-inference upgrade. Neither independent benchmark confirmation nor measured deployment evidence has arrived; comment-only refreshes warrant no promotion or accelerated review.
2026-09-08T13:38:06Z
The refreshed discussion adds no substantive evidence: intended uses and unsupported comparisons still do not establish benchmark leadership or a practical local-inference upgrade. The release remains an evaluation lead, with no new reason to change Scott’s model choices.
2026-09-08T10:27:47Z
The refreshed discussion supplies neither tested local-inference results nor attributable benchmark confirmation; the repository echo remains a pointer, not independent release-artifact verification. MiniCPM5-2B remains an evaluation lead, with no new basis to change Scott’s model choices.
2026-09-07T23:28:29Z
The comment refresh adds no independent benchmark evidence or measured deployment results; the known release remains an evaluation lead, not a demonstrated improvement for Scott’s local stack. Proposed uses and unsupported comparisons do not change that interpretation.
2026-09-07T22:34:04Z
The refreshed discussion remains repetitive amplification, with no measured deployment results or independent support for benchmark leadership. The release is still a local-model evaluation lead rather than evidence to change Scott’s model choices; the repository echo supplies no separate corroboration.
2026-09-07T19:40:14Z
The refreshed discussion adds no tested deployment or credible comparison, so the known release remains an evaluation lead rather than a demonstrated upgrade for resource-constrained inference. Benchmark leadership still lacks independent support; further comment-only refreshes do not warrant hourly review.
2026-09-07T18:25:23Z
The comment refresh adds no substantive evidence beyond the already-known release: proposed uses and unsupported model comparisons do not establish a local-inference upgrade. Keep this as an evaluation lead, with a slower review cadence until benchmark provenance or measured deployment results arrive.
2026-09-07T17:47:03Z
The discussion remains repetitive amplification and prospective use cases, not evidence of a practical local-inference upgrade; the unsupported comparison to a 120B model adds no reliable capability signal. Keep the release as an evaluation lead, with benchmark leadership and suitability for Scott’s stack still unverified.
2026-09-07T16:29:02Z
The refreshed discussion is still prospective use cases and comparison requests, not evidence that MiniCPM5-2B improves local inference quality or deployment economics. The release remains an evaluation lead; neither its claimed benchmark leadership nor an advantage for Scott’s stack gains independent support.
2026-09-07T15:24:12Z
The refreshed comments remain prospective use cases and requests for comparisons, not deployment results or independent benchmark confirmation. The release remains a local-model evaluation lead; nothing new establishes an advantage for Scott’s stack.
2026-09-07T14:35:34Z
The refreshed discussion adds intended uses, not evidence of successful deployment or comparative performance; MiniCPM5-2B remains a local-model evaluation lead rather than a demonstrated upgrade. The repository echo repeats the original report’s links and does not independently corroborate the benchmark claim.
2026-09-07T14:26:48Z
grounded: known/low — As supplied, this is another instance of Scott’s already-held Model Perishability position: model progress warrants replaceable backends and re-evaluation, with
2026-09-07T14:24:12Z
case created — A concrete small open-weight model release and a specific benchmark claim warrant tracking, without importing the scout’s unsupported multimodal claim.