2026-10-11 17:14 UTC

AndroidLife creator East-Muffin-6472 reports that Qwen3.8-27B completed only 56.7% of 60 consecutive phone tasks while the device consumed 69% of its battery, suggesting sustained smartphone-agent deployment faces substantial reliability and device-resource constraints.

state: seedheat: lowuncertainty: highknownscott: lowmobile-agents agent-benchmarks agent-harnessesEast-Muffin-6472

What is this?

The case describes AndroidLife as a real-phone agent evaluation created by East-Muffin-6472, who reportedly ran Qwen3.8-27B through 60 consecutive tasks, recording 34 manually audited successes (56.7%) and 69% battery consumption. None of the supplied web snippets directly corroborates AndroidLife, its creator, or those run measurements. The snippets describe other mobile-agent benchmarks, including AndroidWorld and MobileWorld-Real; AndroidWorld's reported 81.9% for Qwen3.8-27B is a different evaluation, not a directly comparable or conflicting result. The supplied material does not establish the run's duration, device configuration, inference location, or battery baseline, so it cannot isolate the cause of battery drain or establish general deployment limits.

Why it matters to Scott

The report illustrates positions Scott already holds in Model-Plus-Harness Benchmark Unit and Verification Loops: evaluate the configured agent system and check actual outcomes, rather than treating model identity or claimed completion as capability. No supplied radar hit tracks this exact AndroidLife run, but the uncorroborated measurements and missing harness, duration, inference-location and battery-baseline details do not establish a consequential extension of those positions or an actionable constraint on his projects.
ip:concept.model-plus-harness-benchmark-unitip:concept.verification-loopsradar:concept.mobile-agentsradar:concept.agent-evaluationradar:concept.agent-reliability
queries asked of Scott's wikis
  • long-horizon agent reliability consecutive task evaluation
  • agent harness step budgets failure recovery
  • benchmark scores versus real-world agent performance
  • manual audit task success verification
  • mobile GUI automation device resource budgets
  • local versus remote inference deployment economics

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 1082h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

08-27 14:00⭐ origin echo-reconstructedPrimary run report for qwen3.8-27b TEXT on 60 real-phone tasks. It states: β€œ34 / 60 (56.7%)” manual-audit success, 29.25 average steps, abou
Yuvraj Singh on other (echo) Β· attributed from reddit.post.1wis5p0
β€”
09-17 12:06first on r/LocalLLaMA Β· published Β· +502.1hAndroidLife: Can an AI agent survive a day in the life of a real user? Qwen3.8-27b run: 56.7% SR
East-Muffin-6472
β€”
09-18 10:54first on r/artificial Β· published Β· +524.9hAndroidLife: Can an AI agent survive a day in the life of a real user? Qwen3.8-27b run: 56.7% SR
East-Muffin-6472
β€”
09-17 12:06amplified on r/LocalLLaMA πŸ‘‘reddit.post.1wis5p0
East-Muffin-6472
peak 29 Β· 17 comments Β· 92% of case engagement
09-18 10:54amplified on r/artificialreddit.post.1wjmyhc
East-Muffin-6472
peak 4 Β· 0 comments Β· 8% of case engagement
09-17 12:20our radar first saw it Β· +502.3hdiscovery anchor: reddit.post.1wis5p0β€”

Evidence (3) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditAndroidLife: Can an AI agent survive a day in the life of a real user? Qwen3.8-27b run: 56.7% SR
LocalLLaMA
Retrieved article excerpt

Open article Β· Retrieved 2026-09-17T12:22:01.220025+00:00

# Prove your humanity

We’re committed to safety and security. But not for bots. Complete the challenge below and let us know you’re
a real person.

[Reddit, Inc. Β© "2026". All rights reserved.](https://www.redditinc.com/)

[User Agreement](https://www.reddit.com/help/useragreement)
[Privacy Policy](https://www.reddit.com/help/privacypolicy)
[Content Policy](https://www.reddit.com/help/contentpolicy)
[Help](https://support.reddithelp.com/hc/en-us)
East-Muffin-64722917
🟧 echo.other ⭐Primary run report for qwen3.8-27b TEXT on 60 real-phone tasks. It states: β€œ34 / 60 (56.7%)” manual-audit success, 29.25 average steps, abouYuvraj Singhβ€”β€”
🟠 redditAndroidLife: Can an AI agent survive a day in the life of a real user? Qwen3.8-27b run: 56.7% SR
artificial
East-Muffin-647240

Interpretation history

Decision trace