AndroidLife creator East-Muffin-6472 reports that Qwen3.8-27B completed only 56.7% of 60 consecutive phone tasks while the device consumed 69% of its battery, suggesting sustained smartphone-agent deployment faces substantial reliability and device-resource constraints.
state: seedheat: lowuncertainty: highknownscott: lowmobile-agents agent-benchmarks agent-harnessesEast-Muffin-6472
What is this?
The case describes AndroidLife as a real-phone agent evaluation created by East-Muffin-6472, who reportedly ran Qwen3.8-27B through 60 consecutive tasks, recording 34 manually audited successes (56.7%) and 69% battery consumption. None of the supplied web snippets directly corroborates AndroidLife, its creator, or those run measurements. The snippets describe other mobile-agent benchmarks, including AndroidWorld and MobileWorld-Real; AndroidWorld's reported 81.9% for Qwen3.8-27B is a different evaluation, not a directly comparable or conflicting result. The supplied material does not establish the run's duration, device configuration, inference location, or battery baseline, so it cannot isolate the cause of battery drain or establish general deployment limits.
Why it matters to Scott
The report illustrates positions Scott already holds in Model-Plus-Harness Benchmark Unit and Verification Loops: evaluate the configured agent system and check actual outcomes, rather than treating model identity or claimed completion as capability. No supplied radar hit tracks this exact AndroidLife run, but the uncorroborated measurements and missing harness, duration, inference-location and battery-baseline details do not establish a consequential extension of those positions or an actionable constraint on his projects.
ip:concept.model-plus-harness-benchmark-unitip:concept.verification-loopsradar:concept.mobile-agentsradar:concept.agent-evaluationradar:concept.agent-reliability
queries asked of Scott's wikis
- long-horizon agent reliability consecutive task evaluation
- agent harness step budgets failure recovery
- benchmark scores versus real-world agent performance
- manual audit task success verification
- mobile GUI automation device resource budgets
- local versus remote inference deployment economics
Measured heat
now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 1082h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion
How the heat travelled
Evidence (3) β β canonical anchor
Interpretation history
2026-09-18T11:32:44Z
The attached post repeats the same creator's run report; it supplies neither independent corroboration nor a new implementation result. Benchmark validity and resource attribution remain unresolved, so the case still illustrates evaluation concerns rather than establishing general smartphone-agent deployment limits.
2026-09-18T11:22:18Z
evidence attached: reddit.post.1wjmyhc β This directly supplies the reported 56.7% success rate, battery drain, heat, and multi-app failure evidence for the open smartphone-agent reliability case.
2026-09-18T00:28:09Z
Discussion raises specific but unverified questions about task feasibility and claimed completions, making benchmark validity another unresolved issue rather than establishing worse agent reliability. This remains a single-run report, not corroborated evidence of general smartphone-agent deployment limits.
2026-09-17T12:28:18Z
grounded: known/low β The report illustrates positions Scott already holds in Model-Plus-Harness Benchmark Unit and Verification Loops: evaluate the configured agent system and check
2026-09-17T12:24:35Z
origin walked (codex/luna, conf 0.99): anchor reddit.post.1wis5p0 -> echo.other.baf0afe034 by Yuvraj Singh
2026-09-17T12:22:35Z
case created β The reported run provides concrete sustained-workload measurements, but does not establish that model inference ran on the phone or that its results generalize across models.
Decision trace
- 10-07 21:39review_dormantscheduled targets exhausted or 28 quiet days
- 10-07 21:39drop_targetsquiet through full ladder or over cap 8
- 10-07 19:40drop_targetsquiet through full ladder or over cap 8
- 09-18 21:32repriceThe attached post repeats the same creator's run report; it supplies neither independent corroboration nor a new implementation result. Benchmark validity and resource attribution remain unresolv
- 09-18 21:22attachThis directly supplies the reported 56.7% success rate, battery drain, heat, and multi-app failure evidence for the open smartphone-agent reliability case.
- 09-18 21:21propose_attachThis directly supplies the reported 56.7% success rate, battery drain, heat, and multi-app failure evidence for the open smartphone-agent reliability case.
- 09-18 10:28repriceDiscussion raises specific but unverified questions about task feasibility and claimed completions, making benchmark validity another unresolved issue rather than establishing worse agent reliability.
- 09-18 10:27review_screenThe new comments raise potentially consequential concerns that some tasks were impossible or that claimed wins did not occur on-device, but they provide only informal spot-checking rather than suffici
- 09-18 10:21sensor_dirtycomment_update
- 09-18 01:23review_screenThe added comment asks for model quantization details and expresses an opinion, without providing new evidence or changing the assessment.
- 09-17 23:21sensor_dirtycomment_update
- 09-17 22:28groundThe report illustrates positions Scott already holds in Model-Plus-Harness Benchmark Unit and Verification Loops: evaluate the configured agent system and check actual outcomes, rather than treating m
- 09-17 22:24promote_anchororigin walk conf 0.99
- 09-17 22:22createThe reported run provides concrete sustained-workload measurements, but does not establish that model inference ran on the phone or that its results generalize across models.