2026-10-11 17:09 UTC

Google's Pixel-Test-Engineering Fusion team claims its released ARTEMIS framework turns natural-language requests into reliable Android workflows with 99%+ AndroidWorld task completion, potentially letting coding assistants execute device tests and collect diagnostics through MCP.

state: watchingheat: lowuncertainty: highconvergesscott: mediummobile-agents agent-harnesses software-testingGoogleGoogle Pixel-Test-Engineering Fusion team

What is this?

ARTEMIS is an open-source Android automation framework released by Google's Pixel-Test-Engineering Fusion team, according to the google/artemis repository and accompanying coverage. It turns natural-language instructions into cross-app workflows, uses element-based targeting with coordinate and visual fallbacks, and exposes MCP integration so coding assistants can drive test devices and collect Logcat output and screenshots. The repository claims 99%+ success on AndroidWorld, whose own page describes 116 programmatic tasks across 20 Android apps; the supplied snippets do not establish ARTEMIS's evaluation conditions or independently verify that result.

Why it matters to Scott

Google’s released device-action and diagnostic tooling converges with Scott’s “Agent Hands and Eyes” and “MCP as the Tool Belt Standard,” opening a concrete publishing comparison about extending coding assistants’ action-and-observation surfaces to Android testing. This is a new consequential implementation, not an ARTEMIS development already tracked in the supplied radar hits; however, no active Android-testing dependency is established for Scott, and the unverified 99%+ claim does not establish the repeatable quality gates required by his Evaluation-Driven Development position.
ip:concept.agent-hands-and-eyesip:source.mcp-as-the-tool-belt-standard-giving-ai-agents-hands-and-eyes-ebookip:concept.evaluation-driven-developmentradar:mobile-harness-cross-platform-controlradar:concept.mobile-testingradar:concept.mcp
queries asked of Scott's wikis
  • coding agent harnesses end-to-end execution verification loops
  • MCP tool integration coding assistants device diagnostics
  • Android mobile projects automated UI testing
  • agent reliability benchmarks versus production task success
  • multimodal UI automation recovery rollback

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 705h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-12 07:22 (minted)⭐ origin echo-reconstructedARTEMIS provides natural-language Android automation on real devices or emulators, MCP integration with coding assistants, and diagnostic ca
Google Pixel-Test-Engineering Fusion team on github (echo) · attributed from hn.story.49669516 · published time unknown
—
09-12 06:34first on hacker news · published · lag ?Artemis: Google's new AI agent framework for mobile test automation
soltanov
—
09-12 06:34amplified on hacker newshn.story.49669516
soltanov
peak 2 · 0 comments · 25% of case engagement
09-14 02:57amplified on hacker news 👑hn.story.49691379
yarapavan
peak 4 · 0 comments · 49% of case engagement
09-15 07:53amplified on hacker newshn.story.49709184
domhudson
peak 1 · 0 comments · 13% of case engagement
09-16 08:40amplified on hacker newshn.story.49723652
fourfire
peak 1 · 0 comments · 13% of case engagement
09-12 07:21our radar first saw it · lag ?discovery anchor: hn.story.49669516—
pace: p39 vs 1032 stories at the 336h mark (now 705h old) — ahead of agentsec-static-config-auditing (1.2x), behind anthropic-meta-lawsuit (0.8x)

Evidence (5) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnArtemis: Google's new AI agent framework for mobile test automation
Retrieved article excerpt

Open article · Retrieved 2026-09-12T07:22:07.417858+00:00

[ARTEMIS Banner](https://github.com/google/artemis/blob/main/docs/assets/artemis-banner.png?v=7)

**Let AI assistants and test suites use real phones like a human.**

[**English**](https://github.com/google/artemis/blob/main/README.md) •
[中文文档](https://github.com/google/artemis/blob/main/README_CN.md) •
[Workflow Showcase](https://github.com/google/artemis#workflow-showcase) •
[Quick Start](https://github.com/google/artemis#quick-start) •
[MCP for IDEs](https://github.com/google/artemis#mcp-setup) •
[Benchmarks](https://github.com/google/artemis#benchmarks) •
[Discord Community](https://discord.gg/wF2FN4WHGY)

[Python 3.12+](https://www.python.org/downloads/)
[License: Apache-2.0](https://github.com/google/artemis/blob/main/LICENSE)
[MCP Native](https://modelcontextprotocol.io/)
[Multi-Model](https://ai.google.dev/)
[AndroidWorld SOTA](https://github.com/google-research/android_world)

[Artemis in Action](https://github.com/google/artemis/blob/main/docs/assets/demo.gif)
  
*Live Demo: Setup driving routes and calculate total durations in Google Maps, then open YouTube to play a Coldplay song.*

## Key Highlights

- **Cross-App Automation**: Executes testing workflows and everyday tasks on Android from natural language instructions.
- **Multimodal Targeting**: Uses element indices when available, with coordinate and visual locating fallbacks for custom interfaces.
- **IDE Diagnostics**: **Model Context Protocol (MCP)** integration lets **Antigravity, Claude Code, and Windsurf** drive test devices and collect **Logcat** output and screenshots.
- **Flash Execution**: A reactive observe-and-act loop with asynchronous history summaries, typically **3–5s per step**.
- **Pro Exploration**: Checks targets before individual actions and returns blocked actions to the Operator for recovery. Supports long-running exploratory and stability tests.
- **AndroidWorld Results**: **99%+ task completion** on Google Research's **AndroidWorld** benchmark (100+ multi-step tasks).

## Antigravity × ARTEMIS: Autonomous Testing Workflow

**Antigravity** uses **ARTEMIS** through MCP to turn a test request into a plan, device execution, and a diagnostic report:

|  |  |
| --- | --- |
| **1. Prompt Input (Task Dispatch)**  Describe your test scenario and target metrics in Antigravity   [Step 1: Prompt Input in Antigravity](https://github.com/google/artemis/blob/main/docs/assets/workflow-1-prompt.png) | **2. Test Plan Generation**  Formulates a step-by-step test plan & architecture for review   [Step 2: Test Plan Generation](https://github.com/google/artemis/blob/main/docs/assets/workflow-2-plan.png) |
| **3. Autonomous Test Execution**  Drives real device, navigates UI, and profiles performance   [Step 3: Autonomous Test Execution](https://github.com/google/artemis/blob/main/docs/assets/workflow-3-exec.png) | **4. Final Report**  Delivers structured audit findings, metric tables, and raw datasets   [Step 4: Final Report](https://github.com/google/artemis/blob/main/docs/assets/workflow-4-report.png) |

## Quick Start

Ensure an Android device (with **USB Debugging** enabled) or emulator is connected. The one-click startup script will automatically:

- **Install System Toolchains**: Detect and auto-install ADB, scrcpy, FFmpeg, and Python (`uv`) dependencies.
- **Mount Global MCP Server & AI Agent Rules**: Prompt to automatically install global MCP configurations and the **Artemis Mobile Testing Mindset (`rules.md`)** into your AI IDEs (**Antigravity**, **Cursor**, **Claude Code**, **Codex**, **Windsurf**, **VS Code**, **Cline/Roo**, **OpenClaw**).

### macOS and Linux

```
# 1. Clone repo & navigate to directory
git clone https://github.com/google/artemis.git && cd artemis

# 2. One-click launch
./start.sh
```

### Windows PowerShell

```
# 1. Clone repo & navigate to directory
git clone https://github.com/google/artemis.git
cd artemis

# 2. One-click launch
.\start.bat
```

> PowerShell does not search the current directory for executable scripts by default, so use `.\start.bat` without a trailing `\`. In Command Prompt (CMD), use `start.bat` instead.

> **Tip**: Opens `http://localhost:8000` in your default browser with a device connection wizard, live screen mirroring, prompt sandbox, and execution replays. You can also run directly from CLI: `uv run artemis run "Open Settings, find Battery and tell me current level" --profile flash`.

**MCP Setup for Codex / Antigravity / Claude Code / Windsurf (Click to expand)**
  

ARTEMIS includes a native **Model Context Protocol (MCP)** server. Connect your real phone directly into AI IDEs:

### 1. One-Click Auto Install (Recommended)

Running `./start.sh` (macOS/Linux) or `.\start.bat` (Windows PowerShell) will prompt you to configure global MCP and testing rules for detected IDEs (or you can install/update anytime later manually using the commands below):

```
# Auto-install MCP server & global rules for Antigravity / Jetski:
uv run artemis mcp --install antigravity

# Or install for all supported AI IDEs (including Codex):
uv run artemis mcp --install all
```

> **Tip**: You can also configure MCP interactively during first-time setup via `uv run artemis init`.
> **Pro Tip**: If you want to use the `artemis` command globally without `uv run` in any directory, run `uv tool install -e .` once in the project root.

### 2. Manual Configuration (Optional)

If you prefer to configure manually, run `uv run artemis mcp --generate-config <client>` (for example, `codex` or `antigravity`) to output the appropriate TOML or JSON snippet. Replace `/path/to/artemis` with your actual repo path and point `command` to your `.venv` Python executable:

- **Codex** (`~/.codex/config.toml`):

```
[mcp_servers.artemis]
command = "/path/to/artemis/.venv/bin/python"
args = ["-m", "mcp_server"]
cwd = "/path/to/artemis"

[mcp_servers.artemis.env]
PYTHONUNBUFFERED = "1"
PYTHONPATH = "/path/to/artemis"
```

- **Antigravity** (`~/.gemini/jetski/mcp_config.json`):

```
{
  "mcpServers": {
    "artemis": {
      "command": "/path/to/artemis/.venv/bin/python",
      "args": ["-m", "mcp_server"],
      "cwd": "/path/to/artemis",
      "env": {
        "PYTHONUNBUFFERED": "1"
      },
      "tools": {
        "mobile_run_task": { "eager": true },
        "mobile_manage_task": { "eager": true },
        "mobile_get_device_state": { "eager": true },
        "mobile_inspect_trace": { "eager": true },
        "mobile_diagnose": { "eager": true }
      }
    }
  }
}
```

- **Claude Desktop** (`claude_desktop_config.json`):

```
{
  "mcpServers": {
    "artemis": {
      "command": "/path/to/artemis/.venv/bin/python",
      "args": ["-m", "mcp_server"],
      "cwd": "/path/to/artemis"
    }
  }
}
```

### 3. Mount Behavioral Rules for AI Agents (Highly Recommended)

To ensure your AI coding assistant acts with the rigor of a senior mobile test engineer and never hallucinates UI interactions, we provide a dedicated testing mindset rules file at [`mcp_server/rules.md`](https://github.com/google/artemis/blob/main/mcp_server/rules.md) (covering **Active Exploration before coding**, **Flash vs. Pro routing strategy**, **Latency & Timing compensation**, and the **"Dynamic-First, Coordinate-Fallback" locator pattern**).

You can mount or copy [`mcp_server/rules.md`](https://github.com/google/artemis/blob/main/mcp_server/rules.md) into your AI IDE's rule configuration:

- **Antigravity**: Add the contents of `rules.md` to your Workspace Rules, Global Rules settings, or agent instructions.
- **Claude Code**: Run `artemis mcp --install claude` to install the rules to `~/.claude/rules/artemis.md` (install to exactly one location — Claude Code loads both `~/.claude/CLAUDE.md` and `~/.claude/rules/*.md`, so duplicating the rules wastes context).
- **Cursor**: Copy the contents into `.cursorrules` or create a rule file at `.cursor/rules/artemis.mdc`.
- **Codex**: Add the contents to `~/.codex/AGENTS.md` (or the active `AGENTS.override.md`).
- **Windsurf / OpenClaw**: Add the rules to your workspace rules or global system prompts.

> For more details on the testing mindset and MCP architecture, see the [MCP Server README](https://github.com/google/artemis/blob/main/mcp_server/README.md).

### 4. Prompt Your Phone in the IDE Chat

In Codex, Antigravity, or Claude Code, simply prompt:

> *"Build the latest changes into an APK, install it on the connected device, open the login screen with a test account, verify if there are any unexpected popups after login, and return screenshots of the final page."*

**Python SDK Integration (Click to expand)**
  

Install the zero-runtime-dependency client on the development machine. ADB,
agents, models, and image processing remain on the device host:

```
uv add "artemis-client @ git+https://github.com/google/artemis.git#subdirectory=packages/artemis-client"
```

```
import asyncio
from artemis_client import ArtemisClient


async def main():
    client = ArtemisClient(
        "http://artemis-host:8000",
        device_serial="emulator-5554",  # optional: target specific device serial
        default_profile="flash",  # "flash" (fast reactive) or "pro" (deep reasoning)
    )

    result = await client.run(
        "Open System Settings, go to 'Battery', verify battery percentage is displayed, and check for any crash dialogs.",
    )

    assert result.succeeded, f"Test failed: {result.error or result.status}"
    print(f"✅ Test Passed! Device: {result.device_serial} | Trace ID: {result.trace_id}")


if __name__ == "__main__":
    asyncio.run(main())
```

## Usage Modes

[Artemis Web Console](https://github.com/google/artemis/blob/main/docs/assets/artemis-ui-showcase-en.png)
  
**Console Overview**: **① View Switcher** (Home / Workspace) · **② Model & Replay** (Flash/Pro status & video replay) · **③ Live Agent Stream** (Action perception, target coordinates & structured results) · **④ Prompt Dock** (Natural language dispatch) · **⑤ Task Queue & Dashboard** (Lifecycle & history)

- **Web Visual Test Console (`uv run artemis ui`)**: Real-time screen projection and interactive panel, supporting natural language test dispatch, live reasoning telemetry, action trajectories, and execution replay; manage server lifecycle anytime from any terminal using `uv run artemis restart`, `uv run artemis stop`, and `uv run artemis status`;
- **MCP Server**: Connects **Antigravity, Claude Code, Windsurf**, and other MCP clients to real devices for bug reproduction and test execution;
- **Developer CLI (`uv run artemis run`)**: Direct terminal execution for automated test cases, exploratory stability inspection, or AndroidWorld benchmarks with high-fidelity structured terminal output;
- **Python SDK**: Integrates as a standard Python library into existing automated testing frameworks (e.g., pytest) or CI/CD pipelines with strongly typed Pydantic structured outputs and assertion support.

## What ARTEMIS Installs on Your Phone

The first task on a device installs the **Artemis Accessibility Helper**, a small
accessibility service that reads the screen layout without taking the
UiAutomation connection. Tools using UiAutomation can suppress the helper unless
they enable `FLAG_DONT_SUPPRESS_ACCESSIBILITY_SERVICES`. You will see
a collapsed "Artemis test helper is running" notification and a new entry under
Settings > Accessibility; both are that helper. It listens only on the phone
itself and sends nothing elsewhere.

- Pre-install it (avoids the ~3 s delay on the first task): `uv run artemis helper install`
- Inspect it: `uv run artemis helper status` / `uv run artemis doctor`
- Remove it any time: `uv run artemis helper uninstall`
- Use UIAutomator2 instead: `ARTEMIS_HIERARCHY_BACKEND=uiautomator` in `.env`
- Prevent automatic installation: `ARTEMIS_HELPER_AUTO_INSTALL=false` in `.env`

If the helper ever fails mid-task, ARTEMIS falls back to UIAutomator2 and says
so in the task timeline, in `mobile_manage_task` status, and in the final report.

## Benchmarks: AndroidWorld (SOTA 99%+)

Artemis ac
soltanov20
🟧 echo.github ⭐ARTEMIS provides natural-language Android automation on real devices or emulators, MCP integration with coding assistants, and diagnostic caGoogle Pixel-Test-Engineering Fusion team——
🟧 hnGoogle Artemis: Let AI assistants and test suites use real phones like a humanyarapavan40
🟧 hnGoogle Artemis – Missing attribution to mobile-use and its contributorsdomhudson10
🟧 hnGoogle released Artemis for AI-driven Android automationfourfire10

Interpretation history

Decision trace