2026-10-11 16:38 UTC

Castmates demonstrates viable on-device 3B roleplay models on iPhone with Metal acceleration, establishing a consumer pattern for private local agents.

state: seedheat: lowuncertainty: mediumconvergesscott: highon-device-inference iphone-llm roleplay-agents llama-cpp-metal local-agentsLow-Future-9387 (Castmates maker)

What is this?

Castmates is an iPhone app by maker Low-Future-9387 that runs a 3B-parameter roleplay finetune fully on-device via llama.cpp with Metal acceleration, requiring no network after model download. The app ships 32 hand-crafted characters across eight genres plus end-to-end custom character creation, and the maker has published measured memory/performance numbers from their engineering work. The web results confirm the app exists on the App Store and runs offline, but the first-party technical write-up cited in the case (with concrete quantization, latency, and memory figures) is not directly visible in these snippets โ€” its claims about 'what we measured about keeping a small model in character' rest on that unpublished engineering post.

Why it matters to Scott

A shipping iPhone app (Castmates) independently demonstrates the viable consumer pattern Scott's frameworks argue for: a 3B roleplay model running fully on-device via llama.cpp Metal, offline, no account, with measured memory/performance numbers. This validates his sovereign-software-assurance and agent-native-computing positions with a concrete product proof point, and supplies hardware-aware local-inference data (Metal on iPhone, 3B quantization tradeoffs, character fidelity at small scale) that bears directly on his ollama/MLX/hardware-aware-inference work.
ip:framework.sovereign-software-assuranceip:framework.agent-native-computingdev:concept.hardware-aware-local-inferencedev:concept.ai-roleplay-personadev:technology.ollamadev:technology.mlxradar:backburner-iphone-offloadradar:aiope-android-agent-runtimeradar:dfm-mimir-small-model-validationradar:compressed-llm-fidelity-safety-gapradar:adaptive-kv-cache-streamingradar:littlebit-latent-factorization-quantization
queries asked of Scott's wikis
  • on-device inference patterns for consumer mobile
  • llama.cpp Metal optimization and quantization tradeoffs
  • local agent architectures that run entirely on-device
  • private AI consumer product patterns (no cloud, no account)
  • small-model roleplay/character fidelity techniques
  • model sovereignty / local-first AI product strategy

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p25momentum: steady1 platformsage 70h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

10-08 17:46โญ origin directly observedRunning a 3B roleplay finetune fully on iPhone: what we measured about keeping a small model in character (I make the app)
Low-Future-9387 on r/LocalLLaMA
โ€”
10-08 17:46amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1x0xkk7
Low-Future-9387
peak 0 ยท 3 comments ยท 100% of case engagement
10-08 22:30our radar first saw it ยท +4.7hdiscovery anchor: reddit.post.1x0xkk7โ€”
pace: p37 vs 1204 stories at the 48h mark (now 70h old) โ€” ahead of 3jsbench-llm-3d-generation-benchmark (1.5x), behind acs-local-skill-risk-catalog (0.8x)

Evidence (1) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  reddit โญRunning a 3B roleplay finetune fully on iPhone: what we measured about keeping a small model in character (I make the app)
LocalLLaMA
Low-Future-938703

Interpretation history

Decision trace