Castmates demonstrates viable on-device 3B roleplay models on iPhone with Metal acceleration, establishing a consumer pattern for private local agents.
state: seedheat: lowuncertainty: mediumconvergesscott: highon-device-inference iphone-llm roleplay-agents llama-cpp-metal local-agentsLow-Future-9387 (Castmates maker)
What is this?
Castmates is an iPhone app by maker Low-Future-9387 that runs a 3B-parameter roleplay finetune fully on-device via llama.cpp with Metal acceleration, requiring no network after model download. The app ships 32 hand-crafted characters across eight genres plus end-to-end custom character creation, and the maker has published measured memory/performance numbers from their engineering work. The web results confirm the app exists on the App Store and runs offline, but the first-party technical write-up cited in the case (with concrete quantization, latency, and memory figures) is not directly visible in these snippets โ its claims about 'what we measured about keeping a small model in character' rest on that unpublished engineering post.
Why it matters to Scott
A shipping iPhone app (Castmates) independently demonstrates the viable consumer pattern Scott's frameworks argue for: a 3B roleplay model running fully on-device via llama.cpp Metal, offline, no account, with measured memory/performance numbers. This validates his sovereign-software-assurance and agent-native-computing positions with a concrete product proof point, and supplies hardware-aware local-inference data (Metal on iPhone, 3B quantization tradeoffs, character fidelity at small scale) that bears directly on his ollama/MLX/hardware-aware-inference work.
ip:framework.sovereign-software-assuranceip:framework.agent-native-computingdev:concept.hardware-aware-local-inferencedev:concept.ai-roleplay-personadev:technology.ollamadev:technology.mlxradar:backburner-iphone-offloadradar:aiope-android-agent-runtimeradar:dfm-mimir-small-model-validationradar:compressed-llm-fidelity-safety-gapradar:adaptive-kv-cache-streamingradar:littlebit-latent-factorization-quantization
queries asked of Scott's wikis
- on-device inference patterns for consumer mobile
- llama.cpp Metal optimization and quantization tradeoffs
- local agent architectures that run entirely on-device
- private AI consumer product patterns (no cloud, no account)
- small-model roleplay/character fidelity techniques
- model sovereignty / local-first AI product strategy
Measured heat
now 0 pts/hpeak 0 pts/hcomments 0/hpeers p25momentum: steady1 platformsage 70h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion
How the heat travelled
pace: p37 vs 1204 stories at the 48h mark (now 70h old) โ ahead of 3jsbench-llm-3d-generation-benchmark (1.5x), behind acs-local-skill-risk-catalog (0.8x)
Evidence (1) โ โญ canonical anchor
Interpretation history
2026-10-09T00:12:19Z
grounded: converges/high โ A shipping iPhone app (Castmates) independently demonstrates the viable consumer pattern Scott's frameworks argue for: a 3B roleplay model running fully on-devi
2026-10-08T23:59:46Z
case created โ First-party engineering write-up with measured memory/performance numbers for a 3B roleplay model running fully on iPhone via llama.cpp Metal.
Decision trace
- 10-09 18:08attention_communicatedShipping iOS app (Castmates) runs Impish Llama 3B (Q4_K_M, ~2GB) entirely on iPhone via llama.cpp Metal: 4GB phones floor at 4096 context, 6GB get 6144, 8GB get 8192; ~6 tok/s on A18. Key findings: hi
- 10-09 18:08attention_routeFurther reading for 6 PM briefing: concrete on-device validation with measured tradeoffs directly relevant to Scott's ollama/MLX/gamepc stack. No urgency โ the app ships and the engineering notes
- 10-09 13:18attention_routeConcrete on-device validation with measured tradeoffs directly relevant to Scott's ollama/MLX/gamepc stack. No urgency โ the app ships and the engineering notes are published. Fits the 18:00 brie
- 10-09 13:11attention_candidatecreate
- 10-09 11:12groundA shipping iPhone app (Castmates) independently demonstrates the viable consumer pattern Scott's frameworks argue for: a 3B roleplay model running fully on-device via llama.cpp Metal, offline, no
- 10-09 10:59createFirst-party engineering write-up with measured memory/performance numbers for a 3B roleplay model running fully on iPhone via llama.cpp Metal.