Independent testing will determine whether Zeno’s offloading approach makes Qwen3.5-35B-A3B practically usable as a private agentic work tool on 16GB Macs.
state: expiredheat: lowuncertainty: highknownscott: mediumlocal-inference coding-agents agent-harnessesIcosaZenoQwen
What is this?
Icosa claims to have built Zeno, a private local AI work tool that runs Qwen3.5-35B-A3B on a 16GB Mac by offloading model data beyond the machine’s available memory. The supplied sizing estimate says a standard Q4 configuration needs about 28.1GB against roughly 11.5GB available and predicts only about 2 tokens/second with substantial offloading, while a much smaller Q1 quantization may fit. The snippets do not independently document Zeno’s architecture, performance, agent capabilities, or even identify its creators beyond the case, so practical usability remains an unverified claim requiring hands-on testing.
Why it matters to Scott
Scott already holds the operative position in “Hardware-aware local inference”: memory placement, precision, throughput, and usability must be tested as runtime policy rather than inferred from model size. The radar’s open Slipstream case already asks nearly the same practical question about offloading large MoE coding models on constrained Macs; Zeno matters as another hands-on validation target that could inform the Ask agent’s local backend, but its architecture, speed, and agent quality remain unverified.
dev:concept.hardware-aware-local-inferencedev:project.askip:concept.model-plus-harness-benchmark-unitradar:slipstream-ssd-moe-streamingradar:concept.apple-silicon-inferenceradar:concept.local-inference
queries asked of Scott's wikis
- local inference offloading economics and usability thresholds
- private coding agents on constrained Apple Silicon
- agent harness performance versus base-model capability
- quantization tradeoffs for tool-using agents
- local-first AI work tools and data sovereignty
- minimum viable token speed and context for coding agents
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (1) — ⭐ canonical anchor
Interpretation history
2026-08-21T18:32:40Z
After 48 hours, no benchmarks, implementation details, or independent users have emerged to advance the original maker claim. The episode has faded and should reopen only if concrete 16GB-Mac performance or agent-workflow testing appears.
2026-08-19T17:55:58Z
The refreshed discussion only repeats requests for benchmarks; it adds no measurements, implementation details, or independent validation. Zeno remains a relevant but unverified maker claim about constrained-Mac offloading.
2026-08-19T15:48:15Z
The slight engagement increase adds no independent testing or implementation detail, so the case remains a single maker claim rather than evidence that Zeno makes the model practically usable on 16GB Macs.
2026-08-19T15:30:46Z
grounded: known/medium — Scott already holds the operative position in “Hardware-aware local inference”: memory placement, precision, throughput, and usability must be tested as runtime
2026-08-19T15:26:49Z
case created — The claimed release targets a consequential memory constraint for fully local agent workflows, but currently has only one low-engagement observation.
Decision trace
- 08-22 04:32expireAfter 48 hours, no benchmarks, implementation details, or independent users have emerged to advance the original maker claim. The episode has faded and should reopen only if concrete 16GB-Mac performa
- 08-22 04:32alert_silentThere is no new consequential delta to surface; the case remains an unvalidated release claim and can be rediscovered if independent testing appears.
- 08-22 04:32alert_routeThere is no new consequential delta to surface; the case remains an unvalidated release claim and can be rediscovered if independent testing appears.
- 08-20 03:55repriceThe refreshed discussion only repeats requests for benchmarks; it adds no measurements, implementation details, or independent validation. Zeno remains a relevant but unverified maker claim about cons
- 08-20 03:55alert_silentNo consequential new fact has emerged beyond the original release claim, so independent throughput, memory, context, and agent-quality testing can wait for the next briefing.
- 08-20 03:55alert_routeNo consequential new fact has emerged beyond the original release claim, so independent throughput, memory, context, and agent-quality testing can wait for the next briefing.
- 08-20 02:21sensor_dirtycomment_update
- 08-20 01:48repriceThe slight engagement increase adds no independent testing or implementation detail, so the case remains a single maker claim rather than evidence that Zeno makes the model practically usable on 16GB
- 08-20 01:48alert_silentNo consequential new delta occurred; independent benchmarks of throughput, memory behavior, context handling, or agent quality are still needed, so this can wait for routine review.
- 08-20 01:48alert_routeNo consequential new delta occurred; independent benchmarks of throughput, memory behavior, context handling, or agent quality are still needed, so this can wait for routine review.
- 08-20 01:44alert_silentZeno is a relevant hands-on validation target for constrained-Mac offloading, but the only evidence is a low-engagement maker post and its consequential claims—usable speed, agent quality, memory beha
- 08-20 01:44surface_candidateZeno is a relevant hands-on validation target for constrained-Mac offloading, but the only evidence is a low-engagement maker post and its consequential claims—usable speed, agent quality, memory beha
- 08-20 01:44alert_routeZeno is a relevant hands-on validation target for constrained-Mac offloading, but the only evidence is a low-engagement maker post and its consequential claims—usable speed, agent quality, memory beha
- 08-20 01:30groundScott already holds the operative position in “Hardware-aware local inference”: memory placement, precision, throughput, and usability must be tested as runtime policy rather than inferred from model
- 08-20 01:26createThe claimed release targets a consequential memory constraint for fully local agent workflows, but currently has only one low-engagement observation.