Numbat is an open-source, Apache-2.0 agent-security suite released by Perplexity AI to provide endpoint-level visibility across widely used client-side agent harnesses. Its stated capabilities include on-device detection, optional blocking before an action executes, structured activity logging, and forensic reconstruction for investigation and auditing. The supplied evidence establishes Perplexity’s claims and repository release, but provides no independent deployment results showing how reliable, complete, or operationally useful that visibility is in practice.
Perplexity’s Numbat independently converges with Scott’s structured agent-observability, execution-receipt and OS-level containment work, while potentially supplying a cross-harness endpoint layer relevant to SiloOS and trace-backed agent testing. The release creates a dated-receipts and practical evaluation opportunity, but its significance remains limited until independent deployments establish coverage, tamper resistance, privacy properties and blocking reliability.
ip:source.observability-for-agentic-systems-what-to-log-how-to-redact-how-to-debug-ebookip:concept.agent-observabilityip:framework.siloosip:framework.agent-provenance-stackip:concept.agent-receiptsdev:concept.trace-backed-agent-comparisondev:project.silo-osradar:concept.agent-observabilityradar:sentience-governor-agent-execution-trailsradar:concept.agent-harnesses
queries asked of Scott's wikis
- endpoint observability for coding-agent harnesses
- OS-level auditing versus harness-native agent tracing
- pre-action policy enforcement for autonomous agents
- tamper-resistant logs and forensic replay for agent execution
- cross-harness security interfaces and event schemas
- agent monitoring without exposing prompts or sensitive context
2026-08-26T20:35:38Z
Repeated checks have produced no product-specific validation, while adjacent implementations only establish category demand; that no longer justifies keeping a dedicated Numbat reliability episode active. Reopen if an independent deployment, benchmark, or operational report appears.
2026-08-24T19:54:34Z
The added local MCP tracing project provides another independent implementation of the broader execution-auditing pattern, but it does not evaluate Numbat. Repeated observation still leaves Numbat’s endpoint coverage, blocking reliability, and operational usefulness entirely dependent on future deployment evidence.
2026-08-24T19:26:37Z
evidence attached: reddit.post.1vxc25n — The project appears to address post-session tracing of Claude tool calls and exchanged data, directly bearing on practical agent observability.
2026-08-22T18:25:43Z
Agentmetry independently reinforces hook-coverage attestation and tamper-evident local audit trails as useful agent-observability patterns, but it is an adjacent self-reported alpha rather than a Numbat evaluation. Numbat’s endpoint coverage, blocking reliability, and operational usefulness therefore remain unvalidated.
2026-08-22T18:23:23Z
evidence attached: reddit.post.1vvitew — This usable local artifact materially contextualizes agent observability by attesting hook coverage and detecting MCP schema changes.
2026-08-20T17:34:04Z
Repeated reobservation still finds no independent Numbat deployment, benchmark, or operational report; the adjacent agent-security activity does not advance the product-specific reliability question. Keep the case open but move to a slower cadence unless concrete evaluation evidence appears.
2026-08-18T16:59:00Z
Another reobservation produced only negligible engagement and no independent deployment, benchmark, or operational report. Adjacent interest in agent observability is rising, but Numbat’s endpoint coverage and reliability remain entirely unvalidated.
2026-08-16T16:32:46Z
Rungraph adds an adjacent implementation of coding-agent activity visualization, reinforcing demand for execution auditing but not validating Numbat’s endpoint coverage, blocking reliability, or operational usefulness. The case still depends on independent Numbat deployments or benchmarks.
2026-08-16T16:22:47Z
evidence attached: hn.story.49321388 — A first-party coding-agent run-graph artifact provides relevant evidence about practical visibility and auditing of agent activity.
2026-08-14T22:29:40Z
The case remains an unvalidated first-party release: two days of reobservation produced no independent deployment, benchmark, or operational evidence. The slight engagement increase is repetitive amplification and does not change Numbat’s practical credibility.
2026-08-12T21:36:35Z
Independent technical context strengthens the need for runtime visibility beyond sandbox isolation, making Numbat worth watching, but it does not validate Numbat’s actual coverage or reliability. The central question still awaits independent deployment results, benchmarks, or operational reports.
2026-08-12T21:22:57Z
evidence attached: hn.story.49278090 — Independent technical context supports the case that sandbox isolation alone is insufficient without runtime visibility and auditability.
2026-08-12T02:31:23Z
No independent deployment evidence or discussion has emerged; the case remains a first-party release awaiting validation, and the unchanged observation lowers its immediate temperature.
2026-08-12T02:27:45Z
grounded: converges/medium — Perplexity’s Numbat independently converges with Scott’s structured agent-observability, execution-receipt and OS-level containment work, while potentially supp
2026-08-12T02:24:06Z
case created — Numbat is a concrete first-party security artifact addressing the emerging need to observe and audit autonomous agent actions at endpoints.