2026-10-11 16:37 UTC

OntoPrune maintainer vigmarcarlo claims his MIT-licensed middleware — translating code into RDF/SPARQL contract stubs for local SLMs and coding agents via MCP, Python, and CLI — cuts context ~83% and speeds TTFT 6.7x on CPU with zero invalid API calls, and independent replication or builder adoption beyond its single-file self-benchmark would establish ontology-based context pruning as a practical local-inference layer, while quiet fade closes it as another self-benchmarked release.

state: seedheat: lowuncertainty: mediumknownscott: lowcontext-pruning local-inference agent-harnesses inference-economicsvigmarcarlo

What is this?

OntoPrune is a single-maintainer (vigmarcarlo) MIT-licensed middleware project that translates source code into RDF/SPARQL 'contract stubs' and serves them to local small language models and coding agents over MCP, a Python API, and a CLI — its README and two same-author announcement posts claim ~83–86% context-token savings, 6.7x faster time-to-first-token on CPU via Ollama, and zero invalid API calls. Notably, the supplied web search returned no direct coverage of OntoPrune itself — no repo page, no thread, no third-party mention — so those numbers rest entirely on the author's own materials, self-benchmarked on a single file, with zero visible traction. What the snippets do establish is the surrounding lane: a crowded field of local-first context engines for coding agents (archex, vexp, CodeGraph) making comparable 57–86% token-savings claims from their own self-run benchmarks, making OntoPrune one more entrant in an active, self-benchmarked niche rather than an isolated event.

Why it matters to Scott

Scott's canon already holds this position: dev:concept.deterministic-code-skeleton and dev:concept.adaptive-source-context-compilation are his own versions of compiling source into compact structured stubs under a token budget, and the index-is-the-data cluster already argues symbolic structure over embedding retrieval — OntoPrune re-derives that with the semantic-web toolkit. Its 83%/6.7x numbers are seller-class claims that stop exactly where ip:concept.evidence-class-ladder says evidence stops being spendable, so this is the world agreeing with him again from a zero-traction source; it only goes live if the case's own replication fork resolves, at which point the formal-RDF vs natural-language-claims substrate question (ip:concept.relational-arity-boundary, three-substrates) would bear on what he argues.
dev:concept.deterministic-code-skeletondev:concept.adaptive-source-context-compilationip:framework.the-index-is-the-data-self-cleaning-wiki-graphip:concept.evidence-class-ladderip:concept.relational-arity-boundaryradar:ctxfw-ast-pruning-token-firewallradar:cognee-codebase-memory-efficiencyradar:concept.token-efficiencyradar:concept.local-inferenceradar:concept.knowledge-graphs
queries asked of Scott's wikis
  • graph/symbolic code context vs embedding retrieval for coding agents
  • local SLM CPU inference TTFT and cost economics
  • coding agent context-window budgeting and pruning layers
  • MCP middleware and agent harness tooling patterns
  • self-benchmarked performance claims — replication and eval rigor
  • RDF/knowledge-graph agent memory vs prose wiki memory

Measured heat

now 0 pts/hpeak 9 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 148h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

10-05 12:26 (minted)⭐ origin echo-reconstructedRepo README: 'Neuro-Symbolic Context Pruning Middleware for Local SLMs' translating source into RDF/SPARQL contract stubs; self-benchmark on
vigmarcarlo on github (echo) · attributed from reddit.post.1wy6qvf, hn.story.49963837 · published time unknown
—
10-05 12:02first on r/LocalLLaMA · published · lag ?[Project] OntoPrune: 6.7x faster TTFT on local CPU with Ollama by neuro-symbolic context pruning (83-86% token savings + 0% hallucinations)
vigmarcarlo
—
10-05 12:15first on hacker news · published · lag ?OntoPrune – Pruning 85% LLM context tokens and 6.7x TTFT on CPU
vigmarcarlo
—
10-05 12:02amplified on r/LocalLLaMA 👑reddit.post.1wy6qvf
vigmarcarlo
peak 1 · 5 comments · 77% of case engagement
10-05 12:15amplified on hacker newshn.story.49963837
vigmarcarlo
peak 1 · 0 comments · 24% of case engagement
10-05 12:20our radar first saw it · lag ?discovery anchor: reddit.post.1wy6qvf—
pace: p43 vs 1247 stories at the 96h mark (now 148h old) — ahead of acs-local-skill-risk-catalog (1.2x), behind anthropic-blocked-request-billing (0.9x)

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit[Project] OntoPrune: 6.7x faster TTFT on local CPU with Ollama by neuro-symbolic context pruning (83-86% token savings + 0% hallucinations)
LocalLLaMA
vigmarcarlo05
🟧 hnOntoPrune – Pruning 85% LLM context tokens and 6.7x TTFT on CPU
Retrieved article excerpt

Open article · Retrieved 2026-10-05T12:24:49.807649+00:00

# OntoPrune

*Neuro-Symbolic Context Pruning Middleware for Local SLMs*

OntoPrune es un middleware traductor ligero que transforma código fuente en contratos de contexto mínimos utilizando representación ontológica (RDF/SPARQL), reduciendo drásticamente los tokens de entrada y la latencia TTFT (Time to First Token) para modelos de lenguaje pequeños (SLMs) y evitando alucinaciones de API.

---

## Resultados Empíricos (Benchmark en CPU)

Evaluación real con streaming sobre [sample\_service.py](https://github.com/vigmarcarlo/OntoPrune/blob/main/fixtures/sample_service.py) (300+ LOC) en CPU local (12 cores):

| Métrica | Naive (Archivo Completo) | OntoPrune (`stubs`) | Ganancia Real |
| --- | --- | --- | --- |
| **Sobrecarga CPU** | 0.02 ms | **9.9 ms** | $\le 10\text{ ms}$ (Meta: $\le 15\text{ ms}$) |
| **Tokens Entrada** | 2,390 tokens | **406 tokens** | **-83.0%** ($\approx 6\text{x}$ menos) |
| **TTFT (`qwen2.5-coder:3b`)** | 22.4 s | **3.3 s** | **6.7x más rápido** (ahorra 19.1 s) |
| **Tiempo Total (3B)** | 59.9 s | **16.5 s** | **-72.5%** ($3.6\text{x}$ más rápido) |
| **Alucinaciones API** | 1 método inválido | **0 métodos inválidos** | **100% Precisión Contractual** |



---

## Instalación

```
pip install -e .

# Con dependencias para benchmarking:
pip install -e ".[dev,benchmark]"
```

---

## Modos de Uso

### 1. Como Servidor MCP (Model Context Protocol)

OntoPrune incluye un servidor MCP nativo (`ontoprune-mcp`) para integrarse con Cursor, Claude Desktop, Gemini CLI o cualquier agente:

```
# Ejecutar servidor MCP en transporte stdio:
ontoprune-mcp
```

**Configuración en `claude_desktop_config.json` o similar:**

```
{
  "mcpServers": {
    "ontoprune": {
      "command": "ontoprune-mcp"
    }
  }
}
```

**Herramientas MCP expuestas:**

- `prune_context(file_path, target_symbol, format='stubs', include_body=False, project_root=None)`: Extrae el contrato podado mínimo (< 400 tokens) resolviendo dependencias entre múltiples archivos del proyecto.
- `verify_response(response_code, contract_or_file, target_symbol)`: Detecta alucinaciones en código generado comparando contra el contrato o archivo.

---

### 2. Como Librería Python

```
import ontoprune

# 1. Traducir un archivo a contexto compacto podado (formato stubs)
# Resuelve automáticamente imports relativos y absolutos entre módulos del proyecto
context = ontoprune.translate(
    "services/order_service.py",
    target="procesar_orden",
    fmt="stubs",
    multi_module=True,
)
print(context)

# 2. Verificar respuestas del modelo frente al contrato
violations = ontoprune.check(llm_code_response, against=context)
if not violations:
    print("Código 100% válido")
```

---

### 3. Desde la Línea de Comandos (CLI & Pipes)

```
# Traducir función en proyectos multi-módulo:
ontoprune translate services/order_service.py procesar_orden --format stubs

# Pipeline directo con Ollama en proyectos modulares:
ontoprune translate services/order_service.py procesar_orden | ollama run qwen2.5-coder:3b
```

---

### 4. Ejecución del Benchmark Empírico

```
# Correr benchmark comparativo (Naive vs OntoPrune con ablación de formatos):
python -m ontoprune.benchmark --file fixtures/sample_service.py --func procesar_orden --backend ollama --model qwen2.5-coder:3b --compare-formats
```

---

## Modelos de Monetización y Aplicación Comercial

OntoPrune resuelve dos problemas críticos de costo y confiabilidad en ingeniería de IA:

### 1. Token Cost Optimization Gateway (B2B SaaS / Middleware de Ahorro)

- **Problema:** Equipos que operan agentes de código autónomos (Devin, Cursor, Copilot Workspace) gastan miles de dólares mensuales en tokens de entrada donde más del 80% es código irrelevante.
- **Solución:** OntoPrune como proxy o sidecar que poda el contexto antes de enviarlo a APIs comerciales (Gemini, Claude, OpenAI), reduciendo la factura de tokens en un **85%**.
- **Monetización:** Cobro basado en porcentaje de ahorro (*gain-share*: ej. 10% del ahorro mensual generado).

### 2. Local-First Developer Tooling (Edición Profesional / Equipos)

- **Problema:** Desarrolladores y empresas que requieren privacidad estricta ejecutan SLMs en laptops o servidores locales (Ollama/vLLM), sufriendo latencias inaceptables en CPU.
- **Solución:** OntoPrune reduce el TTFT de 22s a 3.3s (6.7x más rápido).
- **Monetización:** Versión open-source para desarrolladores individuales + licencia empresarial para equipos (soporte multi-repositorio, telemetría y reglas de cumplimiento de arquitectura).

### 3. Agent Compliance & CI/CD Security Gate

- **Problema:** Los agentes de código alucinan APIs y rompen contratos de software en producción.
- **Solución:** `ontoprune check` como paso automatizado en GitHub Actions / GitLab CI que bloquea pull requests con llamadas no autorizadas al grafo del sistema.
- **Monetización:** Modelo SaaS de auditoría y seguridad para código generado por agentes.
vigmarcarlo10
🟧 echo.github ⭐Repo README: 'Neuro-Symbolic Context Pruning Middleware for Local SLMs' translating source into RDF/SPARQL contract stubs; self-benchmark onvigmarcarlo——

Interpretation history

Decision trace