2026-10-11 16:37 UTC

Interfaze AI claims its released Apache-2.0 Interfaze-1 Lite โ€” one vision-language reasoning core routing document, speech, and detection specialist architectures on a single 80GB GPU โ€” becomes an adopted unified local multimodal model for developer and agent workloads; sustained third-party adoption confirms it, a quiet post-launch fade closes it.

state: seedheat: lowuncertainty: highconvergesscott: mediumopen-model-releases local-inference mixture-of-architectures multimodalInterfaze AI

What is this?

Interfaze AI (YC-backed seed company, formerly JigsawStack) released Interfaze-1 Lite on Hugging Face under Apache-2.0; the model card describes a 'mixture-of-architectures' design where a single vision-language reasoning core works alongside specialist architectures for document, speech, and detection tasks on one 80GB GPU. The company's arXiv paper (accepted IEEE CAI 2026) argues for fusing task-specific CNN/DNN encoders directly into the transformer decoder โ€” explicitly claiming to avoid routing and tool-calling โ€” targeting deterministic developer tasks (OCR, extraction, scraping, classification), with first-party benchmarks claiming wins over flash-tier frontier models on OCRBench V2, olmOCR, and RefCOCO. The web record otherwise shows low external chatter (a ~7.5k-view YouTube demo, some LinkedIn posts) and no third-party adoption signals, so the case's adoption question is open. One wrinkle: the case and model card use 'routing specialists' language while the paper insists it avoids routing via encoder fusion โ€” the supplied material doesn't fully reconcile the two framings.

Why it matters to Scott

Converges with the gathering/judging split in his Model Barbell and deterministic-first extraction doctrine: Interfaze's paper independently argues that deterministic developer tasks (OCR, extraction, classification) shouldn't pay generalist-orchestration costs, and its explicit anti-routing/anti-tool-calling encoder-fusion stance is a live counterpoint to the routing-heavy pattern across Scott's code-first architecture work and the radar's router cluster โ€” the case tests where routing earns its keep vs where fusion wins. But the flash-tier-beating benchmarks (OCRBench V2, olmOCR, RefCOCO) are first-party with no third-party adoption yet, and the single-80GB-GPU target sits above his consumer gamepc/Mac-mini substrate, so this is a test-and-watch candidate for his OCR/ingestion pipelines (video spike, pytesseract/ocrmac stack) rather than something that changes what he builds today.
ip:concept.model-barbellip:framework.code-first-architecturedev:project.gamepcdev:concept.hardware-aware-local-inferencedev:project.videoradar:concept.open-model-releaseradar:concept.multimodal-modelsradar:concept.document-parsingradar:concept.tool-routing
queries asked of Scott's wikis
  • specialist small models vs generalist LLM structured extraction reliability
  • OCR document ingestion pipeline RAG knowledge base
  • local inference single-GPU multimodal open weights
  • agent tool calling vs fused perception routing harness
  • structured JSON output hallucination production agent workloads
  • open-weights strategy small lab adoption

Measured heat

now 0 pts/hpeak 2 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 135h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

10-06 03:44 (minted)โญ origin echo-reconstructed"Interfaze 1 Lite is a mixture-of-architectures (MoA) model for developer workloads... A single vision-language reasoning core works alongsi
Interfaze AI on github (echo) ยท attributed from reddit.post.1wyp0y6 ยท published time unknown
โ€”
10-06 00:36first on r/LocalLLaMA ยท published ยท lag ?interfaze-ai/interfaze-1-lite ยท Hugging Face
zmarty
โ€”
10-08 07:22first on hacker news ยท published ยท lag ?Interfaze (YC P26) redistribute openRAIL-M model under Apache 2.0
pierre
โ€”
10-06 00:36amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1wyp0y6
zmarty
peak 17 ยท 2 comments ยท 92% of case engagement
10-08 07:22amplified on hacker newshn.story.50002786
pierre
peak 1 ยท 0 comments ยท 9% of case engagement
10-06 03:20our radar first saw it ยท lag ?discovery anchor: reddit.post.1wyp0y6โ€”
pace: p54 vs 1247 stories at the 96h mark (now 135h old) โ€” ahead of agent-iap-credential-brokering (1.1x), behind 37signals-agent-driven-default (0.9x)

Evidence (3) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  redditinterfaze-ai/interfaze-1-lite ยท Hugging Face
LocalLLaMA
Retrieved article excerpt

Open article ยท Retrieved 2026-10-06T03:34:03.517123+00:00

# [interfaze-ai](https://huggingface.co/interfaze-ai) / [interfaze-1-lite](https://huggingface.co/interfaze-ai/interfaze-1-lite) Like 28 Follow Interfaze 49

[Image-Text-to-Image](https://huggingface.co/models?pipeline_tag=image-text-to-image)[Transformers](https://huggingface.co/models?library=transformers)[Safetensors](https://huggingface.co/models?library=safetensors)

15 languages

[interfaze\_lite](https://huggingface.co/models?other=interfaze_lite)[feature-extraction](https://huggingface.co/models?other=feature-extraction)[vlm](https://huggingface.co/models?other=vlm)[multimodal](https://huggingface.co/models?other=multimodal)[multilingual](https://huggingface.co/models?other=multilingual)[mixture-of-architectures](https://huggingface.co/models?other=mixture-of-architectures)[ocr](https://huggingface.co/models?other=ocr)[document-understanding](https://huggingface.co/models?other=document-understanding)[speech-recognition](https://huggingface.co/models?other=speech-recognition)[speaker-diarization](https://huggingface.co/models?other=speaker-diarization)[object-detection](https://huggingface.co/models?other=object-detection)[gui-grounding](https://huggingface.co/models?other=gui-grounding)[image-segmentation](https://huggingface.co/models?other=image-segmentation)[translation](https://huggingface.co/models?other=translation)[time-series-forecasting](https://huggingface.co/models?other=time-series-forecasting)[guardrails](https://huggingface.co/models?other=guardrails)[structured-output](https://huggingface.co/models?other=structured-output)[agent](https://huggingface.co/models?other=agent)[custom\_code](https://huggingface.co/models?other=custom_code)

License: apache-2.0

[Model card](https://huggingface.co/interfaze-ai/interfaze-1-lite)  [Files Files and versions  

xet](https://huggingface.co/interfaze-ai/interfaze-1-lite/tree/main)  [Community](https://huggingface.co/interfaze-ai/interfaze-1-lite/discussions)

 

Deploy

  Copy to bucket new   

Use this model

 

[Interfaze 1 Lite](https://huggingface.co/interfaze-ai/interfaze-1-lite/resolve/main/assets/interfaze-1-lite-banner.png)

# Interfaze 1 Lite

[Website](https://interfaze.ai) ยท [Docs](https://interfaze.ai/docs/models/interfaze-1-lite) ยท [Run tasks](https://interfaze.ai/docs/run-tasks) ยท [Blog](https://interfaze.ai/blog/the-first-open-weight-model-for-deterministic-work-interfaze-1-lite) ยท [GitHub](https://github.com/InterfazeAI/interfaze-1-lite)

## Introduction

Interfaze 1 Lite is a mixture-of-architectures (MoA) model for developer workloads: reading documents, transcribing speech, locating objects and interface elements, and answering over all of it with structured output.

A single vision-language reasoning core works alongside a set of specialist architectures, each built for one kind of perception. The core reads the request, decides which specialists to run, and composes the answer from what they return. The whole model runs on one 80 GB GPU with no external services.

### Key features

- **Document understanding.** Text, reading order, tables and layout from images, PDFs (up to 50 pages a call) and Word files, with a box and confidence for every line and word.
- **Speech.** Transcription with timestamps and speaker diarization. Long recordings are cut at pauses and decoded in batches: a 95-minute recording transcribes in about 90 seconds.
- **Visual grounding.** Open-vocabulary object detection with outlines, and GUI element grounding for computer-use agents.
- **Structured output.** Responses constrained to a JSON schema you supply, reading from any mix of text, images, documents and audio.
- **Translation, forecasting and guardrails.** Translation across 160+ languages, time-series forecasting from CSV or JSON, and safety checks on text and images.
- **Multilingual reasoning.** Science, math, SQL and general knowledge across 14+ languages, with a 131k-token context.
- **Self-contained.** One repository, one GPU, runs offline.

### Model architecture

Interfaze 1 Lite is not one network. It is a reasoning core plus specialists, each chosen for the task it is best at, connected by tool calls.

| Component | Architecture | Role |
| --- | --- | --- |
| Reasoning core | Hybrid-attention decoder with a vision encoder, FP8, 131k context | Plans, calls specialists, grounds boxes on a 0โ€“1000 grid, writes the answer |
| Document reader | Vision-language model trained for page reading | Text, reading order, tables and markdown |
| Line geometry | Text detector and recognizer | Every line's box and confidence |
| Layout | Document layout detector | Titles, paragraphs, tables and figures, with boxes |
| Speech | Encoder-decoder speech recognizer | Transcripts and timestamps in 99 languages |
| Diarization | Speaker segmentation and embedding pipeline | Who spoke when |
| Segmentation | Promptable segmentation model | Object outlines and masks |
| Forecasting | Time-series foundation model | Future values of a numeric series |
| Guardrails | Safety classifier | 14 text safety categories |

How the parts combine:

- **OCR is two views of one page, stitched.** The document reader supplies the text, and the line detector supplies the geometry. Each detected line takes the reader's words for it, so boxes are exact and text is complete.
- **Speakers are attributed per word,** by the largest overlap with each speaker's turns, then grouped into chunks.
- **Detection and GUI grounding run on the reasoning core,** which returns boxes on a 0โ€“1000 grid. Outlines come from the segmentation model.
- **A run task skips planning.** Naming one capability (`task="ocr"`, `"speech_to_text"`, โ€ฆ) runs that specialist directly and returns its raw result.

## Performance

| Benchmark | What it measures | **Interfaze 1 Lite** | Interfaze | GPT-5.4-Mini | Claude-Sonnet-4.6 | Gemini-3-Flash | Grok-4.3 |
| --- | --- | --- | --- | --- | --- | --- | --- |
| GPQA Diamond | Graduate-level science | **85.9** | 92.4 | 82.8 | 89.9 | 88.5 | 73.6 |
| MMMLU | Knowledge in 14 languages | **87.8** | 90.9 | 75.3 | 84.9 | 88.7 | 89.7 |
| MMMU-Pro | Multimodal reasoning | **73.2** | 71.1 | 40.4 | 46.3 | 67.6 | 68.7 |
| olmOCR-Bench | Document OCR | **83.8** | 85.7 | 80.1 | 73.9 | 75.3 | 81.9 |
| OCRBench v2 (English) | Text in images | **60.9** | 70.7 | 52.7 | 54.7 | 55.8 | 54.7 |
| RefCOCO ([email protected]) | Referring-expression grounding | **83.8** | 82.1 | โ€“ | โ€“ | โ€“ | โ€“ |
| VoxPopuli-Cleaned (WER โ†“) | Speech recognition | **3.01** | 2.4 | โ€“ | โ€“ | 4.0 | โ€“ |
| SOB (value accuracy) | Structured output from text, images and audio | **81.5** | 80.5 | โ€“ | 77.9 | 77.3\* | โ€“ |
| Spider 2.0-Lite (SQLite) | Text-to-SQL | **48.9** | 52.9 | 26.7 | 49.6 | 45.2 | 45.9 |

Interfaze 1 Lite was scored by us with each benchmark's official scorer. Every other score is from the [Interfaze leaderboard](https://interfaze.ai/leaderboards). Higher is better except WER. \*Gemini-3-Flash-Preview.

### The benchmarks

- **GPQA Diamond** (all 198 questions). Graduate-level physics, chemistry and biology multiple choice, written to resist search. Lite scores 85.9, ahead of GPT-5.4-Mini and Grok-4.3, and strongest in physics.
- **MMMLU** (MMMLU-lite, all 19,950: 1,425 questions in each of 14 languages). MMLU translated by professional translators. Lite averages 87.8, ahead of Claude-Sonnet-4.6. The low-resource languages (Swahili, Yoruba, Bengali) are where it loses most.
- **MMMU-Pro** (all 1,730 questions per track, mean of standard and vision tracks). College-level questions that need the image, including a track where the question itself is inside the picture. Lite leads the board at 73.2.
- **olmOCR-Bench** (all 1,403 PDFs). Unit tests on real documents: arXiv math, old scans, tables, headers and footers, multi-column pages and long tiny text. Lite scores 83.8, with 91.9 on long tiny text and 88.8 on tables.
- **OCRBench v2, English** (all 7,400 English items). Recognition, referring, spotting, extraction, parsing, calculation, understanding and reasoning over text in images. Lite scores 60.9; its text spotting leads every general-purpose model on the board.
- **RefCOCO** ([email protected]). Find the one object a sentence describes ("the man in red on the left"). Lite's answer box scores 83.8, first on the board.
- **VoxPopuli-Cleaned** (all 628 clips). European Parliament speech, scored by word error rate after the benchmark's standard text normalisation. Lite's WER is 3.01%.
- **SOB, the Structured Output Benchmark** (all 5,324 records). Extract values into a JSON schema from text, images and audio; value accuracy counts exact field matches. Lite scores 81.5, second of 30 models, with 97% of responses valid JSON.
- **Spider 2.0-Lite** (the 135 SQLite tasks). Enterprise text-to-SQL over real schemas, scored by executing the query. Lite solves 48.9%, between the Claude models and Gemini-3-Flash.

## Quickstart

To run Lite as an OpenAI-compatible server with Docker, see [GitHub](https://github.com/InterfazeAI/interfaze-1-lite).

### Requirements

- One 80 GB GPU with compute capability 8.9 or newer (Hopper, Ada). Tested on an H100.
- `ffmpeg` on the system, to decode audio.

```
hf download interfaze-ai/interfaze-1-lite requirements.txt --local-dir .
pip install -r requirements.txt
```

`flash-linear-attention` matters: without it, transformers runs the linear-attention layers as a plain PyTorch loop, and generation slows to minutes per page.

### Using ๐Ÿค— Transformers

```
from transformers import AutoModel

model = AutoModel.from_pretrained("interfaze-ai/interfaze-1-lite", trust_remote_code=True)

answer = model.chat(
    [{"role": "user", "content": "What is the total, and which item is highlighted?"}],
    files=["receipt.jpg"],
)
print(answer["content"])      # the answer
print(answer["precontext"])   # each specialist's full result: [{"name", "result"}]
```

`trust_remote_code=True` is required: the model's routing and specialists are defined in this repository. Components load on first use, so a caller that only runs OCR never loads the others.

Each capability is also a method of its own:

| Method | Returns |
| --- | --- |
| `chat(messages, files=[...])` | `content` (the answer) and `precontext[]` (each specialist's result) |
| `ocr(source, page_range=None, return_markdown=False)` | `text`, `sections[]` (one per page, with `lines[].words[]`, four-corner `bounds`, `average_confidence`), `width`, `height` |
| `transcribe(audio, by_speaker=False, language="auto", word_timestamps=False)` | `text` and `chunks[]` with timestamps; each chunk carries a speaker with `by_speaker` |
| `detect(image, prompts, return_masks=False)` | `detected_objects[]` with `label`, `bounds` and `polygon` (and `mask` when asked) |
| `ground(image, prompts=None)` | `gui_elements[]` with `type` and `bounds`; with no prompts, every interactive element |
| `forecast(series, horizon)` | the next `horizon` points of a `{date: value}` series: `timestamp[]` and `value[]` |
| `moderate(text)` | `output`: `"safe"`, or `"unsafe"` and the violated categories (S1โ€“S14) |

Coordinates are pixels of the input: an image's own size, or a PDF page at 144 DPI.

### Using the Interfaze API

The same model is served behind an OpenAI-compatible API. Get your API key from the [Interfaze dashboard](https://interfaze.ai/dashboard/quick-start), then set `model` to `interfaze-1-lite`:

```
import { Interfaze } from "interfaze";

const interfaze = new Interfaze(); // reads INTERFAZE_API_KEY

const res = await interfaze.chat.completions.create({
  model: "interfaze-1-lite",
  messages: [{ role: "user", content: "Summarise the attached contract in three bullets." }],
});
```

## Examples

### Read a document, with boxes

```
doc = model.ocr("invoice.pdf", page_range=[1, 2])

print(doc["text"])
for page in doc["sections"]:
    for line in page["lines"]:
        box = line["bounds"]
        print(page["page"], line["text"], box["top_left"], box["bottom_right"])
```

### Transcribe a call and split it by speaker

```
call = model.transc
zmarty172
๐ŸŸง echo.github โญ"Interfaze 1 Lite is a mixture-of-architectures (MoA) model for developer workloads... A single vision-language reasoning core works alongsiInterfaze AIโ€”โ€”
๐ŸŸง hnInterfaze (YC P26) redistribute openRAIL-M model under Apache 2.0pierre10

Interpretation history

Decision trace