Retrieved article excerpt
Open article ยท Retrieved 2026-09-23T14:28:00.679948+00:00
# Blink
**A one-pass typed-decision model with an embeddable C runtime that also
builds to WebAssembly.**
---
TypeSafe released [Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev),
a closed model in a class they call *System One Models*: fast structured
decisions for software rather than conversation, with calibrated confidence and
no output tokens to pay for. Blink puts that interface behind a C library you
can link into a service, a daemon or a device, or load as WebAssembly in a web
page.
---
Blink answers questions of the form *"given this state and this criterion,
which of these options?"* in a single forward pass. There is no token
generation, no decoding loop, no JSON to parse and nothing to repair: the
output is one probability per option you declared at call time.
The runtime is C99 with no dependencies beyond libc and libm. Weights are
mapped read-only and never copied. Scoring performs **zero allocations** โ a
property the test suite checks mechanically, not by inspection.
The same sources build to a 66 KB WebAssembly module that runs unchanged in
browsers and in Node, with no file system and no server
([details](https://github.com/sqliteai/blink#webassembly)).
```
#include "blink.h"
blink_model *model = blink_model_open_file("blink-tiny.blink", 1, &status);
blink_session *session = blink_session_init(arena, sizeof arena, model, &limits, &status);
blink_state_set(session, ticket, strlen(ticket)); /* encode once */
blink_menu_set(session, queues, lengths, 4); /* encode once */
blink_score(session, "Which team should handle this?", 29, /* cheap */
probabilities, &result);
```
**What it is good for, in one paragraph.** Blink recognises the *form* of a
decision; it does not read text the way a pretrained language model does.
Where the answer is carried by form โ which queue a ticket's wording points
at, whether a claim's verb agrees with the state โ blink-tiny is near perfect
and well calibrated. Where the answer requires reading โ natural language
inference, binding a name to the right sentence โ it is at or a little above
chance, and a frozen 4B model is far ahead. Its confidence is calibrated
enough to branch on: answer when it is sure, escalate when it is not.
## Performance
blink-tiny on one core of an Apple M5 Pro: one decision over a 256-byte state,
a question and four options, p50.
| build | fresh decision | same state, new question | decisions/s, state reused |
| --- | --- | --- | --- |
| default (C99, NEON) | 406 ยตs | 54 ยตs | 18,211 |
| [W8A8](https://github.com/sqliteai/blink#w8a8-backend-sdot--i8mm) (int8 activations, SDOT) | 221 ยตs | 33 ยตs | 29,481 |
| [Accelerate](https://github.com/sqliteai/blink#accelerate-backend-macos) (macOS) | 172 ยตs | 32 ยตs | 31,290 |
A fresh decision encodes the state, the options and the question. Software
usually asks several questions of the same state, and then only the question
is encoded: the cached state is reused exactly, bit for bit. On x86-64 the
AVX2 kernel does the same decision in 760 ยตs, and 90 ยตs with the state reused,
on a GitHub-hosted runner.
**Memory. Nothing is allocated while scoring: not per decision, not per
question, not per option.** All the working memory a session needs is one
arena, sized before the session exists and placed wherever you choose, a static
array or a stack buffer included, so Blink runs where `malloc` is unavailable
or not allowed. This is checked, not assumed: `tests/c/check_no_malloc.sh`
fails the build if the scoring code references any allocator.
blink-tiny's weights are 452 KiB of int8, mapped read-only and shared by every
session and every process that opens the same file. The arena is 260 KiB with
the model's default limits and 22 KiB with the smallest. One decision from the
command line peaks at 2.5 MiB of resident memory, the whole process included.
The W8A8 build keeps the no-allocation guarantee; the Accelerate build gives it
up, because the framework allocates inside its matrix products
([details](https://github.com/sqliteai/blink#memory)).
A larger preset, blink-small (7.9M parameters), is **experimental**: it is
worse than blink-tiny on every corpus measured and about ten times slower to
encode, so its containers are not shipped
([why](https://github.com/sqliteai/blink/blob/main/docs/RESULTS.md#5-what-does-not-work)).
All the benchmark tables, and a comparison with other systems that expose the
same interface, are in [docs/RESULTS.md](https://github.com/sqliteai/blink/blob/main/docs/RESULTS.md#comparison-with-other-systems).
---
## Quick start
```
git clone https://github.com/sqliteai/blink
cd blink
make # C runtime, tools, benchmarks. No Python needed.
make test # unit tests + the no-allocation check
make bench # latency and memory as JSON
make install # blink, libblink, blink.h, blink.pc, man page under /usr/local
make ACCELERATE=1 test # macOS only: the Accelerate backend, in build-accelerate/
make W8A8=1 test # int8 activations with SDOT/I8MM, in build-w8a8/
```
Training and evaluation need Python:
```
uv venv --python 3.12 .venv
uv pip install --python .venv/bin/python torch numpy pytest
export PYTHONPATH=python
.venv/bin/python scripts/make_data.py # deterministic corpus
.venv/bin/python -m blink_train.train \
data/synthetic/train.jsonl \
--validation data/synthetic/validation.jsonl \
--preset tiny --output artifacts/blink-tiny.blink \
--epochs 60 --learning-rate 3e-3 --device auto
.venv/bin/python eval/evaluate.py \
artifacts/blink-tiny.blink data/synthetic/test.jsonl
```
One decision from the shell, with the committed blink-tiny (no training
needed):
```
./build/blink artifacts/blink-tiny-synthetic-s7.blink \
--state "The parcel left the depot on Monday and has not arrived." \
--question "Which team should handle this ticket?" \
--option "delivery and logistics" \
--option "billing and payments" \
--option "account access and sign-in"
```
It answers *delivery and logistics* with probability 0.56: the right queue, with
a confidence low enough that a caller branching at 0.8 would escalate it.
`./build/blink --version` prints the library version and the numeric
backend it runs on this CPU, for example `blink 0.1.0 (neon)`. Every option,
the output fields and the exit codes are in [docs/CLI.md](https://github.com/sqliteai/blink/blob/main/docs/CLI.md), also
installed as the `blink(1)` man page.
`make install` puts `blink`, `libblink.a`, the shared library, `blink.h`, a
`blink.pc` for pkg-config and the man page under `/usr/local`. Use
`PREFIX=...` to choose another place and `DESTDIR=...` to stage a package;
`make uninstall` removes them again. The shared library is named after the ABI
version, `libblink.so.1` on Linux and `libblink.1.dylib` on macOS.
To reproduce everything โ build, tests, data, training, accuracy, controls,
perturbations, benchmarks, and a manifest of every checksum:
```
bash scripts/run_all.sh # the published protocol, about 2 h on an M-series laptop
bash scripts/run_all.sh --smoke # every stage on a small corpus, a few minutes
```
---
## The architecture in one diagram
```
state bytes question bytes option bytes
โ โ โ
unigram + hashed bigram + position + segment embedding (each separately)
โ โ โ
stem: depthwise conv + pointwise projection, full resolution
โ โ โ
mean-pool by stride 4, then blocks: conv + global mean + FFN, self-attention mixer
โ โ โ
state K, V โโโโโ cross โโโโ question โ
(cached) (the question attends to the state) โ
โ โ โ
โ question summary โโโโ FiLM โโโบ option vectors
โ โ
โโโโโโโโโโโบ multi-head cosine option attention โโโโโ
โ
ฮฃ head cosines ร scale / temperature โ softmax over the options
```
Four decisions carry most of the design, and each is explained with its
rationale in [docs/ARCHITECTURE.md](https://github.com/sqliteai/blink/blob/main/docs/ARCHITECTURE.md):
1. **Byte level, no tokenizer.** Program state is not curated text. Bytes
accept JSON, identifiers, any language and embedded NULs, and there is no
second artefact to keep in sync.
2. **Stem, then pool, then blocks.** Running a feed-forward at every byte is
what makes byte-level encoders expensive. Blink runs one cheap
full-resolution stem, pools by the stride, and only then pays for the
blocks. A 512-byte state is 64 positions before any block executes.
3. **State and question are encoded separately.** Reusing an encoded state is
therefore *exact*, not approximate: `tests/c/test_session.c` checks it with
`memcmp`.
4. **Queries are conditioned on the question.** An option query built from
the option text alone lets the question reach the score by one indirect
route: the keys and values it adds to the shared context. Blink additionally
modulates each option vector with a summary of the question, which gives the
two a direct place to meet. A three-run ablation finds **no demonstrated
effect** at this scale โ between-seed spread exceeds the gap between the
arms โ so this is a motivated design choice, not a measured win. The table
is in [docs/RESULTS.md](https://github.com/sqliteai/blink/blob/main/docs/RESULTS.md#the-film-ablation-still-undecided).
---
## Memory
**Blink allocates no memory while scoring.** Its memory comes in three parts,
and the third is zero:
- **Weights** โ mapped read-only, never copied, shared by every process and
every session.
- **Session arena** โ one contiguous block whose size is a pure function of the
model and the limits you declare. `blink_session_size` tells you the number
and `blink_session_init` places the session in a buffer you own: a static
array, a stack frame, a pool.
- **Scoring** โ **no allocation at all.** `tests/c/check_no_malloc.sh` asserts
that `blink_runtime.o` and `blink_kernels.o` reference no allocator symbol.
`examples/embed_static.c` is the whole pattern in 100 lines with no heap.
These guarantees describe the default build. The W8A8 backend keeps all of
them except exact agreement with the W8A32 definition. The Accelerate backend
trades two of them for speed.
### W8A8 backend (SDOT / I8MM)
`make W8A8=1` builds into `build-w8a8/`, with the activations quantized to
int8 too. Each input row of a projection gets its own scale and the product
becomes an exact integer sum. It runs on SDOT. I8MM is chosen at run time on
non-Apple cores that have it; on Apple silicon it measured level with SDOT
and would need an extra copy of the weights, so SDOT stays the default there
(`BLINK_W8A8_KERNEL=scalar|sdot|i8mm` overrides the choice). On
a CPU without either, a plain integer loop computes the same thing.
| 4 options, p50 | default | W8A8 | speed-up |
| --- | --- | --- | --- |
| blink-tiny, fresh decision, 256-byte state | 406 ยตs | 221 ยตs | 1.8ร |
| blink-tiny, state reused | 54 ยตs | 33 ยตs | 1.6ร |
| blink-small, fresh decision, 512-byte state | 4.29 ms | 1.45 ms | 3.0ร |
| blink-small, state reused | 314 ยตs | 115 ยตs | 2.7ร |
What it keeps: nothing allocates while scoring, a cached state is reused
exactly, a batch is bit-identical to single questions, and the scalar, SDOT
and I8MM kernels are bit-identical to each other. What it changes is the
model, since activations are rounded to 8 bits. On all 15 published models,
top-1 moved by โ0.15 to +0.42 points and calibration did not change. The
decisions that flip are near-ties, almost all in task families the model sits
at chance on