2026-10-11 18:01 UTC

General Instinct claims its released InstinctFlash runtime accelerates robotics-model inference on Jetson Thor by roughly 1.2–7.9 times at matched schedules, with larger gains from reduced sampling steps, potentially enabling responsive edge control without materially degrading task success.

state: watchingheat: lowuncertainty: mediumnovelscott: mediumrobotics-inference edge-ai inference-optimizationGeneral InstinctGuanming

What is this?

General Instinct is a YC P26 startup (founded by Guanming/'Bill') building inference tooling for frontier models on edge devices; in September 2026 it open-sourced InstinctFlash (AGPL-3.0), a serving runtime that runs VLA/world-action robotics policies in real time on NVIDIA's Jetson Thor and consumer RTX 4090/5090, covering eight model families behind an openpi-compatible API with checkpoint-driven serving. Its own blog claims LingBot-VA median latency dropping from 15.51s in native PyTorch to 459ms on Thor — the case's caveat stands: the headline figure requires a distilled sampling schedule, while runtime-only gains at matched schedules are the 1.2–7.9x range — and a Reuters-reported $1B raise (late Sept 2026) appears in the case's evidence but is not independently visible in the current snippets. The surrounding ecosystem is real and moving: Jetson Thor (launched Aug 2025) is explicitly positioned by NVIDIA for real-time VLA/physical-AI inference at the edge, with Agility Robotics, Connect Tech, and others building on it. All performance figures remain vendor-run; the supplied material shows no independent reproduction, third-party adoption, or external verification of quality retention under FP8/few-step schedules on real robots.

Why it matters to Scott

Extends rather than repeats dev:concept.hardware-aware-local-inference: InstinctFlash's distinctive move is putting the sampling schedule itself under runtime policy — the honest split between 1.2–7.9x runtime-only gains and the distilled-schedule 33.78x headline adds an axis (schedule-as-policy) his canon doesn't yet carry, and applies the radar's distillation lever to VLA action chunks rather than tokens or video. The Reuters-reported $1B raise strengthens the actor but not the evidence — every benchmark remains vendor-run — so this stays medium rather than high: it would update his runtime-policy concept and is a live benchmark-integrity specimen, but touches no hardware he actually serves (gamepc/Ollama/MLX) and the load-bearing open question (task-success retention under FP8/few-step schedules on real robots) still awaits third-party reproduction.
dev:concept.hardware-aware-local-inferenceradar:concept.edge-inferenceradar:concept.inference-optimizationradar:concept.distillationradar:concept.benchmark-integrity
queries asked of Scott's wikis
  • hardware-aware local inference optimization
  • FP8 quantization quality tradeoff local models
  • sampling schedule distillation latency speedup
  • edge on-device inference vs cloud latency budget
  • vendor benchmark independent reproduction receipts
  • real-time agent loop latency requirements

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 457h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-22 23:24 (minted)⭐ origin echo-reconstructedFull-source robotics serving runtime with reproduction recipes; the 33.78-fold headline includes changing LingBot-VA's sampling schedule, un
General Instinct on github (echo) · attributed from hn.story.49802789 · published time unknown
—
09-22 15:20first on hacker news · published · lag ?Show HN: InstinctFlash – Run 5B world-action models in real time on Jetson Thor
guanming0717
—
09-22 15:20amplified on hacker news 👑hn.story.49802789
guanming0717
peak 27 · 4 comments · 97% of case engagement
09-28 15:45amplified on hacker newshn.story.49879872
doppp
peak 1 · 0 comments · 3% of case engagement
09-22 17:22our radar first saw it · lag ?discovery anchor: hn.story.49802789—
pace: p58 vs 1032 stories at the 336h mark (now 457h old) — ahead of chatgpt-word-integration (1.0x), behind opencontext-project-local-agent-memory (1.0x)

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: InstinctFlash – Run 5B world-action models in real time on Jetson Thor
Retrieved article excerpt

Open article · Retrieved 2026-09-22T17:28:47.857607+00:00

[InstinctFlash](https://github.com/General-Instinct/InstinctFlash/blob/main/assets/iFlash.png)

**A high-performance serving framework for robotics models.**

[License](https://www.gnu.org/licenses/agpl-3.0)
[Website](https://general-instinct.com/)
[YC](https://www.ycombinator.com/companies/general-instinct)

---

## What's new 🔥

- **[2026/09/17] RTX 5090 support.** Deploy on your workstation with the same Runtime API used on Jetson Thor. [Setup](https://github.com/General-Instinct/InstinctFlash/blob/main/INSTALL.rst#rtx-5090) · [Reproduce](https://github.com/General-Instinct/InstinctFlash/blob/main/REPRODUCE.rst#rtx-5090).
- **[2026/09/16] RTX 4090 support.** Desktop inference and WebSocket serving with dedicated installation profiles. [Setup](https://github.com/General-Instinct/InstinctFlash/blob/main/INSTALL.rst#rtx-4090) · [Reproduce](https://github.com/General-Instinct/InstinctFlash/blob/main/REPRODUCE.rst#rtx-4090).
- **[2026/09/15] Full-source release.** Eight robotics model families, acceleration kernels, and Python / WebSocket serving through one Runtime. [Get started](https://github.com/General-Instinct/InstinctFlash#install).
- **[2026/09/15] Jetson Thor benchmarks.** Up to **33.78×** speedup with LingBot-VA @2V/4A, using FP8 and fewer sampling steps. [Results](https://github.com/General-Instinct/InstinctFlash#results) · [Reproduce](https://github.com/General-Instinct/InstinctFlash/blob/main/REPRODUCE.rst).

## Results

Prediction p50 on **Jetson Thor** (ms), measured September 15, 2026.

We’ve seen up to **33.78× speedup** with no observed loss in task performance in our real-robot tests.

| Model | Acceleration line | PyTorch | InstinctFlash | Speedup |
| --- | --- | --- | --- | --- |
| LingBot-VA | FP8 · 25V/50A | 15506.32 | **2891.74** | **5.36×** |
| ↳ LingBot-VA | FP8 · 2V/4A | 2071.29 | **459.10** | **4.51×** |
| LingBot-VLA-4B | FP8 | 624.22 | **221.53** | **2.82×** |
| LingBot-VLA-V2-6B | FP8 | 734.56 | **394.11** | **1.86×** |
| Cosmos3 Edge DROID | NUMERIC · UniPC4 / CFG3 | 3393.78 | **1048.01** | **3.24×** |
| Cosmos3 Nano DROID | NUMERIC · UniPC4 / CFG3 | 10184.68 | **4772.38** | **2.13×** |
| pi05 | FP8 | 408.58 | **51.85** | **7.88×** |
| GR00T N1.7 | BITEXACT | 139.50 | **117.30** | **1.19×** |
| DreamZero DROID | FP8 · 16 steps · dynamic cache | 23563.08 | **11899.42** | **1.98×** |

VA measures early continuations; each row compares the same schedule.
The 33.78× headline includes 25V/50A → 2V/4A.
FP8 and sampling changes are optional.

[Protocol and raw results](https://github.com/General-Instinct/InstinctFlash/blob/main/eval/public_release_2026-09-15/results.rst) · [Native VA 2V/4A](https://github.com/General-Instinct/InstinctFlash/blob/main/eval/va_native_2v4a_2026-09-15/README.rst) · [Reproduction commands](https://github.com/General-Instinct/InstinctFlash/blob/main/REPRODUCE.rst)

## Install

```
git clone https://github.com/General-Instinct/InstinctFlash && cd InstinctFlash
python3 -m venv .venv-core
source .venv-core/bin/activate
python -m pip install . uv==0.12.5
```

The Python 3.10+ core inspects checkpoints and plans without PyTorch or a GPU.
Inference uses a separate, pinned environment for each model family. For RTX 4090:

```
python3 scripts/bootstrap_vendor.py install pi05 --target rtx4090 \
  --python python3.12 --root ~/ifl-pi05-4090 --ptxas /usr/local/cuda/bin/ptxas
source ~/ifl-pi05-4090/activate.sh
```

Use `va`, `vla4`, `vla2`, `pi05`, `groot`, `edge`, `nano` or `dreamzero`.
Edge and Nano use Python 3.13; the other families use Python 3.12.
The bootstrap installs the upstream source, compatibility patches, core and adapter.
Model weights are downloaded separately. See [RTX 5090 setup](https://github.com/General-Instinct/InstinctFlash/blob/main/INSTALL.rst#rtx-5090),
[RTX 4090 setup](https://github.com/General-Instinct/InstinctFlash/blob/main/INSTALL.rst#rtx-4090)
or [Jetson Thor setup](https://github.com/General-Instinct/InstinctFlash/blob/main/INSTALL.rst#jetson-thor), which selects `--target jetson_thor`
and uses the [Thor CUDA backend build](https://github.com/General-Instinct/InstinctFlash/blob/main/serving/README.rst).

## Load a model

**Your fine-tuned checkpoint** — the expected case. Point `serve` at the training output; it
detects the family, writes the small `instinctflash.json` declaration from what the checkpoint
itself proves, and starts serving. One command:

```
instinctflash serve /path/to/your/checkpoint
```

Anything the checkpoint cannot prove is asked for explicitly, never guessed. Once the
declaration exists (serve writes it on first run), the same directory also loads in Python:

```
from instinctflash import Runtime

runtime = Runtime.from_pretrained("/path/to/your/checkpoint")
```

**A stock release** — use its Hub id after installing the family's environment:

```
runtime = Runtime.from_pretrained("robbyant/lingbot-va-posttrain-robotwin")
```

| family | model id |
| --- | --- |
| LingBot-VA (5B WAM) | `robbyant/lingbot-va-posttrain-robotwin` |
| LingBot-VLA-4B | `robbyant/lingbot-vla-4b-posttrain-robotwin` |
| LingBot-VLA-V2-6B | `robbyant/lingbot-vla-v2-6b-robotwin` |
| pi0.5 | `lerobot/pi05_base` · `lerobot/pi05_libero_finetuned_v044` |
| GR00T-N1.7-3B | `nvidia/GR00T-N1.7-3B` |
| Cosmos3 policies | `nvidia/Cosmos3-Edge-Policy-DROID` · `nvidia/Cosmos3-Nano-Policy-DROID` |
| DreamZero | `GEAR-Dreams/DreamZero-DROID` |

Fine-tunes reuse their family's adapter; quality is evaluated per checkpoint.

The same `Runtime` defaults to `precision="native"` with a BITEXACT transformation ceiling.
Use `tier_ceiling="numeric"` to allow numerical changes, or `precision="fp8"`
(CLI: `--fp8`) to explicitly enable FP8. Step schedules are selected separately.
See [precision policy](https://github.com/General-Instinct/InstinctFlash/blob/main/INSTALL.rst#load-and-predict) and
[FP8 support and validation](https://github.com/General-Instinct/InstinctFlash/blob/main/eval/thor_precision_completion_2026-09-09/README.md).

DreamZero's opt-in [dynamic step cache](https://github.com/General-Instinct/InstinctFlash/blob/main/INSTALL.rst#load-and-predict) requires
`tier_ceiling="behavioral"` with either precision. See the
[Thor measurements](https://github.com/General-Instinct/InstinctFlash/blob/main/eval/dynamic_step_cache_integration_2026-09-14/results.md).

## Get actions

**In process** — this is the whole Python API:

```
with runtime.episode(prompt="put the bottle in the dustbin") as episode:
    while not done:
        result = episode.predict(observation)
        action = result["action"]
```

`observation` is a dict in the model's own format; `result["action"]` contains its action array.
For LingBot-VA, pass `executed_action=...` when the controller changes a predicted action
chunk, so the next prediction uses the actions actually executed.

**Over the network** — the `serve` command above hosts the same runtime behind the
msgpack-over-websocket wire protocol the pi0/openpi ecosystem already speaks, so existing
robot-side clients connect unchanged (`pip install openpi-client`):

```
from openpi_client.websocket_client_policy import WebsocketClientPolicy

client = WebsocketClientPolicy("my-server", 8000)
result = client.infer(observation)
action = result["action"]
```

The prompt rides in the observation; a changed prompt starts a new episode, and a client can
say it explicitly with `{"reset": True, ...}`. Four flags cover the rest:

- `--serve.dry_run` — preflight only: device, declaration, plan. No weights, no GPU.
- `--serve.smoke` — load, produce one action, exit.
- `--serve.seed` — seed native execution for paired comparisons; FP8 serving rejects this option.
- `--serve.viz` — stream observations, actions and latency to a [Rerun](https://rerun.io) viewer.

The second verb, `instinctflash validate <dir>`, checks a checkpoint is publishable; given
`--validate.teacher_outcomes/.student_outcomes/.margin` it also certifies non-inferiority and
stamps the certificate into the package.

## Benchmark acceleration and quantization

After the [vendor and auxiliary-asset preparation](https://github.com/General-Instinct/InstinctFlash/blob/main/REPRODUCE.rst), reproduce paired
eager/default/selected Runtime measurements with the included inputs and fixed
checkpoint revision. Thor also requires its [native backend](https://github.com/General-Instinct/InstinctFlash/blob/main/serving/README.rst).
Keep the model and asset environments activated. For RTX 4090:

```
python -I -m benchmarks.regression.reproduce prepare --target rtx4090 \
  --model pi05 --mode fp8 --output pi05-inputs
python -I -m benchmarks.regression.reproduce run --prepared pi05-inputs --output pi05-results
python -I -m benchmarks.regression.serve_smoke --prepared pi05-inputs --output pi05-serving
```

`run` writes checked JSON/CSV reports and full action arrays. `serve_smoke` tests
the actual CLI and WebSocket pipeline across two episodes. Use `--mode native`
for default precision; FP8, numerical compilation and changed schedules are
explicit selections. [Reproduction guide](https://github.com/General-Instinct/InstinctFlash/blob/main/REPRODUCE.rst).
For additional framework comparisons, use the [pinned comparison recipes](https://github.com/General-Instinct/InstinctFlash/blob/main/benchmarks/regression/FRAMEWORK_COMPARISON.rst).

Compare original and optimized models with `instinctflash eval`. Reports separate
latency, action agreement and simulator task success.

```
instinctflash eval adapters
instinctflash eval coverage --run /path/to/run
instinctflash eval --registry plan.registry.json report --run /path/to/run
```

See the [evaluation guide](https://github.com/General-Instinct/InstinctFlash/blob/main/benchmarks/vla/SIMULATOR_EVALUATION.md) to create and run
paired LIBERO / RoboTwin experiments, or [benchmark details](https://github.com/General-Instinct/InstinctFlash/blob/main/benchmarks/vla/README.md)
for acceleration and quantization protocols. Results: [simulator screening](https://github.com/General-Instinct/InstinctFlash/blob/main/eval/simulator_quality_2026-09-06/README.md)
and [repeatability, checkpoints and edge latency](https://github.com/General-Instinct/InstinctFlash/blob/main/eval/simulator_next_steps_2026-09-06/README.md).
The [expanded V2 evaluation](https://github.com/General-Instinct/InstinctFlash/blob/main/eval/precision_evidence_2026-09-06/README.md) binds
latency and quality evidence to execution profiles and checks explicit control budgets.
The [native qualification workflow](https://github.com/General-Instinct/InstinctFlash/blob/main/benchmarks/regression/README.md)
adds fresh-start admission, retained failures and checkpoint-specific evidence for each device.
LingBot-VA Hub IDs retain native step counts; 2V/4A requires an explicit `nfe` selection.
The [September 9 Thor comparison](https://github.com/General-Instinct/InstinctFlash/blob/main/eval/thor_precision_completion_2026-09-09/COMPARISON.md)
separates native acceleration, FP8 Runtime gains and paired task outcomes;
[historical engine controls](https://github.com/General-Instinct/InstinctFlash/blob/main/eval/fp8_comparison_2026-09-09/README.md) isolate additional implementation effects.

[Shared BF16 fusion](https://github.com/General-Instinct/InstinctFlash/blob/main/instinctflash/native/bf16/README.md) provides an opt-in NUMERIC path, with per-model compatibility and paired Thor regression results.
[Shared tensor caching and prefill separation](https://github.com/General-Instinct/InstinctFlash/blob/main/eval/shared_tensor_cache_2026-09-13/README.md) extend native Cosmos optimization to Edge and Nano; exact caching and NUMERIC compilation remain separate options.

## Framework overview

InstinctFlash keeps model declarations, optimization planning, runtime execution, and evidence in
one inspectable path, whether it is called from Python or the command line.

# Architecture

A checkpoint carries a short declaration of what it is. The runtime reads the declaration, decides
which optim
guanming0717274
🟧 echo.github ⭐Full-source robotics serving runtime with reproduction recipes; the 33.78-fold headline includes changing LingBot-VA's sampling schedule, unGeneral Instinct——
🟧 hnAI agent firm Instinct raises $1B in latest funding rounddoppp10

Interpretation history

Decision trace