2026-10-11 17:11 UTC

AWS engineer Andrey Grehov released Range, a tool that opens container images and Hugging Face repositories via ranged reads without downloading β€” claiming a 1.03TB Kimi K2 'opened' in ~3.4s by moving 9.5MB β€” and sustained adoption in local-inference and eval tooling would establish lazy remote weight-streaming as a practical access pattern, while real full-model-run bandwidth or latency walls would confine it to inspection use.

state: resolvedheat: lowuncertainty: lowconvergesscott: lowlocal-inference weight-streaming model-servingAndrey Grehov

What is this?

Range is a developer tool by AWS engineer Andrey Grehov that 'opens' container images and Hugging Face repositories via ranged reads β€” fetching only index/metadata and requested byte ranges instead of the whole artifact β€” and claims a 1.03TB Kimi K2 checkpoint opened in ~3.4s by transferring just 9.5MB; the supplied web results contain no direct coverage of Range, Grehov, or its benchmarks, so those claims rest entirely on the case's own evidence object. What the snippets do establish is the context that makes the tool interesting: Kimi K2/K2 Thinking are ~1T-parameter open-weight models Moonshot AI publishes on Hugging Face, the July 2026 Kimi K3 release shipped ~1.56TB of weights with commentary explicitly calling a local download 'a multi-day bandwidth exercise with no inference payoff,' and the ecosystem's current default alternative is hosted inference (e.g. Novita Labs) that avoids local weights altogether. The snippets are thin-to-silent on the tool itself; whether lazy remote weight-streaming holds up under real inference loads is untested by this material.

Why it matters to Scott

Range is the network twin of the radar's NVMe/SSD-streaming cluster: it independently implements the read-only-archive-discovery-ladder's 'index first, then bounded exact-source reads on demand' pattern against model weights, attacking the same multi-day-download wall the radar's Kimi K3 cases document for his gamepc/Ollama zoo. But the 9.5MB/3.4s receipt proves opening only, and no canon page holds a position on remote weight-streaming β€” so this converges on his archive-discovery pattern and extends the streaming open-question to the remote axis without yet changing what he builds; it becomes high-relevance the moment ranged reads are shown surviving real inference traffic.
dev:concept.read-only-archive-discovery-ladderdev:project.gamepcradar:kimi-k3-nvme-expert-streamingradar:llama-cpp-lazy-tensor-loadingradar:slipstream-ssd-moe-streamingradar:minirun-kimi-k3-iphone-ssd-inferenceradar:deepseek-v4-nvme-demand-pagingradar:kimi-k3-english-pruned-gguf
queries asked of Scott's wikis
  • lazy loading weight streaming model serving
  • ranged reads sparse fetch object storage
  • TB-scale open weights download bandwidth
  • local inference stack model access pattern
  • hugging face hub tooling eval harness
  • open weights local-first access strategy

Measured heat

no measured readings yet β€” the hourly heat pass fills this in

How the heat travelled

09-25 14:00⭐ origin echo-reconstructedREADME.md: "Use a remote environment before downloading it. Range opens a shell inside a container image, a Hugging Face repository, or an e
Andrey Grehov (software engineer at AWS; personal project) on github (echo) Β· attributed from hn.story.49884918
β€”
09-28 21:50first on hacker news Β· published Β· +79.8hRange – open a 1 TB AI model in 3 seconds without downloading it
andreygrehov
β€”
09-28 21:50amplified on hacker newshn.story.49884918
andreygrehov
peak 4 Β· 0 comments Β· 40% of case engagement
09-30 13:24amplified on hacker news πŸ‘‘hn.story.49908611
andreygrehov
peak 5 Β· 0 comments Β· 50% of case engagement
10-02 14:51amplified on hacker newshn.story.49934167
andreygrehov
peak 1 Β· 0 comments Β· 10% of case engagement
09-28 23:21our radar first saw it Β· +81.3hdiscovery anchor: hn.story.49884918β€”

Evidence (4) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnRange – open a 1 TB AI model in 3 seconds without downloading it
Retrieved article excerpt

Open article Β· Retrieved 2026-09-29T00:32:00.022294+00:00

[Range](https://github.com/andreygrehov/range "Range on GitHub (Alt+R)")

Use a remote environment before downloading it.

README.md

README.md

# Use a remote environment before downloading it.

Range opens a shell in a container image, a Hugging Face repository, or an environment in S3
or on any HTTP server, without downloading it first. Only the bytes your program reads cross the
network.

```
$ range shell python:3.12
$ range shell python:3.12 --mount hf://moonshotai/Kimi-K2-Instruct:/model
$ range shell s3://<your-bucket>/dev.range
```

**2.8 s**to run Python in python:3.12

**48 MB**moved, of a 435 MB image

**9.5 MB**read, of a 1.03 TB model

No Docker, no daemon and no pull. Linux runs it natively. On macOS, Range runs Linux in a
small VM that it manages itself.

Measured on EC2 in us‑east‑1, with the image indexed once. See bench.log.

demo.txt

## A chat model, from nothing, in one line

```
$ range run ghcr.io/ggml-org/llama.cpp:light-b11206 \
    --mount hf://unsloth/gemma-3-270m-it-GGUF:/model -- \
    llama-cli -m /model/gemma-3-270m-it-Q4_K_M.gguf -st \
    -p "Why is the sky blue? Answer in one sentence."

The sky is blue because of a phenomenon called Rayleigh scattering,
where blue light is scattered more than other colors.
```

Range opens the llama.cpp image from its registry and mounts the model repository at
`/model`. The repository holds 6.38 GB in 24 files. Range reads one of them.
With the image indexed, the answer took 6.7 s from an empty cache.
`docker pull` plus `hf download` took 18.3 s. The very first run,
which also indexes the image, took 15.5 s.

## A 1 TB model, open in seconds

```
$ range run python:3.12 --mount hf://moonshotai/Kimi-K2-Instruct:/model -- \
    du -sh --apparent-size /model
959G    /model
```

Kimi K2 is 1.03 TB in 61 shards. A Python script inside read its config, the header of one
shard and one tensor. That took 3.4 s with the image indexed, and moved 9.5 MB of the model. The
other 60 shards never left Hugging Face.

Range reads a file when a program opens it. A program that reads a whole model
still downloads the whole model, once.

bench.log

## The chat demo, against docker pull

docker pull + hf download1.0x18.27 s579 MB

Range, first run1.2x15.46 s590 MB

Range, image indexed2.7x6.73 s317 MB

Range, again4.3x4.25 s0 MB

*0 s**10 s*

| Empty cache to output | docker pull | Range, first run | Range, indexed | Range, again |
| --- | --- | --- | --- | --- |
| Chat demo | 18.3 s 579 MB | 15.5 s 590 MB | 6.7 s 317 MB | 4.2 s 0 MB |
| python:3.12 | 15.8 s 435 MB | 16.5 s 415 MB | 2.8 s 48 MB | 1.1 s 0 MB |
| rust:1.82 | 19.6 s 569 MB | 22.4 s 546 MB | 8.0 s 125 MB | 1.6 s 0 MB |
| eclipse-temurin:21 | 7.9 s 232 MB | 9.2 s 225 MB | 2.9 s 49 MB | 1.0 s 0 MB |
| Kimi K2, 1 TB | not tried 1.03 TB | 17.6 s 433 MB | 3.4 s 56 MB | 1.2 s 0 MB |

The commands: import json and sqlite3, cargo --version, java -version, and a
read of one Kimi K2 tensor. A first run reads each layer once to index it. The python:3.12 index
is 4.1 MB. Medians of three, m6i.large, us‑east‑1, 28 September 2026. Every run starts
empty, except "again". The bars replay at 3x speed.

problem.txt

## What you wait for today

A machine that needs a large environment downloads all of it, every time, to use a small part.

### docker pull downloads a whole image to run one command
:   python:3.12 is a 435 MB download. Starting Python needs 48 MB of it.

### A model repository comes whole
:   unsloth/gemma-3-270m-it-GGUF holds 24 versions of one model in 6.38 GB. The demo
    needs one of them, 253 MB.

### Fifty eval workers download the same thing fifty times
:   Each worker copies it to its own disk before it starts.

Range reads only the bytes each machine touches, and the next run fetches them
before it asks.

design.txt

## One abstraction, and nothing above it

`ReadAt(offset, length) -> bytes`. Range turns a source into a disk, and turns
each read of that disk into a ranged request to the source.

### An image becomes a disk
:   Range reads each layer once to index it. After that a file costs one ranged request to the
    registry, and a few megabytes of gzip or zstd decompression at most. The image stays as it is.

### A model repository becomes a disk
:   The file list comes from the Hugging Face API, pinned to one commit. Reading a file sends a
    ranged request to the Hub.

### A shell is that, with a filesystem on top
:   disk -> read-only EROFS -> writable overlay -> namespaces. Writes stay local.
    Range checks every 64 KiB of an image layer against a SHA-256 in the index.

### It learns the working set
:   Each session records which blocks it needed. The next session fetches them in the background
    as the shell starts. A real read always goes first.

install.txt

## Install

A release archive for macOS or Linux, x86-64 or arm64. On macOS, Range also needs Lima for
its Linux VM. Then open a shell in any image:

```
$ curl -fsSL https://github.com/andreygrehov/range/releases/latest/download/range_$(uname -s)_$(uname -m).tar.gz | tar -xz
$ brew install lima      # macOS only
$ ./range shell python:3.12
```

For your own environments, build once, and every first run reads lazily:

```
$ range build --from-oci python:3.12 -o py.range
$ range publish py.range s3://<your-bucket>/py.range
$ range shell s3://<your-bucket>/py.range
```

A published artifact needs no indexing. The go1.23 demo artifact was ready in
0.41 s on its first run, and moved 6 MB of 1.03 GB.

Or build from source, with Go 1.25 or newer:

```
$ git clone https://github.com/andreygrehov/range && cd range && make install
```

[GitHub](https://github.com/andreygrehov/range)

Linux needs root and the nbd, erofs and overlay kernel modules.
`range doctor` checks them. Windows works through WSL2, untested.

about\_me.txt

A dithered black and white portrait of Andrey Grehov

I am a software engineer at AWS. Range is my personal project.

~/range $ view README.md
andreygrehov40
🟧 echo.github ⭐README.md: "Use a remote environment before downloading it. Range opens a shell inside a container image, a Hugging Face repository, or an eAndrey Grehov (software engineer at AWS; personal project)β€”β€”
🟧 hnShow HN: Range – open a 1TB Linux environment in under a secondandreygrehov50
🟧 hnShell into a remote environment before downloading itandreygrehov10

Interpretation history

Decision trace