2026-10-11 16:38 UTC

Agnes AI claims its released Apache-2.0 Agnes-3.0-Flash supports 262K-context multimodal reasoning and tool use with growing KV caches in only 18 of 72 layers, potentially reducing memory requirements for long-context self-hosted inference.

state: watchingheat: lowuncertainty: highnovelscott: mediumlocal-inference long-context multimodal-modelsAgnes AI

What is this?

Agnes AI’s documentation presents Agnes 3.0 Flash as a model for agentic coding and tool-driven workflows, accepting text and image-URL inputs through Chat Completions, Responses, and Messages APIs. Artificial Analysis describes Agnes AI as a Singapore-based lab training its own foundation models, though that snippet concerns the earlier 2.5 Pro Alpha. The supplied snippets do not establish the case’s central release claims: Apache-2.0 open weights, 33B parameters, 262,144-token context, or growing KV caches in only 18 of 72 layers; the proposed self-hosted memory benefit therefore remains unverified.

Why it matters to Scott

The claimed KV-cache reduction creates a concrete evaluation lead for Scott’s gamepc model-serving substrate and hardware-aware local inference work: whether longer contexts fit within his available GPU memory. However, the supplied grounding does not verify the architecture, licence, context limit or runtime compatibility, so this is not yet an actionable deployment improvement or established convergence with his positions; no supplied radar page tracks this Agnes release.
dev:project.gamepcdev:concept.hardware-aware-local-inferenceradar:concept.local-inferenceradar:concept.kv-cacheradar:qwen38-flash-next-hybrid-inferenceradar:ctx-cliff-local-inference-benchmark
queries asked of Scott's wikis
  • local inference memory budgets KV cache economics
  • hybrid attention long-context inference architectures
  • self-hosted open-weight models deployment licensing
  • coding agent harness model integration tool-use reliability
  • long-context versus RAG agent memory tradeoffs

Measured heat

now 0 pts/hpeak 48 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 704h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-12 08:22 (minted)⭐ origin echo-reconstructedThe model card releases open weights and inference implementations for a 33B multimodal hybrid decoder with a 262,144-token context and sequ
Agnes AI on blog (echo) · attributed from reddit.post.1we6lrn · published time unknown
—
09-12 08:05first on r/LocalLLaMA · published · lag ?Agnes-AI/Agnes-3.0-Flash 33B Multimodal, AA score: 36
Skyline34rGt
—
09-12 08:05amplified on r/LocalLLaMA 👑reddit.post.1we6lrn
Skyline34rGt
peak 255 · 55 comments · 48% of case engagement
09-12 15:33amplified on r/LocalLLaMAreddit.post.1wefo21
Informal-Trouble2183
peak 0 · 34 comments · 5% of case engagement
09-12 23:36amplified on r/LocalLLaMAreddit.post.1wers2v
Aggressive_Aspect436
peak 66 · 11 comments · 12% of case engagement
10-05 12:58amplified on r/LocalLLaMAreddit.post.1wy7wn0
ComfortableKindly507
peak 190 · 40 comments · 35% of case engagement
09-12 08:20our radar first saw it · lag ?discovery anchor: reddit.post.1we6lrn—
pace: p82 vs 1032 stories at the 336h mark (now 704h old) — ahead of cms-wiser-medicare-ai-denials (1.0x), behind deathray-macos-webgpu-hang (1.0x)

Evidence (5) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditAgnes-AI/Agnes-3.0-Flash 33B Multimodal, AA score: 36
LocalLLaMA
Retrieved article excerpt

Open article · Retrieved 2026-09-12T08:21:43.469606+00:00

# [Agnes-AI](https://huggingface.co/Agnes-AI) / [Agnes-3.0-Flash](https://huggingface.co/Agnes-AI/Agnes-3.0-Flash) Like 10

[Image-Text-to-Text](https://huggingface.co/models?pipeline_tag=image-text-to-text)[Transformers](https://huggingface.co/models?library=transformers)[Safetensors](https://huggingface.co/models?library=safetensors)[English](https://huggingface.co/models?language=en)[Chinese](https://huggingface.co/models?language=zh)[agnes](https://huggingface.co/models?other=agnes)[text-generation](https://huggingface.co/models?other=text-generation)[agnes-ai](https://huggingface.co/models?other=agnes-ai)[reasoning](https://huggingface.co/models?other=reasoning)[multimodal](https://huggingface.co/models?other=multimodal)[long-context](https://huggingface.co/models?other=long-context)[hybrid-attention](https://huggingface.co/models?other=hybrid-attention)[conversational](https://huggingface.co/models?other=conversational)[custom\_code](https://huggingface.co/models?other=custom_code)

License: apache-2.0

[Model card](https://huggingface.co/Agnes-AI/Agnes-3.0-Flash)  [Files Files and versions  

xet](https://huggingface.co/Agnes-AI/Agnes-3.0-Flash/tree/main)  [Community 

2](https://huggingface.co/Agnes-AI/Agnes-3.0-Flash/discussions)

 

Deploy

  Copy to bucket new   

Use this model

 

Agnes AI logo

[Agnes AI website](https://agnes-ai.com/)
[Open weights](https://huggingface.co/Agnes-AI/Agnes-3.0-Flash#quickstart)
Apache 2.0

# Agnes-3.0-Flash

Hello! 👋 Today we are introducing **Agnes-3.0-Flash**, an **open-weights multimodal model** built for people who want flagship-class reasoning without flagship-class hardware.

Highlights:

- **Competitive across core capabilities.** Agnes-3.0-Flash posts competitive results across reasoning, coding, and instruction-following evaluations.
- **Built for demanding work.** A **262 144-token context window**, adjustable reasoning effort, tool calling, and **text, image and video** understanding.

Agnes-3.0-Flash benchmark reference results

## Agnes-3.0-Flash

### Benchmarks

Reference results across contemporary models are shown below. The figures were compiled from different sources, harnesses, and model snapshots and do not constitute a controlled head-to-head comparison.

| Benchmark | Agnes-3.0-Flash | Qwen3.6-35B-A3B 35B / 3B active | Kimi K2.5 1T / 32B active | Muse Glimmer 30B | Qwen3.5 27B | DeepSeek V4 Flash 0731 284B / 13B active | Qwen3.8 27B | Gemini 3.5 Flash undisclosed | Qwen3.8 Flash Next 125B / 6B active | MiniMax M3 428B / 23B active |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| IFBench | 74.20 | 64.4 | 43.7 | 77.0 | 75.6 | 75.8 | 79.5 | 76.3 | 81.3 | 82.9 |
| SciCode | 38.08 | 35.8 | 39.6 | 43.6 | 39.5 | 50.3 | 46.6 | 53.1 | 50.6 | 45.4 |
| GPQA Diamond | 85.05 | 84.1 | 78.9 | 83.5 | 85.8 | 90.8 | 90.5 | 92.2 | 92.3 | 92.9 |
| AA-LCR | 68.33 | 66.7 | 59.0 | 80.0 | 72.3 | 79.7 | 82.0 | 81.0 | 79.7 | 74.0 |
| AA-Omniscience Accuracy | 23.00 | 18.8 | 22.9 | 27.0 | 20.7 | 40.4 | 18.4 | 51.4 | 24.5 | 16.7 |

Higher is better for every row. Header parameter figures mix total and active counts, and harnesses and snapshot dates differ across sources, so treat cross-column comparisons as reference values rather than a controlled head-to-head evaluation.

## Architecture

Agnes-3.0-Flash is a hybrid-attention decoder: three of every four layers run a gated delta rule (recurrent, with per-layer state independent of sequence length), and the fourth runs standard global attention. Only 18 of the 72 layers therefore hold a KV cache that grows with context.

|  |  |
| --- | --- |
| Context length | **262 144** tokens |
| Decoder layers | 72 = 54 delta-rule recurrent + 18 global attention, alternating 3 : 1 |
| Hidden size | 5120 |
| Global attention | 24 query heads / 4 KV heads (6 : 1 GQA), head dim 256; RMS-norm on q and k, sigmoid-gated output |
| Delta-rule layers | 16 key heads / 48 value heads, head dim 128; causal conv (kernel 4) in front, gated RMS-norm; recurrent state in fp32 |
| Feed-forward | SwiGLU, intermediate size 17408; plus a parallel SwiGLU 2048 branch in every layer |
| Positions | 3-axis rotary (text / height / width), interleaved mrope sections 11 : 11 : 10, base 1e7, applied to the first 25 % of each head dim (64 dims) |
| Vocabulary | 248 320 |
| Vision tower | 27 layers, hidden 1152, patch 16, 2 × 2 spatial merge, projected to 5120 |

## Quickstart

REMOTE CODE REQUIRED

**Agnes-3.0-Flash** ships its own model implementation. Always load it with `trust_remote_code=True`.

### Requirements

```
pip install "transformers>=5.12" torch torchvision accelerate
```

Tested on transformers 5.12.1. Image and video inputs go through the bundled processor, which needs `torchvision`.

### Transformers

```
from transformers import AutoModelForCausalLM, AutoTokenizer

path = "Agnes-AI/Agnes-3.0-Flash"
tok = AutoTokenizer.from_pretrained(path)
model = AutoModelForCausalLM.from_pretrained(
    path, dtype="bfloat16", device_map="auto", trust_remote_code=True
)

msgs = [{"role": "user", "content": "请用三句话解释什么是人工智能。"}]
ids = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(ids, max_new_tokens=256)
print(tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True))
```

### Images and video

Image and video inputs go through the bundled processor (also remote code):

```
from transformers import AutoProcessor

proc = AutoProcessor.from_pretrained(path, trust_remote_code=True)
msgs = [{"role": "user", "content": [{"type": "image", "image": "photo.jpg"},
                                     {"type": "text", "text": "描述这张图。"}]}]
inputs = proc.apply_chat_template(msgs, add_generation_prompt=True, tokenize=True,
                                  return_dict=True, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=256)
print(proc.batch_decode(out[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True)[0])
```

### Reasoning effort

The chat template exposes three reasoning levels — `high` (default), `medium`, `low` — plus a thinking-off switch:

```
ids = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt",
                              reasoning_effort="medium")   # or enable_thinking=False
```

### Tool calling

The chat template renders tool definitions for you. The model emits calls as `<tool_call><function=…><parameter=…>`, and you feed results back as a `tool` role message:

```
tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Look up current weather for a city",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string", "description": "City name"}},
            "required": ["city"],
        },
    },
}]

msgs = [{"role": "user", "content": "What's the weather in Beijing right now?"}]
ids = tok.apply_chat_template(msgs, tools=tools, add_generation_prompt=True,
                              return_tensors="pt").to(model.device)
out = model.generate(ids, max_new_tokens=256)
reply = tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True)
# <tool_call>
# <function=get_weather>
# <parameter=city>
# Beijing
# </parameter>
# </function>
# </tool_call>

# run the tool, append the result, generate the final answer
msgs += [{"role": "assistant", "content": reply},
         {"role": "tool", "content": "Clear, 26°C, light northeasterly wind"}]
```

Over the OpenAI API pass `tools=` the same way. The server returns the text above verbatim by default; to get structured `tool_calls`, configure sglang with a tool-call parser matching this format (likewise a reasoning parser, if you want the thinking span in `reasoning_content`).

### SGLang

`serve.sh` starts a server from a stock public image, overlaying three files onto the image's `sglang` package and nothing else. See `sglang_patch/README.md`.

```
docker run --gpus all --shm-size 64g -p 30001:8080 \
    -v /path/to/agnes-3.0-flash:/model \
    lmsysorg/sglang:nightly-dev-20260908-20ca564b \
    bash /agnes-3.0-flash/serve.sh --served-model-name Agnes-3.0-Flash
```

`serve.sh` forwards extra command-line arguments to sglang, which is how `--served-model-name` takes effect; `--tp 2` works the same way. The server listens on port 8080 inside the container:

```
from openai import OpenAI

client = OpenAI(api_key="EMPTY", base_url="http://localhost:30001/v1")
response = client.chat.completions.create(
    model="Agnes-3.0-Flash",
    messages=[{"role": "user", "content": "Design a fault-tolerant event processing architecture."}],
    temperature=1.0,
    max_tokens=2000,
)
print(response.choices[0].message.content)
```

Pass `stream=True` for streaming; `tools=` and `reasoning_effort=` are accepted the same way.

## Hardware Requirements

| Resource | Recommendation |
| --- | --- |
| GPUs | 1 × NVIDIA H200 141 GB or NVIDIA H100 80 GB (or equivalent) at bf16 |
| Tensor parallel | `--tp 1`; `--tp 2` for maximum context and concurrency |
| Weights on disk | Approximately 66 GB for the bf16 checkpoint |
| Host memory | 128 GB or more recommended |

Actual context length and concurrency depend on KV-cache allocation, runtime overhead, and tensor-parallel configuration; validate the target workload on the intended hardware.

## Recommended Inference Settings

| Setting | Recommended |
| --- | --- |
| `temperature` | 1.0 |
| `top_p` | 0.95 |
| `top_k` | 20 |
| `reasoning_effort` | `high` for hard reasoning, `low` for latency-sensitive traffic |
| `max_tokens` | 2000 or higher |

These are the checkpoint's own `generation_config.json` defaults.

## Model Capabilities

| Capability | Support |
| --- | --- |
| Advanced reasoning | Yes, with `high` / `medium` / `low` effort levels |
| Coding and debugging | Yes |
| Long-context analysis | 262 144 tokens |
| Image understanding | Yes |
| Video understanding | Yes |
| Tool calling | Yes (`<tool_call>` / `<tool_response>`) |
| Streaming | Yes |
| OpenAI-compatible APIs | Chat Completions via sglang |

## License

Released under the [Apache License 2.0](https://huggingface.co/Agnes-AI/Agnes-3.0-Flash/tree/main/LICENSE).

## Citation

```
@misc{agnes30flash2026,
  title        = {Agnes-3.0-Flash},
  author       = {{Agnes AI}},
  year         = {2026},
  month        = sep,
  howpublished = {Open-weights model},
  url          = {https://agnes-ai.com/}
}
```

Downloads last month
:   -

 

Safetensors

Model size

33B params

Tensor type

BF16

·

Chat template

Files info
Skyline34rGt25555
🟧 echo.blog ⭐The model card releases open weights and inference implementations for a 33B multimodal hybrid decoder with a 262,144-token context and sequAgnes AI——
🟠 redditA mysterious new model tops my efficiency index
LocalLLaMA
Informal-Trouble2183034
🟠 redditWhat's the Story with Agnes-3.0-Flash?
LocalLLaMA
Aggressive_Aspect4366611
🟠 redditAgens Volundr 32B Preview: our small team's first model on our own hybrid architecture. Only 18 of 72 layers keep a KV cache (Apache-2.0)
LocalLLaMA
ComfortableKindly50719040

Interpretation history

Decision trace