2026-10-11 17:15 UTC

Modal and CMU's Full Stack Data Lab claim their released open-source Quail engine — a SQL query planner fused into the inference engine for AI-SQL workloads — processes over one billion tokens per minute per H100, more than 10x their vLLM baseline on a multi-join query at under 6¢ per billion tokens on Modal; independent benchmarking and adoption would establish query-aware serving as a new throughput-and-cost regime for batch LLM workloads.

state: seedheat: lowuncertainty: mediumconvergesscott: highllm-serving-throughput inference-economics ai-sql modal-quailCharles FryeShreya ShankarModalCMU Full Stack Data Lab

What is this?

Quail is an open-source 'query-aware inference layer' for AI-SQL released by CMU's Full Stack Data Lab (Shreya Shankar) together with Modal (Charles Frye): it fuses a SQL query planner into the LLM inference engine so execution order, request batching, and KV-cache reuse are planned per query instead of left to a generic serving stack. The lab's own benchmark shows one multi-join query running in 29.26 minutes versus 6.84 hours for 'stock' vLLM (14.04x, and 1.96x their state-of-the-art estimate) at $1.93 versus $27.03 per query, with reported throughput of ~19M requested input tokens/second — the >1B tokens/minute/GPU headline — and a cited full-run cost of $0.3675 on Modal H100s; against their own hand-tuned and pipelined vLLM baselines the gain narrows to ~1.84x. All figures are first-party from the announcing lab and vendor; the supplied material shows no independent benchmarking or third-party adoption yet.

Why it matters to Scott

Converges with Scott's soft-data BI line: Shankar/Frye independently built the cost regime his BI-for-Soft Data framework and The Soft Join presuppose, and Quail's planner-fused serving moves the exact threshold his Join-Cost Collapse and Cost-of-Cognition concepts name — with open code, per-query costs ($1.93 vs $27.03) and a Modal-denominated bill he could cite or run against his own bulk LLM-extraction pipelines. The grounding's catch — 14x over stock vLLM collapses to ~1.84x over hand-tuned, pipelined baselines — is itself the receipts-hygiene story the radar's vLLM episode family keeps arriving at, and the publishable nuance if he cites the number.
ip:framework.bi-for-soft-dataip:concept.join-cost-collapseip:concept.cost-of-cognitionip:concept.prefix-caching-economicsip:source.the-soft-join-ebookdev:concept.llm-structured-extractionradar:concept.inference-economicsradar:concept.vllmradar:concept.kv-cacheradar:concept.inference-optimizationradar:concept.data-engineeringradar:vllm-h100-config-latency-gainsradar:otlet-postgres-local-inference
queries asked of Scott's wikis
  • batch inference economics cost per billion tokens
  • vLLM serving throughput tuning baselines
  • semantic operators AI-SQL LLM data processing
  • query planning KV cache reuse inference optimization
  • Modal serverless GPU functions inference projects
  • RAG indexing bulk extraction pipeline cost

Measured heat

now 0 pts/hpeak 4 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 434h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-23 14:00⭐ origin echo-reconstructedAnnounces Quail, the 'QUery-Aware Inference Layer' combining a query planner with the inference engine for AI-SQL: 'On one multi-join query
Charles Frye (Modal) and Shreya Shankar (CMU Full Stack Data Lab) on blog (echo) · attributed from hn.story.49861274
—
09-26 22:41first on hacker news · published · +80.7h>1B tokens/minute/GPU by combining query planner and inference engine
charles_irl
—
09-26 22:41amplified on hacker news 👑hn.story.49861274
charles_irl
peak 7 · 2 comments · 50% of case engagement
09-27 14:12amplified on hacker newshn.story.49866809
birdculture
peak 1 · 0 comments · 6% of case engagement
09-29 23:19amplified on hacker newshn.story.49902184
charles_irl
peak 3 · 0 comments · 17% of case engagement
09-30 17:38amplified on hacker newshn.story.49912006
gmays
peak 2 · 0 comments · 11% of case engagement
09-30 18:27amplified on hacker newshn.story.49912536
shreya_shankar
peak 3 · 0 comments · 17% of case engagement
09-26 23:20our radar first saw it · +81.3hdiscovery anchor: hn.story.49861274—
pace: p52 vs 1032 stories at the 336h mark (now 434h old) — ahead of android-editable-graph-agent (1.1x), behind astra-skills-prompt-migration (0.9x)

Evidence (6) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hn>1B tokens/minute/GPU by combining query planner and inference engine
Retrieved article excerpt

Open article · Retrieved 2026-09-26T23:23:50.496408+00:00

[All posts](https://modal.com/blog)

[Back](https://modal.com/blog) 

Research

September 24, 2026 •15 minute read

# Hitting a billion tokens per minute on one GPU by combining a query planner and an inference engine

User avatar

[Charles Frye](https://twitter.com/charles_irl) 

Member of Technical Staff

 [@charles\_irl](https://twitter.com/charles_irl)

User avatar

[Shreya Shankar](https://twitter.com/sh_reya) 

Asst Professor, CMU FSD Lab

 [@sh\_reya](https://twitter.com/sh_reya)

> *I see it as a point on the LLM pareto optimal curve in a regime that had a large revealed latent demand (no thinking, single token, low latency acceptable intelligence) that was under-invested into because of a race to higher intelligence.*  
>   
> - [Karpathy-san, on Jev](https://x.com/karpathy/status/2102124533729955960?s=20)

While everyone and their cousin is loudly building coding agents and chatbots, there’s a quieter inference revolution going on in the backend. Simple LLM transformations of data can be incredibly powerful, provided the cost-performance is good enough — just scroll social media and catch a few of the eye-popping, [hack-inspiring](https://x.com/mattdesl/status/2100899669802963060?s=20) demos of [TypeSafe AI’](https://typesafe.ai/)s [Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev) model.

Jev implements these transformations at what you might call the “JSON layer”, Web-style interfaces between clients and services.

[AI-SQL](https://docs.snowflake.com/en/user-guide/snowflake-cortex/aisql) implements it at the analytic SQL layer, at the interface between business intelligence and the database:

```
-- get hot leads with AI™
SELECT customers.id, products.id
FROM customers JOIN products ON  -- for each row in both tables
AI.IF(  -- run the prompt below and filter by truthiness
		PROMPT("{customers.profile} might buy this: {products.description}")
)
```

Different inference applications produce different [inference workloads](https://modal.com/llm-almanac/workloads), and AI-SQL is no exception. A query like the one above might produce millions of sequences of thousands of tokens — RIP your token budget. These queries often require much less than frontier intelligence, so small open-weights models can crush. But naïvely delivering these sequences directly to an inference engine optimized for agentic inference through interfaces for arbitrary user-controlled requests is inherently and massively inefficient.

So we built an inference engine to fix this: the [QUery-Aware Inference Layer](https://github.com/fsdatalab/quail) (Quail). On one multi-join query where planning is particularly important, Quail hits over a billion tokens processed per minute per H100 GPU (TPM/GPU), >10x faster than our vLLM baseline on the same hardware. On Modal, that comes out to under 6¢ per billion tokens.

 

On [our newly-released benchmark for AI-SQL queries](https://github.com/fsdatalab/quail-bench), Quail runs 1.84x faster than vLLM, geometrically averaged over tasks -- including two queries we designed to demonstrate areas for future improvement in AI-SQL inference.

You can take it for a spin on Modal right now:

```
# uvx modal run try_quail.py

import modal

app = modal.App("try-quail")
image = (
    modal.Image.from_registry("nvidia/cuda:13.0.1-devel-ubuntu24.04", add_python="3.12")
    .entrypoint([])
    .apt_install("git")
    .uv_pip_install("quail-engine==0.1.0")
)


@app.function(gpu="H100!", image=image, timeout=600)
def run(sql=None, documents=None):
    from datasets import load_dataset
    import pyarrow as pa
    import quail

    if sql is None:
        sql = """ // no spoilers!
                SELECT r.id
                FROM reviews r
                WHERE AI_FILTER(PROMPT('Does this review discuss the ending?\n\n{0}', r.review))
                """

    if documents is None:
        imdb = load_dataset("stanfordnlp/imdb")["train"]

        documents = pa.table(
            {
                "id": pa.array(f"review-{i}" for i in range(len(imdb))),
                "review": imdb.data.table.column("text"),
            }
        )

    config = quail.EngineConfig(
        gpus=1,
        model="qwen3-4b-fp8",
        backend="quail",
        device="h100-sxm",
    )

    with quail.Session(config) as session:
        session.register(
            "reviews",
            quail.DocumentProvider.from_table(documents, id_col="id"),
        )
        result = session.sql(sql).run()

        print(result.collect())
        print(result.report)
```

In this blog, we’ll give a quick overview of the problem we’re solving and how Quail works today. Spoilers: the big win is that with a structured query in hand, you can order requests to better cache (and evict) KV. This requires a slight revision of [Hydragen](https://arxiv.org/abs/2402.05099)-style [cascade attention](https://flashinfer.ai/2024/02/02/cascade-inference.html). Large numbers of small requests for small models can also incur lots of [host overhead](https://modal.com/blog/host-overhead-inference-efficiency), aka have low [GPU utilization](https://modal.com/blog/gpu-utilization-guide), which can be avoided when you know the structure of the requests ahead of time.

This was a collaboration between inference researchers at Modal and database researchers Carnegie Mellon University’s [Full Stack Data Lab](https://fsdatalab.github.io/) — call it a “mixture of experts”. We’re sharing what we did because we’d like to make this work more “expert-parallel”, as it were. We believe this is only the beginning for open source performance engineering at the intersection of inference and databases — two of the most important applications of computing.

In this post, we’ll focus more on considerations for inference engineers. You can read more, from a database engineer’s perspective, at [the Full Stack Data Lab blog](https://fsdatalab.github.io/blog/introducing-quail/#43-quail-dominates-vllm-on-bio-4-1404x-faster). You can also check out the code for Quail [here](https://github.com/fsdatalab/quail) or the docs [here](https://fsdatalab.github.io/quail/docs). And if you run AI-SQL queries at scale and are interested in improving performance and cutting costs, [get in touch with us](https://modal.com/blog/quail-billion-tpm).

# What are AI Functions and AI-SQL?

First, a bit more background on the workload.

This is emphatically *not* prompting AI systems to produce SQL based on natural language inputs — that’s [NL2SQL](https://arxiv.org/html/2408.05109v4). That looks a lot like a traditional chatbot or coding agent workload, so existing inference engines work well.

It’s actually the other way around! In AI-SQL, we use an extension of SQL to programmatically produce (and consume) prompts for AI systems. Prompts are constructed from database entries and produce tables.

Like this:

```
-- get hot leads with AI™
SELECT customers.id, products.id
FROM customers JOIN products ON  -- for each row in both tables
AI.IF(  -- run the prompt below and filter by truthiness
		PROMPT("{customers.profile} might buy this: {products.description}")
)
```

AI-SQL is primarily used inside of business intelligence (BI) platforms to help data scientists and stakeholders ask more “fuzzy” questions of their semi-structured data, like documents and free-text fields.

There’s not a standard (yet), but major managed analytical database platforms have their own flavor: [Snowflake Cortex AI-SQL](https://docs.snowflake.com/en/user-guide/snowflake-cortex/aisql), [Databricks AI Functions](https://docs.databricks.com/aws/en/large-language-models/ai-functions), [BigQuery AI functions](https://cloud.google.com/blog/products/data-analytics/sql-reimagined-for-the-ai-era-with-bigquery-ai-functions).

# Unlike the rest of SQL, this problem actually needs GPUs.

Consider the following plan for a query over the [BioDEX dataset](https://github.com/KarelDO/BioDEX), which selects reports of serious adverse events in response to drugs that include both a neurological and a cardiovascular component:

 

If these were normal filters and joins, say based on string matching and logical equality, there’d be no good reason to use a high-throughput numerical accelerator like a GPU, even though this is an analytical query, which seems “throughput-y”. This may be obvious to some, but let’s step through the logic anyway.

Each byte loaded from durable storage to memory (or from memory to registers) would need at most a handful of arithmetic/logic operations to implement comparisons. GPUs are designed for workloads with high [arithmetic intensity](https://modal.com/gpu-glossary/perf/arithmetic-intensity) — many operations per byte loaded. And the latest GPUs have most of their [arithmetic bandwidth](https://modal.com/gpu-glossary/perf/arithmetic-bandwidth) in specialized hardware for large matrix multiplications, aka [Tensor Cores](https://modal.com/gpu-glossary/device-hardware/tensor-core). Normal filtering/joining requires no large matrix multiplications.

But this query plan uses `AI_FILTER` and `AI_JOIN`, which instead pass the inputs through a large language model. An LLM is a sequence of numerical operations, the bottleneck for which is large matrix multiplications. Each byte loaded from durable storage will be subject to on the order of billions of operations before a byte is written to storage.

# Why is this interesting to inference engineers?

Most inference engineering these days is focused on one workload shape in particular: iterative construction of long input sequences by users and tool calls external to the inference service. This is the shape of workloads from chatbots and agents — and of “rollout” inference during the reinforcement learning runs that fine-tune models to be chatbots or agents.

Don’t get us wrong, this is very important work! We’ve written about our approach to it [here](https://modal.com/blog/trillion-tokens-trillion-parameters). But for the hardcore inference engineer, it’s honestly starting to feel a little… played out.

There’s also some work on ultra low-latency inference where speed matters as much as intelligence. We’ve written about our techniques for this [here](https://modal.com/blog/achieve-sota-specdec). In general, these workloads use structured outputs/tool-calling. They end up as something like the “OLTP” of inference, slotting into other computer applications more easily than open-ended agents. The recent popularity of Jev demonstrates the importance of these workloads — and that we are still so early!

AI-SQL workloads haven’t gotten so much attention — yet — but we think they are interesting for inference engineers for a number of fundamental reasons, quite outside their importance to applications. Most intriguingly, they are an incredible fit for transformers (because they enable “perfect” KV cache use) and for transformers-on-GPUs (because they don’t require decode).

## Manage a KV cache without all the regrets.

In typical inference, requests are client-controlled and arbitrary. This causes [no end of pain](https://modal.com/blog/trillion-tokens-trillion-parameters). But in AI-SQL inference, clients only control SQL queries, which create many requests, and the combined query planner/inference engine has substantial control over the processing of those requests.

This makes it particularly easy to operate a cache that amortizes more work. For instance, we know exactly when any cache entry is no longer needed, so we can fearlessly evict it. We also know quite a bit about what the cache demand will look like, since we get an entire query plan’s worth of requests up front.

And we badly need caching for Transformers, because their forward passes are naïvely quadratic in the sequence length. We can exchange that for linear time and linear storage with KV caching.

KV caches can be tricky to operate for agent workloads, because the time between accesses is completely unknown. But for an AI-SQL query, we control the infere
charles_irl72
🟧 echo.blog ⭐Announces Quail, the 'QUery-Aware Inference Layer' combining a query planner with the inference engine for AI-SQL: 'On one multi-join query Charles Frye (Modal) and Shreya Shankar (CMU Full Stack Data Lab)——
🟧 hn>1B tokens/minute/GPU by combining query planner and inference enginebirdculture10
🟧 hnQuail: Speed up AI-SQL by jointly optimizing query planner and inference enginecharles_irl30
🟧 hnHitting 1B tokens/minute on 1 GPU combining a query planner and inference enginegmays20
🟧 hnJointly optimizing SQL queries and LLM inference for up to 14x speedupsshreya_shankar30

Interpretation history

Decision trace