Retrieved article excerpt
Open article ยท Retrieved 2026-09-22T20:24:15.175260+00:00
# shrewd
Turn LLM judgments into a small, fast, local text model for one fixed task.
I built shrewd while making local models for another project. It started with a question:
can GEPA prompt optimization get better labels from an LLM? Then, since Jev-style decisions
have gotten popular recently, I wanted to see how far I could get toward them locally for a
fixed set of questions. This repo contains the pipeline I used and what I measured along
the way.
It is useful when you repeatedly classify similar text, have a few hundred hand-labeled
examples and a larger unlabeled pool, and care about inference cost, latency, or keeping
text on your own hardware after training. It builds either a basic **classifier** (one label per
document) or a **decision panel** type of classifier (a fixed set of choices, yes/no probabilities,
and ratings answered together). Both use an LLM as the teacher and produce a saved model that runs
locally with no LLM in the loop.
Some findings may be useful even if you never use the library:
- Prompt optimization helped some teachers, but gains on the development split often
disappeared on held-out data.
- On the two tasks tested for cost, choosing informative rows from a good teacher was
more effective than buying cheaper labels or escalating uncertain ones.
- Better teacher labels did not always improve the student. Changing the student or
giving it more training data sometimes helped more.
- Calibration helped the probabilities match human labels, but low calibration error
sometimes hid a model that barely distinguished one example from another.
I used established methods. Most experiments weren't repeated with different data splits
and training seeds, so small score differences need more testing. I'm sharing the code
to make trying this on your own task easier. [BENCHMARKS.md](https://github.com/sshah03/shrewd/blob/main/BENCHMARKS.md) has the
measurements, limits, and alternatives.
I'd especially like to hear which findings hold up on other people's data.
Four prebuilt decision panels (SMS spam, email triage, an AI-input guardrail, and a
personal-data gate) are available as demonstrations, with their known failures below.
```
seed CSV (labeled) โโโถ [1] split: dev / locked test set
[2] optimize the teacher prompt on dev (GEPA) optional
pool CSV (text only) โโโถ [3] the teacher labels or judges the pool
[4] fix any of its answers by hand optional
[5] train a small local student on the result
[6] score teacher and student on the locked test set
```
**Part 1** covers using it. **Part 2** covers the findings and their limits, with fuller
tables in [BENCHMARKS.md](https://github.com/sshah03/shrewd/blob/main/BENCHMARKS.md). The repo includes pipeline examples and panel
rebuilds. Some research results came from private experiment scripts that aren't included.
---
# Part 1: Using it
## Install
```
pip install "shrewd[teacher]" # training: LLM client + prompt optimizer, tfidf student
pip install shrewd # inference only: sklearn stack, no LLM libraries
pip install "shrewd[teacher,embed]" # + model2vec static-embedding student (no torch)
pip install "shrewd[teacher,setfit]" # + SetFit student (torch, best few-shot)
pip install "shrewd[teacher,encoder]" # + fine-tuned ModernBERT (torch, strongest at 1k+ rows)
```
## Compile a classifier
One label per document, a locked test set, and a report that says what to fix. Start here
if your problem is "which of these N buckets does this go in".
```
import pandas as pd
from shrewd import Project
proj = Project(
"runs/tickets",
instructions="Classify customer support tickets by the customer's primary intent.",
labels={"billing": "charges, invoices, refunds", "bug": "something is broken",
"cancellation": "wants to cancel, pause, or downgrade", "other": "none of the above"},
teacher="anthropic", # or "openai", or any litellm model string
)
proj.add_seed(pd.read_csv("labeled.csv")) # columns: text, label
proj.optimize(budget=800) # prompt optimization on the dev split
proj.label(pd.read_csv("unlabeled.csv")) # dry_run=True prices it first
result = proj.distill(student="tfidf")
print(result.report())
```
```
from shrewd import load
clf = load("runs/tickets")
clf.predict(["I was double charged last month"]) # -> ["billing"]
clf.predict_proba(["..."]) # ndarray in clf.classes_ order
clf.predict(texts, min_confidence=0.52) # None below the threshold: route to a person
```
To try it without an API key: `python examples/make_demo_data.py`, then
`python examples/quickstart.py`. It replays a bundled cache of a real run, so no API calls.
Options:
- **Label under a budget.** `proj.label(pool, budget_usd=20)` (or `n=2000`) acquires rows in
rounds, picking the rows a quick tfidf probe is least sure about instead of going in file
order. In my tests that took 1.1x to 2.3x fewer teacher calls for the same student
accuracy. Re-running continues where it stopped.
- **Fix labels by hand.** After `label()`, `needs_review.csv` lists the rows the teachers
disagreed on or could not answer, with an empty `human_label` column. Fill in the ones
you care about and the next `distill()` uses them at confidence 1.0. Pass
`teacher=[...]` with two or more models, or `votes=3`, to get disagreements to review.
With one teacher and one vote the queue stays empty.
- **Pick a student.** `"tfidf"` (default, instant), `"embed"` (static embeddings), `"encoder"`
(fine-tuned ModernBERT), `"setfit"` (contrastive fine-tuning), or your own object with
`fit` / `predict` / `predict_proba` / `classes_` / `save`.
```
rows = proj.compare(students=["tfidf", "embed", "encoder"]) # dev F1 vs latency vs size
proj.promote(rows[0]["student"])
result = proj.distill(student=rows[0]["student"]) # the one test-set evaluation
```
`autotune(dir, instructions, labels, seed_df, pool_df, budget_usd=...)` does this for you
under a spending cap. It starts with the cheapest teacher and moves up a tier only while
the dev score misses `target` and the next tier fits what's left of the budget. It scores
the locked test set once, on the winner, and writes every step to `autotune_trail.json`.
`teacher` takes any [litellm](https://docs.litellm.ai/docs/providers) model string.
`"anthropic"` and `"openai"` pick a strong default for that provider. **API keys come from
the environment**, the way litellm reads them: `export ANTHROPIC_API_KEY=...`,
`OPENAI_API_KEY`, `DEEPSEEK_API_KEY`, and so on. shrewd never takes a key as an argument or
writes one to disk, and `load()` doesn't need one.
`optimize()`, `label()` and `judge()` send your seed and pool text to the provider, so if
your text is sensitive, read their data-use terms first. The prompt goes out as a system
message marked for the provider's prompt cache (Claude charges ~10% of the input price for a
cached prefix), so a long prompt or many questions add little per document. Every response
is also cached locally in `cache.db`, so re-running a stage never pays twice.
**When to use something else:** if you won't run many documents through it, calling an
LLM or a hosted decision model may cost less than building and maintaining a student. Where
the line sits depends on your labeling, inference and retraining costs. It's also a poor fit
for tasks that need reasoning or fresh world knowledge per item, or for domains that change
quickly. The saved model only answers the task it was trained for.
## Compile a calibrated decision panel
The second path builds a fixed panel of typed questions (pick one of N, yes/no with a
probability, a rating on an ordered scale), all answered together in one pass. The
probabilities are calibrated against the teacher's answers and then scored against your
hand labels, because calibrating to the teacher alone doesn't guarantee
that a 0.8 means 80% on your task.
```
import pandas as pd
from shrewd import Decisions, Choice, Noul, Score
d = Decisions(
"runs/tickets",
instructions="These are customer support tickets.",
questions={
"department": Choice(
instructions="Which team should handle this?",
criteria={"billing": "charges, invoices, refunds",
"bug": "something is broken",
"other": "none of the above"}),
"angry": Noul(instructions="Does the customer sound angry?"),
"severity": Score(instructions="How severe is this?",
criteria=["cosmetic", "workaround exists", "blocking"]),
},
teacher="anthropic/claude-fable-5-1", # any litellm model string, or "anthropic" / "openai"
)
d.add_seed(pd.read_csv("labeled.csv")) # text + one column per question id (blank where unknown)
d.judge(pd.read_csv("unlabeled.csv")) # ONE teacher call answers every question per document
print(d.distill().report()) # trains, calibrates, scores against your hand labels
```
```
from shrewd import load
dec = load("runs/tickets")
dec.decide("I've been charged twice and nobody will call me back")
# {"department": <Answer department choice='billing'>,
# "angry": <Answer angry noul=0.91>,
# "severity": <Answer severity score=1.4>}
dec.predict_proba(texts) # {question: (n x options) calibrated probabilities}
```
There are three question types. `Choice` picks one of up to 255 named options and returns
a probability for each. `Noul` (borrowed from Jev) answers a yes/no question with one
number, the probability of yes. `Score` rates against 2 to 10 ordered levels and returns
the probability-weighted mean, so a document split between "cosmetic" and "blocking" lands
in the middle. Describe each option in `criteria`. The teacher and the zero-shot stack read
those descriptions, and bare option names don't tell them much.
Options:
- **Fix the teacher's answers.** `judge()` writes `needs_review.csv` listing every
(document, question) where the teacher's winning option is below 0.6, with an empty
`human_answer` column. Write an option name, yes/no, or a level's number or description.
the next `distill()` (or `apply_review()`) uses it as a certain target. Your answers are
ledgered in `reviewed.csv` and survive any re-judging.
- **Optimize the prompt.** `d.optimize(budget=800)` runs GEPA over the free text above the
questions, scored on your dev split by a proper scoring rule (it rewards accurate
probabilities as well as accurate labels).
- **Choose the student.** `distill(features="auto")` fits tf-idf and static embeddings and
keeps whichever wins on your hand-labeled dev split (0.01-0.5 ms/doc). `features="encoder"`
fine-tunes ModernBERT (`pip install "shrewd[encoder]"`, 3-30 ms/doc, ~20-45 min of GPU
or Apple-silicon training, ~11 GB RAM): use it when the question needs the model to
read rather than count words. Part 2 covers where each one wins.
- **Add a zero-shot second opinion.** `distill(zero_shot=True)` stacks an NLI model onto
each head and keeps it per question only where it beats the plain head on your hand-labeled
dev split. It helped on sentiment and emotion and got rejected almost everywhere else.
20-800 ms/doc.
- **Use several teachers.** `backend=EnsembleBackend(["anthropic/claude-fable-5-1", "openai/gpt-6-astra"])` averages their distributions. `backend="logprobs"` reads token
probabilities where the API exposes them (OpenAI-compatible endpoints).
Adding a question later is cheap. Answers are cached per `(model, question, document)`, so
a ninth question reuses the eight you already paid for. Changing a question's text or
options throws away its cached answers. The compiled artifact answers exactly the questions it was
compiled for.
### Try a pre-built panel
Four decision panels for softwa