2026-10-11 17:14 UTC

Shrewd's maintainer claims its released teacher-labeling and student-training pipeline replaces repeated LLM judgments with local fixed-task classifiers, reducing inference cost while showing that better teacher labels and prompt optimization do not reliably improve held-out student accuracy.

state: seedheat: lowuncertainty: mediumconvergesscott: mediummodel-distillation local-inference inference-economicssshah03

What is this?

Shrewd is a Show HN release (news.ycombinator.com item 49807276) by maintainer sshah03: an open pipeline that uses an LLM 'teacher' to label data, then trains small local fixed-task classifiers ('students') on those labels so that repeated LLM judgments at inference time can be replaced by cheap local models. The maintainer reports two Negative-ish findings: using GEPA (a prompt-optimization method) to get better teacher labels gave 'mixed results', and better teacher labels / prompt optimization did not reliably improve held-out student accuracy. The supplied snippets only show the HN post's opening lines plus adjacent academic work on LLM teacher-student distillation (e.g., an arXiv pipeline paper and an ACL paper on teacher-student labeling frameworks); they corroborate that this is an active research area but don't independently verify Shrewd's measurements.

Why it matters to Scott

Two findings pull different directions. The pipeline itself โ€” LLM teacher labels โ†’ local fixed-task students replacing repeated LLM judgments โ€” is territory Scott already occupies (his reddit fine-tuning-data factory does LLM-gated distillation onto local models; Context Arbitrage names exactly this frontier-to-utility spread) and territory the radar already tracks in several open cases (Blink, System One Lite, Verdict, Tracelint), so that half adds nothing. But the negative result โ€” that GEPA-optimized teacher labels do not reliably improve held-out student accuracy โ€” is an independent dated receipt for Perishable Prompt Layer's prompt-optimization-diminishing-returns position, and it quietly complicates his Specification Quality / North Star stance: if prompt craft at the teacher stage doesn't transfer to the student, the value of specification effort in labeling pipelines needs a caveat. All of this rests on one maintainer's unverified Show HN claims, so treat as a lead to check, not a settled measurement.
ip:concept.perishable-prompt-layerip:concept.context-arbitragedev:project.redditip:concept.deterministic-ai-pendulumradar:concept.model-distillationradar:concept.small-modelsradar:concept.llm-judgesradar:concept.inference-economics
queries asked of Scott's wikis
  • LLM-as-judge replacing with small local classifiers โ€” cost/latency notes in dev projects
  • prompt optimization diminishing returns (GEPA, DSPy-style optimizers) โ€” positions in IP frameworks
  • local inference economics and when small models beat API calls โ€” agent harness cost notes
  • distillation from frontier models into task-specific students โ€” prior writing or experiments
  • label quality vs downstream student accuracy โ€” anything in agent-memory or RAG labeling pipelines
  • fixed decision panels / classifier panels in production agent systems

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 452h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-23 12:23 (minted)โญ origin echo-reconstructedReleased a pipeline for turning LLM judgments into local classifiers and fixed decision panels, with measurements and explicit limitations o
sshah03 on github (echo) ยท attributed from hn.story.49807276 ยท published time unknown
โ€”
09-22 20:00first on hacker news ยท published ยท lag ?Show HN: Shrewd โ€“ what I learned distilling LLM labels into local classifiers
sshah03
โ€”
09-22 20:00amplified on hacker news ๐Ÿ‘‘hn.story.49807276
sshah03
peak 3 ยท 0 comments ยท 101% of case engagement
09-22 20:20our radar first saw it ยท lag ?discovery anchor: hn.story.49807276โ€”
pace: p32 vs 1032 stories at the 336h mark (now 452h old) โ€” ahead of addom-local-coding-harness (1.5x), behind agentsec-static-config-auditing (0.8x)

Evidence (2) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸง hnShow HN: Shrewd โ€“ what I learned distilling LLM labels into local classifiers
Retrieved article excerpt

Open article ยท Retrieved 2026-09-22T20:24:15.175260+00:00

# shrewd

Turn LLM judgments into a small, fast, local text model for one fixed task.

I built shrewd while making local models for another project. It started with a question:
can GEPA prompt optimization get better labels from an LLM? Then, since Jev-style decisions
have gotten popular recently, I wanted to see how far I could get toward them locally for a
fixed set of questions. This repo contains the pipeline I used and what I measured along
the way.

It is useful when you repeatedly classify similar text, have a few hundred hand-labeled
examples and a larger unlabeled pool, and care about inference cost, latency, or keeping
text on your own hardware after training. It builds either a basic **classifier** (one label per
document) or a **decision panel** type of classifier (a fixed set of choices, yes/no probabilities,
and ratings answered together). Both use an LLM as the teacher and produce a saved model that runs
locally with no LLM in the loop.

Some findings may be useful even if you never use the library:

- Prompt optimization helped some teachers, but gains on the development split often
  disappeared on held-out data.
- On the two tasks tested for cost, choosing informative rows from a good teacher was
  more effective than buying cheaper labels or escalating uncertain ones.
- Better teacher labels did not always improve the student. Changing the student or
  giving it more training data sometimes helped more.
- Calibration helped the probabilities match human labels, but low calibration error
  sometimes hid a model that barely distinguished one example from another.

I used established methods. Most experiments weren't repeated with different data splits
and training seeds, so small score differences need more testing. I'm sharing the code
to make trying this on your own task easier. [BENCHMARKS.md](https://github.com/sshah03/shrewd/blob/main/BENCHMARKS.md) has the
measurements, limits, and alternatives.

I'd especially like to hear which findings hold up on other people's data.

Four prebuilt decision panels (SMS spam, email triage, an AI-input guardrail, and a
personal-data gate) are available as demonstrations, with their known failures below.

```
seed CSV (labeled)    โ”€โ”€โ–ถ [1] split: dev / locked test set
                          [2] optimize the teacher prompt on dev (GEPA)        optional
pool CSV (text only)  โ”€โ”€โ–ถ [3] the teacher labels or judges the pool
                          [4] fix any of its answers by hand                   optional
                          [5] train a small local student on the result
                          [6] score teacher and student on the locked test set
```

**Part 1** covers using it. **Part 2** covers the findings and their limits, with fuller
tables in [BENCHMARKS.md](https://github.com/sshah03/shrewd/blob/main/BENCHMARKS.md). The repo includes pipeline examples and panel
rebuilds. Some research results came from private experiment scripts that aren't included.

---

# Part 1: Using it

## Install

```
pip install "shrewd[teacher]"      # training: LLM client + prompt optimizer, tfidf student
pip install shrewd                 # inference only: sklearn stack, no LLM libraries
pip install "shrewd[teacher,embed]"     # + model2vec static-embedding student (no torch)
pip install "shrewd[teacher,setfit]"    # + SetFit student (torch, best few-shot)
pip install "shrewd[teacher,encoder]"   # + fine-tuned ModernBERT (torch, strongest at 1k+ rows)
```

## Compile a classifier

One label per document, a locked test set, and a report that says what to fix. Start here
if your problem is "which of these N buckets does this go in".

```
import pandas as pd
from shrewd import Project

proj = Project(
    "runs/tickets",
    instructions="Classify customer support tickets by the customer's primary intent.",
    labels={"billing": "charges, invoices, refunds", "bug": "something is broken",
            "cancellation": "wants to cancel, pause, or downgrade", "other": "none of the above"},
    teacher="anthropic",   # or "openai", or any litellm model string
)
proj.add_seed(pd.read_csv("labeled.csv"))            # columns: text, label
proj.optimize(budget=800)                            # prompt optimization on the dev split
proj.label(pd.read_csv("unlabeled.csv"))             # dry_run=True prices it first
result = proj.distill(student="tfidf")
print(result.report())
```

```
from shrewd import load

clf = load("runs/tickets")
clf.predict(["I was double charged last month"])   # -> ["billing"]
clf.predict_proba(["..."])                         # ndarray in clf.classes_ order
clf.predict(texts, min_confidence=0.52)            # None below the threshold: route to a person
```

To try it without an API key: `python examples/make_demo_data.py`, then
`python examples/quickstart.py`. It replays a bundled cache of a real run, so no API calls.

Options:

- **Label under a budget.** `proj.label(pool, budget_usd=20)` (or `n=2000`) acquires rows in
  rounds, picking the rows a quick tfidf probe is least sure about instead of going in file
  order. In my tests that took 1.1x to 2.3x fewer teacher calls for the same student
  accuracy. Re-running continues where it stopped.
- **Fix labels by hand.** After `label()`, `needs_review.csv` lists the rows the teachers
  disagreed on or could not answer, with an empty `human_label` column. Fill in the ones
  you care about and the next `distill()` uses them at confidence 1.0. Pass
  `teacher=[...]` with two or more models, or `votes=3`, to get disagreements to review.
  With one teacher and one vote the queue stays empty.
- **Pick a student.** `"tfidf"` (default, instant), `"embed"` (static embeddings), `"encoder"`
  (fine-tuned ModernBERT), `"setfit"` (contrastive fine-tuning), or your own object with
  `fit` / `predict` / `predict_proba` / `classes_` / `save`.

```
rows = proj.compare(students=["tfidf", "embed", "encoder"])   # dev F1 vs latency vs size
proj.promote(rows[0]["student"])
result = proj.distill(student=rows[0]["student"])             # the one test-set evaluation
```

`autotune(dir, instructions, labels, seed_df, pool_df, budget_usd=...)` does this for you
under a spending cap. It starts with the cheapest teacher and moves up a tier only while
the dev score misses `target` and the next tier fits what's left of the budget. It scores
the locked test set once, on the winner, and writes every step to `autotune_trail.json`.

`teacher` takes any [litellm](https://docs.litellm.ai/docs/providers) model string.
`"anthropic"` and `"openai"` pick a strong default for that provider. **API keys come from
the environment**, the way litellm reads them: `export ANTHROPIC_API_KEY=...`,
`OPENAI_API_KEY`, `DEEPSEEK_API_KEY`, and so on. shrewd never takes a key as an argument or
writes one to disk, and `load()` doesn't need one.

`optimize()`, `label()` and `judge()` send your seed and pool text to the provider, so if
your text is sensitive, read their data-use terms first. The prompt goes out as a system
message marked for the provider's prompt cache (Claude charges ~10% of the input price for a
cached prefix), so a long prompt or many questions add little per document. Every response
is also cached locally in `cache.db`, so re-running a stage never pays twice.

**When to use something else:** if you won't run many documents through it, calling an
LLM or a hosted decision model may cost less than building and maintaining a student. Where
the line sits depends on your labeling, inference and retraining costs. It's also a poor fit
for tasks that need reasoning or fresh world knowledge per item, or for domains that change
quickly. The saved model only answers the task it was trained for.

## Compile a calibrated decision panel

The second path builds a fixed panel of typed questions (pick one of N, yes/no with a
probability, a rating on an ordered scale), all answered together in one pass. The
probabilities are calibrated against the teacher's answers and then scored against your
hand labels, because calibrating to the teacher alone doesn't guarantee
that a 0.8 means 80% on your task.

```
import pandas as pd
from shrewd import Decisions, Choice, Noul, Score

d = Decisions(
    "runs/tickets",
    instructions="These are customer support tickets.",
    questions={
        "department": Choice(
            instructions="Which team should handle this?",
            criteria={"billing": "charges, invoices, refunds",
                      "bug": "something is broken",
                      "other": "none of the above"}),
        "angry":    Noul(instructions="Does the customer sound angry?"),
        "severity": Score(instructions="How severe is this?",
                          criteria=["cosmetic", "workaround exists", "blocking"]),
    },
    teacher="anthropic/claude-fable-5-1",     # any litellm model string, or "anthropic" / "openai"
)
d.add_seed(pd.read_csv("labeled.csv"))    # text + one column per question id (blank where unknown)
d.judge(pd.read_csv("unlabeled.csv"))     # ONE teacher call answers every question per document
print(d.distill().report())               # trains, calibrates, scores against your hand labels
```

```
from shrewd import load

dec = load("runs/tickets")
dec.decide("I've been charged twice and nobody will call me back")
# {"department": <Answer department choice='billing'>,
#  "angry":      <Answer angry noul=0.91>,
#  "severity":   <Answer severity score=1.4>}
dec.predict_proba(texts)                  # {question: (n x options) calibrated probabilities}
```

There are three question types. `Choice` picks one of up to 255 named options and returns
a probability for each. `Noul` (borrowed from Jev) answers a yes/no question with one
number, the probability of yes. `Score` rates against 2 to 10 ordered levels and returns
the probability-weighted mean, so a document split between "cosmetic" and "blocking" lands
in the middle. Describe each option in `criteria`. The teacher and the zero-shot stack read
those descriptions, and bare option names don't tell them much.

Options:

- **Fix the teacher's answers.** `judge()` writes `needs_review.csv` listing every
  (document, question) where the teacher's winning option is below 0.6, with an empty
  `human_answer` column. Write an option name, yes/no, or a level's number or description.
  the next `distill()` (or `apply_review()`) uses it as a certain target. Your answers are
  ledgered in `reviewed.csv` and survive any re-judging.
- **Optimize the prompt.** `d.optimize(budget=800)` runs GEPA over the free text above the
  questions, scored on your dev split by a proper scoring rule (it rewards accurate
  probabilities as well as accurate labels).
- **Choose the student.** `distill(features="auto")` fits tf-idf and static embeddings and
  keeps whichever wins on your hand-labeled dev split (0.01-0.5 ms/doc). `features="encoder"`
  fine-tunes ModernBERT (`pip install "shrewd[encoder]"`, 3-30 ms/doc, ~20-45 min of GPU
  or Apple-silicon training, ~11 GB RAM): use it when the question needs the model to
  read rather than count words. Part 2 covers where each one wins.
- **Add a zero-shot second opinion.** `distill(zero_shot=True)` stacks an NLI model onto
  each head and keeps it per question only where it beats the plain head on your hand-labeled
  dev split. It helped on sentiment and emotion and got rejected almost everywhere else.
  20-800 ms/doc.
- **Use several teachers.** `backend=EnsembleBackend(["anthropic/claude-fable-5-1", "openai/gpt-6-astra"])` averages their distributions. `backend="logprobs"` reads token
  probabilities where the API exposes them (OpenAI-compatible endpoints).

Adding a question later is cheap. Answers are cached per `(model, question, document)`, so
a ninth question reuses the eight you already paid for. Changing a question's text or
options throws away its cached answers. The compiled artifact answers exactly the questions it was
compiled for.

### Try a pre-built panel

Four decision panels for softwa
sshah0330
๐ŸŸง echo.github โญReleased a pipeline for turning LLM judgments into local classifiers and fixed decision panels, with measurements and explicit limitations osshah03โ€”โ€”

Interpretation history

Decision trace