2026-10-11 16:38 UTC

The authors of the SMITH framework (accepted to NeurIPS 2026) claim that jointly training tool creation and tool use in a single policy via reinforcement learning enables a 4B Qwen3 model to achieve 79.9% macro-average accuracy on held-out procedural reasoning tasks and transfer tools to a 350M student model, outperforming inference-time tool-creation baselines; if replicated, this would establish joint tool-creation/use training as a superior paradigm for agent tool generalization.

state: seedheat: mediumuncertainty: mediumconvergesscott: highagent-harnesses tool-use-generalization rl-trainingZhi Rui TamChieh-Yen LinYun-Nung (Vivian) ChenShao-Hua SunHung-yi Lee

What is this?

The web search results do not surface the SMITH framework paper or its NeurIPS 2026 project page. Results return generic Qwen3 model announcements, NeurIPS workshop/downloads listings, and unrelated tool-use surveys — none mention 'SMITH', 'Schema-grounded Multi-task Iterative Tool Honing', or the claimed 79.9% accuracy on procedural reasoning with a 4B Qwen3 model. The case's evidence titles appear to be first-party project page content, but the web search cannot verify the paper's existence, acceptance, or claims. Without the actual paper or project page in the supplied material, the hypothesis remains ungrounded in external evidence.

Why it matters to Scott

The radar already tracks the core hypothesis as an open question (radar:cross-model-tool-creation-transfer: 'Independent replication will determine whether jointly training language models to create and use tools produces tools that improve agent performance beyond the model that created them'). This SMITH paper claims to be that replication with specific numbers (79.9% on procedural reasoning, 4B→350M transfer). Scott's canon carries the load-bearing frameworks this would validate: epistemic tool forging (tools as disposable cognition apparatus), runtime capability synthesis (agents manufacturing missing capabilities mid-mission), the Three Ingredients Framework (Tools ground), agent-native computing, and code-first architecture. The cross-model transfer claim also bears on model dividend, capability symmetry, and model perishability — if tools trained on a 4B model genuinely improve a 350M student, the compiled tool knowledge becomes the durable asymmetry across model swaps.
ip:concept.epistemic-tool-forgingip:concept.runtime-capability-synthesisip:framework.three-ingredients-frameworkip:framework.agent-native-computingip:framework.code-first-architectureip:concept.model-dividendip:concept.capability-symmetryip:concept.model-perishabilityip:concept.verification-loopsdev:concept.agentic-tool-loopdev:concept.llm-self-play-refinementradar:cross-model-tool-creation-transferradar:concept.tool-useradar:concept.agentic-rlradar:concept.model-distillationradar:concept.agent-benchmarksradar:concept.qwenradar:pcss-zebra-puzzle-reasoning-transferradar:pccg-qwen3-continuation-control
queries asked of Scott's wikis
  • joint tool creation and use training RL paradigm
  • tool generalization across model sizes distillation
  • procedural reasoning benchmarks agent harnesses
  • RL-trained tool use vs inference-time tool creation
  • Qwen3 4B model agent tool use capabilities
  • schema-grounded tool honing iterative training

Measured heat

now 0 pts/hpeak 12 pts/hcomments 0/hpeers p59momentum: steady2 platformsage 6h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

10-11 10:55 (minted)⭐ origin echo-reconstructedSMITH (Schema-grounded Multi-task Iterative Tool Honing) is a reinforcement learning framework that jointly trains tool creation and tool us
Zhi Rui Tam, Chieh-Yen Lin, Yun-Nung (Vivian) Chen, Shao-Hua Sun, Hung-yi Lee on paper (echo) · attributed from hn.story.50041626 · published time unknown
—
10-11 10:29first on hacker news · published · lag ?Training LLMs to write tools generalized beyond self use
blackcat201
—
10-11 10:29amplified on hacker news 👑hn.story.50041626
blackcat201
peak 2 · 0 comments · 98% of case engagement
10-11 10:33our radar first saw it · lag ?discovery anchor: hn.story.50041626—
pace: p41 vs 875 stories at the 3h mark (now 6h old) — ahead of aafp-commons-signed-agent-notebook (2.0x), behind agent-chaperone-jev-tool-screening (0.7x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnTraining LLMs to write tools generalized beyond self use
Retrieved article excerpt

Open article · Retrieved 2026-10-11T10:35:43.110842+00:00

Accepted to NeurIPS 2026 (Poster)

# Joint Optimization of Tool Creation and Use for Large Language Model Agents

SMITH: **S**chema-grounded **M**ulti-task **I**terative **T**ool **H**oning

- [Zhi Rui Tam](https://zrt.wtf)1,2
- Chieh-Yen Lin1
- Yun-Nung (Vivian) Chen2
- Shao-Hua Sun1,2
- Hung-yi Lee2

1Appier AI Research   2National Taiwan University

[Read the Paper](https://arxiv.org/abs/2608.24571)
[View Code](https://github.com/appier-research/smith)

SMITH training loop: the same model invents a tool, uses it to solve problems, learns from the results, and gets better.


SMITH is a training loop, not a fixed pipeline: the *same* policy invents a tool, uses it to solve problems, and is optimized on whether that use succeeds. That feedback is what keeps tool creation improving over time.

## Abstract

Tool-augmented language models are bounded by the APIs humans bothered to write; existing tool-creation
systems patch this by prompting a frozen LLM at inference time, leaving the model that writes a tool
decoupled from the one that uses it, with no signal that the schemas it produces are schemas it can
actually invoke. We propose **SMITH** (Schema-grounded Multi-task Iterative Tool Honing), a
reinforcement learning framework that jointly trains tool creation and tool use inside a single policy.
Each rollout is either a *build* task (write a tool from a few examples) or a *use* task
(invoke a pooled tool on a held-out question). Three separate reward axes catch schema, code, and
outcome failures independently, so each failure mode contributes its own gradient. A 4B Qwen3 trained
with SMITH on 13 procedural reasoning tasks with exact verifiers reaches **79.9**
macro-average accuracy on held-out tasks, the best across all evaluated methods and ahead of an
untrained 30B-A3B tool-writer. It also reaches **40.4** on TabMWP-Hard and
**42.6** on out-of-domain GQA (**+7.6** over the best same-backbone
inference-time baseline), without any visual or tabular training data. When invoked by a frozen 350M
student, tools written by our 4B match those produced by a writer an order of magnitude larger. The
same recipe also lifts Qwen3-8B and Granite-3.3-8B without modification.

79.9RG (Unseen) macro accuracy, best of all methods

32×fewer output tokens than standard CoT

+7.6points on out-of-domain GQA vs. best baseline

42.9held-out accuracy for a 350M model using SMITH's tools

## Why tool creation and tool use need to be trained together

Tool-augmented LLMs are only as capable as the tools someone already wrote for them: a calculator,
a search API, a Python sandbox. When the right tool doesn't exist, the agent is stuck. Recent work
(LATM, CRAFT, TroVE, KTCE) lets a model synthesize new tools on the fly, but almost always with a
powerful model *writing* the tool and a separate, weaker model *using* it. The writer
never finds out whether its interface was actually easy to call.

Decoupled pipelines

1. Large frozen LLM writes a tool at inference time
2. A different, weaker model tries to invoke it
3. Ambiguous schema → wrong call → wrong answer
4. ∇No gradient: the writer never learns its schema failed

→

SMITH: one closed loop

1. Same policy writes the tool (code + JSON schema)
2. Same policy invokes it later from the schema alone
3. Reward is computed from whether that use succeeded
4. ∇Gradient flows straight back to the tool writer

This creates two concrete training problems the paper has to solve. **Reward decomposition:**
a tool can fail because its *code* is wrong, its *schema* is wrong, the two disagree with
each other, or the tool is technically correct but poorly designed. Each needs a different corrective
signal. **Circular evaluation:** scoring tool quality needs a judge, but a model judging
its own live weights is unreliable, and a frozen external judge never improves alongside the policy.

## SMITH: Schema-grounded Multi-task Iterative Tool Honing

SMITH is a multi-task RL framework, trained with [DAPO](https://tool-use-smith.github.io/) (a clip-higher
variant of GRPO), that mixes two rollout types into every batch: **build** and
**use**. Both are optimized inside the *same* policy, so gradients from tool
creation and tool consumption update the same weights every step.

SMITH diagram: a build task produces a Python tool and JSON schema, tools are validated and stored in a pool, and a use task later invokes a pooled tool on a held-out question, feeding an execution-based reward back to the policy.


The build task synthesizes a tool from a handful of examples and validates it against held-out
questions; tools that pass are stored in a shared pool. The use task later receives *only*
the JSON schema, not the code, and must call the right tool to answer a new question. Because
the same model both writes and later invokes each tool, an ambiguous or broken schema is penalized
directly, a feedback loop inference-time prompting cannot provide.

### Task 1Build: write the tool

The policy sees **N = 4** question–answer pairs and must infer the general
procedure behind them, then express it as an OpenAI-compatible `(Python function, JSON
schema)` pair. The tool is then run against **K = 16** held-out questions
drawn from a *harder* difficulty band than the examples it was built from; the model never
sees the ground-truth answers at generation time. A tool that only pattern-matches the easy
induction examples scores near zero; only a genuinely reusable abstraction survives.

Walk through an example

4 in-context examples (easy band)

- **Q:** How many 1-bits in the binary form of 42? **A:** 3
- **Q:** How many 1-bits in the binary form of 255? **A:** 8
- … 2 more

Tool the policy writes

```
def solve(question: str) -> str:
    n = int(re.search(r"\d+", question).group())
    return str(bin(n).count("1"))
```

```
{
  "name": "solve",
  "description": "Counts the 1-bits in the
    binary form of the integer named in
    the question.",
  "parameters": {
    "type": "object",
    "properties": {
      "question": { "type": "string" }
    },
    "required": ["question"]
  }
}
```

**13 / 16** held-out (harder) questions correct → admitted to the tool pool

### Task 2Use: call the tool

The policy receives a target question and a tool pool entry: **one** matching
domain tool plus **two** distractor tools from unrelated categories, forcing it to
identify the right schema. It has up to **T = 5** turns to call the tool and answer;
an efficiency penalty discourages burning through the turn budget. If no pool tool exists yet for
the category, the model must build one first from the same in-context examples.

Try it: pick the right tool

Target question

“How many 1-bits are in the binary form of 92,401?”

Pool entry — every schema is named `solve(question)`; only the description tells them apart

solve(question)
Counts the 1-bits in the binary form of the integer named in the question.

solve(question)
Computes the number of days between two calendar dates named in the question.

solve(question)
Finds the real roots of a quadratic equation described in the question.

### Three reward axes, kept separate on purpose

Rather than collapsing everything into one score, SMITH keeps **format**,
**execution**, and **judge** feedback as independent axes fed to DAPO, so
each failure mode gets its own gradient instead of being averaged away.

#### Format reward

Checks the response contains exactly one Python block and one JSON block whose function name
and parameters actually match. A malformed pair terminates the rollout with zero reward on every
axis.

$$r^{\mathrm{fmt}} \in \{0,\ 0.5\}$$

#### Evaluation reward

The tool is handed to an evaluator model, which must call it to answer each held-out question.
Only answers that came through a real tool call count. A model can't shortcut this by reasoning
the answer out in text.

$$r^{\mathrm{eval}} = \frac{1}{|\mathcal{T}|}\sum\_{j=1}^{|\mathcal{T}|} \mathbf{1}\!\left[\pi^{\mathrm{eval}}(q\_j \mid \mathcal{C}, \mathcal{S}) \approx a\_j\right]$$

#### Judge reward

An LLM judge scores code correctness, schema quality, and overall quality, kept as a
*separate* axis from execution reward, not folded in. A syntax error is penalized directly;
a schema/code mismatch halves the score.

$$r^{\mathrm{judge}} = \begin{cases}-0.5 & \text{syntax error}\\ 0.5\, s\_{\mathrm{overall}} & \text{schema}\neq\text{code}\\ s\_{\mathrm{overall}} & \text{otherwise}\end{cases}$$

**Breaking the circularity of self-judging.** The evaluator \(\pi^{\mathrm{eval}}\)
and judge \(\pi^{\mathrm{judge}}\) both start from the same base checkpoint as the policy, then are
periodically re-synced to the latest policy weights. This lets the evaluator improve alongside the
policy without the instability of scoring against live, still-updating weights.

### Use-task correctness, with an efficiency penalty

Let \(c \in \{0,1\}\) mark whether the final answer matches ground truth. The reward is scaled by an
efficiency multiplier \(\eta(\rho)\) that decays as the turn fraction \(\rho = \min(n/T, 1)\) grows,
with a floor \(\eta\_{\min}=0.3\) so the policy is never indifferent to correctness even at the turn
limit (\(\eta\_{\mathrm{mid}} = 0.7\)):

$$
r^{\mathrm{correct}} = 2c\,\eta(\rho), \qquad
\eta(\rho) = \begin{cases}
1 - 2(1-\eta\_{\mathrm{mid}})\rho & \rho \le 0.5 \\[2pt]
\max\!\bigl(\eta\_{\min},\ \eta\_{\mathrm{mid}}\,(1-2(\rho-0.5))^2\bigr) & \rho > 0.5
\end{cases}
$$

Every training batch is split evenly, \(\mathcal{B} = \mathcal{B}\_{\mathrm{build}} \sqcup
\mathcal{B}\_{\mathrm{use}}\) with \(|\mathcal{B}\_{\mathrm{build}}| = |\mathcal{B}\_{\mathrm{use}}| =
B/2\), and DAPO accumulates both losses in a single backward pass:
\(\mathcal{L} = \mathcal{L}\_{\mathrm{DAPO}}(\mathcal{B}\_{\mathrm{build}}) + \mathcal{L}\_{\mathrm{DAPO}}
(\mathcal{B}\_{\mathrm{use}})\). Build and use gradients therefore update the same parameters
\(\theta\) every step. Any tool with \(r^{\mathrm{eval}} > 0\) is pushed into a shared
**Tool Pool** for reuse by future use-task rollouts. This is the "iterative honing" in
SMITH's name.

## Experimental setup

SMITH is trained on **13 procedural task categories** from
[Reasoning-Gym](https://github.com/open-thought/reasoning-gym), spanning arithmetic, algorithms, algebra,
games, and logical reasoning, chosen because their answers are exact and automatically verifiable and
each exposes a difficulty curriculum. Tools are induced on *easy* examples but graded on the
*hardest* band of the same task family, deliberately separating tool-writing quality from
instance difficulty.

13RG training task categories

10fully held-out RG categories

N=4 / K=16build examples / held-out eval questions

T=5max tool-use turns per rollout

The primary backbone is **Qwen3-4B-Instruct**, fine-tuned with LoRA (r = 64, α = 128)
using DAPO for 60 gradient steps, build/use rollouts mixed 1:1. The recipe is also applied unmodified
to **Qwen3-8B** and **Granite-3.3-8B** to test generality across model
families. Baselines share the Qwen3-4B-Instruct backbone wherever possible: inference-time tool
writers **LATM**, **CRAFT**, **TroVE**, and **KTCE**;
distillation baselines **ReTool** (from Qwen-32B traces) and **LATM**
(distilled from GPT-4.1); and a deliberate scaling probe, **LATM on Qwen3-30B-A3B**.
Transfer is measured on **TabMWP-Hard** (a strengthened tabular-reasoning benchmark) and
**GQA** (visual question answering), neither seen during training, and on
**BFCL v4**, an externally specified function-calling benchmark.

## Results

### Main results on Reasoning-Gym

SMITH is the only method that leads on genuinely held-out tasks while also using the fewest tokens.
Against distillation, a 4B model trained with SMITH's RL objective generalizes more reliably than 4B
models distilled from far larger oracles: ReTool leads on seen tasks (92.2) but drops nearly 30 points
on unseen ones, a sign of overfitting to the demonstrator's distribution rather than learning a
transferable build-and-u
blackcat20120
🟧 echo.paper ⭐SMITH (Schema-grounded Multi-task Iterative Tool Honing) is a reinforcement learning framework that jointly trains tool creation and tool usZhi Rui Tam, Chieh-Yen Lin, Yun-Nung (Vivian) Chen, Shao-Hua Sun, Hung-yi Lee——

Interpretation history

Decision trace