Retrieved article excerpt
Open article · Retrieved 2026-10-11T10:35:43.110842+00:00
Accepted to NeurIPS 2026 (Poster)
# Joint Optimization of Tool Creation and Use for Large Language Model Agents
SMITH: **S**chema-grounded **M**ulti-task **I**terative **T**ool **H**oning
- [Zhi Rui Tam](https://zrt.wtf)1,2
- Chieh-Yen Lin1
- Yun-Nung (Vivian) Chen2
- Shao-Hua Sun1,2
- Hung-yi Lee2
1Appier AI Research 2National Taiwan University
[Read the Paper](https://arxiv.org/abs/2608.24571)
[View Code](https://github.com/appier-research/smith)
SMITH training loop: the same model invents a tool, uses it to solve problems, learns from the results, and gets better.
SMITH is a training loop, not a fixed pipeline: the *same* policy invents a tool, uses it to solve problems, and is optimized on whether that use succeeds. That feedback is what keeps tool creation improving over time.
## Abstract
Tool-augmented language models are bounded by the APIs humans bothered to write; existing tool-creation
systems patch this by prompting a frozen LLM at inference time, leaving the model that writes a tool
decoupled from the one that uses it, with no signal that the schemas it produces are schemas it can
actually invoke. We propose **SMITH** (Schema-grounded Multi-task Iterative Tool Honing), a
reinforcement learning framework that jointly trains tool creation and tool use inside a single policy.
Each rollout is either a *build* task (write a tool from a few examples) or a *use* task
(invoke a pooled tool on a held-out question). Three separate reward axes catch schema, code, and
outcome failures independently, so each failure mode contributes its own gradient. A 4B Qwen3 trained
with SMITH on 13 procedural reasoning tasks with exact verifiers reaches **79.9**
macro-average accuracy on held-out tasks, the best across all evaluated methods and ahead of an
untrained 30B-A3B tool-writer. It also reaches **40.4** on TabMWP-Hard and
**42.6** on out-of-domain GQA (**+7.6** over the best same-backbone
inference-time baseline), without any visual or tabular training data. When invoked by a frozen 350M
student, tools written by our 4B match those produced by a writer an order of magnitude larger. The
same recipe also lifts Qwen3-8B and Granite-3.3-8B without modification.
79.9RG (Unseen) macro accuracy, best of all methods
32×fewer output tokens than standard CoT
+7.6points on out-of-domain GQA vs. best baseline
42.9held-out accuracy for a 350M model using SMITH's tools
## Why tool creation and tool use need to be trained together
Tool-augmented LLMs are only as capable as the tools someone already wrote for them: a calculator,
a search API, a Python sandbox. When the right tool doesn't exist, the agent is stuck. Recent work
(LATM, CRAFT, TroVE, KTCE) lets a model synthesize new tools on the fly, but almost always with a
powerful model *writing* the tool and a separate, weaker model *using* it. The writer
never finds out whether its interface was actually easy to call.
Decoupled pipelines
1. Large frozen LLM writes a tool at inference time
2. A different, weaker model tries to invoke it
3. Ambiguous schema → wrong call → wrong answer
4. ∇No gradient: the writer never learns its schema failed
→
SMITH: one closed loop
1. Same policy writes the tool (code + JSON schema)
2. Same policy invokes it later from the schema alone
3. Reward is computed from whether that use succeeded
4. ∇Gradient flows straight back to the tool writer
This creates two concrete training problems the paper has to solve. **Reward decomposition:**
a tool can fail because its *code* is wrong, its *schema* is wrong, the two disagree with
each other, or the tool is technically correct but poorly designed. Each needs a different corrective
signal. **Circular evaluation:** scoring tool quality needs a judge, but a model judging
its own live weights is unreliable, and a frozen external judge never improves alongside the policy.
## SMITH: Schema-grounded Multi-task Iterative Tool Honing
SMITH is a multi-task RL framework, trained with [DAPO](https://tool-use-smith.github.io/) (a clip-higher
variant of GRPO), that mixes two rollout types into every batch: **build** and
**use**. Both are optimized inside the *same* policy, so gradients from tool
creation and tool consumption update the same weights every step.
SMITH diagram: a build task produces a Python tool and JSON schema, tools are validated and stored in a pool, and a use task later invokes a pooled tool on a held-out question, feeding an execution-based reward back to the policy.
The build task synthesizes a tool from a handful of examples and validates it against held-out
questions; tools that pass are stored in a shared pool. The use task later receives *only*
the JSON schema, not the code, and must call the right tool to answer a new question. Because
the same model both writes and later invokes each tool, an ambiguous or broken schema is penalized
directly, a feedback loop inference-time prompting cannot provide.
### Task 1Build: write the tool
The policy sees **N = 4** question–answer pairs and must infer the general
procedure behind them, then express it as an OpenAI-compatible `(Python function, JSON
schema)` pair. The tool is then run against **K = 16** held-out questions
drawn from a *harder* difficulty band than the examples it was built from; the model never
sees the ground-truth answers at generation time. A tool that only pattern-matches the easy
induction examples scores near zero; only a genuinely reusable abstraction survives.
Walk through an example
4 in-context examples (easy band)
- **Q:** How many 1-bits in the binary form of 42? **A:** 3
- **Q:** How many 1-bits in the binary form of 255? **A:** 8
- … 2 more
Tool the policy writes
```
def solve(question: str) -> str:
n = int(re.search(r"\d+", question).group())
return str(bin(n).count("1"))
```
```
{
"name": "solve",
"description": "Counts the 1-bits in the
binary form of the integer named in
the question.",
"parameters": {
"type": "object",
"properties": {
"question": { "type": "string" }
},
"required": ["question"]
}
}
```
**13 / 16** held-out (harder) questions correct → admitted to the tool pool
### Task 2Use: call the tool
The policy receives a target question and a tool pool entry: **one** matching
domain tool plus **two** distractor tools from unrelated categories, forcing it to
identify the right schema. It has up to **T = 5** turns to call the tool and answer;
an efficiency penalty discourages burning through the turn budget. If no pool tool exists yet for
the category, the model must build one first from the same in-context examples.
Try it: pick the right tool
Target question
“How many 1-bits are in the binary form of 92,401?”
Pool entry — every schema is named `solve(question)`; only the description tells them apart
solve(question)
Counts the 1-bits in the binary form of the integer named in the question.
solve(question)
Computes the number of days between two calendar dates named in the question.
solve(question)
Finds the real roots of a quadratic equation described in the question.
### Three reward axes, kept separate on purpose
Rather than collapsing everything into one score, SMITH keeps **format**,
**execution**, and **judge** feedback as independent axes fed to DAPO, so
each failure mode gets its own gradient instead of being averaged away.
#### Format reward
Checks the response contains exactly one Python block and one JSON block whose function name
and parameters actually match. A malformed pair terminates the rollout with zero reward on every
axis.
$$r^{\mathrm{fmt}} \in \{0,\ 0.5\}$$
#### Evaluation reward
The tool is handed to an evaluator model, which must call it to answer each held-out question.
Only answers that came through a real tool call count. A model can't shortcut this by reasoning
the answer out in text.
$$r^{\mathrm{eval}} = \frac{1}{|\mathcal{T}|}\sum\_{j=1}^{|\mathcal{T}|} \mathbf{1}\!\left[\pi^{\mathrm{eval}}(q\_j \mid \mathcal{C}, \mathcal{S}) \approx a\_j\right]$$
#### Judge reward
An LLM judge scores code correctness, schema quality, and overall quality, kept as a
*separate* axis from execution reward, not folded in. A syntax error is penalized directly;
a schema/code mismatch halves the score.
$$r^{\mathrm{judge}} = \begin{cases}-0.5 & \text{syntax error}\\ 0.5\, s\_{\mathrm{overall}} & \text{schema}\neq\text{code}\\ s\_{\mathrm{overall}} & \text{otherwise}\end{cases}$$
**Breaking the circularity of self-judging.** The evaluator \(\pi^{\mathrm{eval}}\)
and judge \(\pi^{\mathrm{judge}}\) both start from the same base checkpoint as the policy, then are
periodically re-synced to the latest policy weights. This lets the evaluator improve alongside the
policy without the instability of scoring against live, still-updating weights.
### Use-task correctness, with an efficiency penalty
Let \(c \in \{0,1\}\) mark whether the final answer matches ground truth. The reward is scaled by an
efficiency multiplier \(\eta(\rho)\) that decays as the turn fraction \(\rho = \min(n/T, 1)\) grows,
with a floor \(\eta\_{\min}=0.3\) so the policy is never indifferent to correctness even at the turn
limit (\(\eta\_{\mathrm{mid}} = 0.7\)):
$$
r^{\mathrm{correct}} = 2c\,\eta(\rho), \qquad
\eta(\rho) = \begin{cases}
1 - 2(1-\eta\_{\mathrm{mid}})\rho & \rho \le 0.5 \\[2pt]
\max\!\bigl(\eta\_{\min},\ \eta\_{\mathrm{mid}}\,(1-2(\rho-0.5))^2\bigr) & \rho > 0.5
\end{cases}
$$
Every training batch is split evenly, \(\mathcal{B} = \mathcal{B}\_{\mathrm{build}} \sqcup
\mathcal{B}\_{\mathrm{use}}\) with \(|\mathcal{B}\_{\mathrm{build}}| = |\mathcal{B}\_{\mathrm{use}}| =
B/2\), and DAPO accumulates both losses in a single backward pass:
\(\mathcal{L} = \mathcal{L}\_{\mathrm{DAPO}}(\mathcal{B}\_{\mathrm{build}}) + \mathcal{L}\_{\mathrm{DAPO}}
(\mathcal{B}\_{\mathrm{use}})\). Build and use gradients therefore update the same parameters
\(\theta\) every step. Any tool with \(r^{\mathrm{eval}} > 0\) is pushed into a shared
**Tool Pool** for reuse by future use-task rollouts. This is the "iterative honing" in
SMITH's name.
## Experimental setup
SMITH is trained on **13 procedural task categories** from
[Reasoning-Gym](https://github.com/open-thought/reasoning-gym), spanning arithmetic, algorithms, algebra,
games, and logical reasoning, chosen because their answers are exact and automatically verifiable and
each exposes a difficulty curriculum. Tools are induced on *easy* examples but graded on the
*hardest* band of the same task family, deliberately separating tool-writing quality from
instance difficulty.
13RG training task categories
10fully held-out RG categories
N=4 / K=16build examples / held-out eval questions
T=5max tool-use turns per rollout
The primary backbone is **Qwen3-4B-Instruct**, fine-tuned with LoRA (r = 64, α = 128)
using DAPO for 60 gradient steps, build/use rollouts mixed 1:1. The recipe is also applied unmodified
to **Qwen3-8B** and **Granite-3.3-8B** to test generality across model
families. Baselines share the Qwen3-4B-Instruct backbone wherever possible: inference-time tool
writers **LATM**, **CRAFT**, **TroVE**, and **KTCE**;
distillation baselines **ReTool** (from Qwen-32B traces) and **LATM**
(distilled from GPT-4.1); and a deliberate scaling probe, **LATM on Qwen3-30B-A3B**.
Transfer is measured on **TabMWP-Hard** (a strengthened tabular-reasoning benchmark) and
**GQA** (visual question answering), neither seen during training, and on
**BFCL v4**, an externally specified function-calling benchmark.
## Results
### Main results on Reasoning-Gym
SMITH is the only method that leads on genuinely held-out tasks while also using the fewest tokens.
Against distillation, a 4B model trained with SMITH's RL objective generalizes more reliably than 4B
models distilled from far larger oracles: ReTool leads on seen tasks (92.2) but drops nearly 30 points
on unseen ones, a sign of overfitting to the demonstrator's distribution rather than learning a
transferable build-and-u