2026-10-11 16:37 UTC

OpenAI claims its released Programmatic Tool Calling — a hosted Responses API tool where the model writes and runs sandboxed JavaScript to coordinate its own tool calls (parallel calls, loops, intermediate results) in one program instead of sequential tool rounds — becomes a default agent-orchestration pattern; adoption in agent workloads and imitation by competing providers would establish code-orchestration as the standard multi-tool agent mechanism.

state: corroboratedheat: highuncertainty: mediumconvergesscott: highopenai agent-harnesses tool-orchestration agent-orchestration inference-economicsOpenAI

What is this?

OpenAI released Programmatic Tool Calling (PTC) on July 9, 2026 as a hosted Responses API tool where models write and execute JavaScript in an isolated V8 sandbox to orchestrate tool calls — parallel invocations, loops, and intermediate result handling — in a single program turn instead of sequential round-trips. Enabled by default in the Agents API, it supports function/custom/MCP/shell/code_interpreter tools via per-tool `allowed_callers` opt-in and surfaces `program`/`program_output`/`function_call` items in the Responses envelope. Anthropic shipped a GA Python equivalent (code_execution_20260120) in November 2025 citing ~85% token reduction on agentic search benchmarks; Cloudflare Agents independently implemented "Code Mode" with TypeScript in sandboxed JS environments; BFCL v4 evaluation across 14 models finds programmatic beats JSON tool calling. OpenHands SDK has an open issue investigating PTC support. The contested question is adoption at scale as the default multi-tool agent mechanism, with practitioner discussion accelerating around Armin Ronacher's codemode explainer.

Why it matters to Scott

OpenAI shipping Programmatic Tool Calling as a default-on, first-party hosted primitive independently arrives at the exact pattern Scott's Code-First Architecture framework stakes — the framework page explicitly lists 'programmatic tool calling' as a recognized name for the architecture. This is a consequential provider (OpenAI) productizing the position Scott already argued, creating a dated-receipt validation opportunity and a directly testable comparison in his harness work (ask, superlever, all-in-one-software). The case also touches his provider-native-backend-execution concept (delegating to provider-hosted runtime) and the token-economics/inference-economics claims (~85% token reduction, BFCL v4 evidence).
ip:framework.code-first-architectureip:source.why-code-execution-beats-mcpip:concept.code-as-step-between-model-runsip:concept.hybrid-architecturedev:project.askdev:concept.provider-native-backend-executiondev:project.superleverdev:project.all-in-one-softwarework:concept.superrraiip:concept.agent-hands-and-eyesradar:speculative-programmatic-tool-callingradar:openai-agents-apiradar:concept.tool-orchestrationradar:concept.agent-harnessesradar:concept.agent-orchestrationradar:concept.inference-economicsradar:concept.program-synthesis
queries asked of Scott's wikis
  • code-first architecture
  • agent harness
  • tool orchestration
  • inference economics
  • programmatic tool calling
  • model sovereignty

Measured heat

now 0 pts/hpeak 14 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 290h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-29 14:00⭐ origin echo-reconstructedOpenAI's own documentation announces Programmatic Tool Calling: it 'lets a model write and run JavaScript that coordinates its tools. A prog
OpenAI on blog (echo) · attributed from hn.story.49906732
—
09-30 10:05first on hacker news · published · +20.1hProgrammatic Tool Calling
tosh
—
09-30 10:05amplified on hacker newshn.story.49906732
tosh
peak 1 · 0 comments · 0% of case engagement
10-06 13:41amplified on hacker news 👑hn.story.49978333
Tomte
peak 157 · 78 comments · 100% of case engagement
09-30 10:21our radar first saw it · +20.4hdiscovery anchor: hn.story.49906732—
10-11 01:36reached heat=high · +275.6h · via queue+ledger——
pace: p8 vs 1188 stories at the 168h mark (now 290h old) — behind addom-local-coding-harness (0.5x)

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnProgrammatic Tool Calling
Retrieved article excerpt

Open article · Retrieved 2026-09-30T10:25:02.114088+00:00

Programmatic Tool Calling lets a model write and run JavaScript that coordinates its tools. A program can call tools in parallel, use loops and conditions, and keep intermediate results in the hosted runtime. This is useful when a task needs a sequence of related tool calls or needs to process large tool outputs before returning a result.

In the Responses API, your application decides whether Programmatic Tool Calling is available and which eligible tools the model can call directly, from a program, or either way. It continues to run any client-owned tool calls. The [Agents API](https://developers.openai.com/api/docs/guides/tools-programmatic-tool-calling#agents-api) enables Programmatic Tool Calling by default and manages the agent loop for you.

Check the [model page](https://developers.openai.com/api/docs/models) before enabling Programmatic Tool Calling.

## Understand the runtime environment

OpenAI runs each generated program in a fresh, isolated V8 runtime. The runtime supports JavaScript with top-level `await`, but it does not provide Node.js, package installation, direct network access, a general-purpose filesystem, subprocess execution, a console, or persistent JavaScript state between program executions. Programs can interact with external systems only through tools enabled in the request and can emit output with `text(...)` or `image(...)`.

For Responses API requests, Programmatic Tool Calling supports Zero Data Retention (ZDR) workflows without requiring a persistent code-execution container. ZDR must be enabled for the organization or project; setting `store: false` enables stateless continuation but does not enable ZDR by itself. Eligibility and retention depend on the complete request, including its model, tools, and third-party services; see [data controls](https://developers.openai.com/api/docs/guides/your-data).

## Choose when to use Programmatic Tool Calling

Use Programmatic Tool Calling when a stage has predictable control flow and code can return a smaller structured result. Use direct tool calling when one call is sufficient, each result requires fresh model judgment, or the work requires approval or preservation of citations or native artifacts.

| Task shape | Recommended mode |
| --- | --- |
| A single lookup or action | Use direct tool calling. |
| Several results that code can filter, join, rank, remove duplicates from, aggregate, or validate | Use Programmatic Tool Calling when the program can return a smaller structured result. |
| Dependent calls with predictable data flow | Use Programmatic Tool Calling when code can derive later arguments and the limits and failure behavior are explicit. |
| Adaptive search or semantic evaluation | Use direct tool calling when each result should influence the model’s next decision. |
| Writes or approval-sensitive actions | Use direct tool calling by default to preserve a clear authorization boundary. |
| Final citation or native artifact validation | Use direct tool calling unless the program preserves the native output and validates every required item. |

## Configure Programmatic Tool Calling

For the Responses API, add the `programmatic_tool_calling` hosted tool to the request. Then set `allowed_callers` on each eligible tool that the program can invoke.

Enable Programmatic Tool Calling

```
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28[
  {
    "type": "function",
    "name": "get_inventory",
    "description": "Return an object with sku (string) and available_units (number).",
    "parameters": {
      "type": "object",
      "properties": {
        "sku": { "type": "string" }
      },
      "required": ["sku"],
      "additionalProperties": false
    },
    "output_schema": {
      "type": "object",
      "properties": {
        "sku": { "type": "string" },
        "available_units": { "type": "number" }
      },
      "required": ["sku", "available_units"],
      "additionalProperties": false
    },
    "allowed_callers": ["programmatic"]
  },
  {
    "type": "programmatic_tool_calling"
  }
]
```

`allowed_callers` controls how the model can invoke a tool:

| Value | Behavior |
| --- | --- |
| Omitted or `["direct"]` | The model can call the tool directly. |
| `["programmatic"]` | Only code in a `program` item can call the tool. |
| `["direct", "programmatic"]` | The model can call the tool directly or from a program. |

`parameters` describes the function arguments. When a function returns predictable structured data, `output_schema` describes the JSON object encoded in its `function_call_output.output` string. Define both so generated JavaScript can use the returned fields reliably.

### Supported tools

The following tool types support `allowed_callers: ["programmatic"]`:

- `function` and `custom`
- `mcp`
- `apply_patch`
- Local and hosted `shell`
- `code_interpreter`

For MCP tools, the tool’s `require_approval` policy can pause the program until you approve the call.

For OpenAI-hosted tools, review the tool’s data-retention and security guidance before enabling it in a program.

### Combine with tool search

[Tool search](https://developers.openai.com/api/docs/guides/tools-tool-search) runs as a top-level Responses API tool, not from inside generated JavaScript. Function, custom, and MCP tools with `defer_loading: true` are not initially available to a program. After the model loads a matching tool, a later program can invoke it through `tools.*` when its `allowed_callers` includes `"programmatic"`. An already-running program cannot invoke tool search, so the model must load deferred tools before starting a program that needs them.

## Guide routing when both modes are available

When your application lets the model call a function directly or from a program, assign each route to a specific workflow stage. Generic instructions such as “use Programmatic Tool Calling efficiently” don’t identify the intended boundary. For example:

```
<tool_orchestration>
Use Programmatic Tool Calling for [bounded stage] using only [eligible tools].
Run independent calls concurrently when safe. Use only documented tool input
and output fields.

Process and reduce the intermediate results, then emit exactly [program result shape],
including the evidence needed for the final answer.

Stop when [condition] is met. Retry transient failures at most [R] times.
Do not repeat completed calls or perform side-effecting actions. If a required
result is still missing, return a clear structured failure.

Use direct tool calls for [semantic judgment, approval, or final validation].
</tool_orchestration>
```

Here is an example of how to use this template:

```
<tool_orchestration>
Use Programmatic Tool Calling to compare inventory with demand for sku_123
using only get_inventory and get_demand. Run both calls concurrently. Use
only documented tool input and output fields.

Process and reduce the intermediate results, then emit exactly one JSON object
with sku, available_units, requested_units, and shortage_units, where
shortage_units is max(requested_units - available_units, 0). Include
available_units and requested_units as evidence for the calculation.

Stop when both tool results contain the required fields. Retry transient
failures at most 1 time. Do not repeat completed calls or perform
side-effecting actions. If a required result is still missing, return a clear
structured failure.

Use direct tool calls only for approval before any inventory-changing action.
</tool_orchestration>
```

For workflows that need both modes, define one handoff and avoid switching routes or repeating work. If a safe fallback exists, define it once and limit its retries.

## Understand program response items

Each API call still returns the standard [Responses API object](https://developers.openai.com/api/reference/resources/responses/methods/create). Programmatic Tool Calling doesn’t introduce a separate response envelope. When the model uses Programmatic Tool Calling, the response’s `output` array can contain:

- A `program` item containing the generated JavaScript, a `call_id`, and an opaque `fingerprint` used to resume or replay the program.
- A `function_call` item made by the program. It has its own `call_id`, which your application uses to return the function result. Its `caller.caller_id` matches the program’s `call_id`.
- A `program_output` item containing the program’s final result and status. Its `call_id` matches the program’s `call_id`, and its `status` is `completed` or `incomplete`.

These are separate top-level items in `response.output`; the `caller` field records their execution relationship.

For example, a program can pause while your application runs `get_inventory` and `get_demand`:

Program and nested function calls

```
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31[
  {
    "type": "program",
    "id": "prog_123",
    "call_id": "call_prog_123",
    "code": "const [stock, demand] = await Promise.all([tools.get_inventory({ sku: 'sku_123' }), tools.get_demand({ sku: 'sku_123' })]); text(JSON.stringify({ sku: stock.sku, available_units: stock.available_units, requested_units: demand.requested_units, shortage_units: Math.max(demand.requested_units - stock.available_units, 0) }));",
    "fingerprint": "opaque_replay_state"
  },
  {
    "type": "function_call",
    "id": "fc_123",
    "call_id": "call_inventory_123",
    "name": "get_inventory",
    "arguments": "{\"sku\":\"sku_123\"}",
    "caller": {
      "type": "program",
      "caller_id": "call_prog_123"
    }
  },
  {
    "type": "function_call",
    "id": "fc_456",
    "call_id": "call_demand_123",
    "name": "get_demand",
    "arguments": "{\"sku\":\"sku_123\"}",
    "caller": {
      "type": "program",
      "caller_id": "call_prog_123"
    }
  }
]
```

These examples show only the relevant items from `response.output`; they omit the surrounding standard Responses object. After your application returns the nested function results, a later response can contain the complete `program_output` item:

Program output

```
1
2
3
4
5
6
7{
  "type": "program_output",
  "id": "prog_out_123",
  "call_id": "call_prog_123",
  "result": "{\"sku\":\"sku_123\",\"available_units\":42,\"requested_units\":31,\"shortage_units\":0}",
  "status": "completed"
}
```

The JSON string in `program_output.result` follows the program result shape from your instructions. The surrounding `program_output` item follows the API contract shown above. These are separate contracts. A final `message` can arrive with the program output or in a later response, so continue until you receive that message.

OpenAI runs the model-generated JavaScript in the hosted runtime. Your application executes returned client-owned function calls; it does not execute the generated JavaScript.

Return the function result as a `function_call_output`. Copy `caller` from the function call without changing it. The service uses that value to resume the correct program.

## Continue after client-owned function calls

A program can pause more than once as it reaches client-owned tools. Continue until the response contains a final assistant message:

1. Send the request with the hosted tool and functions that allow programmatic calls.
2. Run every returned client-owned function call.
3. Return each function result with the original `call_id` and `caller`.
4. Handle an incomplete response before continuing.
5. If the response contains no pending `function_call` items and no final `message` item, continue from that response. With `store: false`, replay its output items; for a stored response, use `previous_response_id`.
6. Stop when the response contains a final `message` item. Read `response.output_text` or the message’s refusal content.

The following example uses `store: false`, preserves every response item, and returns each function result to the program:

Run a programmatic tool-calling loop

Python

```
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
tosh10
🟧 echo.blog ⭐OpenAI's own documentation announces Programmatic Tool Calling: it 'lets a model write and run JavaScript that coordinates its tools. A progOpenAI——
🟧 hnWhat Is CodemodeTomte15778

Interpretation history

Decision trace