Programmatic Tool Calling

Intermediate

How AI agents write code that calls tools programmatically, reducing latency and token consumption.

Last updated: Sep 13, 2026

What is Programmatic Tool Calling?

Programmatic tool calling lets an AI agent write code that invokes tools inside a sandboxed execution environment — instead of requiring a separate model round-trip for every tool call. The agent writes a script, the runtime executes it, and tool calls happen directly from the code. Only the final result is returned to the model's context window.

Why It Matters

Code can coordinate dependent calls and filter results without another model decision at each step. Independent standard tool calls can already run in parallel; compare against that baseline too.

Fewer Round-Trips

Call multiple tools in a single code execution instead of one model turn per tool.

Lower Token Usage

Intermediate results stay in the sandbox — only the summary enters the context window.

Data Filtering

Process and filter large tool outputs in code before they reach the model.

Native Control Flow

Use loops, conditionals, and error handling — the model writes real code, not just JSON calls.

How It Works

The flow involves four steps between the agent, sandbox, and your tool server.

1

Agent Writes Code

The model generates a Python script that calls your tools as async functions.

2

Sandbox Executes

The code runs in a sandboxed container. When a tool function is called, execution pauses.

3

Tool Runs Externally

Your server receives the tool call, executes it, and returns the result to the sandbox.

4

Result to Model

Once the script finishes, only the final output is added to the model's context.

Traditional vs Programmatic

See how the two approaches differ for a task that queries three database regions.

Traditional Tool Use

With three sequential, dependent calls, the model is invoked before each tool and once more for the final answer: four model calls.

Model → tool call → result → Model → tool call → result → Model → tool call → result → Model → final answer

4 model calls (sequential example)

Programmatic Tool Calling

One model call writes the script; another interprets its returned result and answers. The script coordinates tool calls.

Model → code (3 tool calls + aggregation) → result → Model → final answer

2 model calls

Control: three independent standard tool calls can also run in parallel. That needs one model call to request them and one to summarize, like this PTC example. PTC can additionally filter and aggregate intermediate data in code, reducing what enters model context.

Example: Programmatic Database Query

The agent writes Python that loops over regions, calls a database tool, and aggregates results — all in one execution.

import json

regions = ["West", "East", "Central"]
results = {}

for region in regions:
    raw = await query_database({
        "sql": f"SELECT SUM(revenue) AS total FROM sales WHERE region='{region}'"
    })
    rows = json.loads(raw)
    results[region] = rows[0]["total"] or 0

print(json.dumps({
    "top_region": max(results, key=results.get),
    "total": sum(results.values())
}))

Use Cases

Programmatic tool calling shines when agents need to do more than one-shot tool calls.

Batch Processing

Query a database for each of 50 regions in a loop, aggregate results, and return a summary — all in one execution.

Conditional Logic

Check file size first, then decide whether to read the full file or just a summary. No wasted round-trips.

Data Filtering

Fetch 10,000 log entries, filter to only errors, and return the last 10 — keeping the context window clean.

Early Termination

Check endpoints in sequence and stop as soon as a healthy one is found. No need to check all of them.

The allowed_callers Concept

In Anthropic’s API, allowed_callers specifies the intended direct or versioned code-execution caller. This guides tool use but does not replace runtime authorization.

For clarity, it's best to choose one mode per tool rather than enabling both. This gives the model clearer guidance on how to use each tool.

Anthropic API example, checked September 2026. allowed_callers guides invocation; enforce access control in your tool runtime because it is not a hard API security boundary.

{
  "type": "code_execution_20260120",
  "name": "code_execution"
}

Direct Only

The model calls the tool directly via the standard tool-use flow. This is the default.

"allowed_callers": ["direct"]

Code Execution Only

Guide the model to invoke the tool from the named code-execution version. Validate the caller and permissions in your own tool handler.

"allowed_callers":
  ["code_execution_20260120"]

Both Modes

The tool can be called either directly or from code. Use sparingly — it can confuse tool selection.

"allowed_callers":
  ["direct", "code_execution_20260120"]

Key Takeaways

  • 1Programmatic tool calling lets agents write code that calls tools, eliminating per-tool round-trips.
  • 2Only the script output needs to enter model context; printing raw intermediate data forfeits that saving.
  • 3Use the provider’s versioned caller identifiers and enforce actual access rights in the runtime.
  • 4Best for batch processing, conditional workflows, data filtering, and multi-step tool chains.
  • 5Multiple AI providers are implementing this pattern as a way to make agents faster and more efficient.

Primary sources