Introducing advanced tool use on the Claude Developer Platform
by Anthropic
Three patterns for agents with hundreds of tools: tool search, programmatic tool calling and tool use examples.
Overview
Published on Anthropic's engineering blog on November 24, 2025 by Bin Wu with contributions from the Claude Developer Platform team, this guide addresses three failure modes of tool-heavy agents. Section one covers the Tool Search Tool: tools are marked defer_loading: true and Claude searches for them on demand instead of loading every definition up front; in Anthropic's example, 58 tools that consumed about 55K tokens drop to about 8.7K, and Opus 4.5 accuracy with tool search enabled rose from 79.5% to 88.1%. Section two covers Programmatic Tool Calling: tools flagged with allowed_callers are invoked from Python that Claude writes and runs in the code execution sandbox, so intermediate results never enter the context window. A budget-compliance walkthrough shows 200KB of expense data reduced to 1KB of results, and average tokens on complex research tasks fell from 43,588 to 27,297. Section three covers Tool Use Examples: an input_examples field that shows the model format conventions and parameter correlations JSON Schema cannot express, which lifted complex-parameter accuracy from 72% to 90%. A best-practices section maps each bottleneck to a feature and gives setup rules, such as keeping three to five high-use tools always loaded and providing one to five realistic examples per tool. It links to the official docs and two runnable notebooks in the Claude Cookbooks repository. As of 2026-10-01 the programmatic tool calling docs list the feature as generally available on the Claude API.
At a Glance
- Topic
- Agentic
- Level
- Intermediate
- Format
- Guide
- Cost
- Free
- Duration
- ~25 min read; linked docs and cookbook notebooks are optional hands-on follow-ups
- Provider
- Anthropic
- Hands-on
- Yes — code/exercises
- Certificate
- None
What You’ll Learn
- ✓Defer rarely used tool definitions with defer_loading and let Claude discover them through the Tool Search Tool
- ✓Defer whole MCP servers while keeping the three to five most-used tools always loaded
- ✓Let Claude orchestrate many tool calls from sandboxed Python code using allowed_callers and the code execution tool
- ✓Keep large intermediate tool results out of the context window so only final aggregates reach the model
- ✓Write input_examples on tool definitions to teach format conventions, nested structures and optional-parameter patterns
- ✓Diagnose whether context bloat, large intermediate results or parameter errors is your agent's actual bottleneck
- ✓Judge when each feature does not pay off, such as libraries under ten tools or single-call tasks
- ✓Preserve prompt caching while using deferred tools and track tool calls made from code through the caller field
Highlights
- •Every feature comes with measured before/after numbers (token counts and accuracy on named Claude models), not just API descriptions
- •Each feature has an explicit section on when it is NOT worth using, which most vendor launch posts leave out
- •Drew 673 points and 266 comments on Hacker News; practitioners welcomed programmatic tool calling, and skeptics questioned provider lock-in and compared it to CLI-first and smolagents-style code agents
- •Links to two runnable Claude Cookbooks notebooks (tool search with embeddings, programmatic tool calling) and the API reference docs
- •The patterns carry over to other stacks: deferred tool loading and code-mode orchestration also appear in Cloudflare Code Mode and the code-execution-with-MCP approach the post credits
Who It’s For
Best For
- ✓Engineers building agents that connect to several MCP servers or more than ten tools
- ✓Teams hitting context-window limits or tool-selection errors in production agents
- ✓Developers on the Claude API deciding whether to adopt tool search or programmatic tool calling
- ✓Architects comparing code-execution tool orchestration with classic one-call-per-turn tool use
Prerequisites
- •Working familiarity with LLM tool or function calling and JSON Schema tool definitions
- •Basic Python, to follow the code execution examples and the cookbook notebooks
- •Helpful but optional: experience connecting an agent to MCP servers
FAQ
What is Introducing advanced tool use on the Claude Developer Platform?
An Anthropic engineering guide for developers building agents that connect to large tool libraries or many MCP servers. It explains three Claude API features (Tool Search Tool, Programmatic Tool Calling and Tool Use Examples), when each one helps, and the measured token and accuracy gains, so you can cut context bloat and tool-call errors in your own agents.
Is Introducing advanced tool use on the Claude Developer Platform free?
Introducing advanced tool use on the Claude Developer Platform is free to access.
What level is Introducing advanced tool use on the Claude Developer Platform for?
Introducing advanced tool use on the Claude Developer Platform is aimed at a intermediate audience. Recommended background: Working familiarity with LLM tool or function calling and JSON Schema tool definitions, Basic Python, to follow the code execution examples and the cookbook notebooks, Helpful but optional: experience connecting an agent to MCP servers.
How long does Introducing advanced tool use on the Claude Developer Platform take?
Expect roughly ~25 min read; linked docs and cookbook notebooks are optional hands-on follow-ups. Most learners work through it at their own pace.
What will I learn from Introducing advanced tool use on the Claude Developer Platform?
You'll learn: Defer rarely used tool definitions with defer_loading and let Claude discover them through the Tool Search Tool; Defer whole MCP servers while keeping the three to five most-used tools always loaded; Let Claude orchestrate many tool calls from sandboxed Python code using allowed_callers and the code execution tool; Keep large intermediate tool results out of the context window so only final aggregates reach the model; Write input_examples on tool definitions to teach format conventions, nested structures and optional-parameter patterns; Diagnose whether context bloat, large intermediate results or parameter errors is your agent's actual bottleneck; Judge when each feature does not pay off, such as libraries under ten tools or single-call tasks; Preserve prompt caching while using deferred tools and track tool calls made from code through the caller field.
Topics
Sources
This page was written from 3 sources, 2 on domains other than anthropic.com.