Claude Tool Use Picks the Wrong Tool: Selection Errors vs Argument Errors
Why Claude calls the wrong tool when many similar tools are loaded, how selection errors differ from argument errors, what Anthropic's tool search and strict tool use do, and how to structure tools so selection stays reliable.
Key takeaways
- -Tool-use failures come in two kinds: selection errors, where Claude picks the wrong tool, and argument errors, where it picks the right tool and fills it in wrong. They have different fixes.
- -Anthropic's documentation says Claude's ability to pick the right tool degrades once more than 30 to 50 tools are available.
- -Anthropic names wrong tool selection and incorrect parameters as the most common tool-use failures, especially with similarly named tools such as notification-send-user and notification-send-channel.
- -The tool search tool with defer_loading keeps only a focused set of tools in context. Anthropic recommends it for 10 or more tools or more than 10,000 tokens of definitions.
- -Search adds its own failure mode: public Claude Code issues report deferred tools missed by search ranking even when the search was close to the exact name.
- -Swytchcode exposes a small fixed set of MCP tools no matter how large the API catalog is, finds operations by intent, and only runs operations enabled for the project.
Claude picks the wrong tool when it has too many similar options in context and too little to tell them apart. Anthropic's own documentation says selection accuracy degrades once more than 30 to 50 tools are available, and names wrong tool selection and incorrect parameters as the most common tool-use failures. These are two different problems. A selection error means the right tool was available and Claude chose another. An argument error means Claude chose correctly and filled in the wrong values. Selection errors are fixed by showing fewer, clearer tools. Argument errors are fixed by stricter schemas and validation.
This article covers how to tell the two apart, why selection degrades as tool counts grow, what Anthropic's tool search and strict tool use do, the failure modes they introduce, and how to structure tools so selection stays reliable. It applies to the Claude API, the Claude Agent SDK, and Claude Code with MCP servers. Checked against Anthropic's documentation and public Claude Code issues in October 2026.
Selection errors vs argument errors
| Question | Selection error | Argument error |
|---|---|---|
| What went wrong | Claude called a different tool than the task needed | Claude called the right tool with wrong or missing values |
| Typical example | notification-send-channel called when the user meant one person | notification-send-user called with a channel name in the user field |
| Main cause | Too many tools, overlapping names, vague descriptions | Loose schemas, unclear units and formats, missing examples |
| Main fixes | Fewer tools in context, distinct names, tool search, scoping | Strict schemas, input examples, validation, specific errors |
| How to spot it in logs | The tool name is wrong for the request | The tool name is right and the input is wrong |
Why selection gets worse as tools grow
- Every definition competes for attention. Each tool's name, description, and schema sits in context. With dozens of tools, the relevant one is a small part of a long list.
- Similar names blur together. Anthropic's engineering post on advanced tool use gives notification-send-user and notification-send-channel as an example of names that lead to wrong selection.
- MCP servers add up quickly. Connecting several servers can put hundreds of tools in context. Anthropic recommends tool search for systems that aggregate multiple MCP servers.
- Definitions cost tokens. Large tool sets take up context that would otherwise hold the task and its data.
What Anthropic provides
| Feature | What it does | Limit |
|---|---|---|
| Tool search tool | Claude searches for tools on demand. Tools marked defer_loading: true load only when search finds them | Adds a search step, and ranking can miss the right tool |
| defer_loading | Keeps rarely used tools out of the initial context while keeping them discoverable | Anthropic advises keeping the 3 to 5 most-used tools loaded, and the search tool itself must not be deferred |
| Strict tool use | Guarantees schema validation on tool names and inputs | Checks shape. Values can still be wrong |
| Tool use examples | Example inputs that show Claude how a tool is meant to be called | Guides the model without enforcing anything |
| Subagent tool scoping (Claude Code) | Agent definitions can restrict tools and scope MCP servers to one subagent | Applies to Claude Code subagents |
Anthropic recommends tool search when you have 10 or more tools, when tool definitions exceed 10,000 tokens, or when you see selection accuracy drop. Search returns five matching tools by default. It comes in two variants, one that matches regular expressions and one that uses BM25 keyword ranking.
Search has its own failure mode
Tool search changes the question from "which of 200 tools" to "did the search return the right five". Claude Code, which uses tool search by default, has public issues on exactly this:
- Issue #66488. Claude knew a deferred tool existed but repeatedly failed to load it because search ranked other tools higher, even for a query close to the exact tool name. A commenter found that raising the result limit to around 20 let Claude find it.
- Issue #30466. Long MCP server name prefixes dominated keyword ranking, so a relevant tool fell outside the five results. The workaround was selecting tools by their exact names.
Exact names help. The Claude Code changelog for version 2.1.113 noted that pasted MCP tool names now surface the actual tool ahead of tools that only match on description.
How to keep selection reliable
- Show fewer tools. Load only the tools the current task needs, and scope MCP servers to the agents that use them.
- Make the difference visible in the name. send_direct_message and post_to_channel are harder to confuse than two names that differ only in their last word.
- Say when to use a different tool. "Send a direct message to one user. For channels, use post_to_channel." A description that points to the alternative resolves most near-misses.
- Merge near-duplicates. One tool with a clear target parameter often works better than two nearly identical tools.
- Use tool search above about ten tools. Keep the most-used tools loaded, defer the rest, and use exact names when you know them.
- Log the two errors separately. Track wrong-tool and wrong-argument errors as different metrics, since they have different fixes.
{
"tools": [
{ "type": "tool_search_tool_bm25_20251119", "name": "tool_search_tool_bm25" },
{
"name": "send_direct_message",
"description": "Send a direct message to one user. For channels, use post_to_channel.",
"input_schema": {
"type": "object",
"properties": { "user_id": { "type": "string" }, "text": { "type": "string" } },
"required": ["user_id", "text"]
}
},
{
"name": "post_to_channel",
"description": "Post a message to a channel. For one person, use send_direct_message.",
"input_schema": {
"type": "object",
"properties": { "channel_id": { "type": "string" }, "text": { "type": "string" } },
"required": ["channel_id", "text"]
},
"defer_loading": true
}
]
}The search tool and the most-used tool stay loaded. The rest are deferred and found on demand, and each description names its alternative.
Where Swytchcode fits
Swytchcode is an execution layer between agents and the APIs they call, and it changes the shape of the selection problem. Instead of one tool per API operation, the Swytchcode MCP server's agent profile exposes a small fixed set of tools: discover, info, exec, list, search, add, and policy. Claude does not choose among hundreds of API tools. It describes the job, swytchcode_discover returns ranked operations from the catalog, swytchcode_info shows the chosen operation's real schema, and swytchcode_exec runs it.
Selection is constrained twice. Discovery ranks operations by intent, and only operations added to the project's tooling.json can run. An operation that was never enabled exits with code 2 and no request is sent, so a wrong choice outside the approved set fails closed. Every call that does run goes through the same pipeline: resolve the tool, validate inputs, evaluate policies, resolve credentials, execute, and normalize the response. Credentials never pass through the model's context.
# Claude Code
npm install -g swytchcode
swy init --editor=claude --mode=sandbox
swy doctor
# Enable only the operations this project needs
swy discover "post a message to a channel" --json
swy get <project>
swy add <canonical_id>
swy info <canonical_id>For your own agents on the Claude API, the Swytchcode runtime SDK hands Claude only the tools for the toolkits you choose. It needs the CLI installed and an initialized project on the same machine. The Anthropic SDK quickstart in the Swytchcode docs shows the full loop.
import { Swytchcode } from "@swytchcode/runtime";
import { AnthropicProvider } from "@swytchcode/runtime/providers/anthropic";
const swx = new Swytchcode(new AnthropicProvider());
const tools = await swx.tools.get({ toolkits: ["<project>"] });What happens when Claude picks or fills in the wrong operation:
- Unknown or not-enabled operations are rejected. An invented canonical ID, or a real one that has not been added to the project, exits with code 2 and no request is sent.
- Invalid inputs fail before the request. A missing required field, a wrong type, or a field the API does not define exits with code 1 before the provider sees the request.
- Provider errors come back structured. Each error includes the status code, an error category, and whether it is retryable. A 200 response with an error in the body counts as a failure.
- Changed operations are flagged. swy sync warns when an installed operation's method has changed upstream.
- Writes can be held. dry_run=true previews the request without sending it. For production, allow and deny rules are available from the Pro plan, and approval rules that route a call to Slack or Telegram are available on Business and Enterprise.
Where this does not help
Swytchcode narrows API operations. It does not choose between your own non-API tools, and discovery can rank the wrong operation first for an ambiguous request. Keep requests specific, read the swytchcode_info output before executing, and keep approvals on writes. It also cannot tell whether a valid value is the right one for your task.
Frequently asked questions
Why does Claude call the wrong tool?
Usually because too many similar tools are in context. Anthropic's documentation says selection accuracy degrades beyond 30 to 50 tools, and similar names make it worse.
How many tools can Claude handle?
There is no hard cutoff. Anthropic recommends tool search once you have 10 or more tools or more than 10,000 tokens of definitions, and notes that accuracy degrades beyond 30 to 50 tools.
What is the difference between a selection error and an argument error?
A selection error is calling the wrong tool. An argument error is calling the right tool with wrong values. Selection errors are fixed by fewer, clearer tools, and argument errors by stricter schemas and validation.
Does tool search fix wrong tool selection?
It helps by keeping only relevant tools in context, but search can miss the right tool. Claude Code issues #66488 and #30466 describe ranking misses. Exact tool names and a higher result limit help.
How does Swytchcode reduce wrong tool selection?
Its MCP server exposes a small fixed set of tools instead of one per API operation, finds operations by intent, and only runs operations enabled in the project's tooling.json.
Swytchcode resources
- Why AI agents call the wrong API: https://www.swytchcode.com/content/why-your-ai-agent-calls-the-wrong-api-and-how-to-fix-it
- Why Claude Code guesses API endpoints: https://www.swytchcode.com/content/why-claude-code-guesses-api-endpoints
- Invalid tool arguments in the OpenAI Agents SDK: https://www.swytchcode.com/content/openai-agents-sdk-invalid-tool-arguments
- Anthropic SDK quickstart: https://docs.swytchcode.com/quickstarts/anthropic-sdk/
- Swytchcode MCP server docs: https://docs.swytchcode.com/cli/mcp/
- Supported APIs: https://www.swytchcode.com/apis
More content
Why AI Coding Agents Still Write stripe.charges.create, and How to Stop It
Why AI coding agents still reach for Stripe's legacy Charges API, when current models get it right on their own, how Stripe steers AI tools toward Payment Intents and Checkout Sessions, and how to keep your agent on current Stripe APIs.
Invalid Tool Arguments in the OpenAI Agents SDK: Why the Model Sends the Wrong Parameters
What "Invalid JSON input for tool" means in the OpenAI Agents SDK, what strict mode guarantees and what it leaves out, why schema-valid arguments can still carry wrong values, and how to catch both kinds of error before a call runs.
Lovable and Bolt Apps Break on Third-Party APIs: Wrong Endpoints, Missing Secrets, Deprecated Calls
Why apps built with Lovable and Bolt break when they call third-party APIs, from wrong endpoints and secret name mismatches to keys exposed in the browser and deprecated SDK calls, and how to give the builder the facts it needs to get the integration right.
