LLM tool schema pruning is becoming a practical cost-control problem for teams building multi-tool AI agents. A recent r/SaaS discussion put numbers on the issue: a support agent with 33 available tool schemas accumulated roughly 72,000 schema input tokens across an 18-turn session—even though its router narrowed the likely choices to only four tools per turn.
That finding matters because many teams inspect average cost per request, see a tolerable number, and move on. But a multi-turn support conversation is not a single request. It is a chain of requests in which static instructions, tool descriptions, JSON schemas, conversation history, tool calls, and tool results may all be sent repeatedly. If every turn includes a large tool registry, a component intended to help the model act can cost more than the text the model generates.
The original Reddit poster was considering Braintrust for span-level attribution, schema-set experiments, and routing regressions. The community’s best response was not simply “remove tools.” It was more useful: first determine whether prompt caching is actually working; then replay real sessions against both policies; finally, judge pruning by whether the agent can still complete the user’s step—not merely whether it selected the identical tool name.
This article turns that discussion into a practical framework for reducing tool-definition cost without creating a more expensive failure mode: misrouted requests, unnecessary retries, longer conversations, or unresolved support tickets.
Why tool schemas become an invisible agent tax
A tool schema is not just a function name. Production definitions often contain a natural-language description, parameter descriptions, enums, nested objects, required fields, examples, safety conditions, and formatting constraints. Thirty-three of those definitions can easily become thousands of tokens before a customer has written a second message.
In the r/SaaS example, approximately 72,000 schema tokens across 18 turns works out to about 4,000 schema tokens per turn. The striking point was that the generation bill was reportedly lower than the schema block. That is entirely plausible in support agents, where replies may be brief but the agent repeatedly receives a large operational toolkit.
This is not necessarily a model-provider defect. Tool definitions are part of the context the model uses to decide what it can do and how to format a valid call. On OpenAI’s API, the rendered context eligible for prompt caching includes tool definitions, developer instructions, and conversation history. But cache reuse depends on an identical shared prefix; a changing element before the relevant cache match can reduce reuse for everything after it. (developers.openai.com)
The per-request dashboard trap
A request-level average obscures three important dynamics:
- Conversation multiplication. A 4,000-token tool block looks manageable once. It looks very different when it appears 12, 18, or 30 times in a support workflow.
- Cache asymmetry. The raw input token count may remain high even when a cache discount applies, but a cache miss turns the same static block into fully billed input.
- Failure amplification. A smaller schema set that makes the agent choose the wrong tool may create extra turns, tool errors, handoffs, or customer abandonment—costs that outweigh savings from the missing definitions.
The correct question, then, is not: “How many tools can we remove?” It is: “What is the smallest, most cache-stable tool context that preserves successful task completion?”
A simple session-cost model
For a given session, model its input cost as:
session input cost = Σ (uncached input tokens × uncached rate) + Σ (cached input tokens × cached rate)
Then isolate the portion caused by tool definitions:
schema burden = schema tokens sent per turn × number of turns × effective token rate
The effective rate is the important part. If schemas are consistently cache hits, their marginal cost can be much lower than their list-price input cost. If the cache is repeatedly broken, they may be among the most expensive tokens in the entire agent.
OpenAI’s current guidance emphasizes checking cached-token usage, keeping the beginning of requests stable, and using cache diagnostics to compare requests when expected reuse does not occur. Its documentation also notes that tools and schemas are within the cacheable rendered context, rather than being somehow free or separate from it. (developers.openai.com)
The r/SaaS thread identified the real optimization problem
The original discussion is valuable because the team had already built a router. Its policy reduced 33 possible tools to four candidates, but all 33 schemas were still included in the context window on every turn. In other words, the application had made an intelligent routing decision without translating that decision into a smaller model request.
That distinction is easy to miss. There are two separate systems:
- Candidate selection: deciding which tools are relevant to the current user need.
- Schema injection: deciding which exact definitions the model sees for the next action.
A system can succeed at the first and fail to benefit at the second. If the full registry rides along regardless of the candidate list, the routing layer may improve accuracy or organization while doing nothing for prompt size.
The community reaction surfaced four measurements that should become standard for any LLM tool schema pruning project:
- Inspect cache usage from turns two through the end of the session.
- Replay the same sessions under full and pruned schema policies.
- Measure whether a pruned policy excludes the tool that the full system would ultimately need.
- Evaluate completed workflow quality, not exact tool-call equality alone.
That final point deserves emphasis. A model might choose find_customer_by_email in one variant and search_accounts in another. If both retrieve the required account and the rest of the workflow succeeds, exact-match scoring calls one of them a regression even though the customer experience is unchanged. Conversely, identical tool selection is not enough if different parameters cause a failed call.
Start with cache health, not pruning
Before rewriting tool infrastructure, establish whether the expensive schema prefix is actually being reused. This was the sharpest observation in the thread: if a constant 4,000-token schema block is billed at full input price repeatedly, something earlier in the request may be changing and preventing a useful cache match.
What to inspect on each turn
At a minimum, log these fields for every model request:
| Field | Why it matters |
|---|---|
| Total input tokens | Shows overall context growth. |
| Cached input tokens | Shows how much repeated context was reused. |
| Uncached input tokens | Reveals the portion being recomputed and fully billed. |
| Tool-schema token estimate | Separates tool overhead from history and instructions. |
| Prompt/version hash | Helps locate accidental prefix changes. |
| Tool-set hash and order | Detects changed schemas, ordering, or serialization. |
| Model and request settings | Caching can be affected by changes to the rendered request. |
| Latency | Cache misses usually hurt responsiveness as well as cost. |
On OpenAI, prompt caching requires the relevant prefix to match exactly. The provider specifically advises comparing an affected request with an earlier response to identify changes in the model, tools, settings, or input that stopped reuse. (developers.openai.com)
Common cache breakers in tool-using agents
The schema list itself may not be the culprit. Look for subtle request changes ahead of it:
- Dynamic timestamps or request IDs embedded in developer instructions.
- Per-turn personalization inserted before stable system content.
- Nondeterministic JSON serialization or changing key order.
- Tool lists reordered based on router confidence.
- Tool descriptions modified with live account state.
- A changing model, reasoning setting, structured-output setting, or tool-choice configuration.
- Injected retrieval results placed before the stable tool registry.
- Multiple workers or traffic patterns that reduce cache locality.
The solution is often structural: place content that changes rarely first, and append volatile user, retrieval, and session state afterward. OpenAI’s documentation explicitly recommends stable shared prefixes and monitoring cached tokens while iterating. (developers.openai.com)
Do not confuse token count with token cost
A request can report a large input-token total and still be economically acceptable if most of it is cached. Conversely, a smaller prompt can cost more if it misses cache repeatedly. That means every pruning experiment should report both raw token volume and billed or effective input cost.
This is also why a “schema tokens per session” dashboard should sit next to a “schema cache-hit rate” dashboard. The first tells you the theoretical opportunity; the second tells you whether it is real.
LLM tool schema pruning should be a routing-quality experiment
Once cache behavior is understood, treat pruning as an evaluation problem rather than a prompt-editing project. The experimental unit should be a complete representative support session, not an isolated user query.
A strong setup has two variants:
- Control: the existing full-schema agent.
- Treatment: the router-selected subset, such as the top four tools, plus any mandatory fallback or safety tools.
Both variants should receive the same model, temperature or equivalent generation settings, system instructions, conversation state, test data, and tool environment. The schema policy must be the principal intentional difference.
Build a replay set from real work
The community suggestion to use fixed replays is exactly right. A replay corpus makes cost and quality comparisons meaningful because the workload does not drift between variants.
Create a dataset that includes:
- Resolved support conversations across common intents.
- Complex, multi-step issues that require more than one tool.
- Edge cases involving ambiguous wording, account-specific exceptions, or missing data.
- Historical incidents where the agent chose the wrong tool or retried a call.
- Escalation-worthy cases where the correct behavior is to ask a question or hand off rather than invoke a tool.
Do not build the dataset from only clean, high-volume intents. A router that handles password resets perfectly but hides a rare billing-adjustment tool could look excellent in aggregate while failing the customers with the highest-value or highest-risk requests.
Keep the ground truth practical
For every session or step, capture more than an expected tool name. Useful labels include:
- User intent and required outcome.
- Acceptable tool family or acceptable alternative tools.
- Required parameters and validation constraints.
- Whether a tool call should happen at all.
- Expected terminal state: resolved, clarified, escalated, or blocked safely.
- Maximum acceptable number of tool calls and model turns.
This turns evaluation from “did it imitate the old run?” into “did it do the job safely and efficiently?”
The metrics that determine whether pruning is safe
Exact tool-match rate is a useful diagnostic, but it should not be the shipping criterion. Use a scorecard that combines routing coverage, tool execution, task success, customer experience, and economics.
1. Candidate-set recall
Candidate-set recall asks: did the pruned set contain at least one tool capable of completing the next required action?
candidate-set recall = turns where an acceptable tool is included / turns requiring a tool
This is often the most revealing router metric. If the required tool is absent, the downstream model cannot choose it no matter how good its reasoning is. The thread’s suggestion to track how often the full-system choice falls outside the router’s top four is a practical approximation, but it should be expanded to include acceptable substitutes.
2. Tool-call validity
Track whether the model emits a syntactically valid call with valid required arguments, values within allowed enums, and an executable target. This guards against a misleading win where reduced choice improves tool-name selection but degrades parameter quality because a compressed description lost important constraints.
3. Step completion rate
A tool call is useful only if it advances the workflow. Measure whether the action returned the data or mutation needed to continue. For example, a search tool that succeeds technically but queries the wrong identifier is not a successful step.
4. End-to-end task resolution
This is the north-star quality measure: did the session reach the correct outcome without an unnecessary handoff, unsupported promise, or risky action? A customer-support agent should be scored against operational outcomes, not against aesthetically similar tool traces.
5. Recovery and retry rate
Measure retries, invalid calls, clarification loops, tool failures, and fallback usage. Pruning can appear successful on first-choice accuracy but increase the number of turns needed to recover from uncertainty.
6. Cost per resolved session
This is more decision-useful than cost per request. Calculate all model input and output cost, tool-runtime cost where applicable, and the cost of added turns. Then divide by successfully resolved sessions.
cost per resolved session = total session cost / successful resolutions
A pruned policy that saves 25% on input but reduces resolution by 5% may be a poor trade for a support operation. The right threshold depends on ticket value, support staffing cost, user retention risk, and the severity of failure.
7. Latency to resolution
Tool pruning can improve first-token latency by reducing prefill work, but a wrong route can eliminate that benefit with a retry. Track both per-turn latency and total time until the customer has a useful answer.
Use task-equivalent scoring instead of rigid exact match
The most important methodological upgrade is to define equivalence classes for tools. A route should not be scored wrong solely because it differs from the old agent’s exact function name.
Consider a billing question: “Why was I charged twice?” A full system might call get_invoice_details, while a pruned system might call list_recent_charges followed by get_payment_status. Different traces, same outcome—provided both surface the duplicate authorization and guide the customer correctly.
A practical scoring hierarchy
Use a hierarchy such as this:
- Resolved correctly and safely: full credit.
- Resolved through an acceptable alternate tool path: full or near-full credit.
- Needed a reasonable clarification before acting: partial credit when clarification was necessary.
- Escalated appropriately: credit when the case was outside policy or tool permissions.
- Selected an unavailable or irrelevant tool: failure.
- Made an unsafe mutation or fabricated completion: severe failure.
This approach answers the question raised in the thread: how do you know a smaller schema set preserves routing quality? You know by comparing task-capable behavior on the same conversations, with explicit credit for valid alternate paths and meaningful penalties for harmful behavior.
Why 95% tool-name agreement is not enough
A community commenter suggested a 95%+ same-tool benchmark as a simple baseline. That can be a useful early smoke test, especially where each intent maps to a single canonical tool. But it is not universally safe.
If the missing 5% contains payment cancellations, security incidents, or destructive account changes, the aggregate number hides unacceptable risk. Conversely, a system can score below 95% exact agreement because it uses harmless substitutes while retaining or improving end-to-end resolution.
Segment every result by intent, customer tier, action risk, and tool family. Do not let an excellent score on easy read-only queries conceal failures in high-consequence workflows.
Architecture patterns that reduce schema payload
There is no single correct design. The best pattern depends on the number of tools, how volatile they are, whether tool descriptions are long, how much routing ambiguity exists, and how provider caching works for the chosen model.
Pattern 1: Static full registry with a stable cache
In this approach, every turn receives all tools in a fixed deterministic order. The application prioritizes a stable request prefix and relies on prompt caching to reduce repeated input cost.
Best for: a moderately sized, mostly static registry where cache hits are consistently high and routing mistakes would be expensive.
Advantages: simple behavior, no router recall failures, consistent model awareness.
Risks: a single cache-breaking change can make the full registry costly; schemas still consume context capacity even when discounted.
Pattern 2: Two-stage category routing
The agent first sees a small taxonomy, such as billing, identity, subscription, technical troubleshooting, and account administration. It selects one or more categories, after which the application injects the relevant concrete schemas.
This mirrors the Reddit commenter’s proposal: first expose categories, then load the four real tools after routing. It can sharply reduce schema payload, but it adds a decision point that must be evaluated for category recall and total-turn cost.
Best for: large registries whose tools cluster cleanly into stable domains.
Advantages: major context reduction and clearer ownership boundaries.
Risks: an incorrect category route can make the required tool unreachable; a separate routing call can erase some savings if it is not compact or reliable.
Pattern 3: Deterministic application router plus top-k schemas
Use rules, embeddings, classifiers, metadata filters, or a lightweight model to select the top-k tools. The action model then sees only those schemas, perhaps with a safe fallback such as search_knowledge_base, ask_clarifying_question, or escalate_to_human.
Best for: mature products with clear intent signals and good labeled traffic.
Advantages: fewer tools in the main model context and direct control over candidate policies.
Risks: offline classifier performance may decay as products and vocabulary change; top-k recall must be monitored continuously.
Pattern 4: Tool search or just-in-time loading
The model begins with a meta-tool for discovering tools or capabilities, then receives detailed schemas only after selecting a domain or result. This resembles progressive disclosure in user-interface design.
Best for: very large or fast-changing tool catalogs, including some MCP-style integrations.
Advantages: avoids sending the entire catalog upfront.
Risks: discovery becomes another agentic step; poor search descriptions can make tools effectively invisible; multi-step tool discovery adds latency and evaluation complexity.
Pattern 5: Consolidate overly granular tools
Sometimes schema pruning is a symptom of an API-design problem. Ten separately named read-only customer lookup tools may be better represented by one well-designed get_customer_context tool with constrained fields or modes.
Consolidation should not become a giant “do anything” tool with vague parameters. The goal is a coherent capability boundary with validated inputs, strong descriptions, and predictable behavior.
Reduce schema tokens before you remove tools
Pruning is not the only lever. Teams often discover that descriptions and schemas have grown through copy-paste, generated API specs, redundant examples, and explanatory prose written for humans rather than tool-calling models.
Schema compression checklist
Review each tool definition for:
- Repeated policy text that belongs in stable global instructions instead.
- Long examples that do not change call correctness.
- Redundant field descriptions implied by a clear field name and type.
- Oversized enum descriptions or duplicated allowed values.
- Unused optional parameters exposed “just in case.”
- Multiple near-identical tools that differ only by an obscure implementation detail.
- Deeply nested payloads that could be assembled server-side from a smaller intent-level argument.
Compression must be validated, not assumed. Over-compressing descriptions can remove the distinction between similar tools or omit a safety constraint. Use the same replay suite to compare a full-description baseline with a compact-schema variant.
A useful strategy is to track tokens per successful invocation for each tool. A tool that is expensive to describe but rarely selected may be a better candidate for deferred loading than a frequently used tool with a short schema.
Observability: what to trace at the span level
The original poster mentioned Braintrust because the problem needs trace-level, not invoice-level, visibility. That is a sound direction. Braintrust describes traces as end-to-end executions composed of nested spans, allowing teams to examine the units of work inside an agent interaction. (braintrust.dev)
Whether you use Braintrust or another observability stack, model requests should be connected to the router decision, injected schema set, tool calls, tool results, and final customer outcome.
Recommended trace structure
support_session
├── classify_intent
├── select_tool_candidates
│ ├── candidate_count
│ ├── candidate_tool_ids
│ ├── router_confidence
│ └── required-tool-in-top-k (evaluation label)
├── model_turn_1
│ ├── input_tokens
│ ├── cached_input_tokens
│ ├── schema_tokens
│ ├── schema_set_hash
│ ├── output_tokens
│ └── selected_tool
├── tool_execution
│ ├── tool_name
│ ├── validity
│ ├── latency
│ └── result_status
└── session_outcome
├── resolved
├── escalation
├── turns
└── total_cost
The schema-set hash is especially helpful. It lets an engineer group behavior by the actual definitions that were injected, rather than relying on a deployment version that may hide conditional tool selection.
The dashboard that changes decisions
Build a dashboard that can answer these questions without manually inspecting logs:
- Which intents have the worst candidate-set recall?
- How much schema cost is cached versus uncached?
- Which schemas account for most context volume but least tool usage?
- Does a smaller top-k reduce total session cost after retries?
- Which misses cause customer-visible failure rather than harmless alternate paths?
- Has a prompt, tool, or routing change reduced cache reuse since the previous release?
This is more actionable than a single “average token cost” chart because it connects spend to a concrete system behavior.
A safe rollout plan for schema pruning
Do not deploy an aggressively pruned routing policy based solely on offline agreement. Use a staged release with guardrails.
Phase 1: Instrument the current full-tool baseline
For one to two representative traffic cycles, collect session-level cost, cache data, tool usage, retries, task outcomes, and human escalation outcomes. This establishes the true baseline and identifies expensive but low-value schemas.
Phase 2: Shadow-mode candidate evaluation
Run the proposed router alongside production without changing the live agent’s schemas. Record the top-k set and whether it would have contained the tools actually used by the baseline. This estimates candidate recall safely.
Phase 3: Offline replay with executable tools
Replay fixed historical sessions through control and treatment configurations. Where possible, execute against a sandbox or deterministic tool-result fixture so that parameter validity and subsequent decisions can be evaluated.
Phase 4: Limited live experiment
Route a small, low-risk portion of eligible conversations to the pruned policy. Exclude security, payment mutations, cancellations, and other high-consequence intents until their segmented evaluation is strong.
Phase 5: Progressive expansion with automatic fallback
If the action model cannot complete a step, allow a controlled fallback: broaden from top four to top eight, invoke a tool-search layer, ask a clarifying question, or escalate. Log every fallback as a potential router-quality signal.
A good deployment policy is not “top four forever.” It is “top four by default, expand when confidence is low or the agent detects a missing capability.”
How current prompt-caching guidance changes the decision
Prompt caching changes the economics of tool schemas, but it does not eliminate the need for LLM tool schema pruning. It should change the order of operations.
First, stabilize and verify caching. OpenAI states that full rendered context—including tool definitions—can be cached, but that matching depends on the shared prefix. Its newer diagnostics are intended to identify changes that limited reuse between requests. (developers.openai.com)
Second, measure the remaining uncached schema burden. Even a perfectly cached static block still occupies context window capacity, can affect latency characteristics, and may grow beyond practical limits as a product adds integrations.
Third, use pruning to improve both cost resilience and model focus. A model asked to distinguish among four highly relevant tools faces a simpler action-selection problem than a model confronted with 33 overlapping definitions. That is an engineering hypothesis to test, not a universal truth: sometimes removing tools eliminates useful alternatives or makes the remaining choices deceptively similar.
Anthropic’s tool-caching documentation makes the same broader point from a different implementation angle: tool definitions can be cached across turns, and cache placement and tool-loading behavior matter. It specifically documents placing a cache breakpoint on the final tool definition to cache the preceding tools, while warning that tool configuration patterns affect cache behavior. (platform.claude.com)
The deeper lesson: optimize agent economics at the session level
The r/SaaS post is not really about 33 tools. It is about a broader mistake in agent engineering: optimizing local request metrics while ignoring the economics and reliability of a complete workflow.
Tool schemas, retrieval context, reasoning settings, retries, output verbosity, and model fallback policies all interact. A 10% reduction in prompt tokens is not useful if it creates a 20% rise in retries. A cache hit is not a complete solution if the growing tool registry crowds out relevant conversation context. And an accurate router is not economically valuable if its chosen candidate set is never actually used to shrink the request.
The mature operating metric is therefore successful customer outcome per dollar and per second. It combines quality, cost, and latency in the way customers experience the system.
For teams facing the same issue, the immediate next step is straightforward: inspect cached-token counts across a real multi-turn session, then replay that exact session with a smaller candidate set. If the treatment preserves candidate-set recall, task completion, and safe behavior while lowering total cost per resolved ticket, ship it gradually. If it fails, the trace will tell you whether the issue is cache instability, inadequate routing, missing fallback tools, or schemas that need redesign rather than removal.
FAQ
What is LLM tool schema pruning?
LLM tool schema pruning is the practice of sending an agent only the subset of tool definitions most relevant to the current task, rather than including every available tool schema in every model request.
Does prompt caching make tool schema pruning unnecessary?
No. Caching can substantially reduce the cost and latency of repeated tool definitions when the request prefix remains stable, but cache misses can make them expensive again. Large schemas also consume context space, so pruning can still improve resilience and focus.
What is the best metric for evaluating a smaller tool set?
Use end-to-end task resolution alongside candidate-set recall, tool-call validity, retry rate, latency to resolution, and cost per resolved session. Exact tool-name match is a useful secondary diagnostic, not the only success metric.
How many tools should an agent see per turn?
There is no universal number. Start with the smallest top-k that maintains near-complete candidate-set recall on representative sessions, then add a safe fallback path for low-confidence or ambiguous requests.
Should a support agent use a two-stage router?
Often, yes—especially when tools cluster into clear domains such as billing, identity, and technical troubleshooting. But evaluate the additional routing step against total cost, latency, and the risk that the first-stage classifier excludes the tool needed to solve the case.