Tools
How to expose a 3,000-operation API to agents.
An agent cannot read three thousand tool definitions before it acts. It can search, read one, rehearse it and run it, if the API lets it.
By Himadri Mishra ·
In short: Do not load a large API into an agent as one tool per operation: the definitions alone can fill the context window. Let the agent pull what it needs instead. It searches the catalogue in task words, reads the help and schema for one operation, rehearses the call without credentials, runs it, and trusts an exit code that tells the truth. Cloudflare's cf CLI and Code Mode MCP server, Google's Workspace CLI, gh api and Stainless's dynamic tools each show part of this. Generated descriptions still need human curation, and every layer needs a context budget.
The article in 6 points
- One tool per operation does not scale: Cloudflare estimates that a one-tool-per-endpoint MCP server for its API would need about 1.17 million tokens of definitions, against about a thousand for its Code Mode server built on a search tool and an execute tool (vendor figures).
- The working pattern is pull, not push: search the catalogue in task words, read the schema for one operation, rehearse it without credentials, then run it.
- Exit codes are often the agent's main success signal, so a cancelled delete that exits 0 leads the agent to report a deletion that never happened.
- A raw request command, such as `gh api`, keeps an agent moving when the curated surface has a gap, at the cost of weaker guard rails.
- Descriptions generated from an OpenAPI spec inherit its gaps, so curate names, enums and group descriptions in a reviewed layer on top of the spec.
- Research on catalogue size and code as action is promising but early: mostly preprints, often small or single-vendor model sets, so treat the numbers as dated claims.
Short answer. Do not hand an agent one tool per operation. For an API with thousands of operations, the definitions alone can fill the context window before the agent has read the task. Give the agent a way to pull what it needs instead: search the catalogue in task words, read the details for one operation, rehearse the call without credentials, then run it and trust the exit code. Everything below is about making each of those steps cheap and honest.
This guide follows on from Agent tool catalogues need a pruning strategy, which covers small catalogues. Here the catalogue cannot be small, because the product really does have three thousand things it can do.
Why one tool per operation fails
An MCP server or function-calling setup normally sends every tool definition to the model up front. That is fine for ten tools. It breaks for three thousand.
Cloudflare gives the clearest public numbers, labelled here as vendor figures. Its Code Mode MCP announcement (20 February 2026) says an equivalent MCP server without Code Mode would consume 1.17 million tokens, and the server’s README puts a version with only required parameters at about 244,000. The Code Mode server covers the same API with a search tool and an execute tool in about a thousand tokens; the README now lists two small helper tools as well.
Even the size of the API depends on who is counting: Cloudflare’s own pages say over 2,500 endpoints (February), more than 2,900 commands (the cf docs), over 3,000 operations (the cf launch) and over 3,500 operations (the Forge launch). Endpoints, operations and commands are not the same unit, so say which one you mean before you quote a size.
A PayPal team reports the same shape in a preprint about its own production MCP gateway with more than 2,000 tools: tool definitions fell from 140.2k tokens, 70.1% of the context, to 1.3k with a search tool and an execute tool (Saha et al., arXiv 2608.23992, August 2026). That is the company’s own token count, not a measure of whether tasks succeeded.
So the question “how many tools should an MCP server have” has two answers. For a small product, a handful of tools that each do one job. For a large API, a small, fixed set of tools for finding and calling operations, with the operations themselves kept out of the context until they are needed.
Four ways teams ship a whole API
Four shapes appear in public products. They are not exclusive, and several vendors ship more than one.
| Shape | What the agent sees | Public example |
|---|---|---|
| Generated CLI | Commands with --help, a schema command and a dry run |
Cloudflare’s cf (open beta since 28 September 2026), generated by the open-source Forge pipeline; Google’s Workspace CLI, which its README says is not an officially supported Google product |
| Code mode MCP | A search tool and an execute tool; the agent writes code against the API | Cloudflare’s Code Mode MCP server, to which Cloudflare now points users of its deprecated product-specific MCP servers |
| Dynamic tools | Three meta-tools: list endpoints, get one schema, invoke one | Stainless’s dynamic tools mode for generated MCP servers (May 2025) |
| Tool search in the harness | Deferred tool definitions, found by a search tool when needed | Anthropic’s tool search tool, OpenAI’s tool search and the AI SDK’s toolSearch(), each with a defer_loading or deferLoading flag |
The first three put the discovery layer in the product. The fourth puts it in the agent’s harness, which helps any catalogue but depends on what the agent’s client supports.
Seven patterns for a large surface
Each pattern below has a short name, where it can be seen in public, what it costs, and, where there is one, a failure documented in public.
1. Search-first discovery
Give the agent one cheap way to go from a task in plain words to a few candidate operations. Cloudflare’s cf agent documentation describes cf cli search "<task>", a search that “runs locally and needs no credentials” and returns up to five matches as JSON. When the CLI detects from environment variables that a coding agent is running it, it puts a banner at the top of its help text telling the agent not to explore by chaining nested --help calls, and to search instead. Code Mode MCP does the same job differently: the agent writes a small program that searches the API specification.
Trade-off. Keyword search is cheap and predictable, but it only finds words that appear in the catalogue. A task phrased in the user’s words (“stop sending emails to this address”) may not match an operation named in the vendor’s words (“add to suppression list”). Add synonyms and task examples to the index, and test search with real task phrasings before you rely on it. Search queries are also data: cf collects them in its telemetry by default, and its agent banner asks agents to keep queries free of names, domains and IDs. Decide what you log before you ship the search box.
2. Schema on demand
Once the agent has a candidate, let it read the full contract for that one operation and nothing else. cf schema <command> prints the method, path, parameters and body fields for one command. Google’s Workspace CLI has gws schema, and Stainless’s dynamic mode has a get_api_endpoint_schema tool. The agent pays for detail only on the operation it is about to use.
Trade-off. The schema is only as good as the spec. An open cf issue reports a field typed as a string where the API needs an array (#181), and the docs warn that for some request bodies the schema’s field list is incomplete or empty. Fix it in the spec or the curation layer, not in the agent’s prompt.
3. Credential-free rehearsal
Let the agent see the exact request it is about to send, without sending it and without needing a token. In cf, adding --dry-run prints the request as JSON without sending it, and “dry runs need no credentials”. Google’s Workspace CLI offers --dry-run too. Rehearsal lets an agent check its own work, lets a person approve a concrete request, and lets CI run on a pull request from a fork.
Failure mode. Rehearsal output is still output. A public cf issue reported that --dry-run printed secret values in plain text (#103), which Cloudflare fixed by redacting them (#167); the fix notes that some opaque payloads still need upstream annotations before they can be redacted. Another asks for a before-and-after diff rather than only the request (#202), because a reviewer wants to see what will change, not only what will be sent.
4. Truthful exit codes
For an agent, the exit code is often the main signal that something worked. It must not say 0 when the outcome the command promises was not achieved. Cloudflare’s cf documentation is candid about the beta’s behaviour: in a non-interactive session, a destructive command without --force prints Aborted. and exits with status 0, and “a successful exit does not mean the resource was deleted”. Open issues report the same pattern elsewhere: an aborted delete that exits 0 (#94), a bulk delete that reported thousands of objects deleted while they were all still there (#209), and a bulk secrets command that silently deployed an empty version (#201).
Why it matters. An agent that sees exit 0 will tell its user the record was deleted. The rule is simple: a failed, refused or cancelled change exits non-zero, the reason goes to standard error, and the result goes to standard output. An empty result is not always a failure. A search that finds nothing, or a change that was already in place, can exit 0 if the command’s contract counts it as success and the output says so plainly. This is check RECOVER-06 in the rubric. Design for recovery covers what the error should then say.
5. A raw escape hatch
No generated surface is complete on day one. A raw request command lets an agent reach an operation the curated commands miss, using the CLI’s own authentication. The GitHub CLI’s gh api is a well-known example: an authenticated request to any REST or GraphQL endpoint, with --jq filtering and --paginate. Cloudflare’s Code Mode MCP server has the same thing in its execute tool.
Trade-off. The escape hatch bypasses curated names, confirmations and enums, so it needs its own guard rails: read-only by default, or a confirmation for unsafe methods. Leaving it out has a cost too. A request for a raw cf command like gh api was closed with a reply that it may come eventually, but that the team is focused on generating commands for every endpoint, with “a few gaps today” (#170). Until a gap is filled, it is a dead end for an agent using the CLI, although Code Mode’s execute tool still reaches the API.
6. Generated descriptions need curation
Generating the surface from the API specification is how cf and Google’s Workspace CLI cover thousands of operations, and it keeps the CLI and the API in step. Cloudflare’s Forge pipeline generates cf from OpenAPI schemas annotated with extra information, and the cf repository’s own AGENTS.md notes that the overlay files live in Forge rather than in the CLI. Google’s Workspace CLI goes further: its README says it “reads Google’s own Discovery Service at runtime and builds its entire command surface dynamically”.
Failure mode. Generated text inherits whatever the spec says, including nothing. A command group whose help text is just its own name tells the agent nothing at the moment it has to choose. Constraints written in prose rather than as an enum are another inherited gap: one preprint audited 2,501 OpenAPI documents and reports that only 7.5% declared an enum (Li et al., arXiv 2609.00035, August 2026). The same paper reports that, across twelve models, a vocabulary the description only gave examples of was missed on 88 of 88 attempts, and putting it in the schema took that to 0 of 89. Put the human review in a layer above the spec, and lint it like code. Write tool descriptions an agent can act on has the checklist.
7. Context-footprint budgeting
Give each layer a size limit and measure it. A search should return a short list, one schema should fit comfortably beside the task, and a large result should be cut to a size that is still usable. Cloudflare’s Code Mode server caps each tool result at about 6,000 tokens by default and marks every cut, according to its README. Count the cost of the worst path too: measure what an agent pays if it ignores search and walks every help page, because some will.
Trade-off. Truncation hides data, so tell the agent it happened and how to get the rest, for example with paging or a filter. A budget nobody measures is not a budget, so put the sizes in CI.
What the research says, and what it does not
The research on large catalogues and code as action is new and mostly unreviewed. Each figure below is the paper’s own claim, with its model set and date.
- Programs can beat single tool calls. Patel et al., “The Bitter Lesson of Tool Calling” (August 2026, preprint) compared tools exposed as typed Python stubs with native JSON tool calls on one benchmark (BFCL v4) across 14 models, and report that the programmatic form matched or beat JSON calling for 11 of the 14. It measures call accuracy, not safety or review cost.
- There is a middle ground between one tool and hundreds. MCP-GRANITE (Paschalides et al., September 2026, preprint) varied tool granularity across nine small local models from 268M to 20.9B parameters, and reports that a four-tool interface beat both fine-grained tools and one monolithic tool. Small models only, so do not read it as a rule for frontier models.
- The client matters more than the interface. Alier Forment et al. (August 2026, preprint) ran one git task across seven agent scaffolds and five models, and report that the scaffold changed cost far more than whether the service was reached by MCP or by CLI. Their paired MCP-to-CLI cost ratios ranged from 0.43x to 29x, so the paper does not support “CLI beats MCP” as a general rule.
- Shipping a tool does not mean it gets used. Fan et al. (August 2026, preprint) report that, on the 309 tasks of the OSWorld-MCP computer-use benchmark, a reasoning model called one of the available MCP tools on only 55.
None of these papers measures a real product’s catalogue of thousands of operations with a person in the loop. Treat them as reasons to test your own surface, not as targets.
A worked example
The company and API here are invented for illustration. Tidewater Freight has a public API of about 3,000 operations, and ships a generated CLI called tide. A user asks their coding agent: “Cancel shipment 4471 and tell me when it is done.”
- Search. The agent runs
tide search "cancel a shipment". It gets five matches of one line each. The top match istide shipments cancel; the second istide bookings cancel, which the curated description marks as “for bookings not yet collected; for a shipment in transit, use shipments cancel”. - Schema. The agent runs
tide schema shipments cancel. It learns the method isPOST, the path is/shipments/{id}/cancel, andreasonis an enum of four values. - Rehearse. The agent runs the command with
--dry-runand no token. It prints the exact request, which the agent shows to the user. - Approve and run. Cancellation cannot be undone, so Tidewater’s API does not run it on the agent’s word. The agent runs the command, and the API holds that exact request and sends an approval card, “cancel shipment 4471, reason: customer request”, to the account owner’s Tidewater app, a route the agent cannot operate. The command waits for the decision. The user approves in the app. The API records who approved which operation and when, and only then runs that request; an approval for another shipment, another reason, or one that has expired, does not count. A “yes” in the chat would not count either, and nor would a
--forceflag on its own, because the agent can type both. Tidewater’s CLI keeps--forcefor jobs that run behind the company’s own approval step, but the flag only skips the terminal prompt: the API still checks for a recorded approval. - Check. The API refuses: the shipment is already at the port. The CLI exits 4, not 0, and writes “Shipment 4471 cannot be cancelled after arrival at port; use shipments return” to standard error. The agent tells the user the truth and offers the return.
Without step 4’s recorded approval, nothing would show that the user, rather than the agent, chose to cancel (rubric checks ACT-02 and HANDBACK-05). Without step 5’s honest exit code, the agent would have reported a cancellation that never happened. Without step 1’s curated neighbour description, it might have cancelled the booking instead.
A checklist for a large surface
- Can an agent go from a task in plain words to the right operation in one search, without credentials?
- Can it read the full contract for one operation without loading any other?
- Can it see the exact request before sending it, without a token, and with secrets redacted?
- Does every failed, refused or cancelled change exit non-zero, with the reason on standard error, and does an empty or no-op success say so in its output?
- Is there a raw request route for gaps, with its own guard rails?
- Does every group and operation have a curated description, and is every closed set an enum?
- Do you know, in bytes or tokens, what each step costs, and is that number checked in CI?
- Have you tested search with real task phrasings, and kept the misses as regression cases?
If your instructions to agents name commands or flags, keep them in step with what ships: see Keep agent instructions true.
Frequently asked questions
How many tools should an MCP server have? There is no tested number that holds across models and tasks, and the figures quoted online are rules of thumb. What matters is how many plausible neighbours a request has and how much context the definitions cost. For a small product, a handful of task-shaped tools is usually enough. For an API with thousands of operations, expose a small fixed set of tools for searching, describing and calling the API rather than one tool per operation.
What is code mode for MCP? Code mode is a pattern in which the agent writes a short program against a typed API instead of making one tool call at a time. Cloudflare’s Code Mode MCP server exposes its whole API mainly through a search tool and an execute tool: the agent writes code to search the API specification, then writes code that calls the API inside a sandbox. Its main benefit is that the full catalogue never has to sit in the context window.
Is a CLI or an MCP server better for agents? Neither wins in general. A CLI suits coding agents that already have a shell, and it is easy to script and test. An MCP server suits hosts with no shell, and it can scope credentials per connection. One preprint found that the client scaffolding changed cost far more than the choice of interface. Many vendors ship both from one specification.
Can I generate agent tools straight from an OpenAPI spec? Yes, and that is how the large command-line tools described here cover thousands of operations, but generated text inherits every gap in the spec. Placeholder summaries, constraints written in prose rather than as enums, and typing errors all reach the agent unchanged. Keep a reviewed layer of names, enums, confirmations and descriptions on top of the spec, and test that search finds the right operation for real tasks.
What should a CLI do when an agent cancels a destructive command? Exit with a non-zero code and say on standard error that nothing was done. An agent reads exit code 0 as success. If a refused confirmation exits 0, the agent will report a deletion that never happened. Document the codes, and keep the result on standard output separate from messages on standard error.
The test for a large surface is simple: can an agent that has never seen your API find, check and safely run the one operation it needs, using no more context than the task deserves?
