The problem
The agent wakes up without the catalog.
- A new session does not know what it can call. The task arrives. The MCP server may hold hundreds of tools. Nothing in that cold start is a picture of which tools exist, or how they differ.
- Loading every definition spends the window up front. Each tool brings a name, a description, and a parameter schema. At 457 tools, that block is larger than the task it is supposed to serve, and it sits in context on every turn.
- Search assumes you can already name the target. Lexical and semantic search are the right move once the name, or a close description, is known. They return hits. They do not show the neighboring tools you did not know to ask for.
Why it is worth addressing
A dump and a guess both fail.
- Lookalikes are the normal case. Nine apps in this catalog expose a login. A flat prompt makes the model read every schema to tell them apart. The miss is a near tool with a similar name, called because the catalog was never organized.
- A tool left out of context cannot be chosen. Trimming the prompt to save space also hides the tools that were cut. The trace then shows a path that was never available, not a considered rejection.
- The context cost is paid whether the tool is used. Unused schemas still occupy the window. Turns that needed one definition still carried the other 456.
The advantage
Show the catalog, then load the tool.
-
Progressive disclosure.
Labels and counts come first. Schemas come last. In the demo, “user account” takes the working set from 457 to 69. A few include and exclude decisions leave nine login tools, then
amazon.login. That is the definition worth reading. - More than one taxonomy over the same tools. Domain and function are two ways through one catalog. The nine login tools sit together as a kind of action, and apart as one tool per app. The agent can approach from either question.
- Direct search stays the short path. Once the name is known, look it up and skip the walk. Orchard is how the agent learns the name. Retrieval is how it fetches the tool after that.
Lazy load follows the narrowing. The other definitions stay out of state until a later decision needs them. The repository is the visual search in the demo.
← Back to Tool and skill optimization