A better way to build MCP servers with Laravel

Every MCP tool loads on every request. Laravel MCP 1.0 fixes that with searchable tool catalogs: the same payload for 10 or 100 tools.


We ship an MCP server for Laravel Nightwatch. It does what we built it for. Users like it, use it, and immediately ask for the thing it does not do yet. Every ask is reasonable! But every ask is also another tool that the agent has to sift through.

Now think about a full Laravel Cloud MCP server: the API has a couple hundred endpoints. Deployments, environments, databases, workers, domains, logs. Handing that whole surface to an agent makes logical sense, because that is already how a lot of people use laravel cloud via agents.

The problem starts when an MCP server grows. Every tool you add brings its name, description, and full input schema into the agent's context before it has done any work. With enough tools, those definitions start competing for attention. The agent has more context to sift through, more similar choices to confuse, and less room to focus on the user's actual request. A separate tool for every endpoint makes sense at the API layer, but loading every one of them on every turn does not.

Cloudflare's Code Mode had come out about a year earlier, and it felt like the answer. But what I ended up shipping in Laravel MCP 1.0 is smaller and stranger: searchable tool catalogs, where the size of the tools/list payload stays the same whether the catalog holds 10 tools or 100.

This is the story of the four versions I went through, the three I threw away, and the mistake I made twice before I noticed it.

What every tool costs before the agent starts

Tool calling looks like an API call, but it isn’t quite the same. It is text generation wearing a costume, which is the whole reason any of this matters.

A model generates a stream of tokens. Most map to words or fragments of words. But models are fine-tuned to emit a few special tokens that have no textual meaning at all, and two of them mean "a tool call starts here" and "a tool call ends here."

Between them, the model writes JSON:

The harness watches for that closing token, parses the JSON, and turns it into an MCP request against your Laravel app. Your tool runs, the result comes back through another pair of special tokens, and generation resumes as if the model had simply read it.

A diagram of what tools do before before the agents starts and reaches your Laravel applicationLook at where that note sits. Before the question is even read, the name, description, and full JSON Schema of every tool you expose is already in the context window. Not the ones being used. All of them, on every turn!

The numbers are worse than I expected. Anthropic's advanced tool use docs put a 50-plus tool MCP setup at around 72,000 tokens upfront. About a thousand tokens per tool. Several times bigger than my own guess. The token bill is the obvious cost, but the accuracy is the one that bothered me more.

On Anthropic's MCP evaluations with large tool libraries, deferring definitions and searching instead moved Opus 4 from 49% to 74%, and Opus 4.5 from 79.5% to 88.1%. Forty tool descriptions really is 40 chances to pick the wrong one.

Version one: I handed the agent Tinker

Code Mode's argument is sharp, and I think it is correct: "LLMs are better at writing code to call MCP, than at calling MCP directly." Models have seen millions of real code examples and almost no real tool calls, because tool call syntax exists mostly in synthetic training data.

Cloudflare's line for it is better than anything I would write: "Making an LLM perform tasks with tool calling is like putting Shakespeare through a month-long class in Mandarin and then asking him to write a play in it."

They convert MCP tools into a TypeScript API and run the generated code in a V8 isolate. I did not need to convert anything, because the model already writes PHP, and so does the app. So version one generated a PHP function for each registered tool, straight from its name, description, and input schema:

The agent got those signatures instead of 40 tool definitions, wrote PHP against them, and I ran what came back through eval(). Essentially Tinker, with a language model at the keyboard.

I want to make this part clear, because it changes how you should read everything in this post: I did not type most of this code. The implementation was agent-written across all four versions.

My job was specifying, reading, and deciding what to keep. That is why the versions moved as fast as they did. It’s also how I managed to fool myself… twice.

This is the kind of thing the agent would write:

Two chained calls gave us twelve months of readings filtered down before any of it reached the model. Code Mode worked exactly as advertised, in PHP, in about a day. I was thrilled!

Every example I threw at it came back correct, which is exactly why I kept going as long as I did. A run that succeeds tells you nothing about what a run was allowed to do, and my tests only ever exercised the snippets I expected.

Then I looked at what else that snippet could have been:

Nothing in my design prevented any of it. The generated functions were a suggestion, not a boundary, and there was no other boundary anywhere. The code ran inside the application, in the same process, with the same container, the same environment, and the same database connection my own code uses.

I knew it was eval().

I had specified it and approved the diff, but what took me a day and a half to sit with is that eval() means something different when the author is a language model.

With a human, you are one careless line away from a bad afternoon, and there is a person who can be asked what they were thinking. Here, the full capability surface of PHP was reachable from any line the model chose to write, at whatever rate the agent chose to write them, and none of it would even need to be malicious. Even if it was just slightly wrong or overeager, we’d be in big trouble.

That was when I understood why Code Mode is primarily a method related to sandboxing and not a method strictly about code generation. Getting the model to write good code is the easy half, and I had already done it. All the expensive and dangerous parts are downstream of deciding to run it.

Version two: I added the sandbox

Cloudflare gets its sandbox essentially for free. The entire Workers platform is already V8 isolates, so a fresh one starts in milliseconds using a few megabytes. They create one per snippet and throw it away without thinking about it.

A Laravel app has no such substrate. Your app runs on Cloud, or Forge, or Docker, or somebody's shared host. PHP has had no meaningful in-process sandbox since safe_mode was removed. So every option meant new infrastructure:

Approach

What it would have meant

Container per execution

Seconds of startup, an orchestrator to operate, a new thing to secure

Firecracker microVMs

Real isolation at around 125 ms boot, and a microVM control plane to answer a weather question

Hosted sandboxes like E2B or Daytona

Works well, and your MCP server gains a third-party dependency and a per-second bill

WASM via Extism or php-wasm

Genuinely good isolation, and a whole toolchain to ship and explain

Any of them would have worked. If I were building this for one application, I would have picked one and moved on.

But I was building it into a package, and that changes who absorbs the mistake. In an app, a sandbox you got slightly wrong is your problem. In a package, it is a default that ships to everyone who runs composer require, most of whom will reasonably assume the isolation was somebody else's job to get right. I would be handing tens of thousands of developers’ code execution next to their database and asking them to trust my threat model, which had already been wrong once that week.

Version three: A PHP dialect the model could not speak

So I tried to have it both ways. If the language cannot express anything dangerous, it does not need a sandbox.

The eval()went in the bin and I reached for nikic/PHP-Parser. Parse what the agent writes, walk the tree, reject any node outside a small allowlist: calls to registered tools, literals, simple assignment. Nothing dangerous is even representable, so nothing needs isolating.

It worked, the tests passed, and I made exactly the same mistake as the first time without noticing. The allowlist and the tests came from the same spec I wrote, so naturally every snippet under test stayed inside the subset. It looked complete because nothing had yet been written against it that had not been told what it was.

Then a model that had never read my allowlist started writing:

That is a completely reasonable answer to "check alerts for these three cities." It is also four rejections: a foreach, an if, a function call I had not allowlisted, and an array append. Valid PHP, confidently produced, and unfortunately refused by my interpreter.

The model had no way to know where my subset ended. Nothing in its training data describes a boundary I invented last Tuesday. So it would fail, read the error, and correct itself into a different unsupported construct. I had taken the one advantage of the code approach, the model's fluency with PHP, and turned it into a liability.

It was the worst of the four versions, and it was the one I thought was the most clever.

A partial language is worse than no language. Code Mode works because a V8 isolate runs actual JavaScript, not a dialect that looks like JavaScript and then refuses two thirds of it.

Version four: Switching from PHP to JSON

I dropped PHP-Parser entirely and wrote a JSON format the executor supported completely, with no subset to guess at.

This felt like a downgrade and turned out to be the opposite. A format the model has no priors about outperformed the one it had strong but wrong priors about, because the failure mode became easy to understand. JSON Schema states exactly what is allowed, the model gets that schema in the tool definition, and there is no hidden boundary to trip over.

Then I kept cutting.

We had per-catalog fluent configuration, so you could write ToolSearch::for([...])->maxToolCalls(5)->maxOutputBytes(32_000). It read nicely, but it meant limits could live in two places, and I could not answer simply what should happen when the config file and the fluent call disagreed. I removed it.

Detailed exceptions for out-of-range config values were also removed, replaced by max(1, ...).

Each capability I removed was one less thing to explain to the model and one less thing for it to get subtly wrong.

Somewhere in that last round the goal changed. I stopped trying to port an impressive piece of infrastructure and started trying to make the thing feel like the rest of Laravel, where you use a feature instead of sweating over machinery you never wanted to think about. Once that was the bar, most of my remaining ideas were obviously wrong.

The question I should have asked on day one finally surfaced: how much of the win comes from executing code, and how much comes from not shipping 40 tool definitions on every request?

Mostly the second one. The savings come from progressive disclosure, and progressive disclosure does not require executing anything.

What I shipped

Two meta-tools. search_tools searches a catalog. execute_tools runs a batch of calls from it. You declare a catalog with ToolSearch::class as a key in your server's $tools:

CurrentWeatherTool stays advertised because it gets hit on nearly every request. Everything else becomes discoverable instead of resident, and tools/list returns three tools instead of 40.

A diagram showing the before and after difference in tools/listHere is a real session against that server. The agent searches for "rainfall" and gets back one tool:

That result is the detail I did not expect to matter as much as it did. Search is deterministic, lexical scoring, not embeddings: exact name match scores highest, then terms in the name, the description, and the serialized input schema. Nothing in station_readings says "rainfall" in its name or description. It matched on a parameter called rainfall_mm. Parameter names turn out to be documentation, so search scores them.

Then it executes a batch:

What it costs, measured

I set up a server with identically shaped tools, four parameters each, and dumped the tools/list payload with and without a catalog. These are real bytes of JSON from that run:

Tools

Advertised directly

Through a catalog

Reduction

10

5,843

1,431

75.5%

20

11,703

1,431

87.8%

40

23,423

1,431

93.9%

100

58,585

1,431

97.6%

The right-hand column is the part to pay attention to. The catalog column does not move, because the payload is only ever search_tools and execute_tools no matter how many tools sit behind them. Everything you add after that is free at rest.

Treat the absolute numbers as a snapshot rather than a constant. My test tools are deliberately small, so they come out well under the thousand-tokens-per-tool figures real servers hit. The shape of the table is what transfers: one column grows with your API, the other does not.

What surprised me is where it starts paying off. I expected the crossover to sit somewhere around 30 or 40 tools. At 10 tools, it is already cutting three-quarters of the payload. My rough rule now is that past 10 or 15 tools, advertising all of them is a cost I would have to justify, not a default.

What this looks like in a real app

Weather tools make a tidy example. The case I actually care about is the app that wraps a chunk of its own API surface.

Say you expose your storefront to agents. You have tools for orders, customers, refunds, inventory, shipping, and a few reports. That is twenty tools before you have tried very hard, and an agent asking "where is order 4021" pays for the refund and inventory schemas anyway. Split it:

The two tools that get used constantly stay advertised. The other 18 are still fully available, still validated, still authorised, and cost nothing until something asks for them.

The part I am proudest of is duller than that. Each call in a batch becomes a real JsonRpcRequest and runs ``through ToolInvoker, the exact class a normal tools/call goes through. Conditional registration is re-checked at execution, not just at search, so a search result is never a capability grant.

That is the one thing here genuinely easier in a framework than in a generic runtime. A sandbox has to define and defend every hop from generated code back through an agent loop into your server. I got to skip all of it by not having any.

I did not invent this, and that is reassuring

I went through four versions to get here. Then I went looking to see whether anyone else had landed in the same place, and most of the industry had beaten me to it.

Cloudflare got there first and closest. Their Code Mode MCP server exposes 2,500 API endpoints through exactly two tools, search() and execute(), for about 1,000 tokens instead of 1.17 million. Same two-tool shape I arrived at, down to the names.

The difference is what execute() does. Theirs runs JavaScript in a V8 isolate with no filesystem and no outbound fetch by default. They did not avoid the sandbox; they moved it server-side so their users do not have to run one. That option is available to you when you are the platform. It is not available to a package that installs into everyone else's infrastructure, which is the whole reason my version four is a JSON batch and theirs is a language.

Anthropic shipped the other half as the Tool Search Tool: mark tools with defer_loading: true, and the model pulls full definitions on demand, for a reported 85% token reduction. There is an open MCP SEP for progressive disclosure proposing a standardized searchTools meta-tool. Speakeasy benchmarked the pattern. Gateway projects have done dynamic tool discovery for a while. It even has a name in the pattern catalogs: progressive tool discovery.

Finding all of that afterward was the most reassuring part of the project. Several independent groups converging on search plus deferred loading is much better evidence that the design is right than my own reasoning was.

The objection I keep getting

"Isn't search just more round trips? You added latency to save tokens."

Yes.

A search costs one round trip and returns a handful of schemas. In our server, the agent typically touches three tools in a session, so I pay for three or four schemas plus two meta-tool descriptions once, instead of 40 definitions on every turn of a 10-turn conversation. It clears break-even quickly and gets better the longer the conversation runs, because the definitions cost repeats per turn while the search cost does not.

Where it does not pay is on a small server where every tool gets used. Our own has six tools; I never moved into a catalog, because making the agent search for something it always needs is pure overhead. I actually run a mix of both.

I also do not think this beats Code Mode, and I have stopped thinking of it as competing. There are three points on the line, and what I shipped sits in the middle:

Advertise every tool directly: No infrastructure, no indirection, and you pay for every schema on every request. Fine below 10 tools.

Search and batch: No infrastructure, a flat payload whatever the catalog size, and no control flow. The agent can call five tools in one round trip but cannot loop, branch, or feed one result into the next.

Generated code in a sandbox: Full control flow, intermediate results filtered before they reach the model, and a sandbox you either operate or rent.

Each step buys capability and costs something. I shipped the middle one because it is the only one of the three that a package can hand you with nothing to install, and because most servers today are hurting from schema bloat rather than from an inability to write loops. If your agent genuinely needs to compute over your tools, the third option is the right answer, and I would not pretend otherwise.

What changed for me

I made the same mistake twice. In version one, my tests only ran the snippets I expected, so I never saw what the design permitted. In version three, my tests only used the constructs I had allowlisted, so the subset looked complete right up until something that had not read my allowlist started writing. Both times I was validating a design against my own expectations instead of against a model's, and both times it looked like passing tests.

That is what I actually took away, more than any of the architecture. When the user of your API is a language model, you are not the representative user, and your tests are not evidence.

I went in wanting to port an impressive piece of infrastructure and came out having deleted almost all of it, including an interpreter I was quite pleased with. What shipped is ranking and dispatch, no new dependencies, and nothing new to operate. The version that survived contact with a language model was the one with the least in it, and I have stopped treating the sophisticated option as the default target.

The domain-specific languages (DSLs) are small on purpose, and small enough to grow. Letting one call reference a previous call's output is the obvious next step. I would rather add it after watching people hit the limit than guess at it now, so if you try this and the JSON gets in your way, that is the report I want.

Reading that got me here: Cloudflare on Code Mode and their Code Mode MCP server, and Anthropic on code execution with MCP and advanced tool use.

Laravel is the most productive way to
 build, deploy, and monitor software.

By submitting this form, you agree to our terms. You can opt-out anytime.