DCAINIJul 17

Scalable LLM Agent Tool Access in the Cloud

arXiv:2607.1559317.4h-index: 7
Predicted impact top 4% in DC · last 90 daysOriginality Incremental advance
AI Analysis

For cloud providers deploying LLM agents with tool calling, this work addresses scalability and compatibility bottlenecks in MCP-based systems.

The paper tackles the challenge of scaling LLM agent tool access in the cloud using the Model Context Protocol (MCP). The proposed gateway system achieves 98% Top-15 recall, scales to 3,000+ tools, reduces tool selection time by 8.9× and token usage by 23.8×.

LLM agents increasingly rely on tool calling to act on external systems, and the Model Context Protocol (MCP) has quickly become its de facto interface. Operating MCP at cloud scale, however, becomes difficult. On the tool provider side, legacy services are not directly callable through MCP; the rapid protocol development also creates ongoing compatibility cost. On the agent side, the number of accessible tool is limited by the LLM context window and inference overhead; mounting a large tool set increases token usage and inference latency and can reduce task success rate. Moreover, for stateful MCP backends with multiple replicas, preserving session affinity increases client-side complexity. We present a cloud-scale gateway system for MCP service. It breaks the direct-connect model on the data plane and offloads legacy service integration, consolidating incompatible MCP variants, access control, tool recommendation, and session-aware routing to the gateway. Hybrid retrieval sustains 98% Top-15 recall; it scales agent tool access to 3,000+ with high tool selection accuracy, and reduces tool selection time by $8.9\times$ and token usage by $23.8\times$, with low per-call overhead, stable under scale-out. Finally, we share the lessons learned from deploying the gateway system in production.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes