Out of 7 MCP servers included into a ModelSpecs agent definition and activated, only three are known (partially) by the model #16132
Replies: 9 comments 3 replies
|
Self-reporting isn't a reliable check here: models frequently list only a
If the payload contains all 7 servers' tools, the model is misreporting and |
|
@danny-avila You should be aware, that one part of the problem is the native tool calling problem with ollama. Some modelfiles for ollama simply direct the native calls to the prompt without defining a tool calling scheme. This can differ from model to model and from version to version. It is just a matter of wrapping the model with a proper modelfile for ollama. |
|
@danny-avila : my financial analyst just reported that he has 4 MCP servers and 30 tools, whereas there are 7 MCP servers with around 80 tools. It looks like somewhere 30 tools could be a hard-coded limit. |
|
@danny-avila i have tried to use agent builder instead of modelSpecs for defining my research assistant and it is the same. The context window shows about 11K space used by MCP servers and tools, which seems to be the capped list of about 30 tools and the model itself does not know any of the other tools. Being not able to use MCP servers, completely spoils the fun of having LibreChat installed. |
|
@danny-avila : Ok, 11K of context is about enough for 30 tools, so it is LibreChat capping off the MCP tools and that for good reasons, but then i was right in another idea i had before: that LibreChat would need a prompt engineering system similar to that of SillyTavern, that is aware of context and which actions are possible and thus injects world lore and specialities. In the context of SillyTavern role playing this seems to work reliably to focus on the exact scenario in order to provide e.g. the perfect list of 30 MCP tools necessary for the moment. When does context confusion happen because of too many MCP tools? Is it possible to have 100 tools let us say with 256K Qwen3.6:35b-a3b on Ollama, or let us say 64K to keep the KV Cache and TTFT small. Or asked differently: Can a model in a chat be made practically stateless and the conversation being compactified practically at every step in order to inject some clues to the model what scenario it is about to work through and to break down more complicated workflows, e.g. generate and image and then post-edit it or similar. |
|
@danny-avila I have talked to Google Gemini about it: In LibreChat v0.8.8-rc3, registered MCP tool schemas quickly consume the system prompt context budget—hitting roughly 11K tokens for ~30 tools—which hard-caps the total number of active tools a user can attach before reaching internal schema overhead limits. While static schema injection guarantees prefix caching benefits across turns, loading large tool sets upfront creates severe context bloat and degrades performance; adopting Just-In-Time (JIT) dynamic tool routing or suffix-based injection would decouple total available tools from system prompt size, preserving token space and preventing early registration caps. JIT dynamic tool routing would require to have a prompt compositor, that decomposes the user intent into smaller tasks, which then will be processed in proper sequence (or parallel if applicable), with small tool sets injected at the end of the prompt preserving the fixed part of the prompt in the beginning to have some KV cache hits and low prefill. |
|
@danny-avila I just re-built my general research agent with the agent-builder in v0.8.8-rc3 and checked the "deferred tool load" box for all the mcp servers i have, especially for my ComfyUI MCP server and when i asked the agent to draw a painting for me, it sent me an excuse that it is just a text model and gave the prompt for copy and paste instead of making use of the ComfyUI MCP tools. -> Yes, i really think that there is a serious problem with MCP tool calling in LibreChat. The deferral of tools seems not to work properly and the capping of the number of MCP tools is definitely a killer to any useful application of the chat logic of LibreChat v0.8.8-rc3. Personally, i don't know how exactly the MCP calling mechanism works or how tool deferral would actually work, or for that sake how a prompt compositor would actually be able to provide the proper sequence of actions from the text and choose the proper MCP calls, but looking e.g. at Google Gemini and others, this is a problem that has indeed already been successfully solved (decomposing the user intent). |
|
@danny-avila also, the native Web Search tool with firecrawl yields very bad search results, much worse results than what a SearXNG MCP server produces. |
|
The ~30 tools / ~11K token ceiling you're hitting isn't a tool-count limit, it's a token budget, and it's dominated by argument JSON schemas, not names and descriptions. There's a PR open right now (#16179) from someone who measured their own catalog: 8 servers, 202 tools, 239,613 tokens of tool definitions, and 92.6% of that was argument schemas. One server with 12 tools accounted for 67% of the total, with a single tool's schema running 79,699 tokens on its own. Strip the schemas down to name plus description and the same catalog drops to 17,822 tokens. So your 30-tool ceiling is really "how many schemas fit before the budget runs out," and it moves depending on which servers you pick, which matches what you already suspected. Deferred loading, the mechanism you'd want, already exists in the codebase. It's not something that needs to be invented. A tool can be marked defer_loading: true, in which case the model gets the tool_search tool plus a name-only listing instead of the full schema, and pulls the schema in on demand when it decides to call that tool. But two things currently limit it, both still open as PRs against dev as of today:
So the architecture you were describing from the Gemini conversation, name and description now, schema fetched just-in-time, is exactly what defer_loading already does. What's missing in your setup is most likely the capability flag being off, plus there's currently no way to apply deferral without hand-curating tools one at a time. Worth checking whether the deferred_tools capability is enabled for your agent first, since that alone might unblock the ComfyUI case before either PR merges. |
Uh oh!
There was an error while loading. Please reload this page.
What happened?
I have defined my research assistant agent via modelSpecs: in librechat.yaml. For that i choose 7 MCP servers out of the list of my MCP servers, which initialize nicely and provide the tools they should provides as LibreChat v0.8.8-rc3 reports on startup as its last message. But then when i actually ask my research assistant to use any of the MCP servers, it does not know about the majority of them, just about one MCP server with more tools is there upon having the model self-report, which MCP servers it knows and the others are not listed. It seems to be truncated or unwisely selected.
Version Information
v0.8.8-rc3
Steps to Reproduce
What browsers are you seeing the problem on?
Firefox
Relevant log output
Screenshots
Code of Conduct
All reactions