A local LLM gets web search through an MCP server, but the server's tool list can eat the whole context before the first page arrives. On 2 October 2026 our 28 tool definitions measured 13,947 tokens, more than three times Ollama's 4,096-token default. Loading one tool costs 356 tokens and leaves room to read and answer.
That is the part most guides skip. They show the config file, the model calls the tool once, and the screenshot looks fine. Then the reader runs the same setup on an 8 GB laptop and the model forgets the question, ignores the tools or truncates its own answer. The cause is arithmetic, not the model. This post gives the numbers, then links the step-by-step setups for Ollama, LM Studio, Qwen and Open WebUI.
Why a local model needs an MCP server to reach the web
A model running in Ollama, LM Studio or llama.cpp only knows its training data. It has no browser and no network access of its own. What it can do, if it was trained for tool calling, is ask the application around it to run a function and hand back the result. The Model Context Protocol standardises that handshake: the application (the MCP host) connects to an MCP server, reads the list of tools the server offers, and shows those tool definitions to the model with every request.
A web-data MCP server such as the QuanticData MCP server exposes tools like search (a Google, Bing or DuckDuckGo results page as JSON), scrape (one URL as clean Markdown) and search_and_read (search, open the top pages and return one bounded context). The fetching happens on our side through residential IPs, so the laptop running the model needs nothing more than a network connection to the MCP endpoint.
Ollama itself is a model runner, not an MCP host. To give an Ollama model web access you put a host in front of it: a terminal client, a desktop app or a chat UI. LM Studio became an MCP host in version 0.3.17 and Open WebUI added native MCP in v0.6.31, according to their own documentation.
The tool list is the hidden cost: 13,947 tokens before page one
Every request a host sends to the model carries the tool definitions: name, description and the JSON schema of every argument. Rich tools have rich schemas. We serialised the tool list our MCP server returns and divided the characters by four, the usual estimate for English and JSON.
| Tool set | Tools | Definition tokens | Share of a 4,096 context | Share of 32,768 |
|---|---|---|---|---|
| All tools (default) | 28 | 13,947 | 340% | 43% |
| research preset | 6 | 5,122 | 125% | 16% |
| web preset | 3 | 4,215 | 103% | 13% |
| proxies preset | 4 | 1,809 | 44% | 6% |
| collectors preset | 3 | 829 | 20% | 3% |
| lite preset | 1 | 356 | 9% | 1% |
The two heaviest definitions are scrape at 2,209 tokens and search at 1,650, because they carry dozens of options: content modes, CSS extraction, geo-targeting, search verticals. A cloud model with a 200,000-token window never notices. A local model at 4,096 tokens cannot even hold them.
Ollama's documentation lists its defaults by graphics memory: 4k context below 24 GiB of VRAM, 32k between 24 and 48 GiB, 256k above that. The same page says tasks like web search, agents and coding tools should be set to at least 64,000 tokens. Most laptops are in the first row. LM Studio's MCP page gives the same warning in plain words: servers designed for Claude or ChatGPT can use excessive tokens and trigger frequent context overflows on a local model.
What the tools return, measured on real pages
Definitions are the fixed cost. Results are the variable one, and they depend on the arguments the model sends. These are the exact sizes the MCP tool returned to the model on 2 October 2026.
| Call | Arguments | Tokens returned |
|---|---|---|
| search | one query, defaults | 1,642 |
| search_and_read | defaults: 3 pages, 8,000-token cap | 3,986 |
| search_and_read | max_tokens 2,000 | 2,536 |
| search_and_read | 2 pages, max_tokens 1,500 (hosted, lite) | 1,858 |
| scrape, Wikipedia article on LLMs | defaults (smart Markdown) | 48,291 |
| scrape, same article | links_mode strip | 27,383 |
| scrape, same article | query + max_tokens 1,500 | 2,390 |
| scrape, Python asyncio docs page | defaults | 15,971 |
| scrape, same page | query + max_tokens 1,500 | 1,839 |
A default scrape of one long article is 48,291 tokens, which fits no laptop context at all. The same page with a query (keep only the passages relevant to the question) and a max_tokens cap came back at 2,390. The cap counts the page content; JSON metadata around it adds a few hundred tokens, which is why a 2,000 cap returned 2,536. For more on how content modes change the token bill, see our test of 11 pages in four content modes.
A budget that works at 4K, 8K, 32K and 64K
The rule is simple: definitions plus one result plus your prompt plus the answer must fit. With the measured numbers:
| Context | Tool set | Fixed cost | Safe call | Left for prompt and answer |
|---|---|---|---|---|
| 4,096 | lite | 356 | search_and_read, 2 pages, max_tokens 1,500 (1,858) | 1,882 |
| 8,192 | web | 4,215 | search (1,642) or a scrape with query + max_tokens 1,500 (about 2,400) | 1,500 to 2,300 |
| 32,768 | research | 5,122 | search_and_read with defaults (3,986) plus one capped scrape | about 21,000 |
| 64,000 and up | all 28 | 13,947 | several reads per turn | about 40,000 |
Two settings move the most. First, raise the context if the hardware allows it: in Ollama that is the context slider in the app or OLLAMA_CONTEXT_LENGTH=64000 ollama serve, at the cost of more memory. Second, load fewer tools. Since version 0.11.4 our server accepts a tool list: add ?tools=lite (or web, research, collectors, proxies, or names such as search,scrape) to the hosted URL, or set QUANTICDATA_TOOLS when you run the npm package locally.
https://api.quanticdata.io/mcp?tools=lite
Authorization: Bearer YOUR_QUANTICDATA_API_KEY
Setups for each local stack
The connection is the same everywhere: the hosted endpoint over Streamable HTTP with your API key as a Bearer header, or the npm package over stdio. What changes is where each host keeps its config and how it lets you trim tools.
- Ollama in a terminal: the ollmcp client adds a server in one command and lets you switch tools on and off mid-chat. Step by step in how to give Ollama internet access.
- LM Studio: one block in
mcp.json, the same notation Cursor uses. See LM Studio MCP web search. - Qwen3 with Qwen-Agent: a Python agent on top of a Qwen model served by Ollama, with the MCP server in
function_list. See Qwen web search with MCP. - Open WebUI: an admin adds the server under External Tool Servers as MCP (Streamable HTTP). See Open WebUI MCP server setup.
Pick a model that calls tools reliably
No setting helps a model that was not trained for tool calling. Ollama's library has a tools filter for exactly this; on 2 October 2026 it listed, among others, the Qwen3.6 and Qwen3.8 families, IBM Granite 4.1 in 3B, 8B and 30B sizes, and Nemotron 3.5 Lightning. Ollama's own web-search example uses qwen3:4b, which runs on most laptops. Open WebUI's documentation is blunt about the rest: MCP connects the tools, it does not improve the model's ability to use them, and a weak model may ignore tools or send wrong arguments. A small model also has to pick the right tool out of everything it is shown, so a preset with one or three tools gives it an easier job as well as a cheaper one.
Billing follows the call, not the model: every QuanticData account gets $2 of free API usage each month, and failed requests are never billed. The full tool reference, with every argument, is in the API documentation.
Sources & further reading
- Context length, Ollama documentation (fetched 2 October 2026)
- Models with tool support, Ollama library
- Use MCP Servers, LM Studio documentation
- Model Context Protocol (MCP), Open WebUI documentation
- What is the Model Context Protocol (MCP)?, modelcontextprotocol.io
- What is the most effective way to have your local LLM search the web?, r/LocalLLaMA