Documentation Python quickstart Blog Free tools Enterprise solutions hello@quanticdata.ioLog in

Local LLM Web Search MCP: What Fits in 4K

Two 4,096-token context windows: all 28 MCP tool definitions need 13,947 tokens, 340% of the window, while the lite preset needs 356 tokens plus a 1,858-token search result; the card shows 1,882 tokens left for the question and the answer.
A 4,096-token context window against MCP tool definitions: 28 tools need 13,947 tokens and overflow it, the lite preset needs 356 tokens plus a 1,858-token search result

A local LLM gets web search through an MCP server, but the server's tool list can eat the whole context before the first page arrives. On 2 October 2026 our 28 tool definitions measured 13,947 tokens, more than three times Ollama's 4,096-token default. Loading one tool costs 356 tokens and leaves room to read and answer.

That is the part most guides skip. They show the config file, the model calls the tool once, and the screenshot looks fine. Then the reader runs the same setup on an 8 GB laptop and the model forgets the question, ignores the tools or truncates its own answer. The cause is arithmetic, not the model. This post gives the numbers, then links the step-by-step setups for Ollama, LM Studio, Qwen and Open WebUI.

Why a local model needs an MCP server to reach the web

A model running in Ollama, LM Studio or llama.cpp only knows its training data. It has no browser and no network access of its own. What it can do, if it was trained for tool calling, is ask the application around it to run a function and hand back the result. The Model Context Protocol standardises that handshake: the application (the MCP host) connects to an MCP server, reads the list of tools the server offers, and shows those tool definitions to the model with every request.

A web-data MCP server such as the QuanticData MCP server exposes tools like search (a Google, Bing or DuckDuckGo results page as JSON), scrape (one URL as clean Markdown) and search_and_read (search, open the top pages and return one bounded context). The fetching happens on our side through residential IPs, so the laptop running the model needs nothing more than a network connection to the MCP endpoint.

Ollama itself is a model runner, not an MCP host. To give an Ollama model web access you put a host in front of it: a terminal client, a desktop app or a chat UI. LM Studio became an MCP host in version 0.3.17 and Open WebUI added native MCP in v0.6.31, according to their own documentation.

The tool list is the hidden cost: 13,947 tokens before page one

Every request a host sends to the model carries the tool definitions: name, description and the JSON schema of every argument. Rich tools have rich schemas. We serialised the tool list our MCP server returns and divided the characters by four, the usual estimate for English and JSON.

Tool setToolsDefinition tokensShare of a 4,096 contextShare of 32,768
All tools (default)2813,947340%43%
research preset65,122125%16%
web preset34,215103%13%
proxies preset41,80944%6%
collectors preset382920%3%
lite preset13569%1%

The two heaviest definitions are scrape at 2,209 tokens and search at 1,650, because they carry dozens of options: content modes, CSS extraction, geo-targeting, search verticals. A cloud model with a 200,000-token window never notices. A local model at 4,096 tokens cannot even hold them.

Ollama's documentation lists its defaults by graphics memory: 4k context below 24 GiB of VRAM, 32k between 24 and 48 GiB, 256k above that. The same page says tasks like web search, agents and coding tools should be set to at least 64,000 tokens. Most laptops are in the first row. LM Studio's MCP page gives the same warning in plain words: servers designed for Claude or ChatGPT can use excessive tokens and trigger frequent context overflows on a local model.

What the tools return, measured on real pages

Definitions are the fixed cost. Results are the variable one, and they depend on the arguments the model sends. These are the exact sizes the MCP tool returned to the model on 2 October 2026.

CallArgumentsTokens returned
searchone query, defaults1,642
search_and_readdefaults: 3 pages, 8,000-token cap3,986
search_and_readmax_tokens 2,0002,536
search_and_read2 pages, max_tokens 1,500 (hosted, lite)1,858
scrape, Wikipedia article on LLMsdefaults (smart Markdown)48,291
scrape, same articlelinks_mode strip27,383
scrape, same articlequery + max_tokens 1,5002,390
scrape, Python asyncio docs pagedefaults15,971
scrape, same pagequery + max_tokens 1,5001,839

A default scrape of one long article is 48,291 tokens, which fits no laptop context at all. The same page with a query (keep only the passages relevant to the question) and a max_tokens cap came back at 2,390. The cap counts the page content; JSON metadata around it adds a few hundred tokens, which is why a 2,000 cap returned 2,536. For more on how content modes change the token bill, see our test of 11 pages in four content modes.

A budget that works at 4K, 8K, 32K and 64K

The rule is simple: definitions plus one result plus your prompt plus the answer must fit. With the measured numbers:

ContextTool setFixed costSafe callLeft for prompt and answer
4,096lite356search_and_read, 2 pages, max_tokens 1,500 (1,858)1,882
8,192web4,215search (1,642) or a scrape with query + max_tokens 1,500 (about 2,400)1,500 to 2,300
32,768research5,122search_and_read with defaults (3,986) plus one capped scrapeabout 21,000
64,000 and upall 2813,947several reads per turnabout 40,000

Two settings move the most. First, raise the context if the hardware allows it: in Ollama that is the context slider in the app or OLLAMA_CONTEXT_LENGTH=64000 ollama serve, at the cost of more memory. Second, load fewer tools. Since version 0.11.4 our server accepts a tool list: add ?tools=lite (or web, research, collectors, proxies, or names such as search,scrape) to the hosted URL, or set QUANTICDATA_TOOLS when you run the npm package locally.

https://api.quanticdata.io/mcp?tools=lite
Authorization: Bearer YOUR_QUANTICDATA_API_KEY

Setups for each local stack

The connection is the same everywhere: the hosted endpoint over Streamable HTTP with your API key as a Bearer header, or the npm package over stdio. What changes is where each host keeps its config and how it lets you trim tools.

Pick a model that calls tools reliably

No setting helps a model that was not trained for tool calling. Ollama's library has a tools filter for exactly this; on 2 October 2026 it listed, among others, the Qwen3.6 and Qwen3.8 families, IBM Granite 4.1 in 3B, 8B and 30B sizes, and Nemotron 3.5 Lightning. Ollama's own web-search example uses qwen3:4b, which runs on most laptops. Open WebUI's documentation is blunt about the rest: MCP connects the tools, it does not improve the model's ability to use them, and a weak model may ignore tools or send wrong arguments. A small model also has to pick the right tool out of everything it is shown, so a preset with one or three tools gives it an easier job as well as a cheaper one.

Billing follows the call, not the model: every QuanticData account gets $2 of free API usage each month, and failed requests are never billed. The full tool reference, with every argument, is in the API documentation.

Sources & further reading

FAQ

Quick answers on local llm web search mcp.

Something else? Ask us

Can a local LLM search the web?

Yes, if it supports tool calling and runs inside an MCP host such as LM Studio, Open WebUI or the ollmcp terminal client. The host connects to a web-search MCP server, the model asks for a search, and the host returns the results. The model itself never opens a connection.

How many tokens does an MCP server add to every request?

It depends on the tool list. Our full server sends 28 tool definitions worth 13,947 tokens, measured on 2 October 2026. The lite preset sends one tool for 356 tokens, the web preset three tools for 4,215.

What context length do I need for web search in Ollama?

Ollama recommends at least 64,000 tokens for web search, agents and coding tools, and defaults to 4k below 24 GiB of VRAM. At 4,096 tokens, use one tool (the lite preset) and cap each result with max_tokens.

Do I need to run the MCP server on my own machine?

No. The hosted endpoint at api.quanticdata.io/mcp does the fetching; your host only needs network access to it and your API key. You can also run the npm package locally over stdio if your host prefers a command over a URL.

Which local models are good at calling tools?

Use a model tagged for tools in the Ollama library. On 2 October 2026 that list included Qwen3.6, Qwen3.8 and IBM Granite 4.1, and Ollama uses qwen3:4b in its own web-search example. A short tool list also makes the choice easier for a small model.

Is web search free for local models?

Every QuanticData account gets $2 of free API usage each month, and failed requests are never billed. After that, calls are charged per successful request against the account balance.

Web search that fits a 4K local context

One tool, 356 tokens of definitions, bounded results. Add ?tools=lite to the MCP URL. Every account gets $2 of free API usage each month, and failed requests are never billed.

Related reading