# Local LLM Web Search MCP: What Fits in 4K

> Local LLM web search over MCP, measured: 28 tool definitions cost 13,947 tokens, more than Ollama

[Home](https://quanticdata.io/)/[Blog](https://quanticdata.io/blog/)/Local LLM Web Search MCP: What Fits in 4K

# Local LLM Web Search MCP: What Fits in 4K

MCP & agentsOct 2, 2026·7 min read·By [Aldo Morese](https://quanticdata.io/about/), founder of QuanticData

A 4,096-token context window against MCP tool definitions: 28 tools need 13,947 tokens and overflow it, the lite preset needs 356 tokens plus a 1,858-token search result

On this page [Why a local model needs an MCP server to reach the web](/blog/local-llm-web-search-mcp/#why-a-local-model-needs-an-mcp-server-to-reach-the-web) [The tool list is the hidden cost: 13,947 tokens before page one](/blog/local-llm-web-search-mcp/#the-tool-list-is-the-hidden-cost-13-947-tokens-before-page-o) [What the tools return, measured on real pages](/blog/local-llm-web-search-mcp/#what-the-tools-return-measured-on-real-pages) [A budget that works at 4K, 8K, 32K and 64K](/blog/local-llm-web-search-mcp/#a-budget-that-works-at-4k-8k-32k-and-64k) [Setups for each local stack](/blog/local-llm-web-search-mcp/#setups-for-each-local-stack) [Pick a model that calls tools reliably](/blog/local-llm-web-search-mcp/#pick-a-model-that-calls-tools-reliably)

A local LLM gets web search through an MCP server, but the server's tool list can eat the whole context before the first page arrives. On 2 October 2026 our 28 tool definitions measured 13,947 tokens, more than three times Ollama's 4,096-token default. Loading one tool costs 356 tokens and leaves room to read and answer.

That is the part most guides skip. They show the config file, the model calls the tool once, and the screenshot looks fine. Then the reader runs the same setup on an 8 GB laptop and the model forgets the question, ignores the tools or truncates its own answer. The cause is arithmetic, not the model. This post gives the numbers, then links the step-by-step setups for Ollama, LM Studio, Qwen and Open WebUI.

## Why a local model needs an MCP server to reach the web

A model running in Ollama, LM Studio or llama.cpp only knows its training data. It has no browser and no network access of its own. What it can do, if it was trained for tool calling, is ask the application around it to run a function and hand back the result. The Model Context Protocol standardises that handshake: the application (the MCP host) connects to an MCP server, reads the list of tools the server offers, and shows those tool definitions to the model with every request.

A web-data MCP server such as the [QuanticData MCP server](https://quanticdata.io/mcp-server/) exposes tools like `search` (a Google, Bing or DuckDuckGo results page as JSON), `scrape` (one URL as clean Markdown) and `search_and_read` (search, open the top pages and return one bounded context). The fetching happens on our side through residential IPs, so the laptop running the model needs nothing more than a network connection to the MCP endpoint.

Ollama itself is a model runner, not an MCP host. To give an Ollama model web access you put a host in front of it: a terminal client, a desktop app or a chat UI. LM Studio became an MCP host in version 0.3.17 and Open WebUI added native MCP in v0.6.31, according to their own documentation.

## The tool list is the hidden cost: 13,947 tokens before page one

Every request a host sends to the model carries the tool definitions: name, description and the JSON schema of every argument. Rich tools have rich schemas. We serialised the tool list our MCP server returns and divided the characters by four, the usual estimate for English and JSON.

| Tool set | Tools | Definition tokens | Share of a 4,096 context | Share of 32,768 |
| --- | --- | --- | --- | --- |
| All tools (default) | 28 | 13,947 | 340% | 43% |
| research preset | 6 | 5,122 | 125% | 16% |
| web preset | 3 | 4,215 | 103% | 13% |
| proxies preset | 4 | 1,809 | 44% | 6% |
| collectors preset | 3 | 829 | 20% | 3% |
| **lite preset** | **1** | **356** | **9%** | **1%** |

The two heaviest definitions are `scrape` at 2,209 tokens and `search` at 1,650, because they carry dozens of options: content modes, CSS extraction, geo-targeting, search verticals. A cloud model with a 200,000-token window never notices. A local model at 4,096 tokens cannot even hold them.

Ollama's documentation lists its defaults by graphics memory: 4k context below 24 GiB of VRAM, 32k between 24 and 48 GiB, 256k above that. The same page says tasks like web search, agents and coding tools should be set to at least 64,000 tokens. Most laptops are in the first row. LM Studio's MCP page gives the same warning in plain words: servers designed for Claude or ChatGPT can use excessive tokens and trigger frequent context overflows on a local model.

## What the tools return, measured on real pages

Definitions are the fixed cost. Results are the variable one, and they depend on the arguments the model sends. These are the exact sizes the MCP tool returned to the model on 2 October 2026.

| Call | Arguments | Tokens returned |
| --- | --- | --- |
| search | one query, defaults | 1,642 |
| search_and_read | defaults: 3 pages, 8,000-token cap | 3,986 |
| search_and_read | max_tokens 2,000 | 2,536 |
| search_and_read | 2 pages, max_tokens 1,500 (hosted, lite) | 1,858 |
| scrape, Wikipedia article on LLMs | defaults (smart Markdown) | 48,291 |
| scrape, same article | links_mode strip | 27,383 |
| scrape, same article | query + max_tokens 1,500 | 2,390 |
| scrape, Python asyncio docs page | defaults | 15,971 |
| scrape, same page | query + max_tokens 1,500 | 1,839 |

A default scrape of one long article is 48,291 tokens, which fits no laptop context at all. The same page with a `query` (keep only the passages relevant to the question) and a `max_tokens` cap came back at 2,390. The cap counts the page content; JSON metadata around it adds a few hundred tokens, which is why a 2,000 cap returned 2,536. For more on how content modes change the token bill, see [our test of 11 pages in four content modes](https://quanticdata.io/blog/mcp-web-scraper/).

## A budget that works at 4K, 8K, 32K and 64K

The rule is simple: definitions plus one result plus your prompt plus the answer must fit. With the measured numbers:

| Context | Tool set | Fixed cost | Safe call | Left for prompt and answer |
| --- | --- | --- | --- | --- |
| 4,096 | lite | 356 | search_and_read, 2 pages, max_tokens 1,500 (1,858) | 1,882 |
| 8,192 | web | 4,215 | search (1,642) or a scrape with query + max_tokens 1,500 (about 2,400) | 1,500 to 2,300 |
| 32,768 | research | 5,122 | search_and_read with defaults (3,986) plus one capped scrape | about 21,000 |
| 64,000 and up | all 28 | 13,947 | several reads per turn | about 40,000 |

Two settings move the most. First, raise the context if the hardware allows it: in Ollama that is the context slider in the app or `OLLAMA_CONTEXT_LENGTH=64000 ollama serve`, at the cost of more memory. Second, load fewer tools. Since version 0.11.4 our server accepts a tool list: add `?tools=lite` (or `web`, `research`, `collectors`, `proxies`, or names such as `search,scrape`) to the hosted URL, or set `QUANTICDATA_TOOLS` when you run the npm package locally.

```
https://api.quanticdata.io/mcp?tools=lite
Authorization: Bearer YOUR_QUANTICDATA_API_KEY
```

## Setups for each local stack

The connection is the same everywhere: the hosted endpoint over Streamable HTTP with your API key as a Bearer header, or the npm package over stdio. What changes is where each host keeps its config and how it lets you trim tools.

- **Ollama in a terminal**: the ollmcp client adds a server in one command and lets you switch tools on and off mid-chat. Step by step in [how to give Ollama internet access](https://quanticdata.io/blog/give-ollama-internet-access/).

- **LM Studio**: one block in `mcp.json`, the same notation Cursor uses. See [LM Studio MCP web search](https://quanticdata.io/blog/lm-studio-mcp-web-search/).

- **Qwen3 with Qwen-Agent**: a Python agent on top of a Qwen model served by Ollama, with the MCP server in `function_list`. See [Qwen web search with MCP](https://quanticdata.io/blog/qwen-web-search-mcp/).

- **Open WebUI**: an admin adds the server under External Tool Servers as MCP (Streamable HTTP). See [Open WebUI MCP server setup](https://quanticdata.io/blog/open-webui-mcp-server/).

## Pick a model that calls tools reliably

No setting helps a model that was not trained for tool calling. Ollama's library has a tools filter for exactly this; on 2 October 2026 it listed, among others, the Qwen3.6 and Qwen3.8 families, IBM Granite 4.1 in 3B, 8B and 30B sizes, and Nemotron 3.5 Lightning. Ollama's own web-search example uses `qwen3:4b`, which runs on most laptops. Open WebUI's documentation is blunt about the rest: MCP connects the tools, it does not improve the model's ability to use them, and a weak model may ignore tools or send wrong arguments. A small model also has to pick the right tool out of everything it is shown, so a preset with one or three tools gives it an easier job as well as a cheaper one.

Billing follows the call, not the model: every QuanticData account gets $2 of free API usage each month, and failed requests are never billed. The full tool reference, with every argument, is in the [API documentation](https://quanticdata.io/docs/).

### Sources & further reading

- [Context length, Ollama documentation (fetched 2 October 2026)](https://docs.ollama.com/context-length)

- [Models with tool support, Ollama library](https://ollama.com/search?c=tools)

- [Use MCP Servers, LM Studio documentation](https://lmstudio.ai/docs/app/mcp)

- [Model Context Protocol (MCP), Open WebUI documentation](https://docs.openwebui.com/features/extensibility/mcp/)

- [What is the Model Context Protocol (MCP)?, modelcontextprotocol.io](https://modelcontextprotocol.io/docs/2026-07-28/getting-started/intro)

- [What is the most effective way to have your local LLM search the web?, r/LocalLLaMA](https://www.reddit.com/r/LocalLLaMA/comments/1n9sod7/what_is_the_most_effective_way_to_have_your_local/)

## FAQ

Quick answers on local llm web search mcp.

[Something else? Ask us](mailto:hello@quanticdata.io)

### Can a local LLM search the web?

Yes, if it supports tool calling and runs inside an MCP host such as LM Studio, Open WebUI or the ollmcp terminal client. The host connects to a web-search MCP server, the model asks for a search, and the host returns the results. The model itself never opens a connection.

### How many tokens does an MCP server add to every request?

It depends on the tool list. Our full server sends 28 tool definitions worth 13,947 tokens, measured on 2 October 2026. The lite preset sends one tool for 356 tokens, the web preset three tools for 4,215.

### What context length do I need for web search in Ollama?

Ollama recommends at least 64,000 tokens for web search, agents and coding tools, and defaults to 4k below 24 GiB of VRAM. At 4,096 tokens, use one tool (the lite preset) and cap each result with max_tokens.

### Do I need to run the MCP server on my own machine?

No. The hosted endpoint at api.quanticdata.io/mcp does the fetching; your host only needs network access to it and your API key. You can also run the npm package locally over stdio if your host prefers a command over a URL.

### Which local models are good at calling tools?

Use a model tagged for tools in the Ollama library. On 2 October 2026 that list included Qwen3.6, Qwen3.8 and IBM Granite 4.1, and Ollama uses qwen3:4b in its own web-search example. A short tool list also makes the choice easier for a small model.

### Is web search free for local models?

Every QuanticData account gets $2 of free API usage each month, and failed requests are never billed. After that, calls are charged per successful request against the account balance.

## Web search that fits a 4K local context

One tool, 356 tokens of definitions, bounded results. Add ?tools=lite to the MCP URL. Every account gets $2 of free API usage each month, and failed requests are never billed.

[Start free — $2/month included](https://quanticdata.io/signup/)[Explore Web Scraping MCP Server for AI Agents](https://quanticdata.io/mcp-server/)

## Related reading

[MCP & agents LM Studio MCP Web Search: 14K vs 360 Tokens LM Studio has been an MCP host since version 0.3.17, so web search is one block in mcp.json. The catch is in LM Studio's own docs: servers built for cloud models can flood a local context. We measured ours on 2 October 2026: 13,947 tokens for all 28 tools, 4,215 for the three web tools, 356 for one. Read more](https://quanticdata.io/blog/lm-studio-mcp-web-search/) [MCP & agents Qwen Web Search MCP: Qwen3 Online via Ollama A local Qwen3 model gets web search through Qwen-Agent, which speaks MCP. The config has one trap: a server with a URL is treated as SSE unless you set the type to streamable-http. On 2 October 2026 our full tool list measured 13,947 tokens; one tool measured 356, small enough for a 4K Qwen on a laptop. Read more](https://quanticdata.io/blog/qwen-web-search-mcp/) [MCP & agents Open WebUI MCP Server: Web Search in 5 Steps Open WebUI has spoken MCP natively since v0.6.31, so an admin can give every Ollama model in the instance web search from one form. We measured what that costs the model on 2 October 2026: 13,947 tokens of definitions for all 28 of our tools, 4,215 for the three web tools, 356 for one. Read more](https://quanticdata.io/blog/open-webui-mcp-server/)

---

Source: https://quanticdata.io/blog/local-llm-web-search-mcp/ · Site index for AI: https://quanticdata.io/llms.txt
