# LM Studio MCP Web Search: 14K vs 360 Tokens

> LM Studio MCP for web search: the mcp.json block, a key in headers, and a tool list that fits. 28 tools cost 13,947 tokens; the lite set costs 356.

[Home](https://quanticdata.io/)/[Blog](https://quanticdata.io/blog/)/LM Studio MCP Web Search: 14K vs 360 Tokens

# LM Studio MCP Web Search: 14K vs 360 Tokens

MCP & agentsOct 2, 2026·5 min read·By [Aldo Morese](https://quanticdata.io/about/), founder of QuanticData

LM Studio mcp.json with the QuanticData web search server: url plus Authorization header, web tool set at 4,215 tokens instead of 13,947

On this page [What LM Studio needs for web search](/blog/lm-studio-mcp-web-search/#what-lm-studio-needs-for-web-search) [Step by step: the mcp.json block](/blog/lm-studio-mcp-web-search/#step-by-step-the-mcp-json-block) [Choose the tool set by context length](/blog/lm-studio-mcp-web-search/#choose-the-tool-set-by-context-length) [Keep each result small, too](/blog/lm-studio-mcp-web-search/#keep-each-result-small-too) [Which web tool answers which question](/blog/lm-studio-mcp-web-search/#which-web-tool-answers-which-question) [Check that it works](/blog/lm-studio-mcp-web-search/#check-that-it-works)

LM Studio supports MCP for web search since version 0.3.17: add the server to `mcp.json` with its URL and an Authorization header, then turn on tool use for a tool-calling model. Choose the tool list deliberately. Our 28 tools measured 13,947 tokens on 2 October 2026; the three web tools 4,215; one tool 356.

LM Studio's documentation says it plainly: some MCP servers were designed for Claude, ChatGPT or Gemini, use excessive amounts of tokens, and can bog down a local model with frequent context overflows. That is true of ours too if you load everything. This guide sets up web search in LM Studio and sizes the tool list to the context you actually run.

## What LM Studio needs for web search

Three things, all on your machine except the last:

- **LM Studio 0.3.17 or newer.** From that build LM Studio acts as an MCP host for both local (stdio) and remote (HTTP) servers.

- **A model trained for tool calling.** LM Studio passes tool definitions to any model, but only models trained for function calling use them reliably. Qwen, Granite and Nemotron families with tool support are a good start.

- **A web MCP server.** Here, the hosted QuanticData endpoint: it runs the searches and page fetches through residential IPs and returns Markdown or JSON, so nothing scrapes from your laptop.

## Step by step: the mcp.json block

Open the **Program** tab in the right-hand sidebar, click **Install**, then **Edit mcp.json**. LM Studio follows Cursor's notation, so a remote server is a `url` plus optional `headers`. Paste this, with your API key from the QuanticData dashboard:

```
{
  "mcpServers": {
    "quanticdata": {
      "url": "https://api.quanticdata.io/mcp?tools=web",
      "headers": {
        "Authorization": "Bearer YOUR_QUANTICDATA_API_KEY"
      }
    }
  }
}
```

If `mcp.json` already has other servers, copy only the `"quanticdata": { ... }` entry inside the existing `mcpServers` object, as LM Studio's docs advise. Save the file; LM Studio loads every server listed in it.

To run the server on your own machine instead, use the npm package over stdio; the tool filter is an environment variable (version 0.11.4 or later):

```
"quanticdata": {
  "command": "npx",
  "args": ["-y", "quanticdata-mcp"],
  "env": {
    "QUANTICDATA_API_KEY": "YOUR_QUANTICDATA_API_KEY",
    "QUANTICDATA_TOOLS": "web"
  }
}
```

## Choose the tool set by context length

Every request LM Studio sends to the model includes the definitions of every enabled tool. We serialised what our server returns and estimated tokens as characters divided by four:

| URL suffix | Tools the model sees | Definition tokens | Smallest sensible context |
| --- | --- | --- | --- |
| none (all tools) | 28, including crawl, datasets, collectors, proxies | 13,947 | 32K, better 64K |
| ?tools=research | search, search_and_read, scrape, map, batch, batch_status | 5,122 | 16K |
| ?tools=web | search, search_and_read, scrape | 4,215 | 8K |
| ?tools=lite | search_and_read | 356 | 4K |

Set the context when you load the model in LM Studio; a bigger window costs memory, so match it to the preset rather than the other way round. You can also combine names, for example `?tools=lite,collectors` adds the ready-made collectors (829 tokens) to the single search tool.

## Keep each result small, too

Definitions are paid on every turn; results only when the model calls a tool, but a single result can be larger than everything else. Measured on 2 October 2026, as returned to the model:

| Call | Tokens |
| --- | --- |
| search, one query | 1,642 |
| search_and_read, defaults (3 pages) | 3,986 |
| search_and_read, max_tokens 2,000 | 2,536 |
| scrape, one Wikipedia article, defaults | 48,291 |
| scrape, same article, query + max_tokens 1,500 | 2,390 |

Put the limits in LM Studio's system prompt, for example: "When you use search_and_read, set top_n to 2 and max_tokens to 1500. When you use scrape, always pass a query describing what you need and max_tokens 1500." The arguments are documented in the [API documentation](https://quanticdata.io/docs/). Why the content mode matters so much is measured across 11 pages in [our MCP web scraper token test](https://quanticdata.io/blog/mcp-web-scraper/).

## Which web tool answers which question

With the `web` set the model chooses between three tools. They overlap, so it helps to know what each one is for, and to say so in the system prompt:

| Question type | Tool | What comes back |
| --- | --- | --- |
| "What is the latest version of X and what changed?" | search_and_read | The top pages for a query, cleaned and merged into one bounded context with numbered sources |
| "Summarise this page" with a URL in the message | scrape | That one page as Markdown; pass a query to keep only the relevant sections |
| "Which sites rank for X?" or "find me five sources" | search | The results page as JSON: titles, links, snippets, and blocks such as People Also Ask when Google shows them |

For a small model, `search_and_read` alone answers most questions: it does the search and the reading in one call, so the model has only one decision to make. That is why the 4K preset contains nothing else. Add `scrape` when people paste links, and `search` when they want lists of sources rather than an answer.

## Check that it works

Ask a question the model cannot answer from training data, such as what a page published this week says. You should see a tool call to `search_and_read` or `search` in the chat, then an answer that cites the pages it read. If the model answers from memory instead:

- check that the model supports tool calling and that the server and its tools are listed in the Program tab;

- switch to a smaller tool set, because a long tool list can push your question out of a small context;

- an HTTP 401 error means the Authorization header is missing or the key is wrong.

The same server works in other local stacks: [giving Ollama internet access](https://quanticdata.io/blog/give-ollama-internet-access/) covers the terminal route with ollmcp, and [local LLM web search with MCP](https://quanticdata.io/blog/local-llm-web-search-mcp/) has the full budget for 4K to 64K contexts. Billing is per successful call; every account gets $2 of free API usage each month, and failed requests are never billed. All clients are listed on the [QuanticData MCP server](https://quanticdata.io/mcp-server/) page.

### Sources & further reading

- [Use MCP Servers, LM Studio documentation (fetched 2 October 2026)](https://lmstudio.ai/docs/app/mcp)

- [Web search for LMStudio?, r/LocalLLM](https://www.reddit.com/r/LocalLLM/comments/1ou52jp/web_search_for_lmstudio/)

- [Models with tool support, Ollama library](https://ollama.com/search?c=tools)

- [What is the Model Context Protocol (MCP)?, modelcontextprotocol.io](https://modelcontextprotocol.io/docs/2026-07-28/getting-started/intro)

## FAQ

Quick answers on lm studio mcp for web search.

[Something else? Ask us](mailto:hello@quanticdata.io)

### Does LM Studio support MCP?

Yes. Since version 0.3.17 LM Studio is an MCP host for local and remote servers. You add servers in mcp.json from the Program tab, using the same notation as Cursor.

### How do I add web search to LM Studio?

Add a web MCP server to mcp.json: a url such as https://api.quanticdata.io/mcp?tools=web and a headers object with Authorization set to Bearer and your API key. Then chat with a tool-calling model.

### Where is mcp.json in LM Studio?

Open the Program tab in the right-hand sidebar and choose Install, then Edit mcp.json. LM Studio opens the file in its own editor.

### Why does my LM Studio model run out of context with MCP?

Tool definitions are sent on every turn. Our full server measured 13,947 tokens of definitions on 2 October 2026. Load a preset such as ?tools=web (4,215 tokens) or ?tools=lite (356) and cap results with max_tokens.

### Which tool should LM Studio use for a URL I paste in the chat?

scrape. It returns that one page as Markdown. Ask the model to pass a query describing what you need and max_tokens 1500, so a long page does not fill the context: one Wikipedia article came back at 48,291 tokens by default and 2,390 with those two arguments.

### Can I run the MCP server locally instead of the hosted URL?

Yes. Use command npx with args -y quanticdata-mcp and put QUANTICDATA_API_KEY and QUANTICDATA_TOOLS in env. The pages are still fetched by QuanticData, not by your machine.

## Web search for LM Studio in one mcp.json block

Hosted MCP endpoint, a key in the headers, a tool list sized to your context. Every account gets $2 of free API usage each month, and failed requests are never billed.

[Start free — $2/month included](https://quanticdata.io/signup/)[Explore Web Scraping MCP Server for AI Agents](https://quanticdata.io/mcp-server/)

## Related reading

[MCP & agents Give Ollama Internet Access: 360-Token Setup Ollama runs models, it does not browse. To give an Ollama model internet access you put an MCP client in front of it and connect a web-search server. On 2 October 2026 our full tool list measured 13,947 tokens, more than Ollama's 4,096-token default, so the setup below loads one tool for 356 tokens and caps every result. Read more](https://quanticdata.io/blog/give-ollama-internet-access/) [MCP & agents Qwen Web Search MCP: Qwen3 Online via Ollama A local Qwen3 model gets web search through Qwen-Agent, which speaks MCP. The config has one trap: a server with a URL is treated as SSE unless you set the type to streamable-http. On 2 October 2026 our full tool list measured 13,947 tokens; one tool measured 356, small enough for a 4K Qwen on a laptop. Read more](https://quanticdata.io/blog/qwen-web-search-mcp/) [MCP & agents Open WebUI MCP Server: Web Search in 5 Steps Open WebUI has spoken MCP natively since v0.6.31, so an admin can give every Ollama model in the instance web search from one form. We measured what that costs the model on 2 October 2026: 13,947 tokens of definitions for all 28 of our tools, 4,215 for the three web tools, 356 for one. Read more](https://quanticdata.io/blog/open-webui-mcp-server/)

---

Source: https://quanticdata.io/blog/lm-studio-mcp-web-search/ · Site index for AI: https://quanticdata.io/llms.txt
