LM Studio supports MCP for web search since version 0.3.17: add the server to mcp.json with its URL and an Authorization header, then turn on tool use for a tool-calling model. Choose the tool list deliberately. Our 28 tools measured 13,947 tokens on 2 October 2026; the three web tools 4,215; one tool 356.
LM Studio's documentation says it plainly: some MCP servers were designed for Claude, ChatGPT or Gemini, use excessive amounts of tokens, and can bog down a local model with frequent context overflows. That is true of ours too if you load everything. This guide sets up web search in LM Studio and sizes the tool list to the context you actually run.
What LM Studio needs for web search
Three things, all on your machine except the last:
- LM Studio 0.3.17 or newer. From that build LM Studio acts as an MCP host for both local (stdio) and remote (HTTP) servers.
- A model trained for tool calling. LM Studio passes tool definitions to any model, but only models trained for function calling use them reliably. Qwen, Granite and Nemotron families with tool support are a good start.
- A web MCP server. Here, the hosted QuanticData endpoint: it runs the searches and page fetches through residential IPs and returns Markdown or JSON, so nothing scrapes from your laptop.
Step by step: the mcp.json block
Open the Program tab in the right-hand sidebar, click Install, then Edit mcp.json. LM Studio follows Cursor's notation, so a remote server is a url plus optional headers. Paste this, with your API key from the QuanticData dashboard:
{
"mcpServers": {
"quanticdata": {
"url": "https://api.quanticdata.io/mcp?tools=web",
"headers": {
"Authorization": "Bearer YOUR_QUANTICDATA_API_KEY"
}
}
}
}
If mcp.json already has other servers, copy only the "quanticdata": { ... } entry inside the existing mcpServers object, as LM Studio's docs advise. Save the file; LM Studio loads every server listed in it.
To run the server on your own machine instead, use the npm package over stdio; the tool filter is an environment variable (version 0.11.4 or later):
"quanticdata": {
"command": "npx",
"args": ["-y", "quanticdata-mcp"],
"env": {
"QUANTICDATA_API_KEY": "YOUR_QUANTICDATA_API_KEY",
"QUANTICDATA_TOOLS": "web"
}
}
Choose the tool set by context length
Every request LM Studio sends to the model includes the definitions of every enabled tool. We serialised what our server returns and estimated tokens as characters divided by four:
| URL suffix | Tools the model sees | Definition tokens | Smallest sensible context |
|---|---|---|---|
| none (all tools) | 28, including crawl, datasets, collectors, proxies | 13,947 | 32K, better 64K |
| ?tools=research | search, search_and_read, scrape, map, batch, batch_status | 5,122 | 16K |
| ?tools=web | search, search_and_read, scrape | 4,215 | 8K |
| ?tools=lite | search_and_read | 356 | 4K |
Set the context when you load the model in LM Studio; a bigger window costs memory, so match it to the preset rather than the other way round. You can also combine names, for example ?tools=lite,collectors adds the ready-made collectors (829 tokens) to the single search tool.
Keep each result small, too
Definitions are paid on every turn; results only when the model calls a tool, but a single result can be larger than everything else. Measured on 2 October 2026, as returned to the model:
| Call | Tokens |
|---|---|
| search, one query | 1,642 |
| search_and_read, defaults (3 pages) | 3,986 |
| search_and_read, max_tokens 2,000 | 2,536 |
| scrape, one Wikipedia article, defaults | 48,291 |
| scrape, same article, query + max_tokens 1,500 | 2,390 |
Put the limits in LM Studio's system prompt, for example: "When you use search_and_read, set top_n to 2 and max_tokens to 1500. When you use scrape, always pass a query describing what you need and max_tokens 1500." The arguments are documented in the API documentation. Why the content mode matters so much is measured across 11 pages in our MCP web scraper token test.
Which web tool answers which question
With the web set the model chooses between three tools. They overlap, so it helps to know what each one is for, and to say so in the system prompt:
| Question type | Tool | What comes back |
|---|---|---|
| "What is the latest version of X and what changed?" | search_and_read | The top pages for a query, cleaned and merged into one bounded context with numbered sources |
| "Summarise this page" with a URL in the message | scrape | That one page as Markdown; pass a query to keep only the relevant sections |
| "Which sites rank for X?" or "find me five sources" | search | The results page as JSON: titles, links, snippets, and blocks such as People Also Ask when Google shows them |
For a small model, search_and_read alone answers most questions: it does the search and the reading in one call, so the model has only one decision to make. That is why the 4K preset contains nothing else. Add scrape when people paste links, and search when they want lists of sources rather than an answer.
Check that it works
Ask a question the model cannot answer from training data, such as what a page published this week says. You should see a tool call to search_and_read or search in the chat, then an answer that cites the pages it read. If the model answers from memory instead:
- check that the model supports tool calling and that the server and its tools are listed in the Program tab;
- switch to a smaller tool set, because a long tool list can push your question out of a small context;
- an HTTP 401 error means the Authorization header is missing or the key is wrong.
The same server works in other local stacks: giving Ollama internet access covers the terminal route with ollmcp, and local LLM web search with MCP has the full budget for 4K to 64K contexts. Billing is per successful call; every account gets $2 of free API usage each month, and failed requests are never billed. All clients are listed on the QuanticData MCP server page.