Documentation Python quickstart Blog Free tools Enterprise solutions hello@quanticdata.ioLog in

LM Studio MCP Web Search: 14K vs 360 Tokens

An LM Studio mcp.json block with a remote web search server: a url ending in ?tools=web plus an Authorization Bearer header; the card compares the tool definitions sent with every request: all 28 tools 13,947 tokens, research 5,122, web 4,215, lite 356.
LM Studio mcp.json with the QuanticData web search server: url plus Authorization header, web tool set at 4,215 tokens instead of 13,947

LM Studio supports MCP for web search since version 0.3.17: add the server to mcp.json with its URL and an Authorization header, then turn on tool use for a tool-calling model. Choose the tool list deliberately. Our 28 tools measured 13,947 tokens on 2 October 2026; the three web tools 4,215; one tool 356.

LM Studio's documentation says it plainly: some MCP servers were designed for Claude, ChatGPT or Gemini, use excessive amounts of tokens, and can bog down a local model with frequent context overflows. That is true of ours too if you load everything. This guide sets up web search in LM Studio and sizes the tool list to the context you actually run.

Three things, all on your machine except the last:

  • LM Studio 0.3.17 or newer. From that build LM Studio acts as an MCP host for both local (stdio) and remote (HTTP) servers.
  • A model trained for tool calling. LM Studio passes tool definitions to any model, but only models trained for function calling use them reliably. Qwen, Granite and Nemotron families with tool support are a good start.
  • A web MCP server. Here, the hosted QuanticData endpoint: it runs the searches and page fetches through residential IPs and returns Markdown or JSON, so nothing scrapes from your laptop.

Step by step: the mcp.json block

Open the Program tab in the right-hand sidebar, click Install, then Edit mcp.json. LM Studio follows Cursor's notation, so a remote server is a url plus optional headers. Paste this, with your API key from the QuanticData dashboard:

{
  "mcpServers": {
    "quanticdata": {
      "url": "https://api.quanticdata.io/mcp?tools=web",
      "headers": {
        "Authorization": "Bearer YOUR_QUANTICDATA_API_KEY"
      }
    }
  }
}

If mcp.json already has other servers, copy only the "quanticdata": { ... } entry inside the existing mcpServers object, as LM Studio's docs advise. Save the file; LM Studio loads every server listed in it.

To run the server on your own machine instead, use the npm package over stdio; the tool filter is an environment variable (version 0.11.4 or later):

"quanticdata": {
  "command": "npx",
  "args": ["-y", "quanticdata-mcp"],
  "env": {
    "QUANTICDATA_API_KEY": "YOUR_QUANTICDATA_API_KEY",
    "QUANTICDATA_TOOLS": "web"
  }
}

Choose the tool set by context length

Every request LM Studio sends to the model includes the definitions of every enabled tool. We serialised what our server returns and estimated tokens as characters divided by four:

URL suffixTools the model seesDefinition tokensSmallest sensible context
none (all tools)28, including crawl, datasets, collectors, proxies13,94732K, better 64K
?tools=researchsearch, search_and_read, scrape, map, batch, batch_status5,12216K
?tools=websearch, search_and_read, scrape4,2158K
?tools=litesearch_and_read3564K

Set the context when you load the model in LM Studio; a bigger window costs memory, so match it to the preset rather than the other way round. You can also combine names, for example ?tools=lite,collectors adds the ready-made collectors (829 tokens) to the single search tool.

Keep each result small, too

Definitions are paid on every turn; results only when the model calls a tool, but a single result can be larger than everything else. Measured on 2 October 2026, as returned to the model:

CallTokens
search, one query1,642
search_and_read, defaults (3 pages)3,986
search_and_read, max_tokens 2,0002,536
scrape, one Wikipedia article, defaults48,291
scrape, same article, query + max_tokens 1,5002,390

Put the limits in LM Studio's system prompt, for example: "When you use search_and_read, set top_n to 2 and max_tokens to 1500. When you use scrape, always pass a query describing what you need and max_tokens 1500." The arguments are documented in the API documentation. Why the content mode matters so much is measured across 11 pages in our MCP web scraper token test.

Which web tool answers which question

With the web set the model chooses between three tools. They overlap, so it helps to know what each one is for, and to say so in the system prompt:

Question typeToolWhat comes back
"What is the latest version of X and what changed?"search_and_readThe top pages for a query, cleaned and merged into one bounded context with numbered sources
"Summarise this page" with a URL in the messagescrapeThat one page as Markdown; pass a query to keep only the relevant sections
"Which sites rank for X?" or "find me five sources"searchThe results page as JSON: titles, links, snippets, and blocks such as People Also Ask when Google shows them

For a small model, search_and_read alone answers most questions: it does the search and the reading in one call, so the model has only one decision to make. That is why the 4K preset contains nothing else. Add scrape when people paste links, and search when they want lists of sources rather than an answer.

Check that it works

Ask a question the model cannot answer from training data, such as what a page published this week says. You should see a tool call to search_and_read or search in the chat, then an answer that cites the pages it read. If the model answers from memory instead:

  • check that the model supports tool calling and that the server and its tools are listed in the Program tab;
  • switch to a smaller tool set, because a long tool list can push your question out of a small context;
  • an HTTP 401 error means the Authorization header is missing or the key is wrong.

The same server works in other local stacks: giving Ollama internet access covers the terminal route with ollmcp, and local LLM web search with MCP has the full budget for 4K to 64K contexts. Billing is per successful call; every account gets $2 of free API usage each month, and failed requests are never billed. All clients are listed on the QuanticData MCP server page.

Sources & further reading

FAQ

Quick answers on lm studio mcp for web search.

Something else? Ask us

Does LM Studio support MCP?

Yes. Since version 0.3.17 LM Studio is an MCP host for local and remote servers. You add servers in mcp.json from the Program tab, using the same notation as Cursor.

How do I add web search to LM Studio?

Add a web MCP server to mcp.json: a url such as https://api.quanticdata.io/mcp?tools=web and a headers object with Authorization set to Bearer and your API key. Then chat with a tool-calling model.

Where is mcp.json in LM Studio?

Open the Program tab in the right-hand sidebar and choose Install, then Edit mcp.json. LM Studio opens the file in its own editor.

Why does my LM Studio model run out of context with MCP?

Tool definitions are sent on every turn. Our full server measured 13,947 tokens of definitions on 2 October 2026. Load a preset such as ?tools=web (4,215 tokens) or ?tools=lite (356) and cap results with max_tokens.

Which tool should LM Studio use for a URL I paste in the chat?

scrape. It returns that one page as Markdown. Ask the model to pass a query describing what you need and max_tokens 1500, so a long page does not fill the context: one Wikipedia article came back at 48,291 tokens by default and 2,390 with those two arguments.

Can I run the MCP server locally instead of the hosted URL?

Yes. Use command npx with args -y quanticdata-mcp and put QUANTICDATA_API_KEY and QUANTICDATA_TOOLS in env. The pages are still fetched by QuanticData, not by your machine.

Web search for LM Studio in one mcp.json block

Hosted MCP endpoint, a key in the headers, a tool list sized to your context. Every account gets $2 of free API usage each month, and failed requests are never billed.

Related reading