Documentation Python quickstart Blog Free tools Enterprise solutions hello@quanticdata.ioLog in

Qwen Web Search MCP: Qwen3 Online via Ollama

A scale of how much of a 4,096-token context Qwen receives: all 28 tool definitions 340%, the web preset 103%, a 2-page search_and_read result 45%, the lite preset 9%; the card shows the local setup: qwen3:8b served by Ollama, Qwen-Agent as MCP client, a web MCP server.
Qwen web search over MCP: Qwen3 in Ollama, Qwen-Agent as the MCP client, QuanticData MCP server reading the web and returning a bounded context

To give a local Qwen model web search, serve it with Ollama, run Qwen-Agent as the MCP client and add a web MCP server with type streamable-http, its URL and an Authorization header. Load one tool: our full list measured 13,947 tokens on 2 October 2026, the single search tool 356.

Qwen is one of the most pulled open model families for tool calling: on 2 October 2026 Ollama's tools page showed 6.9 million pulls for Qwen3.6 alone. Search for "qwen web search" and most results are about Alibaba Cloud's hosted search, which needs a cloud key and does not apply to a model on your own machine. This guide is the local version.

What you need

PieceRoleExample
Ollama (or vLLM)Serves the Qwen model on an OpenAI-compatible endpointhttp://localhost:11434/v1
A Qwen model with tool callingDecides when to search and writes the answerqwen3:4b or qwen3:8b on a laptop; qwen3.6:27b with 24 GB or more
Qwen-AgentThe agent framework from the Qwen team; acts as MCP clientpip install -U "qwen-agent[mcp]"
A web MCP serverRuns searches and reads pages, returns texthttps://api.quanticdata.io/mcp

Qwen-Agent's README recommends Ollama for local CPU and GPU deployment and vLLM for high-throughput GPU serving. For Qwen3 on vLLM it advises not adding --enable-auto-tool-choice and --tool-call-parser hermes, because Qwen-Agent parses tool calls itself.

Step 1: pull the model and size the context

ollama pull qwen3:8b

Ollama gives a model 4k tokens of context on GPUs under 24 GiB of VRAM, 32k between 24 and 48 GiB and 256k above, and recommends at least 64,000 for web search and agents. If you have the memory, start the server with a bigger window:

OLLAMA_CONTEXT_LENGTH=32000 ollama serve

If not, the tool list in step 3 is what keeps a 4K Qwen usable.

Step 2: install Qwen-Agent with MCP support

pip install -U "qwen-agent[mcp]"

The mcp extra pulls in the MCP Python SDK. Qwen-Agent then accepts MCP servers inside function_list, in the same mcpServers shape Claude Desktop and Cursor use.

Step 3: connect the web MCP server

This is the part that trips people up. In Qwen-Agent's MCP manager, a server with a url is opened as SSE unless its type is streamable-http. Our endpoint, like most current MCP servers, speaks Streamable HTTP, so the type line is required:

from qwen_agent.agents import Assistant

llm_cfg = {
    "model": "qwen3:8b",
    "model_server": "http://localhost:11434/v1",
    "api_key": "EMPTY",
}

tools = [{
    "mcpServers": {
        "quanticdata": {
            "type": "streamable-http",
            "url": "https://api.quanticdata.io/mcp?tools=lite",
            "headers": {"Authorization": "Bearer YOUR_QUANTICDATA_API_KEY"},
        }
    }
}]

bot = Assistant(
    llm=llm_cfg,
    function_list=tools,
    system_message=(
        "For anything recent or outside your training data, call search_and_read "
        "with top_n 2 and max_tokens 1500, then answer and cite the pages."
    ),
)

messages = [{"role": "user", "content": "What changed in the latest Ollama release?"}]
for response in bot.run(messages=messages):
    pass
print(response[-1]["content"])

?tools=lite exposes only search_and_read. With a 16K or larger context, switch to ?tools=web to add search (raw results pages) and scrape (one URL as Markdown). To run the server locally instead of the hosted URL, use "command": "npx", "args": ["-y", "quanticdata-mcp"] and put QUANTICDATA_API_KEY and QUANTICDATA_TOOLS in "env".

How much context the web tools take from Qwen

Measured on 2 October 2026 from the exact text our MCP server returns, tokens estimated as characters divided by four:

What Qwen receivesTokensShare of 4,096Share of 32,000
All 28 tool definitions13,947340%44%
web preset, 3 definitions4,215103%13%
lite preset, 1 definition3569%1%
search_and_read result, 2 pages, max_tokens 1,5001,85845%6%
search_and_read result, defaults (3 pages)3,98697%12%
scrape of one Wikipedia article, defaults48,2911,179%151%

On a 4K Qwen the working combination is the lite preset plus a capped call: 356 plus 1,858 tokens, leaving 1,882 for the system prompt, the question and the answer. The longer reasoning traces of thinking models eat into that too, so on 4K turn thinking off for search questions or move to 8K. The budget for every context size is in our local LLM web search budget, and why a default scrape is so large is measured in the MCP web scraper test.

Using it day to day

Once the agent runs, the questions that benefit most are the ones a model cannot answer from training data: release notes from this month, a price on a page today, what a document published last week says. A few patterns that keep a local Qwen accurate and inside its context:

  • Ask for sources. search_and_read returns numbered sources with their URLs; telling Qwen to cite them makes it quote the pages instead of filling gaps from memory.
  • One question per turn. Each search result stays in the conversation history. On 4K, two capped searches already use most of the window, so start a new conversation for an unrelated question.
  • Name the site when you know it. A query such as "site:docs.ollama.com context length" sends the search straight to the right page and returns less noise to read.
  • Keep thinking short on small windows. Qwen3's reasoning text counts against the same context as the search results; on 4K, switch it off for lookups and keep it for questions that need reasoning over what was found.

Whichever you use, the server side is the same QuanticData MCP server, with every argument in the API documentation. Billing is per successful call: every account gets $2 of free API usage each month, and failed requests are never billed.

Sources & further reading

FAQ

Quick answers on qwen web search mcp.

Something else? Ask us

Can Qwen search the web when it runs locally?

Yes, through tool calling. Run Qwen in Ollama, use Qwen-Agent as the MCP client and connect a web MCP server. The model asks for a search, the server runs it and returns the text.

Why does Qwen-Agent fail to connect to my MCP URL?

Qwen-Agent treats a server with a url as SSE unless type is set to streamable-http. Most current MCP servers, ours included, speak Streamable HTTP, so add "type": "streamable-http" next to the url and headers.

Which Qwen model should I use for web search on a laptop?

qwen3:4b or qwen3:8b fit most laptops; Ollama uses qwen3:4b in its own web-search example. With 24 GB of GPU memory or more, the Qwen3.6 27B model is an option.

How much context does web search need for Qwen?

Ollama recommends 64,000 tokens for web search. At its 4K default, load one tool (356 tokens) and cap results at about 1,500 tokens: our measured total was 2,214, leaving 1,882 for the question and answer.

Does Qwen thinking mode work with web search?

Yes, but the reasoning text uses the same context window as the search results. On a 4K model, turn thinking off for simple lookups; with 16K or more there is room for both.

Is this the same as Alibaba Cloud web search for Qwen?

No. That is a hosted search option for Qwen models in Alibaba Cloud Model Studio. This setup keeps the model on your machine and adds search through an MCP server you choose.

Web search for a local Qwen

Streamable HTTP endpoint, one tool for 4K models, three for 8K and up. Every account gets $2 of free API usage each month, and failed requests are never billed.

Related reading