To give a local Qwen model web search, serve it with Ollama, run Qwen-Agent as the MCP client and add a web MCP server with type streamable-http, its URL and an Authorization header. Load one tool: our full list measured 13,947 tokens on 2 October 2026, the single search tool 356.
Qwen is one of the most pulled open model families for tool calling: on 2 October 2026 Ollama's tools page showed 6.9 million pulls for Qwen3.6 alone. Search for "qwen web search" and most results are about Alibaba Cloud's hosted search, which needs a cloud key and does not apply to a model on your own machine. This guide is the local version.
What you need
| Piece | Role | Example |
|---|---|---|
| Ollama (or vLLM) | Serves the Qwen model on an OpenAI-compatible endpoint | http://localhost:11434/v1 |
| A Qwen model with tool calling | Decides when to search and writes the answer | qwen3:4b or qwen3:8b on a laptop; qwen3.6:27b with 24 GB or more |
| Qwen-Agent | The agent framework from the Qwen team; acts as MCP client | pip install -U "qwen-agent[mcp]" |
| A web MCP server | Runs searches and reads pages, returns text | https://api.quanticdata.io/mcp |
Qwen-Agent's README recommends Ollama for local CPU and GPU deployment and vLLM for high-throughput GPU serving. For Qwen3 on vLLM it advises not adding --enable-auto-tool-choice and --tool-call-parser hermes, because Qwen-Agent parses tool calls itself.
Step 1: pull the model and size the context
ollama pull qwen3:8b
Ollama gives a model 4k tokens of context on GPUs under 24 GiB of VRAM, 32k between 24 and 48 GiB and 256k above, and recommends at least 64,000 for web search and agents. If you have the memory, start the server with a bigger window:
OLLAMA_CONTEXT_LENGTH=32000 ollama serve
If not, the tool list in step 3 is what keeps a 4K Qwen usable.
Step 2: install Qwen-Agent with MCP support
pip install -U "qwen-agent[mcp]"
The mcp extra pulls in the MCP Python SDK. Qwen-Agent then accepts MCP servers inside function_list, in the same mcpServers shape Claude Desktop and Cursor use.
Step 3: connect the web MCP server
This is the part that trips people up. In Qwen-Agent's MCP manager, a server with a url is opened as SSE unless its type is streamable-http. Our endpoint, like most current MCP servers, speaks Streamable HTTP, so the type line is required:
from qwen_agent.agents import Assistant
llm_cfg = {
"model": "qwen3:8b",
"model_server": "http://localhost:11434/v1",
"api_key": "EMPTY",
}
tools = [{
"mcpServers": {
"quanticdata": {
"type": "streamable-http",
"url": "https://api.quanticdata.io/mcp?tools=lite",
"headers": {"Authorization": "Bearer YOUR_QUANTICDATA_API_KEY"},
}
}
}]
bot = Assistant(
llm=llm_cfg,
function_list=tools,
system_message=(
"For anything recent or outside your training data, call search_and_read "
"with top_n 2 and max_tokens 1500, then answer and cite the pages."
),
)
messages = [{"role": "user", "content": "What changed in the latest Ollama release?"}]
for response in bot.run(messages=messages):
pass
print(response[-1]["content"])
?tools=lite exposes only search_and_read. With a 16K or larger context, switch to ?tools=web to add search (raw results pages) and scrape (one URL as Markdown). To run the server locally instead of the hosted URL, use "command": "npx", "args": ["-y", "quanticdata-mcp"] and put QUANTICDATA_API_KEY and QUANTICDATA_TOOLS in "env".
How much context the web tools take from Qwen
Measured on 2 October 2026 from the exact text our MCP server returns, tokens estimated as characters divided by four:
| What Qwen receives | Tokens | Share of 4,096 | Share of 32,000 |
|---|---|---|---|
| All 28 tool definitions | 13,947 | 340% | 44% |
| web preset, 3 definitions | 4,215 | 103% | 13% |
| lite preset, 1 definition | 356 | 9% | 1% |
| search_and_read result, 2 pages, max_tokens 1,500 | 1,858 | 45% | 6% |
| search_and_read result, defaults (3 pages) | 3,986 | 97% | 12% |
| scrape of one Wikipedia article, defaults | 48,291 | 1,179% | 151% |
On a 4K Qwen the working combination is the lite preset plus a capped call: 356 plus 1,858 tokens, leaving 1,882 for the system prompt, the question and the answer. The longer reasoning traces of thinking models eat into that too, so on 4K turn thinking off for search questions or move to 8K. The budget for every context size is in our local LLM web search budget, and why a default scrape is so large is measured in the MCP web scraper test.
Using it day to day
Once the agent runs, the questions that benefit most are the ones a model cannot answer from training data: release notes from this month, a price on a page today, what a document published last week says. A few patterns that keep a local Qwen accurate and inside its context:
- Ask for sources.
search_and_readreturns numbered sources with their URLs; telling Qwen to cite them makes it quote the pages instead of filling gaps from memory. - One question per turn. Each search result stays in the conversation history. On 4K, two capped searches already use most of the window, so start a new conversation for an unrelated question.
- Name the site when you know it. A query such as "site:docs.ollama.com context length" sends the search straight to the right page and returns less noise to read.
- Keep thinking short on small windows. Qwen3's reasoning text counts against the same context as the search results; on 4K, switch it off for lookups and keep it for questions that need reasoning over what was found.
Other ways to run Qwen with web search
- LM Studio runs Qwen models and is an MCP host itself, so the same server goes in its mcp.json with no Python: LM Studio MCP web search.
- Open WebUI in front of Ollama gives a chat interface with the MCP server added once by an admin: Open WebUI MCP server setup.
- A terminal client such as ollmcp talks to any Ollama model, Qwen included: give Ollama internet access.
Whichever you use, the server side is the same QuanticData MCP server, with every argument in the API documentation. Billing is per successful call: every account gets $2 of free API usage each month, and failed requests are never billed.