To give Ollama internet access, run the model behind an MCP client and connect a web-search MCP server; the model then calls search and fetch tools on its own. Keep the tool list short: our 28 tool definitions measured 13,947 tokens on 2 October 2026, while one tool costs 356 and fits Ollama's 4,096-token default.
Ollama is a model runner. It serves the model, keeps it in memory and exposes an API, and that is all. The model inside has no network access, which is why the most common question about it on Reddit is still some version of "how do I enable internet access". The answer is a tool: something outside the model that fetches the page and hands the text back. This guide wires that up with an MCP client and explains the one setting that decides whether it works on a laptop.
Can Ollama access the internet on its own?
No. A model in Ollama answers from its weights. It can, if it was trained for tool calling, reply with a structured request ("call web_search with this query") instead of text. Something has to execute that request and send the result back as the next message. Ollama's API supports this loop, but Ollama does not run the tools for you.
There are three ways to close the loop:
| Approach | What you write | Who fetches the web | Fits 4K context |
|---|---|---|---|
| Your own Python loop with a search function | The agent loop, tool schemas, parsing | Whatever API you call | Depends on your code |
| Ollama's web search API | An Ollama account key and the loop from their docs | Ollama's service, search snippets and single-page fetch | Their example truncates results to 8,000 characters |
| An MCP client plus a web MCP server | One command | The MCP server: search, scrape, search-and-read | Yes, with a short tool list |
The MCP route is the only one where you write no code and can swap the model, the client or the server independently. The rest of this guide uses it.
Step 1: pick a tool-calling model and raise the context
Only models trained for tool calling can use a web tool. The Ollama library has a tools filter; on 2 October 2026 it listed the Qwen3.6 and Qwen3.8 families, IBM Granite 4.1 (3B, 8B and 30B) and Nemotron 3.5 Lightning among others, and Ollama's own web-search example uses qwen3:4b. Pull one that fits your memory:
ollama pull qwen3:4b
Then check the context. Ollama defaults to 4k tokens below 24 GiB of VRAM, 32k between 24 and 48 GiB and 256k above, and its documentation recommends at least 64,000 tokens for web search and agents. If your machine has the memory, raise it:
OLLAMA_CONTEXT_LENGTH=32000 ollama serve
Run ollama ps afterwards: the CONTEXT column shows what the loaded model actually got, and PROCESSOR shows whether it spilled to CPU. If it did, a smaller context is faster than a bigger one on CPU.
Step 2: install an MCP client for Ollama
Ollama is not an MCP host, so you need a client that speaks both: MCP to the tool server, the Ollama API to the model. In a terminal the simplest is ollmcp (MCP Client for Ollama), an open-source Python TUI. It needs Python 3.11 or newer:
uv tool install --upgrade ollmcp
# or
pip install --upgrade ollmcp
If you prefer a desktop app, LM Studio is an MCP host since version 0.3.17 and loads the same server from its mcp.json: see LM Studio MCP web search. For a browser chat in front of Ollama, Open WebUI does the same from its admin panel: see Open WebUI MCP server setup.
Step 3: connect the web-search server with a short tool list
Create a QuanticData API key in the dashboard, then add the hosted endpoint. The ?tools=lite part is what makes this work on a 4K model: it exposes only search_and_read, which searches, opens the top pages and returns one bounded context.
ollmcp mcp add --transport http quanticdata \
"https://api.quanticdata.io/mcp?tools=lite" \
--header "Authorization: Bearer YOUR_QUANTICDATA_API_KEY"
ollmcp --model qwen3:4b
Prefer to run the server locally? The npm package works over stdio with the same tool filter (version 0.11.4 or later):
ollmcp mcp add \
--env QUANTICDATA_API_KEY=YOUR_QUANTICDATA_API_KEY \
--env QUANTICDATA_TOOLS=lite \
quanticdata -- npx -y quanticdata-mcp
Inside the client, /tools shows what is enabled and lets you switch tools on and off during the chat, and the model settings include num_ctx if you want to set the context per session.
What each choice costs a 4K model
We measured the exact text the MCP server hands to the model, tokens estimated as characters divided by four. Definitions are paid on every request; results only when the model calls a tool.
| Setting | Definitions | One typical result | Total | Fits 4,096? |
|---|---|---|---|---|
| All 28 tools | 13,947 | search, 1,642 | 15,589 | No |
| web preset (search, search_and_read, scrape) | 4,215 | search, 1,642 | 5,857 | No, needs 8K |
| lite preset, default call | 356 | search_and_read, 3,986 | 4,342 | No |
| lite preset, 2 pages, max_tokens 1,500 | 356 | 1,858 | 2,214 | Yes, 1,882 left |
The model chooses the arguments, so tell it in the system prompt to call search_and_read with top_n 2 and max_tokens 1,500. With a single tool on the list there is only one call to get right. At 8K and above, the web preset adds scrape for reading a URL the user pastes; give it a query and a max_tokens too, because a default scrape of one long Wikipedia article came back at 48,291 tokens and the same page with a query and a 1,500 cap at 2,390. The full budget table for 4K, 8K, 32K and 64K is in what fits in a local context.
Troubleshooting the usual failures
- The model answers without searching. It either does not support tools or the tool list pushed your question out of the context. Check the model has the tools tag, then switch to the lite preset.
- The answer stops mid-sentence. The result filled the context. Lower
max_tokenson the call or raise the context. - 401 from the server. The Authorization header is missing or the key is wrong; the header must read
Bearerfollowed by the key. - It is slow. In our test a
search_and_readcall took 4.5 to 7.8 seconds on our side; on a laptop the model then has to read the result, which takes longer the bigger it is. A smallermax_tokensshortens both.
Every tool and argument is listed in the API documentation, and the MCP server page covers the other clients. Billing is per successful call: every account gets $2 of free API usage each month, and failed requests are never billed.