LlamaIndex web scraping costs you twice: once to fetch the page and once to embed whatever the reader puts in the Document. On 2 October 2026 we sent 20 pages a RAG app would ingest through one batch job of our web scraping API: 19 came back in 49.8 seconds, all over plain HTTP, for $0.0038. As raw HTML those 19 pages are 1,904,058 tokens; as smart Markdown with link targets stripped they are 103,333, 94.6% fewer. Feed LlamaIndex the second one.
The default web reader puts raw HTML into your Documents
LlamaIndex lists 23 web readers on its reference page. The one most tutorials start with, SimpleWebPageReader, takes an html_to_text flag, and in the current source that flag defaults to False: the Document text is the HTML the server sent, scripts, inline styles and navigation included. Every one of those characters is embedded, stored in the vector index and can be retrieved into the prompt.
That is why we measured the page in the form LlamaIndex receives it, not the form a browser shows. A Document is just text plus a metadata dictionary, and the LlamaIndex guide notes that metadata is injected into the text for both the embedding and the LLM call by default. Whatever you put in text is what you pay for, twice.
19 of 20 pages came back from one batch job in 49.8 seconds
The 20 URLs are the kind of sources a documentation or research assistant ingests: Python, MDN, FastAPI, Docker, Kubernetes, PostgreSQL and LlamaIndex docs, two Wikipedia articles, two GitHub repositories, an arXiv abstract, three pricing pages, two news front pages, a long-form article, a personal blog and a changelog. We submitted them as one batch job with format: markdown and content_mode: smart, no country, default concurrency of 5.
The job finished in 49.8 seconds of wall time, about 2.5 seconds per URL. All 19 returned pages went over plain HTTP with a browser TLS fingerprint; none needed a headless browser, so none was billed at the rendered rate. The twentieth, the LlamaIndex changelog page on GitHub, came back without its content at that minute and was not billed: the batch charges only pages that come back. For changelogs and READMEs, fetch the raw file instead: the MCP tools README from raw.githubusercontent.com weighed 7,224 bytes and came back in 2.4 seconds, against 317,198 to 382,079 bytes for a GitHub repository page.
Batch charges $0.0002 per URL up front and refunds the share that does not come back, so the job cost $0.0038. The same 19 pages forced through a browser would have cost $0.019.
Smart Markdown cut 92.7% of the tokens, stripping links another 25.6%
We then fetched each of the 19 pages again in three forms: format: html (what SimpleWebPageReader stores by default), smart Markdown (the batch default: the page minus navigation, footer and cookie chrome, tables kept as GFM), and smart Markdown with links_mode: strip, which keeps the link text and drops the URL. Tokens are estimated as characters divided by 4, the same rule for every column.
| Page | HTML bytes | HTML tokens | Smart | Links stripped | Seconds |
|---|---|---|---|---|---|
| Python asyncio docs | 177,639 | 44,370 | 15,099 | 11,469 | 9.3 |
| Wikipedia: Retrieval-augmented generation | 251,488 | 62,672 | 6,402 | 4,243 | 2.5 |
| Wikipedia: Vector database | 323,569 | 80,753 | 11,445 | 4,712 | 4.7 |
| GitHub: run-llama/llama_index | 382,079 | 95,470 | 4,236 | 2,829 | 8.4 |
| GitHub: psf/requests | 317,198 | 79,280 | 1,834 | 1,079 | 11.6 |
| Neon pricing | 657,733 | 164,419 | 5,469 | 4,867 | 10.4 |
| BBC technology | 526,742 | 131,565 | 4,767 | 4,557 | 3.5 |
| OpenAI pricing | 837,446 | 209,325 | 5,693 | 4,872 | 8.8 |
| Simon Willison's blog home | 85,681 | 21,394 | 13,782 | 8,069 | 2.0 |
| Martin Fowler: Microservices | 86,071 | 21,488 | 12,787 | 11,183 | 2.3 |
| AP News technology | 1,526,563 | 381,550 | 17,585 | 10,315 | 27.6 |
| MDN: HTTP caching | 260,972 | 65,187 | 9,807 | 8,628 | 12.9 |
| FastAPI: first steps | 143,815 | 35,923 | 3,760 | 3,323 | 3.3 |
| Docker overview | 198,996 | 49,748 | 2,476 | 2,316 | 3.7 |
| LlamaIndex: Documents guide | 257,662 | 64,399 | 2,260 | 2,083 | 5.4 |
| Kubernetes overview | 505,930 | 126,445 | 2,911 | 2,631 | 3.2 |
| PostgreSQL: JSON types | 60,284 | 15,062 | 8,937 | 7,701 | 1.7 |
| arXiv: RAG paper abstract | 45,835 | 11,458 | 2,214 | 1,273 | 1.4 |
| Stripe pricing | 974,548 | 243,550 | 7,508 | 7,183 | 5.5 |
| Total, 19 pages | 7,620,251 | 1,904,058 | 138,972 | 103,333 | median 4.7 |
The median page is 260,972 bytes of HTML, 65,187 tokens raw, 5,693 as smart Markdown and 4,712 with links stripped. Smart Markdown removes 92.7% of the raw tokens; stripping link targets removes a further 25.6% of what is left, for 94.6% in total. The seconds column is a single fetch of the links-stripped version from a US exit.
Stripping links pays most where a page is mostly references. The Vector database article on Wikipedia went from 11,445 to 4,712 tokens, the requests repository from 1,834 to 1,079, the RAG paper abstract from 2,214 to 1,273. On prose-heavy pages the gain is small: Martin Fowler's essay dropped from 12,787 to 11,183. If your app has to cite or follow links, keep them in metadata, not in the text you embed.
Chunk by heading on docs, by tokens on listings and pricing pages
The scrape call can also cut the Markdown into chunks before LlamaIndex sees it, each with its heading path and token count, and it never splits a table or a code fence. Chunked by heading, the 19 pages gave 687 chunks: a median of 24 per page and a median chunk of 73 tokens, with the largest at 688. On documentation and long articles that is what you want: the Fowler essay gave 34 chunks with a median of 355 tokens, the PostgreSQL page 24 chunks at a median of 306.5.
On pages built from cards and short sections, heading chunks shatter. 285 of the 687 chunks, 41%, are under 50 tokens: a headline, a price tier name, a button label. A 30-token vector retrieves well for its headline and carries nothing to answer with. On those pages chunk by tokens instead. We re-ran four of them with 400-token chunks and 40 tokens of overlap:
| Page | Heading chunks | Under 50 tokens | 400-token chunks | Median tokens | Under 50 tokens |
|---|---|---|---|---|---|
| BBC technology | 80 | 58 | 12 | 389.5 | 0 |
| Stripe pricing | 110 | 51 | 19 | 395 | 0 |
| GitHub: psf/requests | 16 | 12 | 4 | 267.5 | 0 |
| Martin Fowler: Microservices | 34 | 3 | 33 | 349 | 1 |
On the three card-built pages, token chunks cut the vector count from 206 to 35 and left no fragment under 50 tokens. On the essay the two settings agree, 34 against 33 chunks. The rule: by: heading for docs, references and articles; by: tokens with a size around 400 for news fronts, pricing and repository pages. Chunk token counts here are the API's own count per chunk.
The code: one batch job into LlamaIndex Documents
The batch endpoint takes up to 5,000 URLs with shared options and returns a job id; you poll it with include_content=true and pass since to receive only the items finished after the last poll. Each item becomes one Document with the URL as its id, so a re-run updates the page instead of duplicating it.
import time, requests
from llama_index.core import Document, VectorStoreIndex
API = "https://api.quanticdata.io/v1"
H = {"Authorization": "Bearer qd_live_YOUR_KEY"}
urls = ["https://docs.python.org/3/library/asyncio-task.html",
"https://en.wikipedia.org/wiki/Vector_database"] # up to 5,000
job = requests.post(API + "/batch", headers=H, timeout=60, json={
"urls": urls, "format": "markdown", "contentMode": "smart"}).json()["payload"]
docs, since = [], 0
while True:
r = requests.get(API + "/batch/" + job["id"], headers=H, timeout=60,
params={"include_content": "true", "since": since}).json()["payload"]
for item in r["items"]:
if item["status"] == 200 and item.get("content"):
docs.append(Document(text=item["content"], id_=item["url"],
metadata={"url": item["url"], "title": item.get("title") or ""}))
since = r["nextCursor"]
if r["status"] != "running" and not r["hasMore"]:
break
if r["status"] == "running":
time.sleep(5)
index = VectorStoreIndex.from_documents(docs)
For a single page you want pre-chunked and without link targets, call scrape with links_mode and chunk and build nodes directly, so LlamaIndex does not split the text a second time:
from llama_index.core import VectorStoreIndex
from llama_index.core.schema import TextNode
url = "https://stripe.com/pricing"
p = requests.post(API + "/scrape", headers=H, timeout=90, json={
"url": url, "format": "markdown", "content_mode": "smart", "links_mode": "strip",
"chunk": {"by": "tokens", "size": 400, "overlap": 40}}).json()["payload"]
nodes = [TextNode(text=c["text"], metadata={"url": url, "section": " / ".join(c.get("path") or [])})
for c in p["chunks"]]
index = VectorStoreIndex(nodes)
Pages that do not come back are not billed, so the loop needs no retry logic for cost reasons. For a whole documentation site rather than a list of URLs, the crawl endpoint follows links for you at $0.0003 per page; ../how-to-web-crawl-python/ has the math for when that beats your own loop.
Agents that read live pages connect through MCP in five lines
Indexing is for pages you will query many times. When an agent needs a page it has not seen, give it the fetch as a tool. LlamaIndex's llama-index-tools-mcp package turns any MCP server into FunctionTool objects through McpToolSpec, and its BasicMCPClient speaks Streamable HTTP and accepts a custom httpx client, which is where the API key goes:
import httpx
from llama_index.tools.mcp import BasicMCPClient, McpToolSpec
from llama_index.core.agent.workflow import FunctionAgent
client = BasicMCPClient("https://api.quanticdata.io/mcp", http_client=httpx.AsyncClient(
headers={"Authorization": "Bearer qd_live_YOUR_KEY"}))
spec = McpToolSpec(client=client, allowed_tools=["search", "scrape", "batch", "batch_status"])
tools = await spec.to_tool_list_async()
agent = FunctionAgent(name="researcher", description="Answers from live web pages",
llm=llm, tools=tools,
system_prompt="Call scrape with links_mode strip before answering.")
Limit allowed_tools to what the agent needs: every tool schema you expose is read on every turn. The same MCP server also signs in with OAuth 2.1, which BasicMCPClient supports through with_oauth. We measured the agent side separately in ../mcp-web-scraper/: 11 pages, 96% fewer tokens with smart Markdown and no link targets, the same direction as the 94.6% here.
What 94.6% fewer tokens does to the embedding bill
OpenAI's pricing page, fetched on 2 October 2026, lists text-embedding-3-small at $0.02 and text-embedding-3-large at $0.13 per million tokens. Our 19 pages average 100,213 tokens as raw HTML, 7,314 as smart Markdown and 5,439 with links stripped. Scaled to 10,000 pages of the same mix:
| Document text | Tokens, 10,000 pages | Embed, 3-small | Embed, 3-large | Fetch, plain HTTP |
|---|---|---|---|---|
| Raw HTML | 1,002 million | $20.04 | $130.28 | you run it |
| Smart Markdown | 73.1 million | $1.46 | $9.51 | $2.00 |
| Smart, links stripped | 54.4 million | $1.09 | $7.07 | $2.00 |
With the small model, fetching 10,000 pages through the API at $0.0002 and embedding them clean costs $3.09; embedding the same pages as raw HTML costs $20.04 before you have paid for a single fetch. The vector store shrinks by the same factor, and every retrieved chunk that reaches the LLM is prose instead of markup. The token figures are character estimates, so treat the dollar amounts as the right order, not the cent.
Skip the browser unless a page proves it needs one: all 19 pages here returned full content over plain HTTP at $0.0002, and the rendered rate is five times that. Our ../how-to-build-an-ai-web-scraper/ covers where the LLM belongs once the text is clean.
The setting that works on LlamaIndex
- Endpoint:
POST /v1/batchfor URL lists (up to 5,000 per job),POST /v1/scrapefor single pages, both on our web scraping API at $0.0002 per page, $0.001 rendered, failed pages never billed. On our 20 URLs: 19 back in 49.8 s for $0.0038. - Fetch mode:
engine: auto(the default). It stayed on plain HTTP for all 19 pages and only escalates to a browser when a page needs it. - Document text:
format: markdown,content_mode: smart, pluslinks_mode: stripon single scrapes. 103,333 tokens for 19 pages instead of 1,904,058 as raw HTML. - Chunking:
chunk: {"by": "heading"}on docs and articles (median 306 to 355 tokens there);by: tokens, size 400, overlap 40 on news, pricing and repository pages, where heading chunks left 121 of 206 pieces under 50 tokens. - Country: unpinned worked for all 19 sources; pin
countryonly for pages that serve prices or news by region. - Agents: McpToolSpec on
https://api.quanticdata.io/mcpwithallowed_toolslimited to search, scrape, batch and batch_status. See the web data API for AI for the full tool list. - When pages are not enough: for places, jobs, products or reviews, the collectors return rows from a keyword and a location instead of pages to chunk.
- Pricing: every account gets $2 of free API usage per month, which is 10,000 plain page fetches, enough to index this test 526 times over.
Sources & further reading
- LlamaIndex API reference: Web readers
- LlamaIndex source: SimpleWebPageReader (html_to_text default)
- LlamaIndex docs: Defining and Customizing Documents
- LlamaIndex API reference: McpToolSpec
- llama-index-tools-mcp README (BasicMCPClient, Streamable HTTP, custom httpx client)
- OpenAI API pricing (embedding models)
- Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks