Documentation Python quickstart Blog Free tools Enterprise solutions hello@quanticdata.ioLog in

LlamaIndex Web Scraping: 95% Fewer Tokens

Nineteen pages from one batch job on 2 October 2026 in three forms: raw HTML at 1,904,058 tokens, smart Markdown at 138,972 and Markdown with link targets stripped at 103,333; the card shows 94.6% fewer tokens for LlamaIndex, all 19 pages over plain HTTP in 49.8 seconds for $0.0038.
Token count of 19 scraped pages in three forms: raw HTML at 1.9 million tokens, smart Markdown at 138,972 and Markdown without link targets at 103,333, from one batch job that took 49.8 seconds

LlamaIndex web scraping costs you twice: once to fetch the page and once to embed whatever the reader puts in the Document. On 2 October 2026 we sent 20 pages a RAG app would ingest through one batch job of our web scraping API: 19 came back in 49.8 seconds, all over plain HTTP, for $0.0038. As raw HTML those 19 pages are 1,904,058 tokens; as smart Markdown with link targets stripped they are 103,333, 94.6% fewer. Feed LlamaIndex the second one.

The default web reader puts raw HTML into your Documents

LlamaIndex lists 23 web readers on its reference page. The one most tutorials start with, SimpleWebPageReader, takes an html_to_text flag, and in the current source that flag defaults to False: the Document text is the HTML the server sent, scripts, inline styles and navigation included. Every one of those characters is embedded, stored in the vector index and can be retrieved into the prompt.

That is why we measured the page in the form LlamaIndex receives it, not the form a browser shows. A Document is just text plus a metadata dictionary, and the LlamaIndex guide notes that metadata is injected into the text for both the embedding and the LLM call by default. Whatever you put in text is what you pay for, twice.

19 of 20 pages came back from one batch job in 49.8 seconds

The 20 URLs are the kind of sources a documentation or research assistant ingests: Python, MDN, FastAPI, Docker, Kubernetes, PostgreSQL and LlamaIndex docs, two Wikipedia articles, two GitHub repositories, an arXiv abstract, three pricing pages, two news front pages, a long-form article, a personal blog and a changelog. We submitted them as one batch job with format: markdown and content_mode: smart, no country, default concurrency of 5.

The job finished in 49.8 seconds of wall time, about 2.5 seconds per URL. All 19 returned pages went over plain HTTP with a browser TLS fingerprint; none needed a headless browser, so none was billed at the rendered rate. The twentieth, the LlamaIndex changelog page on GitHub, came back without its content at that minute and was not billed: the batch charges only pages that come back. For changelogs and READMEs, fetch the raw file instead: the MCP tools README from raw.githubusercontent.com weighed 7,224 bytes and came back in 2.4 seconds, against 317,198 to 382,079 bytes for a GitHub repository page.

Batch charges $0.0002 per URL up front and refunds the share that does not come back, so the job cost $0.0038. The same 19 pages forced through a browser would have cost $0.019.

We then fetched each of the 19 pages again in three forms: format: html (what SimpleWebPageReader stores by default), smart Markdown (the batch default: the page minus navigation, footer and cookie chrome, tables kept as GFM), and smart Markdown with links_mode: strip, which keeps the link text and drops the URL. Tokens are estimated as characters divided by 4, the same rule for every column.

PageHTML bytesHTML tokensSmartLinks strippedSeconds
Python asyncio docs177,63944,37015,09911,4699.3
Wikipedia: Retrieval-augmented generation251,48862,6726,4024,2432.5
Wikipedia: Vector database323,56980,75311,4454,7124.7
GitHub: run-llama/llama_index382,07995,4704,2362,8298.4
GitHub: psf/requests317,19879,2801,8341,07911.6
Neon pricing657,733164,4195,4694,86710.4
BBC technology526,742131,5654,7674,5573.5
OpenAI pricing837,446209,3255,6934,8728.8
Simon Willison's blog home85,68121,39413,7828,0692.0
Martin Fowler: Microservices86,07121,48812,78711,1832.3
AP News technology1,526,563381,55017,58510,31527.6
MDN: HTTP caching260,97265,1879,8078,62812.9
FastAPI: first steps143,81535,9233,7603,3233.3
Docker overview198,99649,7482,4762,3163.7
LlamaIndex: Documents guide257,66264,3992,2602,0835.4
Kubernetes overview505,930126,4452,9112,6313.2
PostgreSQL: JSON types60,28415,0628,9377,7011.7
arXiv: RAG paper abstract45,83511,4582,2141,2731.4
Stripe pricing974,548243,5507,5087,1835.5
Total, 19 pages7,620,2511,904,058138,972103,333median 4.7

The median page is 260,972 bytes of HTML, 65,187 tokens raw, 5,693 as smart Markdown and 4,712 with links stripped. Smart Markdown removes 92.7% of the raw tokens; stripping link targets removes a further 25.6% of what is left, for 94.6% in total. The seconds column is a single fetch of the links-stripped version from a US exit.

Stripping links pays most where a page is mostly references. The Vector database article on Wikipedia went from 11,445 to 4,712 tokens, the requests repository from 1,834 to 1,079, the RAG paper abstract from 2,214 to 1,273. On prose-heavy pages the gain is small: Martin Fowler's essay dropped from 12,787 to 11,183. If your app has to cite or follow links, keep them in metadata, not in the text you embed.

Chunk by heading on docs, by tokens on listings and pricing pages

The scrape call can also cut the Markdown into chunks before LlamaIndex sees it, each with its heading path and token count, and it never splits a table or a code fence. Chunked by heading, the 19 pages gave 687 chunks: a median of 24 per page and a median chunk of 73 tokens, with the largest at 688. On documentation and long articles that is what you want: the Fowler essay gave 34 chunks with a median of 355 tokens, the PostgreSQL page 24 chunks at a median of 306.5.

On pages built from cards and short sections, heading chunks shatter. 285 of the 687 chunks, 41%, are under 50 tokens: a headline, a price tier name, a button label. A 30-token vector retrieves well for its headline and carries nothing to answer with. On those pages chunk by tokens instead. We re-ran four of them with 400-token chunks and 40 tokens of overlap:

PageHeading chunksUnder 50 tokens400-token chunksMedian tokensUnder 50 tokens
BBC technology805812389.50
Stripe pricing11051193950
GitHub: psf/requests16124267.50
Martin Fowler: Microservices343333491

On the three card-built pages, token chunks cut the vector count from 206 to 35 and left no fragment under 50 tokens. On the essay the two settings agree, 34 against 33 chunks. The rule: by: heading for docs, references and articles; by: tokens with a size around 400 for news fronts, pricing and repository pages. Chunk token counts here are the API's own count per chunk.

The code: one batch job into LlamaIndex Documents

The batch endpoint takes up to 5,000 URLs with shared options and returns a job id; you poll it with include_content=true and pass since to receive only the items finished after the last poll. Each item becomes one Document with the URL as its id, so a re-run updates the page instead of duplicating it.

import time, requests
from llama_index.core import Document, VectorStoreIndex

API = "https://api.quanticdata.io/v1"
H = {"Authorization": "Bearer qd_live_YOUR_KEY"}
urls = ["https://docs.python.org/3/library/asyncio-task.html",
        "https://en.wikipedia.org/wiki/Vector_database"]  # up to 5,000

job = requests.post(API + "/batch", headers=H, timeout=60, json={
    "urls": urls, "format": "markdown", "contentMode": "smart"}).json()["payload"]

docs, since = [], 0
while True:
    r = requests.get(API + "/batch/" + job["id"], headers=H, timeout=60,
                     params={"include_content": "true", "since": since}).json()["payload"]
    for item in r["items"]:
        if item["status"] == 200 and item.get("content"):
            docs.append(Document(text=item["content"], id_=item["url"],
                                 metadata={"url": item["url"], "title": item.get("title") or ""}))
    since = r["nextCursor"]
    if r["status"] != "running" and not r["hasMore"]:
        break
    if r["status"] == "running":
        time.sleep(5)

index = VectorStoreIndex.from_documents(docs)

For a single page you want pre-chunked and without link targets, call scrape with links_mode and chunk and build nodes directly, so LlamaIndex does not split the text a second time:

from llama_index.core import VectorStoreIndex
from llama_index.core.schema import TextNode

url = "https://stripe.com/pricing"
p = requests.post(API + "/scrape", headers=H, timeout=90, json={
    "url": url, "format": "markdown", "content_mode": "smart", "links_mode": "strip",
    "chunk": {"by": "tokens", "size": 400, "overlap": 40}}).json()["payload"]

nodes = [TextNode(text=c["text"], metadata={"url": url, "section": " / ".join(c.get("path") or [])})
         for c in p["chunks"]]
index = VectorStoreIndex(nodes)

Pages that do not come back are not billed, so the loop needs no retry logic for cost reasons. For a whole documentation site rather than a list of URLs, the crawl endpoint follows links for you at $0.0003 per page; ../how-to-web-crawl-python/ has the math for when that beats your own loop.

Agents that read live pages connect through MCP in five lines

Indexing is for pages you will query many times. When an agent needs a page it has not seen, give it the fetch as a tool. LlamaIndex's llama-index-tools-mcp package turns any MCP server into FunctionTool objects through McpToolSpec, and its BasicMCPClient speaks Streamable HTTP and accepts a custom httpx client, which is where the API key goes:

import httpx
from llama_index.tools.mcp import BasicMCPClient, McpToolSpec
from llama_index.core.agent.workflow import FunctionAgent

client = BasicMCPClient("https://api.quanticdata.io/mcp", http_client=httpx.AsyncClient(
    headers={"Authorization": "Bearer qd_live_YOUR_KEY"}))
spec = McpToolSpec(client=client, allowed_tools=["search", "scrape", "batch", "batch_status"])
tools = await spec.to_tool_list_async()

agent = FunctionAgent(name="researcher", description="Answers from live web pages",
                      llm=llm, tools=tools,
                      system_prompt="Call scrape with links_mode strip before answering.")

Limit allowed_tools to what the agent needs: every tool schema you expose is read on every turn. The same MCP server also signs in with OAuth 2.1, which BasicMCPClient supports through with_oauth. We measured the agent side separately in ../mcp-web-scraper/: 11 pages, 96% fewer tokens with smart Markdown and no link targets, the same direction as the 94.6% here.

What 94.6% fewer tokens does to the embedding bill

OpenAI's pricing page, fetched on 2 October 2026, lists text-embedding-3-small at $0.02 and text-embedding-3-large at $0.13 per million tokens. Our 19 pages average 100,213 tokens as raw HTML, 7,314 as smart Markdown and 5,439 with links stripped. Scaled to 10,000 pages of the same mix:

Document textTokens, 10,000 pagesEmbed, 3-smallEmbed, 3-largeFetch, plain HTTP
Raw HTML1,002 million$20.04$130.28you run it
Smart Markdown73.1 million$1.46$9.51$2.00
Smart, links stripped54.4 million$1.09$7.07$2.00

With the small model, fetching 10,000 pages through the API at $0.0002 and embedding them clean costs $3.09; embedding the same pages as raw HTML costs $20.04 before you have paid for a single fetch. The vector store shrinks by the same factor, and every retrieved chunk that reaches the LLM is prose instead of markup. The token figures are character estimates, so treat the dollar amounts as the right order, not the cent.

Skip the browser unless a page proves it needs one: all 19 pages here returned full content over plain HTTP at $0.0002, and the rendered rate is five times that. Our ../how-to-build-an-ai-web-scraper/ covers where the LLM belongs once the text is clean.

The setting that works on LlamaIndex

  • Endpoint: POST /v1/batch for URL lists (up to 5,000 per job), POST /v1/scrape for single pages, both on our web scraping API at $0.0002 per page, $0.001 rendered, failed pages never billed. On our 20 URLs: 19 back in 49.8 s for $0.0038.
  • Fetch mode: engine: auto (the default). It stayed on plain HTTP for all 19 pages and only escalates to a browser when a page needs it.
  • Document text: format: markdown, content_mode: smart, plus links_mode: strip on single scrapes. 103,333 tokens for 19 pages instead of 1,904,058 as raw HTML.
  • Chunking: chunk: {"by": "heading"} on docs and articles (median 306 to 355 tokens there); by: tokens, size 400, overlap 40 on news, pricing and repository pages, where heading chunks left 121 of 206 pieces under 50 tokens.
  • Country: unpinned worked for all 19 sources; pin country only for pages that serve prices or news by region.
  • Agents: McpToolSpec on https://api.quanticdata.io/mcp with allowed_tools limited to search, scrape, batch and batch_status. See the web data API for AI for the full tool list.
  • When pages are not enough: for places, jobs, products or reviews, the collectors return rows from a keyword and a location instead of pages to chunk.
  • Pricing: every account gets $2 of free API usage per month, which is 10,000 plain page fetches, enough to index this test 526 times over.

Sources & further reading

FAQ

Quick answers on llamaindex web scraping.

Something else? Ask us

What is the best way to do web scraping with LlamaIndex?

Fetch pages as smart Markdown with link targets stripped and build Documents from that text. On 19 pages measured on 2 October 2026 that was 103,333 tokens, against 1,904,058 for the raw HTML that SimpleWebPageReader stores with its default html_to_text=False, a 94.6% reduction.

Does SimpleWebPageReader return HTML or text?

HTML by default. In the current LlamaIndex source the html_to_text argument defaults to False, so Document.text is the page markup. On our 19 pages that was a median of 65,187 tokens per page, against 4,712 as Markdown without link targets.

How many URLs can I scrape in one batch for a LlamaIndex index?

Up to 5,000 URLs per batch job at $0.0002 per URL, charged up front with the failed share refunded. Our 20-URL job finished in 49.8 seconds at the default concurrency of 5 and billed $0.0038 for the 19 pages that came back.

What chunk size should I use for scraped web pages in LlamaIndex?

Chunk docs and articles by heading: a median of 306 to 355 tokens per chunk on our long pages. On news, pricing and repository pages use 400-token chunks with 40 tokens of overlap: heading chunks there left 121 of 206 pieces under 50 tokens, token chunks left none.

Can a LlamaIndex agent call a web scraping MCP server?

Yes. McpToolSpec from llama-index-tools-mcp converts MCP tools into FunctionTools; point BasicMCPClient at the hosted endpoint with your key in an httpx client header and limit allowed_tools to the 4 you need: search, scrape, batch and batch_status.

How much does it cost to embed scraped pages for RAG?

At $0.02 per million tokens for text-embedding-3-small, 10,000 pages of our mix cost $1.09 to embed as links-stripped Markdown and $20.04 as raw HTML. Fetching them through the API costs $2.00 at $0.0002 per page.

Give your LlamaIndex app clean pages

Twenty URLs in one batch job came back as 103,333 tokens of clean Markdown instead of 1.9 million tokens of HTML, in 49.8 seconds for $0.0038. Every account gets $2 of free API usage per month, enough for 10,000 pages.

Related reading