Web Data and Document Extraction
Tools that turn web pages and documents into clean, structured input an agent or LLM can use.
How to read this category
These profiles cover the step before an agent can reason at all: turning web pages and files into clean input an agent or LLM can use.
Firecrawl and Crawl4AI fetch and clean live sites, at different points on the hosted-to-self-run spectrum. Jina Reader converts one URL at a time with almost no setup. MarkItDown converts files you already have on disk and never fetches the web at all.
Choose by where the content lives, not by whichever demo looks fastest. Live pages need fetching, rendering, and boilerplate removal. Local PDFs, Office documents, and exports need conversion that preserves headings, tables, and lists. The wrong starting point wastes more time than any speed difference saves.
The verified profiles
Firecrawl
Open source, AGPL-3.0A web data API that searches, scrapes, and crawls sites into clean markdown and structured output for AI agents.
Best for: Agents and pipelines that need clean markdown or structured JSON from many pages without running browsers themselves.
By Firecrawl. Checked 2026-10-03.
View profile →Crawl4AI
Open source, Apache-2.0An open-source Python crawler that turns any website into clean, LLM-ready markdown for agents and RAG pipelines.
Best for: Python builders who want a free, self-run crawler that outputs LLM-ready markdown, with Docker and cloud options when they outgrow a laptop.
By Crawl4AI. Checked 2026-10-03.
View profile →Jina Reader
Open source, Apache-2.0A hosted API that converts any URL into LLM-friendly markdown by adding a prefix, with web search through the same pattern.
Best for: Agents and RAG pipelines that need readable page and search content with almost no setup: one URL prefix, no SDK required.
By Jina AI. Checked 2026-10-03.
View profile →MarkItDown
Open source, MITA lightweight Python utility that converts files and Office documents into markdown for LLMs and text pipelines.
Best for: Pipelines that already have files on disk and need clean markdown for an LLM, with fine-grained optional installs.
By Microsoft. Checked 2026-10-03.
View profile →Side by side on sourced facts
| Tool | Pricing model | License | Deployment | MCP support | Checked |
|---|---|---|---|---|---|
| Firecrawl | Freemium (hosted paid tiers) | AGPL-3.0 | Hosted API, Self-hosted (open source) | firecrawl-mcp package | 2026-10-03 |
| Crawl4AI | Open source | Apache-2.0 | Self-hosted library, Docker server, Hosted cloud (optional) | Documented for Docker and cloud | 2026-10-03 |
| Jina Reader | Freemium (hosted paid tiers) | Apache-2.0 | Hosted API, Self-hosted Docker | No MCP claim verified | 2026-10-03 |
| MarkItDown | Open source | MIT | Self-hosted library, Command line, Docker | No MCP claim verified | 2026-10-03 |
Every cell traces to the tool profile, which traces to the official repository, package page, and documentation checked on the date shown. Pricing cells name the model, not a price: vendor prices change, so check the profile for the sourced pricing summary and the official pricing link.
What actually decides the choice
- Input first: live URLs point at Firecrawl, Crawl4AI, or Jina Reader; local files point at MarkItDown.
- Operations decide second: a hosted API trades an account and a metered key for zero browser maintenance, while the self-run libraries trade your machines for no per-call account.
- Depth splits the fetchers: one page or one search is Jina Reader territory, whole-site crawls and schema-driven extraction belong to Firecrawl and Crawl4AI.
- License terms differ materially (Apache-2.0, MIT, and AGPL-3.0 all appear here). Check what a license means for your deployment before standardizing on one.
How this archive is ordered, and why it is short
Profiles appear in the order they passed the AgentsUse review gate, not in a tested quality ranking. AgentsUse has not run head-to-head benchmarks for this category, so this page does not claim a winner. The comparison table is the ranking tool: sort by the dimension your job cannot compromise on, then read the full profiles for limitations.
This archive currently lists 4 verified profiles. AgentsUse normally asks for eight published profiles before opening a category archive. This one opens early, on the record, because every listed profile passed the full tool gate and readers comparing real choices need the side-by-side surface now, not a longer unvetted list later. The archive grows as more tools pass review. See how verification works.
Read before you give these tools real work
How to Evaluate Agent Output Quality
It seems to work is not an evaluation strategy. How to define quality concretely, build runnable evals, and track agent performance over time.
Read guide →Building Agent Error Recovery
Agents fail; the question is whether failure is graceful or catastrophic. A taxonomy of agent failures and the recovery patterns - retries, fallbacks, circuit breakers - that contain them.
Read guide →Back to the tools directory, the category index, or the guides.