Web Data and Document Extraction
Crawl4AI
An open-source Python crawler that turns any website into clean, LLM-ready markdown for agents and RAG pipelines.
Quick decision
- Best for
- Python builders who want a free, self-run crawler that outputs LLM-ready markdown, with Docker and cloud options when they outgrow a laptop.
- Not ideal for
- Teams that want a managed API with no infrastructure decisions on day one. The library path puts browsers, proxies, and scaling on you.
- Pricing model
- The library and Docker server are open source under Apache-2.0 and free to run. The optional hosted cloud is pay as you go. See crawl4ai.com for current cloud pricing.
- Deployment
- Self-hosted library. Docker server. Hosted cloud (optional).
- Authentication
- No key needed for the library. The Docker server uses your own CRAWL4AI_API_TOKEN, and the cloud uses a Crawl4AI key. LLM extraction needs your LLM key unless you use the cloud.
- Review state
- Source verified. Facts checked 2026-10-03. Not locally tested by AgentsUse.
What it does
Crawl4AI is an open-source web crawler and scraper for LLMs and AI agents. Its AsyncWebCrawler returns clean markdown from a URL, with content filters that strip menus and boilerplate and extraction strategies based on CSS, XPath, regex, or an LLM.
You can run it as a Python library, as a Docker server with a REST API and monitoring dashboard, or through the hosted Crawl4AI Cloud with one key. The library path needs no account and installs from PyPI.
Crawling features include deep crawl strategies, adaptive crawling that stops when it has learned enough, sessions, proxies, screenshots, and PDF capture. An MCP connection is documented for both the Docker server and the cloud.
Verified capabilities
Data access
LLM-ready markdown Source verified
Clean markdown with headings, tables, and code, plus fit-markdown filters that remove boilerplate.
Structured extraction Source verified
CSS, XPath, and regex strategies with no LLM, plus LLM extraction into a typed JSON schema.
Deep and adaptive crawling Source verified
Breadth-first, depth-first, and best-first crawls, plus adaptive crawling that stops when enough is learned.
Browser control
Full browser control Source verified
Sessions, proxies, cookies, stealth options, and Chromium, Firefox, and WebKit engines.
Deployment
Docker server with REST API Source verified
Self-host with endpoints for markdown, HTML, crawl, screenshots, PDFs, and JavaScript execution.
Agent integration
CLI and MCP paths Source verified
A crwl command line, plus documented MCP connections for the server and cloud.
Quick start
Install from PyPI
From the official README. The setup command installs the browser once, and a doctor command checks the installation.
pip install -U crawl4ai
crawl4ai-setup
crawl4ai-doctorSource: https://github.com/unclecode/crawl4ai. Examples use placeholders only. Never paste a real key into a profile, config file you share, or a ticket.
MCP support: Documented for the Docker server and the hosted cloud. The README shows an MCP connection for Claude Code, and the docs cover other clients.
Works with
MCP
Documented MCP connection for the Docker server and cloud.
Packages
PyPI: crawl4ai. Docker image published by the project.
Browsers
Chromium, Firefox, and WebKit through the bundled browser stack.
Deployment
Library, Docker server, or optional hosted cloud.
Only sourced support is listed. A missing framework means AgentsUse has not verified it yet, not that it cannot work.
Health and maintenance
- GitHub stars
- 84,708 (checked 2026-10-03)
- GitHub forks
- 8,765 (checked 2026-10-03)
- Package
- PyPI: crawl4ai, version 0.9.4
- License
- Open source, Apache-2.0
- Maintainer
- Crawl4AI
Stars and forks from the GitHub repository page. Version 0.9.4 is the latest release noted in the README (Sep 23, 2026). PyPI page checks were blocked by a client challenge in this pass, so no PyPI download figure is shown. Maintenance signals only, not a quality rating.
Pricing and license
The library and Docker server are open source under Apache-2.0 and free to run. The optional hosted cloud is pay as you go. See crawl4ai.com for current cloud pricing.
Limitations and safety
- Self-running means you own browser installs, proxies, and rate limits. Blocked or JS-heavy sites need your configuration, or the cloud path.
- LLM extraction quality depends on the model and schema you supply. Validate structured output before trusting it downstream.
- No independently checked benchmark is published by AgentsUse for this tool yet.
Browser and data tools can read pages, fill forms, and download files. Start with a test account or read-only access, keep credentials in environment variables, and review agent actions before connecting anything that can spend money, send messages, or delete data.
Alternatives to Crawl4AI
Choose Firecrawl when you want a hosted API for search, scrape, crawl, and batch work without running browsers.
Tradeoff: Hosted use needs an account and metered key. The open-source license is AGPL-3.0, not Apache-2.0.
Choose Jina Reader when you need one page or one search turned into markdown with the least setup.
Tradeoff: No deep crawl engine or extraction schema system. It is a fetch and search API first.
Choose Playwright when you need full browser control and deterministic steps rather than crawl output.
Tradeoff: You build the extraction and markdown conversion yourself.
Related tools
Common questions
What does Crawl4AI do for an AI agent?
An open-source Python crawler that turns any website into clean, LLM-ready markdown for agents and RAG pipelines.
Is Crawl4AI open source?
Yes. This profile records the license as Apache-2.0 from the official repository.
Does Crawl4AI support MCP?
Documented for the Docker server and the hosted cloud. The README shows an MCP connection for Claude Code, and the docs cover other clients.
Sources and freshness
- GitHub repository: https://github.com/unclecode/crawl4ai
- Documentation: https://docs.crawl4ai.com
- Product site: https://crawl4ai.com
Last checked 2026-10-03. Verification label: source verified, which means public claims trace to the sources above. It does not mean AgentsUse ran the tool. Spotted an error? Send a correction. Back to the tools directory.