AgentsUse

Web Data and Document Extraction

Crawl4AI

An open-source Python crawler that turns any website into clean, LLM-ready markdown for agents and RAG pipelines.

Quick decision

Best for
Python builders who want a free, self-run crawler that outputs LLM-ready markdown, with Docker and cloud options when they outgrow a laptop.
Not ideal for
Teams that want a managed API with no infrastructure decisions on day one. The library path puts browsers, proxies, and scaling on you.
Pricing model
The library and Docker server are open source under Apache-2.0 and free to run. The optional hosted cloud is pay as you go. See crawl4ai.com for current cloud pricing.
Deployment
Self-hosted library. Docker server. Hosted cloud (optional).
Authentication
No key needed for the library. The Docker server uses your own CRAWL4AI_API_TOKEN, and the cloud uses a Crawl4AI key. LLM extraction needs your LLM key unless you use the cloud.
Review state
Source verified. Facts checked 2026-10-03. Not locally tested by AgentsUse.

What it does

Crawl4AI is an open-source web crawler and scraper for LLMs and AI agents. Its AsyncWebCrawler returns clean markdown from a URL, with content filters that strip menus and boilerplate and extraction strategies based on CSS, XPath, regex, or an LLM.

You can run it as a Python library, as a Docker server with a REST API and monitoring dashboard, or through the hosted Crawl4AI Cloud with one key. The library path needs no account and installs from PyPI.

Crawling features include deep crawl strategies, adaptive crawling that stops when it has learned enough, sessions, proxies, screenshots, and PDF capture. An MCP connection is documented for both the Docker server and the cloud.

Verified capabilities

Data access

  • LLM-ready markdown Source verified

    Clean markdown with headings, tables, and code, plus fit-markdown filters that remove boilerplate.

  • Structured extraction Source verified

    CSS, XPath, and regex strategies with no LLM, plus LLM extraction into a typed JSON schema.

  • Deep and adaptive crawling Source verified

    Breadth-first, depth-first, and best-first crawls, plus adaptive crawling that stops when enough is learned.

Browser control

  • Full browser control Source verified

    Sessions, proxies, cookies, stealth options, and Chromium, Firefox, and WebKit engines.

Deployment

  • Docker server with REST API Source verified

    Self-host with endpoints for markdown, HTML, crawl, screenshots, PDFs, and JavaScript execution.

Agent integration

  • CLI and MCP paths Source verified

    A crwl command line, plus documented MCP connections for the server and cloud.

Quick start

Install from PyPI

From the official README. The setup command installs the browser once, and a doctor command checks the installation.

pip install -U crawl4ai
crawl4ai-setup
crawl4ai-doctor

Source: https://github.com/unclecode/crawl4ai. Examples use placeholders only. Never paste a real key into a profile, config file you share, or a ticket.

MCP support: Documented for the Docker server and the hosted cloud. The README shows an MCP connection for Claude Code, and the docs cover other clients.

Works with

MCP

Documented MCP connection for the Docker server and cloud.

Packages

PyPI: crawl4ai. Docker image published by the project.

Browsers

Chromium, Firefox, and WebKit through the bundled browser stack.

Deployment

Library, Docker server, or optional hosted cloud.

Only sourced support is listed. A missing framework means AgentsUse has not verified it yet, not that it cannot work.

Health and maintenance

GitHub stars
84,708 (checked 2026-10-03)
GitHub forks
8,765 (checked 2026-10-03)
Package
PyPI: crawl4ai, version 0.9.4
License
Open source, Apache-2.0
Maintainer
Crawl4AI

Stars and forks from the GitHub repository page. Version 0.9.4 is the latest release noted in the README (Sep 23, 2026). PyPI page checks were blocked by a client challenge in this pass, so no PyPI download figure is shown. Maintenance signals only, not a quality rating.

Pricing and license

The library and Docker server are open source under Apache-2.0 and free to run. The optional hosted cloud is pay as you go. See crawl4ai.com for current cloud pricing.

Official pricing or docs →

Limitations and safety

  • Self-running means you own browser installs, proxies, and rate limits. Blocked or JS-heavy sites need your configuration, or the cloud path.
  • LLM extraction quality depends on the model and schema you supply. Validate structured output before trusting it downstream.
  • No independently checked benchmark is published by AgentsUse for this tool yet.

Browser and data tools can read pages, fill forms, and download files. Start with a test account or read-only access, keep credentials in environment variables, and review agent actions before connecting anything that can spend money, send messages, or delete data.

Alternatives to Crawl4AI

Firecrawl

Choose Firecrawl when you want a hosted API for search, scrape, crawl, and batch work without running browsers.

Tradeoff: Hosted use needs an account and metered key. The open-source license is AGPL-3.0, not Apache-2.0.

Jina Reader

Choose Jina Reader when you need one page or one search turned into markdown with the least setup.

Tradeoff: No deep crawl engine or extraction schema system. It is a fetch and search API first.

Playwright

Choose Playwright when you need full browser control and deterministic steps rather than crawl output.

Tradeoff: You build the extraction and markdown conversion yourself.

Related tools

Common questions

What does Crawl4AI do for an AI agent?

An open-source Python crawler that turns any website into clean, LLM-ready markdown for agents and RAG pipelines.

Is Crawl4AI open source?

Yes. This profile records the license as Apache-2.0 from the official repository.

Does Crawl4AI support MCP?

Documented for the Docker server and the hosted cloud. The README shows an MCP connection for Claude Code, and the docs cover other clients.

Sources and freshness

Last checked 2026-10-03. Verification label: source verified, which means public claims trace to the sources above. It does not mean AgentsUse ran the tool. Spotted an error? Send a correction. Back to the tools directory.