crawl4ai
View on GitHub🚀🤖 Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper. Don't be shy, join here: https://discord.gg/jP8KfhDhyN
Async Python web crawler that converts pages into clean, LLM-ready Markdown and structured JSON via CSS/XPath or LLM extraction. Includes deep-crawl, caching, sessions/proxies, and a Docker FastAPI server, aimed at RAG pipelines and agent data ingestion.
Use Cases
Turn websites into LLM-ready Markdown for RAG ingestionExtract structured JSON from pages with LLM or CSS/XPath schemasDeep crawl documentation sites with BFS/DFS strategiesFeed agents and pipelines with fresh web dataScrape product prices and catalogs at scaleBypass bot detection using managed browsers and profilesCapture full-page screenshots and page metadata during crawlsBuild custom AI training or evaluation datasets from the webRun headless crawling as a Dockerized FastAPI serviceChunk and filter page content by topic, regex, or BM25 relevance
Built With
- Language
- Python
- Frameworks
- Playwright · Patchright · LiteLLM · FastAPI · Pydantic · asyncio · aiohttp · httpx · BeautifulSoup · lxml · Docker · transformers · sentence-transformers · PyTorch · scikit-learn · NLTK
Tags
web-scraping · web-crawler · llm · markdown · rag · data-extraction · structured-extraction · deep-crawl · browser-automation · async · docker · api-server · chunking · playwright · anti-bot · caching