Top open source web scraping & crawling projects for your next project

Top Picks

  • Playwright

    96.6k GitHub stars

    Lets your code open a real browser and click, type and read pages like a person. Used to test websites, collect data and give AI assistants a browser.

    • Browser automation & testing
    • Web scraping & crawling
    • Apache 2.0
  • Crawl4AI

    84.2k GitHub stars

    Reads websites for you and turns them into clean text or tidy data your AI can use, even pages that only load in a browser.

    • Web scraping & crawling
    • RAG & agents
    • Apache 2.0 with attribution clause
  • Scrapy

    64.5k GitHub stars

    A fast, long-trusted Python tool for collecting information from large websites and saving it as tidy data you can use.

    • Web scraping & crawling
    • Command line tools
    • BSD 3-Clause
  • Crawlee

    25.9k GitHub stars

    Build web crawlers that keep going when things go wrong. Reads simple pages quickly and opens a real browser only when a page needs one.

    • Web scraping & crawling
    • RAG & agents
    • Apache 2.0