The Open-Source Scraping Index 2026: GitHub Stars, Downloads and the Rise of AI-Native Scrapers

Home › Guides › The Open-Source Scraping Index

✓ data: GitHub, PyPI, npm, Stack Overflow and Hacker News, collected 2026-10-02; star history from the Wayback Machine, October 2024 to April 2026 · the index · method · download the data

Which open-source scraping tools are developers actually adopting? We tracked 20 scraping, crawling, browser-automation and anti-detection tools across five public sources: GitHub stars, forks and commits; monthly downloads from PyPI and npm; Stack Overflow questions; and Hacker News mentions. For star history we read archived copies of each GitHub repository page in the Wayback Machine.

One trend runs through the data. Tools built to feed web data to large language models and AI agents went from a small share of developer attention two years ago to the top of the GitHub rankings. The classic tools, Scrapy, Beautiful Soup, Selenium, Puppeteer and Cheerio, still account for far more downloads.

AI-native: Firecrawl, Browser Use, Crawl4AI, Stagehand Stealth / anti-detect: Scrapling, curl_cffi, undetected-chromedriver, nodriver, playwright-stealth, puppeteer-extra-plugin-stealth Browser automation: Playwright, Puppeteer, Selenium Classic scraping & HTTP: Scrapy, Beautiful Soup, Requests-HTML, HTTPX, Cheerio, Crawlee, Colly

Headline numbers

  1. Firecrawl is now the most-starred scraping project on GitHub, with 187,783 stars on 2 October 2026. That is 2.9× Scrapy’s 64,550. In October 2024 it ranked #7 with 15,727 stars; today’s count is 11.94× that.
  2. AI-native tools took 63.5% of all new stars. Of the 334,539 stars gained by the 19 GitHub projects we track over the last 12 months, Firecrawl, Browser Use, Crawl4AI and Stagehand gained 212,336. Their share of all stars rose from 6.5% in October 2024 to 44.3% today.
  3. The top six now includes four projects created in 2024. Firecrawl, Browser Use, Scrapling and Crawl4AI rank #1, #2, #5 and #6, around Playwright (#3) and Puppeteer (#4). Puppeteer was #1 in October 2024 and Scrapy #3; Scrapy is now #7.
  4. Scrapling had the fastest growth of any tool tracked. It went from 7,376 to 85,148 stars in 12 months (11.54×). Its PyPI downloads rose from 21,668 in September 2025 to 901,834 in September 2026 (41.62×).
  5. Classic parsers still dominate downloads. Beautiful Soup was downloaded 297.1 million times from PyPI in September 2026: 38× Browser Use (7.8 million) and 47× Firecrawl’s Python SDKs (6.3 million). On npm, Cheerio had 113.2 million downloads, against 5.1 million for Firecrawl’s JavaScript SDK.
  6. Playwright is pulling away from Puppeteer and Selenium. Its npm package was downloaded 400.9 million times in September 2026, 8.4× Puppeteer’s 47.6 million; in October 2023 Puppeteer was ahead (23.2 million to 10.3 million). On PyPI, Playwright passed Selenium in April 2026 (85.1 million to 57.2 million) and grew 3.46× over the year to September 2026. Selenium fell from 38.6 million to 28.5 million, most of that in a single step in late August (see caveats).
  7. Some stealth tools are still downloaded heavily without updates. undetected-chromedriver has had no commits in the last 52 weeks and no PyPI release since February 2024, yet it was downloaded 2.0 million times in September 2026. puppeteer-extra-plugin-stealth, last published to npm in March 2023, was downloaded 4.8 million times, 3.31× the figure a year earlier. Newer alternatives are growing faster: nodriver, which describes itself as the “Successor of Undetected-Chromedriver”, grew 3.35× on PyPI, and curl_cffi grew 2.9× to 33.1 million downloads.
  8. Stack Overflow is no longer where scraping questions get asked. Questions tagged web-scraping fell from 7,949 in 2020 to 268 in 2025 (96.6% lower), and to 31 in January to September 2026. The whole site fell 94.1% over the same period. The newer tools barely appear: the firecrawl tag has 5 questions and browser-use 9; Crawl4AI, Stagehand, Scrapling and curl_cffi have no tag at all.

GitHub stars over two years

Line chart of GitHub stars October 2024 to October 2026: Firecrawl rises from about 16k to 188k, Browser Use to 117k, Scrapling and Crawl4AI to about 85k, while Playwright reaches 97k, Puppeteer 96k, Scrapy 65k and Selenium 35k
Stars at the Wayback Machine capture nearest each date, plus the GitHub API on 2 October 2026.

Over the last 12 months the four AI-native projects together went from 202,609 to 414,945 stars (2.05×), and the six stealth and anti-detection projects from 34,257 to 117,832 (3.44×, most of it Scrapling). Browser-automation frameworks grew 1.12× and classic scraping and HTTP libraries 1.09×.

Bar chart of GitHub stars gained October 2025 to October 2026: Firecrawl +127,035, Scrapling +77,772, Browser Use +46,914, Crawl4AI +30,476, Playwright +19,365; classic tools gained under 7,000 each

Downloads: September 2026 against September 2025

Bar chart of download growth, September 2026 vs September 2025: Scrapling 41.6x on PyPI, playwright-stealth 6.8x, Crawlee PyPI 6.7x, Firecrawl npm 6.4x, Browser Use 5.9x, Playwright npm 4.5x; Selenium PyPI 0.74x and Requests-HTML 0.53x declined
Downloads are file fetches, not users; see what downloads measure.

Most packages grew. For scale, the general-purpose HTTP library requests went from 874.9 million to 1.21 billion PyPI downloads (1.38×). Packages that grew less than that lost ground. Three declined in absolute terms: Selenium on PyPI (0.74×), undetected-chromedriver (0.93×) and Requests-HTML (0.53×), whose last PyPI release was in February 2019.

The index

One row per tool, sorted by GitHub stars. “12-mo” compares stars on 2 October 2026 with the Wayback capture nearest 1 October 2025. Downloads are for the calendar month of September 2026, with the change from September 2025 in brackets.

ToolGroupGitHub stars (rank)Rank Oct 202412-mo star growthCommits, last 52 wksPyPI downloads, Sep 2026npm downloads, Sep 2026Stack Overflow questions, 2025HN mentions, Jan–Sep 2026
Firecrawl
firecrawl/firecrawl
AI-native187,783 (#1)#7+127,035 (3.09×)2,1436.3 million (3×)5.1 million (6.43×)384
Browser Use
browser-use/browser-use
AI-native116,999 (#2)not yet created+46,914 (1.67×)3,5697.8 million (5.9×)–8320
Playwright
microsoft/playwright
Browser automation97,000 (#3)#2+19,365 (1.25×)2,61695.0 million (3.46×)400.9 million (4.51×)2591,134
Puppeteer
puppeteer/puppeteer
Browser automation95,647 (#4)#1+3,139 (1.03×)792–47.6 million (1.94×)94190
Scrapling
D4Vinci/Scrapling
Stealth / anti-detect85,148 (#5)not yet created+77,772 (11.54×)819901,834 (41.62×)–no tag7
Crawl4AI
unclecode/crawl4ai
AI-native84,645 (#6)#12+30,476 (1.56×)5871.3 million (1.58×)–no tag11
Scrapy
scrapy/scrapy
Classic scraping & HTTP64,550 (#7)#3+6,050 (1.1×)6404.4 million (2.17×)–2933
Selenium
SeleniumHQ/selenium
Browser automation34,518 (#8)#4+1,147 (1.03×)1,59828.5 million (0.74×)8.7 million (1.23×)589181
Cheerio
cheeriojs/cheerio
Classic scraping & HTTP30,517 (#9)#5+736 (1.02×)561–113.2 million (2.18×)348
Crawlee
apify/crawlee
Classic scraping & HTTP25,973 (#10)#8+6,410 (1.33×)653889,060 (6.69×)663,113 (3.94×)03
Colly
gocolly/colly
Classic scraping & HTTP25,542 (#11)#6+839 (1.03×)47––042
Stagehand
browserbase/stagehand
AI-native25,518 (#12)#17+7,911 (1.45×)8301.3 million (2.92×)6.5 million (3.46×)no tag48
HTTPX
encode/httpx
Classic scraping & HTTP15,525 (#13)#10+984 (1.07×)5616.7 million (2.18×)–15103
Requests-HTML
psf/requests-html
Classic scraping & HTTP13,810 (#14)#9-420344,992 (0.53×)–20
undetected-chromedriver
ultrafunkamsterdam/undetected-chromedriver
Stealth / anti-detect12,857 (#15)#11+1,061 (1.09×)02.0 million (0.93×)–161
puppeteer-extra (stealth)
berstend/puppeteer-extra
Stealth / anti-detect7,404 (#16)#13+348 (1.05×)0–4.8 million (3.31×)no tag3
curl_cffi
lexiforest/curl_cffi
Stealth / anti-detect6,637 (#17)#14+2,300 (1.53×)13833.1 million (2.9×)–no tag18
nodriver
ultrafunkamsterdam/nodriver
Stealth / anti-detect4,799 (#18)#15+1,863 (1.63×)15296,152 (3.35×)–121
playwright-stealth
AtuboDad/playwright_stealth
Stealth / anti-detect987 (#19)#16+231 (1.31×)02.8 million (6.82×)–no tag4
Beautiful Soup
launchpad.net/beautifulsoup
Classic scraping & HTTPnot on GitHub–––297.1 million (1.36×)–7328

Packages counted: Firecrawl = firecrawl-py + firecrawl on PyPI (Firecrawl’s Python SDK is published under both names) and @mendable/firecrawl-js on npm; Stagehand = stagehand on PyPI and @browserbasehq/stagehand on npm; Crawlee = crawlee on PyPI (from apify/crawlee-python, 9,574 stars, not included in the star totals) and on npm; Selenium = selenium on PyPI and selenium-webdriver on npm; puppeteer-extra = puppeteer-extra-plugin-stealth on npm. The playwright-stealth package on PyPI is now published from a fork (Mattwmaster58/playwright_stealth, 268 stars); the star figures are for the original AtuboDad repository. Colly is a Go library and has no comparable download count.

Hacker News mentions

Stories and comments that mention each name, by year, from the Hacker News search API (exact word, no typo matching). Names that are also ordinary words pick up unrelated matches: “browser use” appeared 128 times in 2022 and 143 in 2023, before the Browser Use project existed, and Selenium, Cheerio and Colly also have other meanings. The second figure counts only stories whose link contains the project’s name.

Tool202320242025Jan–Sep 2026
Playwright401 (35 links)543 (30 links)1,055 (34 links)1,134 (32 links)
Selenium453 (12 links)440 (10 links)361 (4 links)181 (2 links)
Puppeteer282 (15 links)273 (12 links)325 (5 links)190 (3 links)
Browser Use143 (0 links)122 (1 link)370 (17 links)320 (27 links)
Firecrawl0 (0 links)30 (13 links)75 (4 links)84 (12 links)
HTTPX23 (3 links)48 (6 links)56 (2 links)103 (3 links)
Scrapy58 (4 links)59 (4 links)29 (3 links)33 (1 link)
BeautifulSoup38 (2 links)55 (2 links)45 (0 links)28 (0 links)
Stagehand4 (0 links)9 (1 link)55 (4 links)48 (7 links)
Cheerio75 (1 link)47 (2 links)54 (1 link)48 (1 link)
Crawl4AI0 (0 links)4 (2 links)18 (4 links)11 (3 links)
curl_cffi0 (0 links)5 (0 links)11 (1 link)18 (3 links)

On Hacker News, Playwright was mentioned 1,055 times in 2025, up from 543 in 2024 and well ahead of Selenium (361) and Puppeteer (325). Firecrawl mentions went from 30 in 2024 to 84 in the first nine months of 2026.

Stack Overflow

Line chart indexed to 2020: questions tagged web-scraping fell from 7,949 in 2020 to 268 in 2025, tracking the fall in all Stack Overflow questions
Tag2020202220242025Jan–Sep 2026
web-scraping7,9496,1981,40126831
selenium-webdriver17,58913,2822,52958952
beautifulsoup5,5133,353465737
scrapy2,1111,157188293
puppeteer1,8521,081407947
playwright12191081025959
undetected-chromedriver08449160
cheerio2481792830
(all questions)1,854,9791,334,835398,390109,26819,332

Web-scraping’s share of all new questions rose from 0.337% in 2019 to 0.464% in 2022, then fell to 0.245% in 2025. Playwright questions kept rising until 2023, when they peaked at 954, as did undetected-chromedriver (peak 114). The older tags were already falling by then.

Methodology

  • Tools and groups. 20 open-source tools used for scraping, crawling and browser automation. We grouped them by how each project describes itself. AI-native: built to supply web data to LLMs or agents (Crawl4AI: “Open-source web crawler and scraper for LLMs and AI agents”; Firecrawl: “Supercharge your AI agents with data from the web”; Browser Use: “Agents that use the browser”; Stagehand on PyPI: “SDK for building browser agents”). Stealth / anti-detect: tools whose main purpose is avoiding bot detection, or that advertise it (curl_cffi “can impersonate browser tls/ja3/http2 fingerprints”; Scrapling’s PyPI summary calls it “an undetectable … Python library”). The groups are our classification. Several tools fit more than one; Crawlee, for example, also advertises data “for AI, LLMs, RAG, or GPTs”.
  • GitHub. Stars, forks, creation date and last push come from the GitHub search API on 2 October 2026. Commits are the sum of the 52 weekly counts from GitHub’s participation statistics, covering the year to 2 October 2026. GitHub’s stargazer timestamps need an authenticated API, so for history we used the Wayback Machine. For each repository and each target date (1 October 2024, 1 April 2025, 1 October 2025, 1 April 2026) we took the closest archived capture of the repository page within 45 days and read the star counter from the saved HTML. 68 of the 82 captures used are within 7 days of the target and the furthest is 45 days away; the exact capture used for each figure is in the CSV. Firecrawl’s earlier captures are under its previous address, github.com/mendableai/firecrawl. A repository created after a target date counts as 0. Growth multiples from a base below 1,000 stars are not reported.
  • PyPI. Monthly downloads by calendar month, October 2023 to September 2026, from ClickHouse’s public copy of the PyPI download logs (the same file-download records that feed Google BigQuery and pypistats.org). These counts include every installer, mirrors included. As a cross-check, our September 2026 figures for 6 packages are 0.3% to 3.4% above pypistats.org’s “without mirrors” numbers.
  • npm. Daily download counts from the npm registry’s downloads API, October 2023 to September 2026, summed by calendar month.
  • Stack Overflow. Stack Exchange API, fetched 2 October 2026: all-time question count per tag, and questions created per calendar year that are still on the site (deleted questions are not counted).
  • Hacker News. Algolia’s Hacker News search API: stories plus comments matching each name as an exact word, by year of posting, plus stories whose URL contains the project’s name.

What these numbers do and don’t measure

  • Downloads are not users. Every CI run, Docker build and fresh virtual environment counts as a download, and so does every install of a package that depends on the tool. Fast-growing installers such as uv make repeated fresh installs cheap: uv accounted for 64.4% of Crawl4AI’s and 50.7% of Browser Use’s September 2026 PyPI downloads, against 44.3% for Beautiful Soup. HTTPX (616.7 million downloads, 5 commits in 52 weeks) is a dependency of many widely installed packages, so most of its downloads are unlikely to be for scraping. Compare download counts within a tool over time, not as head counts across tools.
  • Step changes. A single large automated user can move a package’s numbers on its own. Selenium’s daily PyPI downloads (pypistats.org, without mirrors) averaged 2,100,609 from 17 to 24 August 2026 and 881,780 from 25 to 31 August, a step down that the whole September figure inherits.
  • Playwright and Puppeteer downloads include testing. Both are used heavily for browser testing and are installed as dependencies of other packages. Their download counts are not scraping counts.
  • Stars measure attention. They can be bought or driven by a single viral post. A project with fast star growth is not necessarily more used than one with flat stars; for example, Beautiful Soup has no GitHub stars at all because it is hosted on Launchpad.
  • Wayback sampling. Star counts are read from archived pages, not from GitHub’s own history, and the capture can be up to 45 days from the target date.
  • Stack Overflow’s decline is site-wide. Because the fall affects the whole site, it says little about interest in any one tool.
  • Name collisions. Hacker News counts for ordinary words (browser use, selenium, cheerio, colly, stagehand, playwright) include unrelated uses.
  • Selection. These are popular open-source tools as of 2026, not a complete census; commercial scraping APIs are not covered except through their open-source SDKs (Firecrawl’s SDK downloads largely represent clients of its hosted API).

Download the data

Free to use with attribution (CC BY 4.0).

Cite or embed

Suggested citation: ProxyPicker (2026). “The Open-Source Scraping Index: GitHub stars, PyPI and npm downloads, Stack Overflow and Hacker News activity for 20 scraping and browser-automation tools.” Data collected 2 October 2026. https://proxypicker.com/open-source-scraping-index/

Embed the star chart:

<a href="https://proxypicker.com/open-source-scraping-index/"><img src="https://proxypicker.com/wp-content/uploads/2026/10/proxypicker-oss-index-2026-10-github-stars.png" alt="GitHub stars of open-source scraping tools, 2024 to 2026 (ProxyPicker)" width="800" /></a><br />Source: <a href="https://proxypicker.com/open-source-scraping-index/">ProxyPicker, The Open-Source Scraping Index</a> (2 October 2026)

Related: AI Crawler Census (which AI bots sites block) · Anti-Bot Census (which bot-protection vendors sites use) · Proxies for web scraping · Best scraping APIs · Scraping Law Hub · All guides.

Press & data requests: email [email protected] for interviews, the raw data behind this page, or corrections. See also our contact page.