Home › Guides › The Open-Source Scraping Index
✓ data: GitHub, PyPI, npm, Stack Overflow and Hacker News, collected 2026-10-02; star history from the Wayback Machine, October 2024 to April 2026 · the index · method · download the data
Which open-source scraping tools are developers actually adopting? We tracked 20 scraping, crawling, browser-automation and anti-detection tools across five public sources: GitHub stars, forks and commits; monthly downloads from PyPI and npm; Stack Overflow questions; and Hacker News mentions. For star history we read archived copies of each GitHub repository page in the Wayback Machine.
One trend runs through the data. Tools built to feed web data to large language models and AI agents went from a small share of developer attention two years ago to the top of the GitHub rankings. The classic tools, Scrapy, Beautiful Soup, Selenium, Puppeteer and Cheerio, still account for far more downloads.
AI-native: Firecrawl, Browser Use, Crawl4AI, Stagehand Stealth / anti-detect: Scrapling, curl_cffi, undetected-chromedriver, nodriver, playwright-stealth, puppeteer-extra-plugin-stealth Browser automation: Playwright, Puppeteer, Selenium Classic scraping & HTTP: Scrapy, Beautiful Soup, Requests-HTML, HTTPX, Cheerio, Crawlee, Colly
Headline numbers
- Firecrawl is now the most-starred scraping project on GitHub, with 187,783 stars on 2 October 2026. That is 2.9× Scrapy’s 64,550. In October 2024 it ranked #7 with 15,727 stars; today’s count is 11.94× that.
- AI-native tools took 63.5% of all new stars. Of the 334,539 stars gained by the 19 GitHub projects we track over the last 12 months, Firecrawl, Browser Use, Crawl4AI and Stagehand gained 212,336. Their share of all stars rose from 6.5% in October 2024 to 44.3% today.
- The top six now includes four projects created in 2024. Firecrawl, Browser Use, Scrapling and Crawl4AI rank #1, #2, #5 and #6, around Playwright (#3) and Puppeteer (#4). Puppeteer was #1 in October 2024 and Scrapy #3; Scrapy is now #7.
- Scrapling had the fastest growth of any tool tracked. It went from 7,376 to 85,148 stars in 12 months (11.54×). Its PyPI downloads rose from 21,668 in September 2025 to 901,834 in September 2026 (41.62×).
- Classic parsers still dominate downloads. Beautiful Soup was downloaded 297.1 million times from PyPI in September 2026: 38× Browser Use (7.8 million) and 47× Firecrawl’s Python SDKs (6.3 million). On npm, Cheerio had 113.2 million downloads, against 5.1 million for Firecrawl’s JavaScript SDK.
- Playwright is pulling away from Puppeteer and Selenium. Its npm package was downloaded 400.9 million times in September 2026, 8.4× Puppeteer’s 47.6 million; in October 2023 Puppeteer was ahead (23.2 million to 10.3 million). On PyPI, Playwright passed Selenium in April 2026 (85.1 million to 57.2 million) and grew 3.46× over the year to September 2026. Selenium fell from 38.6 million to 28.5 million, most of that in a single step in late August (see caveats).
- Some stealth tools are still downloaded heavily without updates. undetected-chromedriver has had no commits in the last 52 weeks and no PyPI release since February 2024, yet it was downloaded 2.0 million times in September 2026. puppeteer-extra-plugin-stealth, last published to npm in March 2023, was downloaded 4.8 million times, 3.31× the figure a year earlier. Newer alternatives are growing faster: nodriver, which describes itself as the “Successor of Undetected-Chromedriver”, grew 3.35× on PyPI, and curl_cffi grew 2.9× to 33.1 million downloads.
- Stack Overflow is no longer where scraping questions get asked. Questions tagged web-scraping fell from 7,949 in 2020 to 268 in 2025 (96.6% lower), and to 31 in January to September 2026. The whole site fell 94.1% over the same period. The newer tools barely appear: the firecrawl tag has 5 questions and browser-use 9; Crawl4AI, Stagehand, Scrapling and curl_cffi have no tag at all.
GitHub stars over two years

Over the last 12 months the four AI-native projects together went from 202,609 to 414,945 stars (2.05×), and the six stealth and anti-detection projects from 34,257 to 117,832 (3.44×, most of it Scrapling). Browser-automation frameworks grew 1.12× and classic scraping and HTTP libraries 1.09×.

Downloads: September 2026 against September 2025

Most packages grew. For scale, the general-purpose HTTP library requests went from 874.9 million to 1.21 billion PyPI downloads (1.38×). Packages that grew less than that lost ground. Three declined in absolute terms: Selenium on PyPI (0.74×), undetected-chromedriver (0.93×) and Requests-HTML (0.53×), whose last PyPI release was in February 2019.
The index
One row per tool, sorted by GitHub stars. “12-mo” compares stars on 2 October 2026 with the Wayback capture nearest 1 October 2025. Downloads are for the calendar month of September 2026, with the change from September 2025 in brackets.
| Tool | Group | GitHub stars (rank) | Rank Oct 2024 | 12-mo star growth | Commits, last 52 wks | PyPI downloads, Sep 2026 | npm downloads, Sep 2026 | Stack Overflow questions, 2025 | HN mentions, Jan–Sep 2026 |
|---|---|---|---|---|---|---|---|---|---|
| Firecrawl firecrawl/firecrawl | AI-native | 187,783 (#1) | #7 | +127,035 (3.09×) | 2,143 | 6.3 million (3×) | 5.1 million (6.43×) | 3 | 84 |
| Browser Use browser-use/browser-use | AI-native | 116,999 (#2) | not yet created | +46,914 (1.67×) | 3,569 | 7.8 million (5.9×) | – | 8 | 320 |
| Playwright microsoft/playwright | Browser automation | 97,000 (#3) | #2 | +19,365 (1.25×) | 2,616 | 95.0 million (3.46×) | 400.9 million (4.51×) | 259 | 1,134 |
| Puppeteer puppeteer/puppeteer | Browser automation | 95,647 (#4) | #1 | +3,139 (1.03×) | 792 | – | 47.6 million (1.94×) | 94 | 190 |
| Scrapling D4Vinci/Scrapling | Stealth / anti-detect | 85,148 (#5) | not yet created | +77,772 (11.54×) | 819 | 901,834 (41.62×) | – | no tag | 7 |
| Crawl4AI unclecode/crawl4ai | AI-native | 84,645 (#6) | #12 | +30,476 (1.56×) | 587 | 1.3 million (1.58×) | – | no tag | 11 |
| Scrapy scrapy/scrapy | Classic scraping & HTTP | 64,550 (#7) | #3 | +6,050 (1.1×) | 640 | 4.4 million (2.17×) | – | 29 | 33 |
| Selenium SeleniumHQ/selenium | Browser automation | 34,518 (#8) | #4 | +1,147 (1.03×) | 1,598 | 28.5 million (0.74×) | 8.7 million (1.23×) | 589 | 181 |
| Cheerio cheeriojs/cheerio | Classic scraping & HTTP | 30,517 (#9) | #5 | +736 (1.02×) | 561 | – | 113.2 million (2.18×) | 3 | 48 |
| Crawlee apify/crawlee | Classic scraping & HTTP | 25,973 (#10) | #8 | +6,410 (1.33×) | 653 | 889,060 (6.69×) | 663,113 (3.94×) | 0 | 3 |
| Colly gocolly/colly | Classic scraping & HTTP | 25,542 (#11) | #6 | +839 (1.03×) | 47 | – | – | 0 | 42 |
| Stagehand browserbase/stagehand | AI-native | 25,518 (#12) | #17 | +7,911 (1.45×) | 830 | 1.3 million (2.92×) | 6.5 million (3.46×) | no tag | 48 |
| HTTPX encode/httpx | Classic scraping & HTTP | 15,525 (#13) | #10 | +984 (1.07×) | 5 | 616.7 million (2.18×) | – | 15 | 103 |
| Requests-HTML psf/requests-html | Classic scraping & HTTP | 13,810 (#14) | #9 | -42 | 0 | 344,992 (0.53×) | – | 2 | 0 |
| undetected-chromedriver ultrafunkamsterdam/undetected-chromedriver | Stealth / anti-detect | 12,857 (#15) | #11 | +1,061 (1.09×) | 0 | 2.0 million (0.93×) | – | 16 | 1 |
| puppeteer-extra (stealth) berstend/puppeteer-extra | Stealth / anti-detect | 7,404 (#16) | #13 | +348 (1.05×) | 0 | – | 4.8 million (3.31×) | no tag | 3 |
| curl_cffi lexiforest/curl_cffi | Stealth / anti-detect | 6,637 (#17) | #14 | +2,300 (1.53×) | 138 | 33.1 million (2.9×) | – | no tag | 18 |
| nodriver ultrafunkamsterdam/nodriver | Stealth / anti-detect | 4,799 (#18) | #15 | +1,863 (1.63×) | 15 | 296,152 (3.35×) | – | 12 | 1 |
| playwright-stealth AtuboDad/playwright_stealth | Stealth / anti-detect | 987 (#19) | #16 | +231 (1.31×) | 0 | 2.8 million (6.82×) | – | no tag | 4 |
| Beautiful Soup launchpad.net/beautifulsoup | Classic scraping & HTTP | not on GitHub | – | – | – | 297.1 million (1.36×) | – | 73 | 28 |
Packages counted: Firecrawl = firecrawl-py + firecrawl on PyPI (Firecrawl’s Python SDK is published under both names) and @mendable/firecrawl-js on npm; Stagehand = stagehand on PyPI and @browserbasehq/stagehand on npm; Crawlee = crawlee on PyPI (from apify/crawlee-python, 9,574 stars, not included in the star totals) and on npm; Selenium = selenium on PyPI and selenium-webdriver on npm; puppeteer-extra = puppeteer-extra-plugin-stealth on npm. The playwright-stealth package on PyPI is now published from a fork (Mattwmaster58/playwright_stealth, 268 stars); the star figures are for the original AtuboDad repository. Colly is a Go library and has no comparable download count.
Hacker News mentions
Stories and comments that mention each name, by year, from the Hacker News search API (exact word, no typo matching). Names that are also ordinary words pick up unrelated matches: “browser use” appeared 128 times in 2022 and 143 in 2023, before the Browser Use project existed, and Selenium, Cheerio and Colly also have other meanings. The second figure counts only stories whose link contains the project’s name.
| Tool | 2023 | 2024 | 2025 | Jan–Sep 2026 |
|---|---|---|---|---|
| Playwright | 401 (35 links) | 543 (30 links) | 1,055 (34 links) | 1,134 (32 links) |
| Selenium | 453 (12 links) | 440 (10 links) | 361 (4 links) | 181 (2 links) |
| Puppeteer | 282 (15 links) | 273 (12 links) | 325 (5 links) | 190 (3 links) |
| Browser Use | 143 (0 links) | 122 (1 link) | 370 (17 links) | 320 (27 links) |
| Firecrawl | 0 (0 links) | 30 (13 links) | 75 (4 links) | 84 (12 links) |
| HTTPX | 23 (3 links) | 48 (6 links) | 56 (2 links) | 103 (3 links) |
| Scrapy | 58 (4 links) | 59 (4 links) | 29 (3 links) | 33 (1 link) |
| BeautifulSoup | 38 (2 links) | 55 (2 links) | 45 (0 links) | 28 (0 links) |
| Stagehand | 4 (0 links) | 9 (1 link) | 55 (4 links) | 48 (7 links) |
| Cheerio | 75 (1 link) | 47 (2 links) | 54 (1 link) | 48 (1 link) |
| Crawl4AI | 0 (0 links) | 4 (2 links) | 18 (4 links) | 11 (3 links) |
| curl_cffi | 0 (0 links) | 5 (0 links) | 11 (1 link) | 18 (3 links) |
On Hacker News, Playwright was mentioned 1,055 times in 2025, up from 543 in 2024 and well ahead of Selenium (361) and Puppeteer (325). Firecrawl mentions went from 30 in 2024 to 84 in the first nine months of 2026.
Stack Overflow

| Tag | 2020 | 2022 | 2024 | 2025 | Jan–Sep 2026 |
|---|---|---|---|---|---|
| web-scraping | 7,949 | 6,198 | 1,401 | 268 | 31 |
| selenium-webdriver | 17,589 | 13,282 | 2,529 | 589 | 52 |
| beautifulsoup | 5,513 | 3,353 | 465 | 73 | 7 |
| scrapy | 2,111 | 1,157 | 188 | 29 | 3 |
| puppeteer | 1,852 | 1,081 | 407 | 94 | 7 |
| playwright | 121 | 910 | 810 | 259 | 59 |
| undetected-chromedriver | 0 | 84 | 49 | 16 | 0 |
| cheerio | 248 | 179 | 28 | 3 | 0 |
| (all questions) | 1,854,979 | 1,334,835 | 398,390 | 109,268 | 19,332 |
Web-scraping’s share of all new questions rose from 0.337% in 2019 to 0.464% in 2022, then fell to 0.245% in 2025. Playwright questions kept rising until 2023, when they peaked at 954, as did undetected-chromedriver (peak 114). The older tags were already falling by then.
Methodology
- Tools and groups. 20 open-source tools used for scraping, crawling and browser automation. We grouped them by how each project describes itself. AI-native: built to supply web data to LLMs or agents (Crawl4AI: “Open-source web crawler and scraper for LLMs and AI agents”; Firecrawl: “Supercharge your AI agents with data from the web”; Browser Use: “Agents that use the browser”; Stagehand on PyPI: “SDK for building browser agents”). Stealth / anti-detect: tools whose main purpose is avoiding bot detection, or that advertise it (curl_cffi “can impersonate browser tls/ja3/http2 fingerprints”; Scrapling’s PyPI summary calls it “an undetectable … Python library”). The groups are our classification. Several tools fit more than one; Crawlee, for example, also advertises data “for AI, LLMs, RAG, or GPTs”.
- GitHub. Stars, forks, creation date and last push come from the GitHub search API on 2 October 2026. Commits are the sum of the 52 weekly counts from GitHub’s participation statistics, covering the year to 2 October 2026. GitHub’s stargazer timestamps need an authenticated API, so for history we used the Wayback Machine. For each repository and each target date (1 October 2024, 1 April 2025, 1 October 2025, 1 April 2026) we took the closest archived capture of the repository page within 45 days and read the star counter from the saved HTML. 68 of the 82 captures used are within 7 days of the target and the furthest is 45 days away; the exact capture used for each figure is in the CSV. Firecrawl’s earlier captures are under its previous address, github.com/mendableai/firecrawl. A repository created after a target date counts as 0. Growth multiples from a base below 1,000 stars are not reported.
- PyPI. Monthly downloads by calendar month, October 2023 to September 2026, from ClickHouse’s public copy of the PyPI download logs (the same file-download records that feed Google BigQuery and pypistats.org). These counts include every installer, mirrors included. As a cross-check, our September 2026 figures for 6 packages are 0.3% to 3.4% above pypistats.org’s “without mirrors” numbers.
- npm. Daily download counts from the npm registry’s downloads API, October 2023 to September 2026, summed by calendar month.
- Stack Overflow. Stack Exchange API, fetched 2 October 2026: all-time question count per tag, and questions created per calendar year that are still on the site (deleted questions are not counted).
- Hacker News. Algolia’s Hacker News search API: stories plus comments matching each name as an exact word, by year of posting, plus stories whose URL contains the project’s name.
What these numbers do and don’t measure
- Downloads are not users. Every CI run, Docker build and fresh virtual environment counts as a download, and so does every install of a package that depends on the tool. Fast-growing installers such as uv make repeated fresh installs cheap: uv accounted for 64.4% of Crawl4AI’s and 50.7% of Browser Use’s September 2026 PyPI downloads, against 44.3% for Beautiful Soup. HTTPX (616.7 million downloads, 5 commits in 52 weeks) is a dependency of many widely installed packages, so most of its downloads are unlikely to be for scraping. Compare download counts within a tool over time, not as head counts across tools.
- Step changes. A single large automated user can move a package’s numbers on its own. Selenium’s daily PyPI downloads (pypistats.org, without mirrors) averaged 2,100,609 from 17 to 24 August 2026 and 881,780 from 25 to 31 August, a step down that the whole September figure inherits.
- Playwright and Puppeteer downloads include testing. Both are used heavily for browser testing and are installed as dependencies of other packages. Their download counts are not scraping counts.
- Stars measure attention. They can be bought or driven by a single viral post. A project with fast star growth is not necessarily more used than one with flat stars; for example, Beautiful Soup has no GitHub stars at all because it is hosted on Launchpad.
- Wayback sampling. Star counts are read from archived pages, not from GitHub’s own history, and the capture can be up to 45 days from the target date.
- Stack Overflow’s decline is site-wide. Because the fall affects the whole site, it says little about interest in any one tool.
- Name collisions. Hacker News counts for ordinary words (browser use, selenium, cheerio, colly, stagehand, playwright) include unrelated uses.
- Selection. These are popular open-source tools as of 2026, not a complete census; commercial scraping APIs are not covered except through their open-source SDKs (Firecrawl’s SDK downloads largely represent clients of its hosted API).
Download the data
- The index (CSV): one row per tool, with every metric on this page.
- GitHub star history (CSV): stars at each target date, with the Wayback capture used.
- PyPI monthly downloads (CSV) and npm monthly downloads (CSV), October 2023 to September 2026.
- Stack Overflow questions per tag and year (CSV) and Hacker News mentions per year (CSV).
Free to use with attribution (CC BY 4.0).
Cite or embed
Suggested citation: ProxyPicker (2026). “The Open-Source Scraping Index: GitHub stars, PyPI and npm downloads, Stack Overflow and Hacker News activity for 20 scraping and browser-automation tools.” Data collected 2 October 2026. https://proxypicker.com/open-source-scraping-index/
Embed the star chart:
<a href="https://proxypicker.com/open-source-scraping-index/"><img src="https://proxypicker.com/wp-content/uploads/2026/10/proxypicker-oss-index-2026-10-github-stars.png" alt="GitHub stars of open-source scraping tools, 2024 to 2026 (ProxyPicker)" width="800" /></a><br />Source: <a href="https://proxypicker.com/open-source-scraping-index/">ProxyPicker, The Open-Source Scraping Index</a> (2 October 2026)Related: AI Crawler Census (which AI bots sites block) · Anti-Bot Census (which bot-protection vendors sites use) · Proxies for web scraping · Best scraping APIs · Scraping Law Hub · All guides.
Press & data requests: email [email protected] for interviews, the raw data behind this page, or corrections. See also our contact page.