Python Webscraping

Open-source Python projects categorized as Webscraping

Top 23 Python Webscraping Projects

Webscraping
  1. Scrapling

    🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!

    Project mention: Scrapling MCP: Giving Claude Code a Real Web Scraper | dev.to | 2026-07-07

    Claude Code can read my files and run my shell, but out of the box it can't actually go get a page from the live web in a way that survives modern anti-bot defenses. A curl from the Bash tool gets you a 403 from anything behind Cloudflare. So I took an existing open-source scraper — D4Vinci/Scrapling — installed it locally, and registered its built-in MCP server with Claude Code. Now the agent has ten tools for pulling the real web: plain fetches, headless-browser fetches, stealth fetches that solve Cloudflare, and screenshots. The scraper isn't mine — I want to be clear about that. What I built is the local install and the MCP integration that hands those capabilities to the agent.

  2. AppSignal

    Monitoring that respects your time & budget. APM, error tracking, and dashboards for modern web apps. Ten-minute setup, transparent flat pricing, and support from engineers who actually use the product.

    AppSignal logo
  3. Scrapegraph-ai

    Python scraper based on AI

  4. gpt-researcher

    An autonomous agent that conducts deep research on any data using any LLM providers

    Project mention: 1500 Lines of Markdown vs 15000 Lines of Python. | dev.to | 2025-12-31

    I spent a weekend studying GPT-Researcher, an open-source project with 24,000+ GitHub stars. It builds an autonomous research agent that generates comprehensive reports with citations. The architecture is elegant: multiple specialized agents coordinate through LangGraph, parallel execution speeds up research, and quality gates ensure reliable output.

  5. autoscraper

    A Smart, Automatic, Fast and Lightweight Web Scraper for Python

  6. CyberScraper-2077

    A Powerful web scraper powered by LLM | OpenAI, Gemini & Ollama

  7. CloakBrowser

    Stealth Chromium that passes every bot detection test. Drop-in Playwright replacement with source-level fingerprint patches. 30/30 tests passed.

    Project mention: A Green Run Is Not a Green Light: Engineering Responsibility in Browser Automation | dev.to | 2026-09-05

    CloakBrowser belongs primarily in the runtime layer. It provides a source-modified Chromium environment for Playwright, Puppeteer, Selenium, and CDP workflows where conventional headless setups or JavaScript-level patches can become fragile. A more consistent browser gives teams a cleaner foundation for QA, monitoring, browser agents, brand verification, and public-web research.

  8. CrossLinked

    LinkedIn enumeration tool to extract valid employee names from an organization through search engine scraping

  9. Kargo

    Stop Scripting Promotions. Start Shipping with Kargo. Kargo automates promotion across dev, staging, and prod with approval gates and verification. Open source, built by the team behind Argo CD. Download now.

    Kargo logo
  10. requests-cache

    Transparent persistent cache for python requests

  11. scrapeghost

    👻 Experimental library for scraping websites using OpenAI's GPT API.

  12. parsera

    Lightweight library for scraping web-sites with LLMs

  13. mov-cli

    Watch everything from your terminal.

  14. scrapfly-scrapers

    Scalable Python web scraping scripts for +40 popular domains

    Project mention: Precursor | news.ycombinator.com | 2026-07-13

    happy to offer a counter of some great products for anti-bot defeat:

    https://brightdata.com/

    https://www.zenrows.com/

    https://www.capsolver.com/

    https://scrapfly.io/

    hundreds of millions of residential ips, human browser fingerprints, custom browser binaries, auto solve of turnstyle, recaptcha v3, kasada, datadome, AWS WAF, etc if they come up.

  15. zimit

    Make a ZIM file from any Web site and surf offline!

  16. EasyApplyJobsBot

    A python bot to automatically apply all Linkedin,Glassdoor, etc Easy Apply jobs based on your preferences. Auto login, auto fill additional questions, apply automatically!

  17. gazpacho

    The simple, fast, and modern web scraping library

  18. AutoScraper

    Official implement of paper "AutoScraper: A Progressive Understanding Web Agent for Web Scraper Generation" [EMNLP 24'] (by EZ-hwh)

  19. fortress

    Stealth Chromium engine that stops scrapers and browser agents from getting blocked, with one line of code change. (by tiliondev)

    Project mention: One Wikipedia page costs your AI agent 68,000 tokens | news.ycombinator.com | 2026-07-10

    repo and the reproducible benchmark: https://github.com/tiliondev/fortress/tree/main/mcp

  20. TikTokBot

    A TikTokBot that downloads trending tiktok videos and compiles them using FFmpeg

  21. ChatGPT-OpenAI-Smart-Speaker

    This AI Smart Speaker uses speech recognition, TTS (text-to-speech), and STT (speech-to-text) to enable voice and vision-driven conversations, with additional web search capabilities via OpenAI and Langchain agents.

  22. llm_osint

    LLM OSINT is a proof-of-concept method of using LLMs to gather information from the internet and then perform a task with this information.

  23. ebayScraper

    Scrape all eBay sold listings to determine average/median pricing, plot listings over time with trend lines, and extract to excel

  24. yars

    Yet Another Reddit Scrapper (without API keys) | Scrap search results, posts and images from subreddits filtered by hot, new etc and bulk download any user's data.

  25. par_scrape

    AI assisted web scraping and data extraction

  26. SaaSHub

    SaaSHub - Software Alternatives and Reviews. SaaSHub helps you find the best software and product alternatives

    SaaSHub logo
NOTE: The open source projects on this list are ordered by number of github stars. The number of mentions indicates repo mentiontions in the last 12 Months or since we started tracking (Dec 2020).

Python Webscraping discussion

Log in or Post with

Python Webscraping related posts

Index

What are some of the best open-source Webscraping projects in Python? This list will help you:

# Project Stars
1 Scrapling 78,237
2 Scrapegraph-ai 30,546
3 gpt-researcher 29,280
4 autoscraper 7,936
5 CyberScraper-2077 3,252
6 CloakBrowser 2,249
7 CrossLinked 1,580
8 requests-cache 1,501
9 scrapeghost 1,442
10 parsera 1,354
11 mov-cli 1,298
12 scrapfly-scrapers 1,071
13 zimit 840
14 EasyApplyJobsBot 816
15 gazpacho 769
16 AutoScraper 488
17 fortress 480
18 TikTokBot 417
19 ChatGPT-OpenAI-Smart-Speaker 318
20 llm_osint 316
21 ebayScraper 267
22 yars 220
23 par_scrape 218

Sponsored
Monitoring that respects your time & budget
APM, error tracking, and dashboards for modern web apps. Ten-minute setup, transparent flat pricing, and support from engineers who actually use the product.
www.appsignal.com