wi-page
polite
wi-page | polite | |
---|---|---|
2 | 2 | |
1 | 322 | |
- | - | |
0.0 | 5.3 | |
about 3 years ago | 8 months ago | |
Python | R | |
- | GNU General Public License v3.0 or later |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
wi-page
-
Ask HN: What are the best tools for web scraping in 2022?
[4] https://github.com/altilunium/wi-page (Scrap wikipedia to get most active contributors that contribute to a certain article)
- Show HN: Wi-Page – Rank Wikipedia Article's Contributors by Byte Counts
polite
-
Is it legal to scrape data from RedFin using Selenium?
found the github for you: https://github.com/dmi3kno/polite
-
Ask HN: What are the best tools for web scraping in 2022?
The polite package using R is intended to be a friendly way of scraping content from the owner. "The three pillars of a polite session are seeking permission, taking slowly and never asking twice."
https://github.com/dmi3kno/polite
What are some alternatives?
kiwix-hotspot - Kiwix Hotspot Image Creator (Desktop) for Windows/macOS/Linux
scrapyd - A service daemon to run Scrapy spiders
estela - estela, an elastic web scraping cluster 🕸
undetected-chromedriver - Custom Selenium Chromedriver | Zero-Config | Passes ALL bot mitigation systems (like Distil / Imperva/ Datadadome / CloudFlare IUAM)
danker - Compute PageRank on >3 billion Wikipedia links on off-the-shelf hardware.
scrapy-redis - Redis-based components for Scrapy.
powerpage-web-crawler - a portable, lightweight web crawler using Powerpage.
curl-impersonate - curl-impersonate: A special build of curl that can impersonate Chrome & Firefox
chrome-aws-lambda - Chromium Binary for AWS Lambda and Google Cloud Functions
r-web-scraping-cheat-sheet - Guide, reference and cheatsheet on web scraping using rvest, httr and Rselenium.