RedditExtractor
polite
Our great sponsors
RedditExtractor | polite | |
---|---|---|
5 | 2 | |
82 | 322 | |
- | - | |
3.3 | 5.3 | |
8 months ago | 8 months ago | |
R | R | |
GNU General Public License v3.0 only | GNU General Public License v3.0 or later |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
RedditExtractor
-
Will RedditExtractoR be impacted by API changes?
IIRC RedditExtractor doesn't use OAuth2 so I think the 10reqs/min ratelimit will be applied to the library/client.
-
bulk subreddit datasets?
Sounds like this package might help you reach your objective. The timeframe you can capture will depend on the amount of activity within the subreddit of interest.
- Has anyone here used the Reddit API before (in R)?
-
Using RedditExtractoR to scrape flairs?
My apologies if the Reddit API flair is inappropriate here - RedditExtractoR does use the Reddit API, but it's technically distinct as a simplified package for R (see: https://github.com/ivan-rivera/RedditExtractor)
-
H3 Podcast YouTube Views Analysis
Great idea, yeah Reddit has an API too, and it looks like there are R & Python packages to access it - https://github.com/ivan-rivera/RedditExtractor
polite
-
Is it legal to scrape data from RedFin using Selenium?
found the github for you: https://github.com/dmi3kno/polite
-
Ask HN: What are the best tools for web scraping in 2022?
The polite package using R is intended to be a friendly way of scraping content from the owner. "The three pillars of a polite session are seeking permission, taking slowly and never asking twice."
What are some alternatives?
Pushshift API - Pushshift API
scrapyd - A service daemon to run Scrapy spiders
police-settlements - A FiveThirtyEight/The Marshall Project effort to collect comprehensive data on police misconduct settlements from 2010-19.
undetected-chromedriver - Custom Selenium Chromedriver | Zero-Config | Passes ALL bot mitigation systems (like Distil / Imperva/ Datadadome / CloudFlare IUAM)
reddit-awards-data - Dataset and visualizations of the most popular Reddit Awards, using the PRAW API.
powerpage-web-crawler - a portable, lightweight web crawler using Powerpage.
Rcrawler - An R web crawler and scraper
scrapy-redis - Redis-based components for Scrapy.
tuber - :sweet_potato: Access YouTube from R
chrome-aws-lambda - Chromium Binary for AWS Lambda and Google Cloud Functions
wi-page - Rank Wikipedia Article's Contributors by Byte Counts.
r-web-scraping-cheat-sheet - Guide, reference and cheatsheet on web scraping using rvest, httr and Rselenium.