mlscraper
sneakpeek
mlscraper | sneakpeek | |
---|---|---|
10 | 3 | |
1,229 | 36 | |
- | - | |
0.6 | 7.5 | |
about 2 months ago | 9 months ago | |
Python | Python | |
- | BSD 3-clause "New" or "Revised" License |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
mlscraper
-
What are the best tools for web scraping and analysis of natural language to populate a dataset?
See if something like autoscraper or mlscraper suits your needs.
-
Experimental library for scraping websites using OpenAI's GPT API
Why GPT-based then? There are libraries that do this: You give examples, they generate the rules for you and give you a scraper object that takes any html and returns the scraped data.
Mine: https://github.com/lorey/mlscraper
-
Could someone recommend me a library for c# like one of these two (they are for python) : mlscraper and autoscraper
GitHub - lorey/mlscraper: 🤖 Scrape data from HTML websites automatically by just providing examples
-
Smart Scraper
Check it out here: https://github.com/lorey/mlscraper Example: https://github.com/lorey/mlscraper/blob/master/examples/quotes\_to\_scrape.py
- Pre-trained Webscraping Models
- 🤖 Scrape data from HTML websites automatically by just providing examples
- mlscraper: Scrape data from HTML pages automatically with Machine Learning
-
Show HN: RSS feeds for arbitrary websites using CSS selectors
In case anyone wants to detect the selectors automatically, here's a small python library I wrote that does it for you: https://github.com/lorey/mlscraper
sneakpeek
- Sneakpeek is a framework that helps to quickly and conveniently develop scrapers. It’s the best choice for scrapers that have some specific complex scraping logic that needs to be run on a constant basis
- Sneakpeek is a framework that helps to quickly and conviniently develop scrapers. It’s the best choice for scrapers that have some specific complex scraping logic that needs to be run on a constant basis (useful for ChatGPT plugins)
- flulemon/sneakpeek: Sneakpeek is a framework that helps to quickly and conviniently develop scrapers. It’s the best choice for scrapers that have some specific complex scraping logic that needs to be run on a constant basis
What are some alternatives?
scrapingant-client-python - ScrapingAnt API client for Python.
meta-spy - 👾 CLI MetaSpy (Facebook, Instagram) scraper and crawler - instagram account, facebook accounts, pages and search
ttrss_plugin-feediron - Evolution of ttrss_plugin-af_feedmod
custom-crawler - 🌌 High productivity semi-automatic crawler generator 🛠️🧰
furss - Fix Up RSS (and atom): Make full-text versions of rss/atom feeds
ZvgPortalScraper - Python-based scraper for zvg-portal.de
feed-me-up-scotty
Grab - Web Scraping Framework
rssify - Tool that generates an rss feed out of websites that don't have one
outlook-account-generator - Outlook Account Generator helps you create outlook accounts.
RSSHub - 🧡 Everything is RSSible
feedgen - Generates RSS/ATOM/JSON feeds. Can be reasonably extended or create a feed using the CSS generator.