scrapy-redis vs colly

Our great sponsors

WorkOS - The modern identity platform for B2B SaaS

InfluxDB - Power Real-Time Data Analytics at Scale

SaaSHub - Software Alternatives and Reviews

Our great sponsors

scrapy-redis		colly
	Project
4	Mentions	39
5,451	Stars	22,165
-	Growth	1.8%
5.0	Activity	6.0
5 months ago	Latest Commit	7 days ago
Python	Language	Go
MIT License	License	Apache License 2.0

The number of mentions indicates the total number of mentions that we've tracked plus the number of user suggested alternatives.
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.

scrapy-redis

Posts with mentions or reviews of scrapy-redis. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2023-06-26.

How to make scrapy run multiple times on the same URLs?
2 projects | /r/scrapy | 26 Jun 2023
Ask HN: What are the best tools for web scraping in 2022?
33 projects | news.ycombinator.com | 10 Aug 2022

11. With some work, you can use Scrapy for distributed projects that are scraping thousands (millions) of domains. We are using https://github.com/rmax/scrapy-redis.
How can I clone a github project to offline machine ?
1 project | /r/learnpython | 2 Apr 2022

git clone https://github.com/darkrho/scrapy-redis.git cd scrapy-redis python setup.py install

colly

Posts with mentions or reviews of colly. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2024-01-01.

Scraping the full snippet from Google search result
3 projects | dev.to | 1 Jan 2024

SerpApi focuses on scraping search results. That's why we need extra help to scrape individual sites. We'll use GoColly package.
Show HN: Flyscrape – A standalone and scriptable web scraper in Go
6 projects | news.ycombinator.com | 11 Nov 2023

Interesting. Can you compare it to colly? [0]
Last time I looked it was the most popular choice for scraping in Go and I have some projects using it.
Is it similar? Does it have more/less features or is it more suited for a different use case? (Which one?)
[0] https://github.com/gocolly/colly
Colly: Elegant Scraper and Crawler Framework for Golang
1 project | news.ycombinator.com | 23 Aug 2023
New modern web crawling tool
2 projects | news.ycombinator.com | 30 Apr 2023

Sounds cool, but how is this different from Colly: https://github.com/gocolly/colly?
colly VS scrapemate - a user suggested alternative
2 projects | 15 Apr 2023
Web Scraping in Python: Avoid Detection Like a Ninja
2 projects | dev.to | 5 Apr 2023

We could write some snippets mixing all these, but the best option in real life is to use a tool with it all, like Scrapy, pyspider, node-crawler (Node.js), or Colly (Go).
Web scraping with Go
5 projects | /r/golang | 2 Apr 2023
Web scraper help
1 project | /r/golang | 1 Mar 2023

Unless you're specifically trying to do it using net/http, I recommend using colly. I've used it in a few scrappers and I love it!
Web Scraping in Golang
2 projects | dev.to | 7 Feb 2023

In this blog, we will be covering the basics of web scraping in Go using the Fiber and Colly frameworks. Colly is an open-source web scraping framework written in Go. It provides a simple and flexible API for performing web scraping tasks, making it a popular choice among Go developers. Colly uses Go's concurrency features to efficiently handle multiple requests and extract data from websites. It offers a wide range of customization options, including the ability to set request headers, handle cookies, follow redirects, and more
Learn how to scrape Trustpilot reviews using Go
4 projects | dev.to | 4 Feb 2023

github.com/gocolly/colly - popular and widely-used library for web scraping in Go. It provides a higher-level API than net/http and makes it easier to extract information from websites. It also provides features such as concurrency, automatic request retries, and support for cookies and sessions.

What are some alternatives?

When comparing scrapy-redis and colly you can also consider the following projects:

polite - Be nice on the web

GoQuery - A little like that j-thing, only in Go.

wi-page - Rank Wikipedia Article's Contributors by Byte Counts.

Scrapy - Scrapy, a fast high-level web crawling & scraping framework for Python.

browserless - Deploy headless browsers in Docker. Run on our cloud or bring your own. Free for non-commercial uses.

xpath - XPath package for Golang, supports HTML, XML, JSON document query.

powerpage-web-crawler - a portable, lightweight web crawler using Powerpage.

rod - A Devtools driver for web automation and scraping

scrapyd - A service daemon to run Scrapy spiders

Geziyor - Geziyor, blazing fast web crawling & scraping framework for Go. Supports JS rendering.

Ferret - Declarative web scraping

chromedp - A faster, simpler way to drive browsers supporting the Chrome DevTools Protocol.

scrapy-redis vs polite colly vs GoQuery scrapy-redis vs wi-page colly vs Scrapy scrapy-redis vs browserless colly vs xpath scrapy-redis vs powerpage-web-crawler colly vs rod scrapy-redis vs scrapyd colly vs Geziyor colly vs Ferret colly vs chromedp

Compare scrapy-redis vs colly and see what are their differences.

scrapy-redis

colly

scrapy-redis

colly

What are some alternatives?