|7 days ago||5 days ago|
|MIT License||Apache License 2.0|
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
MDS Newsletter #12
1 project | reddit.com/r/ModernDataStack | 15 Dec 2021
2/ Featured tools this week - Transform and RudderStack
How To Event Stream From Your Gatsby Website Using Open Source RudderStack
3 projects | dev.to | 8 Dec 2021
RudderStack is an open-source Customer Data Pipeline that allows you to track and send real-time events from your web, mobile, and server-side sources to your entire customer data stack. Our primary repository - rudder-server - is open-sourced on GitHub.
Customer Data Pipelines Play a Key Role in Data Privacy
2 projects | dev.to | 1 Dec 2021
This post will explain how your customer data pipeline can help improve your data privacy and how to ensure your data privacy with RudderStack.
Open Source Analytics Stack: Bringing Control, Flexibility, and Data-Privacy to Your Analytics
15 projects | dev.to | 25 Nov 2021
However, limitations to traditional CDPs, especially around connecting to best-of-breed customer tooling and exposing data for use across an organization have driven a new generation of non-CDPs. Solutions like Snowplow's (website, GitHub) data delivery platform and RudderStack's (website, GitHub) customer data platform for developers ingest data from a multitude of sources, apply in-stream transformations, and route data to your data warehouse, like Snowplow, or your warehouse plus your preferred customer tooling destinations for activation, like RudderStack.
RudderStack + Blendo: Better Together
1 project | dev.to | 25 Nov 2021
I learned many lessons from this journey - lessons that deserve a post of their own - but there's one lesson that I learned early on that stands out. In this blog, I talk about why we merged Blendo with RudderStack, building the team and working together to build a great product.
The Open Source Story - Open Sourcing RudderStack Blog and Docs
5 projects | dev.to | 18 Nov 2021
In fact, developers have already started contributing to our documentation. Recently, Benedikt from the Userlist team created the docs for the Userlist destination for RudderStack (see the pull request here). They also built the Userlist integration, submitted a pull request, and it is now live on our platform! This is the beauty of open source!
How to plan and implement a customer data tracking strategy for your Micro-SaaS
1 project | reddit.com/r/ShopifyAppDev | 16 Nov 2021
TLDR: general steps/starting point for setting up an app with Rudderstack ( or Segment) to track customer events
Developing a Custom Plugin using Flutter
5 projects | dev.to | 11 Nov 2021
As a part of our SDK roadmap at RudderStack, we wanted to develop a Flutter SDK. Our existing SDKs include features such as storing event details and persisting user details on the database, and much more. However, these features are already implemented in our Android and iOS SDKs.
Visualize Stripe Payments Data in Postgres using SQL
2 projects | dev.to | 2 Nov 2021
To load Stripe data into Postgres, you can use platforms such as Stitch Data and Rudderstack. In this guide, we will use Stitch Data because it is a cheap and a fast solution.
Dogfooding at RudderStack: Tracking Plans Part 1
1 project | dev.to | 21 Oct 2021
With your Tracking Plans in place, you can use the existing Data Governance API's to evaluate your inbound events, payload samples and metadata to compare them against your plans. You can also use the RudderTyper tool we're releasing alongside Tracking Plans. RudderTyper is a tool for generating strongly-typed RudderStack analytics library wrappers based on your published tracking plan specs, meaning your data will conform to your defined schema upon capture.
AWS Summit London 2022 day recap
6 projects | dev.to | 29 Apr 2022
Amazon Redshift still remains a bit of a mystery to me, even after a whole session on it unpacking loads of its features, possibilities and use cases. Trying to draw some parallels in my head with BigQuery - Google Cloud Platform's own cloud data wharehouse service, which I know well - also didn't help much. So it remains one of those things that now I know a little bit more about than yesterday, but still feels like I haven't even begun to scratch its surface. So one to watch for me, and learn more about. The use case presented by National Rail was great and full of detail too.
A modern data stack for startups
4 projects | dev.to | 21 Apr 2022
With data warehouse solutions (BigQuery, Snowflake, Redshift) going mainstream, modern data stacks are becoming increasingly boring - great news if you're starting from scratch!
How To Start Your Next Data Engineering Project
6 projects | dev.to | 16 Apr 2022
If you wanted to upgrade that idea, track down articles relating to that swing for discussion and post those. There is definite value in that data, and it is a pretty simple thing to do. You are just using a Cloud Composer to ingest the data and storing it in a data warehouse like BigQuery or Snowflake, creating a Twitter bot to post outputs to Twitter using something like Airflow.
How do you handle large spreadsheets?
1 project | reddit.com/r/googlesheets | 7 Apr 2022
look into BigQuery. that's the best way to handle it. Especially if you're familiar with sql. It's super smooth to integrate with sheets using the bigQuery API which makes it great.
Completed my first Data Engineering project with Kafka, Spark, GCP, Airflow, dbt, Terraform, Docker and more!
13 projects | reddit.com/r/dataengineering | 2 Apr 2022
Data Warehouse - BigQuery
dbt for Data Quality Testing & Alerting at FINN
1 project | dev.to | 21 Mar 2022
(Our data stack: Fivetran, BigQuery, dbt, dbt Cloud, Looker, Datafold)
Analytics Stacks for Startups
8 projects | dev.to | 21 Feb 2022
The main DWH offerings that meet the above expectations are Snowflake, Google Bigquery, and Amazon Redshift. Featurewise, these three have similar functionalities but there are differences, e.g. how long it takes to spin up new compute resources or how much maintenance work they need. Costwise, it seems they end up with similar numbers on your bill, depending on which blog post you read.
Pandas and SQL side by side
3 projects | dev.to | 12 Jan 2022
Migrating to Snowflake, Redshift, or BigQuery? Use Datafold to Avoid these Common Pitfalls
2 projects | dev.to | 15 Dec 2021
Although we’ve focused on Snowflake in this article, the same features of Datafold can be used for other cloud products like Redshift or BigQuery.
spent 1 hour trying to work out why DBT on macOS could not find schema.yml! WTF macOS!
2 projects | reddit.com/r/programminghorror | 17 Nov 2021
Working on a new project using DBT, Google BigQuery & MongoDB
What are some alternatives?
airbyte - Airbyte is an open-source EL(T) platform that helps you replicate your data in your warehouses, lakes and databases.
dbt-spark - dbt-spark contains all of the code enabling dbt to work with Apache Spark and Databricks
dbt - dbt enables data analysts and engineers to transform their data using the same practices that software engineers use to build applications. [Moved to: https://github.com/dbt-labs/dbt-core]
pub-dev - Pub.dev, the Dart package repository, written in Dart
dbt-core - dbt enables data analysts and engineers to transform their data using the same practices that software engineers use to build applications.
WordPress - WordPress, Git-ified. This repository is just a mirror of the WordPress subversion repository. Please do not send pull requests. Submit pull requests to https://github.com/WordPress/wordpress-develop and patches to https://core.trac.wordpress.org/ instead.
cube.js - 📊 Cube — Headless Business Intelligence for Building Data Applications
PostgreSQL - Mirror of the official PostgreSQL GIT repository. Note that this is just a *mirror* - we don't work with pull requests on github. To contribute, please see https://wiki.postgresql.org/wiki/Submitting_a_Patch
BigQuery-Python - Simple Python client for interacting with Google BigQuery.
MongoDB - The MongoDB Database
Prophet - Tool for producing high quality forecasts for time series data that has multiple seasonality with linear or non-linear growth.