analytics
jitsu
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
analytics
-
I'm not getting it...what's the point of DBT?
Take a look at gitlab's dbt project: https://gitlab.com/gitlab-data/analytics/-/blob/master/transform/snowflake-dbt/models/common/schema.yml
-
How would you structure a repo with 10+ ETL pipelines and shared code?
A good reference is the Gitlab data team repo. https://gitlab.com/gitlab-data/analytics
- What are your favourite GitHub repos that shows how data engineering should be done?
-
Are there any open corporate Data Team repositories / projects besides GitLab?
For example, their Data Team have a public repository, with a bunch of information on how they organize DAGs, machine learning projects, system configuration, etc.
- Kimball Dim Modelling Code Examples
- Can someone help me, an absolute newbie, understand the usage and benefit of dbt with practical example ?
-
Is jinja templating right for DBT?
So I've run through the DBT tutorial stuff and looked over some fairly complex uses of it i.e. GitLab Data and I was wondering if anyone has any opinions or insights into the use of jinja templating in the sql?
-
Where can I find free data engineering ( big data) projects online?
Gitlab has their DBT repo open source and is very useful for seeing how to structure a project at scale. https://gitlab.com/gitlab-data/analytics/-/tree/master/transform/snowflake-dbt
-
Gitlab's Data Team Platform (in depth look at their stack)
Currently the team is working hard on this: https://gitlab.com/gitlab-data/analytics/-/issues/9508
-
Can someone explain the big deal with dbt?
GitLab's dbt project is an excellent example of a mature project at scale. They also have a comprehensive guide to their methodology.
jitsu
- Jitsu
- Any examples of working activist, socialist, or community-organizing software?
-
Lesser Known Features of ClickHouse
you may check: https://github.com/jitsucom/jitsu. "Jitsu is an open-source Segment alternative. Fully-scriptable data ingestion engine for modern data teams. Set-up a real-time data pipeline in minutes, not days"
You can create an API endpoint, and send those JSON to it. In the "destination" part, it can sync to clickhouse (one of many choices, like redshift, snowflake,besides clickhouse) very quickly, and flatten the JSON into columns. If there is new key found in JSON, it will create a new column in clickhouse.
-
Reference Data Stack for Data-Driven Startups
We also have telemetry set up on our Monosi product which is collected through Snowplow,. As with Airbyte, we chose Snowplow because of its open source offering and because of their scalable event ingestion framework. There are other open source options to consider including Jitsu and RudderStack or closed source options like Segment. Since we started building our product with just a CLI offering, we didn’t need a full CDP solution so we chose Snowplow.
-
Data pipeline suggestions
Ingestion / Extraction: Airbyte, Singer, Jitsu
-
Where can I find free data engineering ( big data) projects online?
Ingestion / ETL: Airbyte, Singer, Jitsu Transformation: dbt Orchestration: Airflow, Dagster Testing: GreatExpectations Observability: Monosi Reverse ETL: Grouparoo, Castled Visualization: Lightdash, Superset
- Ask HN: Good open source alternatives to Google Analytics?
- Jitsu is a FOSS data integration platform that gathers events from several data sources (alternative to Segment)
-
Launch HN: Jitsu (YC S20) – Open-Source Segment Alternative
I’m just saying this is better:
We are building Jitsu, (https://github.com/jitsucom/jitsu, https://jitsu.com/) We help companies collect events from their apps, websites, and APIs and send them to databases.
Think of us as an open-source Segment alternative.
What are some alternatives?
dbt-synapse - dbt adapter for Azure Synapse Dedicated SQL Pools
airbyte - The leading data integration platform for ETL / ELT data pipelines from APIs, databases & files to data warehouses, data lakes & data lakehouses. Both self-hosted and Cloud-hosted.
dagster - An orchestration platform for the development, production, and observation of data assets.
Snowplow - The enterprise-grade behavioral data engine (web, mobile, server-side, webhooks), running cloud-natively on AWS and GCP
castled - Castled is an open source reverse ETL solution that helps you to periodically sync the data in your db/warehouse into sales, marketing, support or custom apps without any help from engineering teams
Airflow - Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
datahub - The Metadata Platform for your Data Stack
posthog-ios - PostHog iOS SDK
AdvancedSQLPuzzles - Welcome to my GitHub repository. I hope you enjoy solving these puzzles as much as I have enjoyed creating them.
sqlpad - Web-based SQL editor. Legacy project in maintenance mode.
lightdash - Self-serve BI to 10x your data team ⚡️
superset - Apache Superset is a Data Visualization and Data Exploration Platform