versatile-data-kit
Benthos
Our great sponsors
versatile-data-kit | Benthos | |
---|---|---|
52 | 76 | |
410 | 7,559 | |
2.4% | 4.5% | |
9.7 | 9.6 | |
2 days ago | 7 days ago | |
Python | Go | |
Apache License 2.0 | MIT License |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
versatile-data-kit
-
Looking for a data blogger
Here's the project: https://github.com/vmware/versatile-data-kit
-
Need advice on ETL tool
I don't really know if this would work for you because the UI is not functional yet, but a very simple REST API ingestion example here, there's one for csv too https://github.com/vmware/versatile-data-kit/wiki/Ingesting-data-from-REST-API-into-Database I can't imagine a simpler way unless it's really drag and drop.
-
If dbt is the "T" part of an "ELT", what do you use for "EL"?
I work at VMware and we use one tool for the whole ELT, it was made internally as there was no good alternative at the time and now we opensourced it, here it is: https://github.com/vmware/versatile-data-kit
-
Best way to fix errors in my data?
With my team we created csv ingestion plugin described here, maybe you want to try it out: https://github.com/vmware/versatile-data-kit/wiki/Ingesting-local-CSV-file-into-Database
-
What Orchestration Tool do you use for batch ETL/ELT?
We use Versatile Data Kit for batch data job orchestration (https://github.com/vmware/versatile-data-kit)
-
Dear, pipeline builders! Which step in your role is the most time consuming?
"suggestions on how to reduce the time spent on initially generating and adjusting the code" is using some tools that automate ELT. Here's one open-source tool I'm working on with my team: https://github.com/vmware/versatile-data-kit
-
Problem definition / vibe check for a repo
here's the repo: https://github.com/vmware/versatile-data-kit
-
Can we take a moment to appreciate how much of dataengineering is open source?
If you wish to contribute, projects usually have good first issues: https://github.com/vmware/versatile-data-kit/labels/good%20first%20issue If you wish to learn, check out examples: https://github.com/vmware/versatile-data-kit/tree/main/examples
-
ETL question (noob)
Have you heard about versatile data kit (https://github.com/vmware/versatile-data-kit)? I think it meets your needs perfectly:
-
DE Open Source
Versatile Data Kit is a framework to bBuild, run and manage your data pipelines with Python or SQL on any cloud https://github.com/vmware/versatile-data-kit here's a list of good first issues: https://github.com/vmware/versatile-data-kit/issues?q=is%3Aissue+is%3Aopen+label%3A%22good+first+issue%22 Join our slack channel to connect with our team: https://cloud-native.slack.com/archives/C033PSLKCPR
Benthos
- Ask HN: Who is hiring? (December 2023)
- Structured Logging with Slog
- Fancy stream processing made operationally mundane
- Benthos: Fancy stream processing made operationally mundane
-
Any golang library to batch process a queue ?
I’ve used https://www.benthos.dev/ and it’s really easy and well implemented. The author is also very responsive
-
Show HN: Arroyo – Write SQL on streaming data
Looks cool. What is the difference between this tools and benthos (https://www.benthos.dev/)?
- Benthos: Open-source stream processing tool
-
book about golang and kafka
You might want to gradually replace that one with https://github.com/twmb/franz-go because Shopify is looking to find a new owner for Sarama and, until or if they do, it seems to be falling behind with maintenance: https://github.com/Shopify/sarama/issues/2461 For example, they still haven’t addressed this breaking change https://github.com/Shopify/sarama/issues/2358. franz-go has worked well so far in Benthos https://github.com/benthosdev/benthos/tree/main/internal/impl/kafka and it will likely end up as the only implementation once the Sarama-based one will be deprecated
- Show HN: Open-source Auth0 alternative Ory Kratos v0.13 released – nearing v1.0
-
Go in depth youtube channels?
I upload a mix of code reviews and live streams on https://www.youtube.com/@Jeffail, mostly building https://www.benthos.dev out in the open so the content ranges from beginner friendly stuff to more advanced things like stream processing, parser combinators, etc.
What are some alternatives?
data-engineering-zoomcamp - Free Data Engineering course!
Confluent Kafka Golang Client - Confluent's Apache Kafka Golang client
Mage - 🧙 The modern replacement for Airflow. Mage is an open-source data pipeline tool for transforming and integrating data. https://github.com/mage-ai/mage-ai
appsmith - Platform to build admin panels, internal tools, and dashboards. Integrates with 25+ databases and any API.
quadratic - Quadratic | Data Science Spreadsheet with Python & SQL
watermill - Building event-driven applications the easy way in Go.
pyramid-jsonapi - Auto-build JSON API from sqlalchemy models using the pyramid framework
sarama - Sarama is a Go library for Apache Kafka. [Moved to: https://github.com/IBM/sarama]
dbt-data-reliability - dbt package that is part of Elementary, the dbt-native data observability solution for data & analytics engineers. Monitor your data pipelines in minutes. Available as self-hosted or cloud service with premium features.
salsa - A generic framework for on-demand, incrementalized computation. Inspired by adapton, glimmer, and rustc's query system.
hamilton - A scalable general purpose micro-framework for defining dataflows. THIS REPOSITORY HAS BEEN MOVED TO www.github.com/dagworks-inc/hamilton
azure-event-hubs-go - Golang client library for Azure Event Hubs https://azure.microsoft.com/services/event-hubs