memphis
versatile-data-kit
memphis | versatile-data-kit | |
---|---|---|
52 | 52 | |
3,158 | 412 | |
0.9% | 1.2% | |
9.9 | 9.7 | |
10 days ago | 1 day ago | |
Go | Python | |
GNU General Public License v3.0 or later | Apache License 2.0 |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
memphis
- Memphis
-
What type of open source contributions can I make that improve my core data engineering skills? Are there any projects that require help of that nature? What are they?
Hey check out the first good issues in the Memphis.dev open source
-
I want to create beginners Data Pipeline with SQL, Python etc. Any expert suggestions on like (Tools, Processes, Sources).
Try Memphis.dev blog or you can check out Github
-
What's an ideal project structure for a Golang web service?
- https://github.com/memphisdev/memphis
-
Connect Memphis as an Argo event source
Argo is a collection of open-source tools for Kubernetes to run workflows, manage clusters, and do GitOps easily. Memphis is an open-source next-generation alternative to traditional message brokers.
-
Creating a brand new data infrastructure for a small company
I have dealt with the same problem. It depends what are the cycle intervals. If for example, it's every few minutes maybe it's worth keeping a machine on the cloud and the DB on that machine. in my opinion, it's a little bit expensive. The thing that worked for me is to run once a day a lambda function and store the data on a message broker. You should take a look at memphis.dev which is open-source and very easy to work with.
- Iām looking for a suggestion for a queuing library
-
Memphis: Low-code real-time data processing platform
Image Source
-
Memphis.dev v0.4.2
Join our 2K stargazers on Github and try Memphis out, I am sure you are going to be surprised š· https://github.com/memphisdev/memphis-broker
- Memphis.dev v0.2.4 is out!
versatile-data-kit
-
Looking for a data blogger
Here's the project: https://github.com/vmware/versatile-data-kit
-
Need advice on ETL tool
I don't really know if this would work for you because the UI is not functional yet, but a very simple REST API ingestion example here, there's one for csv too https://github.com/vmware/versatile-data-kit/wiki/Ingesting-data-from-REST-API-into-Database I can't imagine a simpler way unless it's really drag and drop.
-
If dbt is the "T" part of an "ELT", what do you use for "EL"?
I work at VMware and we use one tool for the whole ELT, it was made internally as there was no good alternative at the time and now we opensourced it, here it is: https://github.com/vmware/versatile-data-kit
-
Best way to fix errors in my data?
With my team we created csv ingestion plugin described here, maybe you want to try it out: https://github.com/vmware/versatile-data-kit/wiki/Ingesting-local-CSV-file-into-Database
-
What Orchestration Tool do you use for batch ETL/ELT?
We use Versatile Data Kit for batch data job orchestration (https://github.com/vmware/versatile-data-kit)
-
Dear, pipeline builders! Which step in your role is the most time consuming?
"suggestions on how to reduce the time spent on initially generating and adjusting the code" is using some tools that automate ELT. Here's one open-source tool I'm working on with my team: https://github.com/vmware/versatile-data-kit
-
Problem definition / vibe check for a repo
here's the repo: https://github.com/vmware/versatile-data-kit
-
Can we take a moment to appreciate how much of dataengineering is open source?
If you wish to contribute, projects usually have good first issues: https://github.com/vmware/versatile-data-kit/labels/good%20first%20issue If you wish to learn, check out examples: https://github.com/vmware/versatile-data-kit/tree/main/examples
-
ETL question (noob)
Have you heard about versatile data kit (https://github.com/vmware/versatile-data-kit)? I think it meets your needs perfectly:
-
DE Open Source
Versatile Data Kit is a framework to bBuild, run and manage your data pipelines with Python or SQL on any cloud https://github.com/vmware/versatile-data-kit here's a list of good first issues: https://github.com/vmware/versatile-data-kit/issues?q=is%3Aissue+is%3Aopen+label%3A%22good+first+issue%22 Join our slack channel to connect with our team: https://cloud-native.slack.com/archives/C033PSLKCPR
What are some alternatives?
gRPC - The C based gRPC (C++, Python, Ruby, Objective-C, PHP, C#)
data-engineering-zoomcamp - Free Data Engineering course!
ApacheKafka - A curated re-sources list for awesome Apache Kafka
Mage - š§ The modern replacement for Airflow. Mage is an open-source data pipeline tool for transforming and integrating data. https://github.com/mage-ai/mage-ai
Nodejs-Developer-Roadmap - A Developer Roadmap to becoming a Node.js developer in 2019
quadratic - Quadratic | Data Science Spreadsheet with Python & SQL
kubernetes - Production-Grade Container Scheduling and Management
pyramid-jsonapi - Auto-build JSON API from sqlalchemy models using the pyramid framework
compression - Node.js compression middleware
dbt-data-reliability - dbt package that is part of Elementary, the dbt-native data observability solution for data & analytics engineers. Monitor your data pipelines in minutes. Available as self-hosted or cloud service with premium features.
v8.dev - The source code of v8.dev, the official website of the V8 project.
hamilton - A scalable general purpose micro-framework for defining dataflows. THIS REPOSITORY HAS BEEN MOVED TO www.github.com/dagworks-inc/hamilton