data-engineering-wiki
The best place to learn data engineering. Built and maintained by the data engineering community. (by data-engineering-community)
versatile-data-kit
One framework to develop, deploy and operate data workflows with Python and SQL. (by vmware)
data-engineering-wiki | versatile-data-kit | |
---|---|---|
15 | 52 | |
1,177 | 414 | |
12.4% | 0.7% | |
7.1 | 9.7 | |
21 days ago | about 14 hours ago | |
CSS | Python | |
Creative Commons Zero v1.0 Universal | Apache License 2.0 |
The number of mentions indicates the total number of mentions that we've tracked plus the number of user suggested alternatives.
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
data-engineering-wiki
Posts with mentions or reviews of data-engineering-wiki.
We have used some of these posts to build our list of alternatives
and similar projects. The last one was on 2023-07-17.
- Data Engineering Glossary
-
ETL practice
My suggestions: 1. Browse https://dataengineering.wiki/ and overall go over r/dataengineering 2. In mid-sized companies, the trend is to outsource Extract and Load to providers like Fivetran or Airbyte (open-source). Then Transform it with dbt in a data warehouse with SQL. 3. In big companies, you won't touch much ETL design. Just need to be proficient in Python / Spark / SQL... 4. Make sure you know what a star schema, fact tables, and dimension tables are.
- Anything else to read
-
Looking for blogs for backend development
Hi everyone! As mentioned in title I recently came across great blogs for data engineering: startdataengineering.com and dataengineering.wiki
-
DE- How to get my foot in the door?
The data engineering subreddit maintains a wiki of advice, resources, and recommendations at https://dataengineering.wiki/. Your question is answered in their FAQ here
- Getting into Data Engineering and more!
-
Are there avenues into sports science as a software engineer or web dev?
Data engineering
-
Switching to something more technical
r/dataengineering has a wiki at https://dataengineering.wiki and also a Discord server which is pretty active.
-
Data Engineering Concepts: Definitions, Backlinks, and Graph View
Almost the same as the wiki https://dataengineering.wiki/
-
dataengineering.wiki Bug
Hi, would you mind opening an issue on GitHub? We can help you debug the issue there.
versatile-data-kit
Posts with mentions or reviews of versatile-data-kit.
We have used some of these posts to build our list of alternatives
and similar projects. The last one was on 2022-11-23.
-
Looking for a data blogger
Here's the project: https://github.com/vmware/versatile-data-kit
-
Need advice on ETL tool
I don't really know if this would work for you because the UI is not functional yet, but a very simple REST API ingestion example here, there's one for csv too https://github.com/vmware/versatile-data-kit/wiki/Ingesting-data-from-REST-API-into-Database I can't imagine a simpler way unless it's really drag and drop.
-
If dbt is the "T" part of an "ELT", what do you use for "EL"?
I work at VMware and we use one tool for the whole ELT, it was made internally as there was no good alternative at the time and now we opensourced it, here it is: https://github.com/vmware/versatile-data-kit
-
Best way to fix errors in my data?
With my team we created csv ingestion plugin described here, maybe you want to try it out: https://github.com/vmware/versatile-data-kit/wiki/Ingesting-local-CSV-file-into-Database
-
What Orchestration Tool do you use for batch ETL/ELT?
We use Versatile Data Kit for batch data job orchestration (https://github.com/vmware/versatile-data-kit)
-
Dear, pipeline builders! Which step in your role is the most time consuming?
"suggestions on how to reduce the time spent on initially generating and adjusting the code" is using some tools that automate ELT. Here's one open-source tool I'm working on with my team: https://github.com/vmware/versatile-data-kit
-
Problem definition / vibe check for a repo
here's the repo: https://github.com/vmware/versatile-data-kit
-
Can we take a moment to appreciate how much of dataengineering is open source?
If you wish to contribute, projects usually have good first issues: https://github.com/vmware/versatile-data-kit/labels/good%20first%20issue If you wish to learn, check out examples: https://github.com/vmware/versatile-data-kit/tree/main/examples
-
ETL question (noob)
Have you heard about versatile data kit (https://github.com/vmware/versatile-data-kit)? I think it meets your needs perfectly:
-
DE Open Source
Versatile Data Kit is a framework to bBuild, run and manage your data pipelines with Python or SQL on any cloud https://github.com/vmware/versatile-data-kit here's a list of good first issues: https://github.com/vmware/versatile-data-kit/issues?q=is%3Aissue+is%3Aopen+label%3A%22good+first+issue%22 Join our slack channel to connect with our team: https://cloud-native.slack.com/archives/C033PSLKCPR
What are some alternatives?
When comparing data-engineering-wiki and versatile-data-kit you can also consider the following projects:
glossary - Data Glossary 🧠: An interactive digital garden for deeper data exploration. Learn through a graph and backlinks, enabling layered knowledge discovery.
data-engineering-zoomcamp - Free Data Engineering course!