versatile-data-kit
quadratic
versatile-data-kit | quadratic | |
---|---|---|
52 | 9 | |
411 | 2,725 | |
1.2% | 2.1% | |
9.7 | 10.0 | |
3 days ago | 7 days ago | |
Python | TypeScript | |
Apache License 2.0 | GNU General Public License v3.0 or later |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
versatile-data-kit
-
Looking for a data blogger
Here's the project: https://github.com/vmware/versatile-data-kit
-
Need advice on ETL tool
I don't really know if this would work for you because the UI is not functional yet, but a very simple REST API ingestion example here, there's one for csv too https://github.com/vmware/versatile-data-kit/wiki/Ingesting-data-from-REST-API-into-Database I can't imagine a simpler way unless it's really drag and drop.
-
If dbt is the "T" part of an "ELT", what do you use for "EL"?
I work at VMware and we use one tool for the whole ELT, it was made internally as there was no good alternative at the time and now we opensourced it, here it is: https://github.com/vmware/versatile-data-kit
-
Best way to fix errors in my data?
With my team we created csv ingestion plugin described here, maybe you want to try it out: https://github.com/vmware/versatile-data-kit/wiki/Ingesting-local-CSV-file-into-Database
-
What Orchestration Tool do you use for batch ETL/ELT?
We use Versatile Data Kit for batch data job orchestration (https://github.com/vmware/versatile-data-kit)
-
Dear, pipeline builders! Which step in your role is the most time consuming?
"suggestions on how to reduce the time spent on initially generating and adjusting the code" is using some tools that automate ELT. Here's one open-source tool I'm working on with my team: https://github.com/vmware/versatile-data-kit
-
Problem definition / vibe check for a repo
here's the repo: https://github.com/vmware/versatile-data-kit
-
Can we take a moment to appreciate how much of dataengineering is open source?
If you wish to contribute, projects usually have good first issues: https://github.com/vmware/versatile-data-kit/labels/good%20first%20issue If you wish to learn, check out examples: https://github.com/vmware/versatile-data-kit/tree/main/examples
-
ETL question (noob)
Have you heard about versatile data kit (https://github.com/vmware/versatile-data-kit)? I think it meets your needs perfectly:
-
DE Open Source
Versatile Data Kit is a framework to bBuild, run and manage your data pipelines with Python or SQL on any cloud https://github.com/vmware/versatile-data-kit here's a list of good first issues: https://github.com/vmware/versatile-data-kit/issues?q=is%3Aissue+is%3Aopen+label%3A%22good+first+issue%22 Join our slack channel to connect with our team: https://cloud-native.slack.com/archives/C033PSLKCPR
quadratic
-
Quadratic – Open-Source Spreadsheet Is Now Multiplayer
https://github.com/quadratichq/quadratic/issues
-
suggestions for a free spreadsheet library (like excel or google spreadsheets)
Hey guys, as the title says I am looking for some suggestions on useful free spreadsheet libraries, similar to excel or google sheets. I saw quadratic (https://github.com/quadratichq/quadratic), which looks really interesting but is in very early alpha, to the point I cant even embed it in my own project yet.
-
The coordinate system for an infinite spreadsheet
I am glad they started brainstorming how to implement my suggestions for named cells and named ranges [1]. Having complex spreadsheets than span through infinite will be unmaintainable relying only on coordinates
[1] https://github.com/quadratichq/quadratic/issues/408
-
Show HN: Quadratic – Open-Source Spreadsheet with Python, & AI (WASM and WebGL)
Yes the bundle is huge, we have made no effort yet to optimize it. Feel free to create a PR :)
Here is how we manage cell dependencies https://github.com/quadratichq/quadratic/blob/main/src/grid/...
- TIL: The autocorrect feature in Excel, which converts certain combinations into dates, has mangled up to 30% of published papers, causing significant issues. As a result, at least 27 gene symbols have been forced to change to prevent further errors from occurring.
- Quadratic: Open-Source Data Science Spreadsheet with Python, JavaScript and SQL
What are some alternatives?
data-engineering-zoomcamp - Free Data Engineering course!
astro-sdk - Astro SDK allows rapid and clean development of {Extract, Load, Transform} workflows using Python and SQL, powered by Apache Airflow.
Mage - 🧙 The modern replacement for Airflow. Mage is an open-source data pipeline tool for transforming and integrating data. https://github.com/mage-ai/mage-ai
perfcalc
pyramid-jsonapi - Auto-build JSON API from sqlalchemy models using the pyramid framework
superset - Apache Superset is a Data Visualization and Data Exploration Platform
dbt-data-reliability - dbt package that is part of Elementary, the dbt-native data observability solution for data & analytics engineers. Monitor your data pipelines in minutes. Available as self-hosted or cloud service with premium features.
mito - The mitosheet package, trymito.io, and other public Mito code.
hamilton - A scalable general purpose micro-framework for defining dataflows. THIS REPOSITORY HAS BEEN MOVED TO www.github.com/dagworks-inc/hamilton
airbyte - The leading data integration platform for ETL / ELT data pipelines from APIs, databases & files to data warehouses, data lakes & data lakehouses. Both self-hosted and Cloud-hosted.
Reddit-API-Pipeline
AWS Data Wrangler - pandas on AWS - Easy integration with Athena, Glue, Redshift, Timestream, Neptune, OpenSearch, QuickSight, Chime, CloudWatchLogs, DynamoDB, EMR, SecretManager, PostgreSQL, MySQL, SQLServer and S3 (Parquet, CSV, JSON and EXCEL).