soda-spark
Soda Spark is a PySpark library that helps you with testing your data in Spark Dataframes (by sodadata)
monosi
Open source data observability platform (by monosidev)
soda-spark | monosi | |
---|---|---|
1 | 20 | |
60 | 320 | |
- | 0.0% | |
0.0 | 0.0 | |
almost 2 years ago | over 1 year ago | |
Python | Python | |
Apache License 2.0 | Apache License 2.0 |
The number of mentions indicates the total number of mentions that we've tracked plus the number of user suggested alternatives.
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
soda-spark
Posts with mentions or reviews of soda-spark.
We have used some of these posts to build our list of alternatives
and similar projects. The last one was on 2022-01-23.
-
How do you test your pipelines?
Since you already have Spark setup, perhaps it would be easier to build a DataFrames by loading data from different tables and validate it in one go ? You can give soda-spark a try (disclosure: I'm one of the developers), using which you can specify your checks using YAML declaratively and run the validations in spark jobs.
monosi
Posts with mentions or reviews of monosi.
We have used some of these posts to build our list of alternatives
and similar projects. The last one was on 2023-03-18.
-
Open source data observability tools with UI?
I also found https://github.com/monosidev/monosi but it seems there are no activities in the repository from last year.
-
Databricks monitoring/observability
I'm building an open source data observability platform - https://github.com/monosidev/monosi that visualizes metadata collected from data warehouses. Databricks is currently not supported (contributions welcome!), but it may help to take a look at how we approach the anomaly detection & visualization aspects.
-
Monitor PostgreSQL for anomalies in ingested data
Building an open source tool that lets you monitor PostgreSQL instances form anomalies in data coming in - https://github.com/monosidev/monosi
- Open Source Data Observability for BigQuery
-
Metadata extraction and management
It’s open source, check out the repository here - https://github.com/monosidev/monosi
-
How to Monitor Supabase with Monosi
🎉 Congratulations, you've just set up and scheduled a data monitor on your Supabase instance. You can now add more monitors to other tables in your database. Find more information on how to use Monosi here.
-
Setting up data monitoring for PostgreSQL
Now that you’ve worked through an example using a public PostgreSQL instance, you can further extend this to your own data store. For more information, get started here.
- Monosi v0.0.3 Released! Open source Data Observability now with a Web UI, Postgres Support, & more.
-
Sunday Daily Thread: What's everyone working on this week?
Continuing to build out & stabilize Monosi (open source data observability) - https://github.com/monosidev/monosi
-
Data pipeline suggestions
Observability: Monosi
What are some alternatives?
When comparing soda-spark and monosi you can also consider the following projects:
great_expectations - Always know what to expect from your data.
datahub - The Metadata Platform for your Data Stack