nessie vs dolt

nessie

Nessie: Transactional Catalog for Data Lakes with Git-like semantics (by projectnessie)

Source Code

projectnessie.org

Suggest alternative

Edit details

InfluxDB - Power Real-Time Data Analytics at Scale

Get real-time insights from all types of time series data with InfluxDB. Ingest, query, and analyze billions of data points in real-time with unbounded cardinality.

www.influxdata.com

featured

SaaSHub - Software Alternatives and Reviews

SaaSHub helps you find the best software and product alternatives

www.saashub.com

featured

nessie		dolt
	Project
13	Mentions	93
834	Stars	16,993
3.6%	Growth	1.7%
9.9	Activity	10.0
4 days ago	Latest Commit	about 7 hours ago
Java	Language	Go
Apache License 2.0	License	Apache License 2.0

The number of mentions indicates the total number of mentions that we've tracked plus the number of user suggested alternatives.
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.

nessie

Posts with mentions or reviews of nessie. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2024-01-22.

A deep dive into the concept and world of Apache Iceberg Catalogs
1 project | dev.to | 1 Mar 2024

Nessie is an innovative open-source catalog that extends beyond the traditional catalog capabilities in the Apache Iceberg ecosystem, introducing git-like features to data management. This catalog not only tracks table metadata but also allows users to capture commits at a holistic level, enabling advanced operations such as multi-table transactions, rollbacks, branching, and tagging. These features provide a new layer of flexibility and control over data changes, resembling version control systems in software development.
FLaNK Stack Weekly 22 January 2024
37 projects | dev.to | 22 Jan 2024
Why is Hive Metastore everywhere? (Especially Iceberg)
1 project | /r/dataengineering | 30 Jun 2023

Try Nessie https://github.com/projectnessie/nessie - it recently got trino support as well ..
What are the main things I need to know to be hired as a Java developer?
4 projects | /r/java | 1 Feb 2023
Is learning and mastering Spring & Spring boot worth it in 2023 ?
3 projects | /r/java | 30 Jan 2023
Which lakehouse table format do you expect your organization will be using by the end of 2023?
2 projects | /r/dataengineering | 25 Dec 2022

Project Nessie (https://projectnessie.org/) will be the catalog that eventually decouples Iceberg from Hive. At that point, I think it will be a no brainer to go Iceberg over Delta.
5 Reasons Your Data Lakehouse should Embrace Dremio Cloud
2 projects | dev.to | 9 Aug 2022

The Dremio Sonar query engine can query your data where it exists whether it's AWS Glue, S3, Nessie Catalogs, MySQL, Postgres, RedShift and an ever growing list of sources.
Project Nessie: Transactional Catalog for Data Lakes with Git-Like Semantics
1 project | news.ycombinator.com | 23 Jun 2022
Introduction to The World of Data - (OLTP, OLAP, Data Warehouses, Data Lakes and more)
2 projects | dev.to | 20 Jun 2022

We will also need a catalog to track all of these tables, with the open source Project Nessie we can do just that, and also get great versioning features similar to using Git when developing applications allowing data engineers to practice "data as code" and "write-audit-publish" patterns on their data.
DoltLab v0.2.0
5 projects | news.ycombinator.com | 11 Feb 2022

dolt

Posts with mentions or reviews of dolt. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2024-04-09.

A MySQL compatible database engine written in pure Go
10 projects | news.ycombinator.com | 9 Apr 2024

Hi, this is my project :)
For us this package is most important as the query engine that powers Dolt:
https://github.com/dolthub/dolt
We aren't the original authors but have contributed the vast majority of its code at this point. Here's the origin story if you're interested:
https://www.dolthub.com/blog/2020-05-04-adopting-go-mysql-se...
The Great Migration from MongoDB to PostgreSQL
1 project | news.ycombinator.com | 29 Mar 2024

It's a pretty good default stance, yeah.
We have been trying to convince people to use our new database [1] for several years and it's an uphill battle, because Postgres really is the best choice for most people. They really have to need our unique feature (version control) to even consider it over Postgres, and I don't blame them.
[1] https://github.com/dolthub/dolt
What I Talk About When I Talk About Query Optimizer (Part 1): IR Design
7 projects | news.ycombinator.com | 29 Jan 2024

We implemented a query optimizer with a flexible intermediate representation in pure Go:
https://github.com/dolthub/go-mysql-server
Getting the IR correct so that it's both easy to use and flexible enough to be useful is a really interesting design challenge. Our primary abstraction in the query plan is called a Node, and is way more general than the IR type described in the article from OP. This has probably hurt us: we only recently separated the responsibility to fetch rows into its own part of the runtime, out of the IR -- originally row fetching was coupled to the Node type directly.
This is also the query engine that Dolt uses:
https://github.com/dolthub/dolt
But it has a plug-in architecture, so you can use the engine on any data source that implements a handful of Go interface.
Dolt – Git for Data
1 project | news.ycombinator.com | 18 Jan 2024
Dolt: A version-controlled SQL database
1 project | news.ycombinator.com | 5 Jan 2024
Show HN: DoltgreSQL – Version-Controlled Database, Like Git and PostgreSQL
7 projects | news.ycombinator.com | 1 Nov 2023

Just want to point out that we're announcing development on the project. It's absolutely not ready for mainstream use yet! We have Dolt (https://github.com/dolthub/dolt) which is production-ready and widely in use, but it uses MySQL's syntax and wire protocol. We are building the Dolt equivalent for PostgreSQL, which is DoltgreSQL, but it's only pre-alpha.
Pg_branch: Pre-alpha Postgres extension brings Neon-like branching
6 projects | news.ycombinator.com | 1 Oct 2023

Interesting that branching is now better supported and almost free. I wonder if merging can be simplified or whether it already is as simple and as fast as it can be?
I guess I am inspired by Dolt’s ability to branch and merge: https://github.com/dolthub/dolt
SQLedge: Replicate Postgres to SQLite on the Edge
9 projects | news.ycombinator.com | 9 Aug 2023

#. SQLite WAL mode
From https://www.sqlite.org/isolation.html https://news.ycombinator.com/item?id=32247085 :
> [sqlite] WAL mode permits simultaneous readers and writers. It can do this because changes do not overwrite the original database file, but rather go into the separate write-ahead log file. That means that readers can continue to read the old, original, unaltered content from the original database file at the same time that the writer is appending to the write-ahead log
#. superfly/litefs: aFUSE-based file system for replicating SQLite https://github.com/superfly/litefs
#. sqldiff: https://www.sqlite.org/sqldiff.html https://news.ycombinator.com/item?id=31265005
#. dolthub/dolt: https://github.com/dolthub/dolt
> Dolt can be set up as a replica of your existing MySQL or MariaDB database using standard MySQL binlog replication. Every write becomes a Dolt commit. This is a great way to get the version control benefits of Dolt and keep an existing MySQL or MariaDB database.
#. pganalyze/libpg_query: https://github.com/pganalyze/libpg_query :
> C library for accessing the PostgreSQL parser outside of the server environment
#. Ibis + Substrait [ + DuckDB ]
> ibis strives to provide a consistent interface for interacting with a multitude of different analytical execution engines, most of which (but not all) speak some dialect of SQL.
> Today, Ibis accomplishes this with a lot of help from `sqlalchemy` and `sqlglot` to handle differences in dialect, or we interact directly with available Python bindings (for instance with the pandas, datafusion, and polars backends).
> [...] `Substrait` is a new cross-language serialization format for communicating (among other things) query plans. It's still in its early days, but there is already nascent support for Substrait in Apache Arrow, DuckDB, and Velox.
#. benbjohnson/postlite: https://github.com/benbjohnson/postlite
> postlite is a network proxy to allow access to remote SQLite databases over the Postgres wire protocol. This allows GUI tools to be used on remote SQLite databases which can make administration easier.
> The proxy works by translating Postgres frontend wire messages into SQLite transactions and converting results back into Postgres response wire messages. Many Postgres clients also inspect the pg_catalog to determine system information so Postlite mirrors this catalog by using an attached in-memory database with virtual tables. The proxy also performs minor rewriting on these system queries to convert them to usable SQLite syntax.
> Note: This software is in alpha. Please report bugs. Postlite doesn't alter your database unless you issue INSERT, UPDATE, DELETE commands so it's probably safe. If anything, the Postlite process may die but it shouldn't affect your database.
#. > "Hosting SQLite Databases on GitHub Pages" (2021) re: sql.js-httpvfs, DuckDB https://news.ycombinator.com/item?id=28021766
#. awesome-db-tools https://github.com/mgramin/awesome-db-tools
How do you sync dev databases across multiple devices?
2 projects | /r/PHP | 9 May 2023
Ask HN: Data Management for AI Training
3 projects | news.ycombinator.com | 30 Apr 2023

If you are just looking for data versioning there is Dolt:
https://github.com/dolthub/dolt
And that has a user-friendly UI in DoltHub:
https://www.dolthub.com/
You wouldn't store the images themselves in Dolt, those would likely be links to S3 but al the labels and surrounding metadata could be stored in Dolt?
DISCLAIMER: I'm the CEO of DoltHub so this is self-promotion.

What are some alternatives?

When comparing nessie and dolt you can also consider the following projects:

git-bug - Distributed, offline-first bug tracker embedded in git, with bridges

liquibase - Main Liquibase Source

dvc - 🦉 ML Experiments and Data Management with Git

absurd-sql - sqlite3 in ur indexeddb (hopefully a better backend soon)

hiveberg - Demonstration of a Hive Input Format for Iceberg

noms - The versioned, forkable, syncable database

dremio-oss - Dremio - the missing link in modern data

TimescaleDB - An open-source time-series SQL database optimized for fast ingest and complex queries. Packaged as a PostgreSQL extension.

vitess - Vitess is a database clustering system for horizontal scaling of MySQL.

Flyway - Flyway by Redgate • Database Migrations Made Easy.

temporal_tables - Temporal Tables PostgreSQL Extension

nessie vs git-bug dolt vs liquibase nessie vs dvc dolt vs absurd-sql nessie vs hiveberg dolt vs noms nessie vs dremio-oss dolt vs TimescaleDB nessie vs noms dolt vs vitess nessie vs Flyway dolt vs temporal_tables

Compare nessie vs dolt and see what are their differences.

nessie

dolt

nessie

dolt

What are some alternatives?