splink vs duckdb

splink

Fast, accurate and scalable probabilistic data linkage with support for multiple SQL backends (by moj-analytical-services)

Source Code

moj-analytical-services.github.io

Suggest alternative

Edit details

duckdb

DuckDB is an in-process SQL OLAP Database Management System (by duckdb)

SQL Database Olap Analytics embedded-database

Source Code

duckdb.org

Suggest alternative

Edit details

Scout Monitoring - Free Django app performance insights with Scout Monitoring

Get Scout setup in minutes, and let us sweat the small stuff. A couple lines in settings.py is all you need to start monitoring your apps. Sign up for our free tier today.

www.scoutapm.com

featured

InfluxDB - Power Real-Time Data Analytics at Scale

Get real-time insights from all types of time series data with InfluxDB. Ingest, query, and analyze billions of data points in real-time with unbounded cardinality.

www.influxdata.com

featured

splink		duckdb
	Project
16	Mentions	52
1,116	Stars	17,924
2.7%	Growth	6.6%
9.9	Activity	10.0
12 days ago	Latest Commit	1 day ago
Python	Language	C++
MIT License	License	MIT License

The number of mentions indicates the total number of mentions that we've tracked plus the number of user suggested alternatives.
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.

splink

Posts with mentions or reviews of splink. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2023-05-05.

Splink: Fast, accurate, scalable probabilistic data linkage
1 project | news.ycombinator.com | 13 Mar 2024
Ask HN: What projects are you working on?
1 project | news.ycombinator.com | 1 Jun 2023

https://github.com/moj-analytical-services/splink
Record linkage/Entity linkage
2 projects | /r/datascience | 5 May 2023

Record linkage has been a big part of a project I've been working on for 6 months now. I personally think a great and free solution be using the splink package in Python which can handle 10+m rows which implements the Fellegi-Sunter model (equivalent to a naive-Bayes model) is the classical model in record linkage. It can be trained in an unsupervised manner using some initial parameter estimation (these are quite intuitive) and then expectation maximisation. The features in the model will be different pairwise string comparisons on your field of interest. These can include exact equality; edit distance comparisons like Levensthein distance and Jaro-Winkler; and phonetic comparisons like soundex and double metaphone. The splink pacakge will handle training the model and then all the graph theory at the end to connect all your links into clusters. All the details you'll need are in the links. https://www.robinlinacre.com/probabilistic\_linkage/ https://moj-analytical-services.github.io/splink/
What is the best approach to removing duplicate person records if the only identifier is person firstname middle name and last name? These names are entered in varying ways to the DB, thus they are free-fromatted.
2 projects | /r/SQL | 25 Mar 2023

https://moj-analytical-services.github.io/splink/ is a FOSS python package (but it runs against your db using SQL).
DuckDB – in-process SQL OLAP database management system
4 projects | news.ycombinator.com | 10 Feb 2023

If you're curious, I've written a FOSS record linkage library that executes everything as SQL. It supports multiple SQL backends including DuckDB and Spark for scale, and runs faster than most competitors because it's able to leverage the speed of these backends: https://github.com/moj-analytical-services/splink
Ask HN: What have you created that deserves a second chance on HN?
44 projects | news.ycombinator.com | 26 Jan 2023

Splink - a python library for probabilistic record linkage (fuzzy matching/entity resolution).
Splink is dramatically faster and works on much larger datasets than other open source libraries. I'm particularly proud of the fact we support multiple execution backends (at the moment, DuckDb Spark Athena and Sqlite, but additional adaptors are relatively straightforward to write).
We've had >4 million pypi downloads and it's used in government, academia and the private sector, often replacing extremely expensive proprietary solutions.
https://github.com/moj-analytical-services/splink
More info in blog posts here:
Conformed Dimensions problem that keeps recurring on every project
1 project | /r/dataengineering | 17 Jan 2023

Splink is a SQL tool that can do this https://github.com/moj-analytical-services/splink
How do you join two sources with attributes that aren't identical?
1 project | /r/dataengineering | 12 Nov 2022

Probabilistic record matching model such as a Fellegi-Sunter. Check out the splink package in Python.
Splink 3: Fast, accurate and scalable record linkage (entity resolution) in Python
1 project | /r/dataengineering | 29 Sep 2022

Main docs here: https://moj-analytical-services.github.io/splink
Splink 3: Fast, accurate and scalable fuzzy record linkage in Python with support for multiple backends (FOSS)
2 projects | /r/datascience | 8 Aug 2022

It'd be great to see Splink add value in this area! Do give us a shout if you have any questions. The best place to post is on the Github discussions: https://github.com/moj-analytical-services/splink/discussions

duckdb

Posts with mentions or reviews of duckdb. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2023-11-06.

🪄 DuckDB sql hack : get things SORTED w/ constraint CHECK
1 project | dev.to | 4 Apr 2024
DuckDB: Move to push-based execution model (2021)
1 project | news.ycombinator.com | 15 Mar 2024
DuckDB performance improvements with the latest release
8 projects | news.ycombinator.com | 6 Nov 2023

I'm not sure if the fix is reassuring or not: https://github.com/duckdb/duckdb/pull/9411/files
Building a Distributed Data Warehouse Without Data Lakes
3 projects | news.ycombinator.com | 2 Nov 2023

It's an interesting question!
The problem is that the data is spread everywhere - no choice about that. So with that in mind, how do you query that data? Today, the idea is that you HAVE to put it into a central location. With tools like Bacalhau[1] and DuckDB [2], you no longer have to - a single query can be sharded amongst all your data - EFFECTIVELY giving you a lot of what you want from a data lake.
It's not a replacement, but if you can do a few of these items WITHOUT moving the data, you will be able to see really significant cost and time savings.
[1] https://github.com/bacalhau-project/bacalhau
[2] https://github.com/duckdb/duckdb
DuckDB 0.9.0
3 projects | news.ycombinator.com | 26 Sep 2023
Push or Pull, is this a question?
2 projects | dev.to | 9 Aug 2023

[4] Switch to Push-Based Execution Model by Mytherin · Pull Request #2393 · duckdb/duckdb (github.com)
Show HN: Hydra 1.0 – open-source column-oriented Postgres
12 projects | news.ycombinator.com | 3 Aug 2023

it depends on your query obviously.
In general, I did very deep benchmarking of pg, clickhouse and duckdb, and I sure didn't make stupid mistakes like this: https://news.ycombinator.com/item?id=36990831
My dataset has 50B rows and 2tb of data, and I think columnar dbs are very overhiped and I chose pg because:
- pg performance is acceptable, maybe 2-3x times slower than clickhouse and duckdb on some queries if pg is configured correctly and run on compressed storage
- clickhouse and duckdb start falling apart very fast because they specialized on very narrow type of queries: https://github.com/ClickHouse/ClickHouse/issues/47520 https://github.com/ClickHouse/ClickHouse/issues/47521 https://github.com/duckdb/duckdb/discussions/6696
🦆 Effortless Data Quality w/duckdb on GitHub ♾️
3 projects | dev.to | 25 Jul 2023

This action installs duckdb with the version provided in input.
Using SQL inside Python pipelines with Duckdb, Glaredb (and others?)
6 projects | /r/dataengineering | 30 Jun 2023

Duckdb: https://github.com/duckdb/duckdb - seems pretty popular, been keeping an eye on this for close to a year now.
CSV or Parquet File Format
3 projects | /r/Python | 1 Jun 2023

The Parquet-Go library is very complex, not yet success to use it. So I ask whether DuckDB can provide API https://github.com/duckdb/duckdb/issues/7776

What are some alternatives?

When comparing splink and duckdb you can also consider the following projects:

zingg - Scalable identity resolution, entity resolution, data mastering and deduplication using ML

ClickHouse - ClickHouse® is a real-time analytics DBMS

dedupe - :id: A python library for accurate and scalable fuzzy matching, record deduplication and entity-resolution.

sqlite-worker - A simple, and persistent, SQLite database for Web and Workers.

libpostal - A C library for parsing/normalizing street addresses around the world. Powered by statistical NLP and open geo data.

datasette - An open source multi-tool for exploring and publishing data

sqlglot - Python SQL Parser and Transpiler

octosql - OctoSQL is a query tool that allows you to join, analyse and transform data from multiple databases and file formats using SQL.

entity-embed - PyTorch library for transforming entities like companies, products, etc. into vectors to support scalable Record Linkage / Entity Resolution using Approximate Nearest Neighbors.

metabase-clickhouse-driver - ClickHouse database driver for the Metabase business intelligence front-end

dblink - Distributed Bayesian Entity Resolution in Apache Spark

datafusion - Apache DataFusion SQL Query Engine

splink vs zingg duckdb vs ClickHouse splink vs dedupe duckdb vs sqlite-worker splink vs libpostal duckdb vs datasette splink vs sqlglot duckdb vs octosql splink vs entity-embed duckdb vs metabase-clickhouse-driver splink vs dblink duckdb vs datafusion

Compare splink vs duckdb and see what are their differences.

splink

duckdb

splink

duckdb

What are some alternatives?