Splink Alternatives

Similar projects and alternatives to splink

osxphotos

96 1,699 9.4 Python splink VS osxphotos

Python app to work with pictures and associated metadata from Apple Photos on macOS. Also includes a package to provide programmatic access to the Photos library, pictures, and metadata.
sqlglot

56 5,511 9.9 Python splink VS sqlglot

Python SQL Parser and Transpiler
InfluxDB

www.influxdata.com featured

Power Real-Time Data Analytics at Scale. Get real-time insights from all types of time series data with InfluxDB. Ingest, query, and analyze billions of data points in real-time with unbounded cardinality.
duckdb

52 16,749 10.0 C++ splink VS duckdb

DuckDB is an in-process SQL OLAP Database Management System
formkiq-core

50 91 6.6 Java splink VS formkiq-core

A full-featured Document Layer for your application, providing the functionality of a flexible document management system, including storage, discovery, processing, and retrieval. Deploys directly into your Amazon Web Services Cloud. 🌟 Star to support our work!
scheme-for-max

34 181 2.8 C splink VS scheme-for-max

Max/MSP external for scripting and live coding Max with s7 Scheme Lisp
zingg

23 880 9.2 Java splink VS zingg

Scalable identity resolution, entity resolution, data mastering and deduplication using ML
KaithemAutomation

17 45 9.8 Python splink VS KaithemAutomation

Pure Python, GUI-focused home automation/consumer grade SCADA
SaaSHub

www.saashub.com featured

SaaSHub - Software Alternatives and Reviews. SaaSHub helps you find the best software and product alternatives
ipyflow

20 1,079 9.5 Python splink VS ipyflow

A reactive Python kernel for Jupyter notebooks.
subtls

20 349 8.0 JavaScript splink VS subtls

A proof-of-concept TypeScript TLS 1.3 client
zetasql

15 2,135 0.0 C++ splink VS zetasql

ZetaSQL - Analyzer Framework for SQL
ultra-weather

15 70 1.5 Svelte splink VS ultra-weather

UltraWeather gives user-friendly, actionable weather forecasts.
codebase-visualizer-action

11 61 0.0 splink VS codebase-visualizer-action

Visualize your codebase during CI.
notabase

10 680 7.7 TypeScript splink VS notabase

A second brain for your knowledge, thoughts, and ideas.
enu

10 447 9.0 Nim splink VS enu

A Logo-like 3D environment, implemented in Nim
dotfile

9 99 4.8 Go splink VS dotfile

Simple version control made for tracking single files
dedupe

9 3,979 7.1 Python splink VS dedupe

:id: A python library for accurate and scalable fuzzy matching, record deduplication and entity-resolution.
libpostal

5 3,951 5.9 C splink VS libpostal

A C library for parsing/normalizing street addresses around the world. Powered by statistical NLP and open geo data.
entity-embed

2 138 0.0 Jupyter Notebook splink VS entity-embed

PyTorch library for transforming entities like companies, products, etc. into vectors to support scalable Record Linkage / Entity Resolution using Approximate Nearest Neighbors.
jetson-nano-image

5 249 0.0 Shell splink VS jetson-nano-image

Discontinued Create minimalist, Ubuntu based images for the Nvidia jetson boards [Moved to: https://github.com/pythops/jetson-image]
dblink

1 54 0.0 Scala splink VS dblink

Distributed Bayesian Entity Resolution in Apache Spark
SaaSHub

www.saashub.com featured

SaaSHub - Software Alternatives and Reviews. SaaSHub helps you find the best software and product alternatives

NOTE: The number of mentions on this list indicates mentions on common posts plus user suggested alternatives. Hence, a higher number means a better splink alternative or higher similarity.

Suggest an alternative to splink

splink reviews and mentions

Posts with mentions or reviews of splink. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2023-05-05.

Splink: Fast, accurate, scalable probabilistic data linkage
1 project | news.ycombinator.com | 13 Mar 2024
Ask HN: What projects are you working on?
1 project | news.ycombinator.com | 1 Jun 2023

https://github.com/moj-analytical-services/splink
Record linkage/Entity linkage
2 projects | /r/datascience | 5 May 2023

Record linkage has been a big part of a project I've been working on for 6 months now. I personally think a great and free solution be using the splink package in Python which can handle 10+m rows which implements the Fellegi-Sunter model (equivalent to a naive-Bayes model) is the classical model in record linkage. It can be trained in an unsupervised manner using some initial parameter estimation (these are quite intuitive) and then expectation maximisation. The features in the model will be different pairwise string comparisons on your field of interest. These can include exact equality; edit distance comparisons like Levensthein distance and Jaro-Winkler; and phonetic comparisons like soundex and double metaphone. The splink pacakge will handle training the model and then all the graph theory at the end to connect all your links into clusters. All the details you'll need are in the links. https://www.robinlinacre.com/probabilistic\_linkage/ https://moj-analytical-services.github.io/splink/
What is the best approach to removing duplicate person records if the only identifier is person firstname middle name and last name? These names are entered in varying ways to the DB, thus they are free-fromatted.
2 projects | /r/SQL | 25 Mar 2023

https://moj-analytical-services.github.io/splink/ is a FOSS python package (but it runs against your db using SQL).
DuckDB – in-process SQL OLAP database management system
4 projects | news.ycombinator.com | 10 Feb 2023

If you're curious, I've written a FOSS record linkage library that executes everything as SQL. It supports multiple SQL backends including DuckDB and Spark for scale, and runs faster than most competitors because it's able to leverage the speed of these backends: https://github.com/moj-analytical-services/splink
Ask HN: What have you created that deserves a second chance on HN?
44 projects | news.ycombinator.com | 26 Jan 2023

Splink - a python library for probabilistic record linkage (fuzzy matching/entity resolution).
Splink is dramatically faster and works on much larger datasets than other open source libraries. I'm particularly proud of the fact we support multiple execution backends (at the moment, DuckDb Spark Athena and Sqlite, but additional adaptors are relatively straightforward to write).
We've had >4 million pypi downloads and it's used in government, academia and the private sector, often replacing extremely expensive proprietary solutions.
https://github.com/moj-analytical-services/splink
More info in blog posts here:
Conformed Dimensions problem that keeps recurring on every project
1 project | /r/dataengineering | 17 Jan 2023

Splink is a SQL tool that can do this https://github.com/moj-analytical-services/splink
How do you join two sources with attributes that aren't identical?
1 project | /r/dataengineering | 12 Nov 2022

Probabilistic record matching model such as a Fellegi-Sunter. Check out the splink package in Python.
Splink 3: Fast, accurate and scalable record linkage (entity resolution) in Python
1 project | /r/dataengineering | 29 Sep 2022

Main docs here: https://moj-analytical-services.github.io/splink
Splink 3: Fast, accurate and scalable fuzzy record linkage in Python with support for multiple backends (FOSS)
2 projects | /r/datascience | 8 Aug 2022

It'd be great to see Splink add value in this area! Do give us a shout if you have any questions. The best place to post is on the Github discussions: https://github.com/moj-analytical-services/splink/discussions
A note from our sponsor - InfluxDB
www.influxdata.com | 4 May 2024

Get real-time insights from all types of time series data with InfluxDB. Ingest, query, and analyze billions of data points in real-time with unbounded cardinality. Learn more →

Stats

Basic splink repo stats

Mentions

Stars

1,091

Activity

9.9

Last Commit

3 days ago

moj-analytical-services/splink is an open source project licensed under MIT License which is an OSI approved license.

The primary programming language of splink is Python.

Popular Comparisons