zingg
Optimus
zingg | Optimus | |
---|---|---|
23 | - | |
886 | 1,446 | |
1.4% | 0.8% | |
9.2 | 0.6 | |
3 days ago | 23 days ago | |
Java | Python | |
GNU Affero General Public License v3.0 | Apache License 2.0 |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
zingg
-
Ask HN: What is the most impactful thing you've ever built?
As part of my data consulting, I struggled with identity resolution and started working on scalable no code identity resolution - https://github.com/zinggAI/zingg/ . It has pushed my limits as a software engineer and product builder, and I had to do a lot of learning to build it. Its cool to see people use Zingg in their workflows and save months of working on custom solutions. Big highlight has been North Carolina Open Campaign Data https://crossroads-cx.medium.com/building-open-access-to-nc-...
-
How to find open source data science python projects to contribute to?
Check https://github.com/zinggAI/zingg/. We recently added Python to our stack and are looking for help with building dbt-zingg python models, databricks-zingg python notebooks, python api, building a python based front end etc.
- Merging datasets
-
is it possible to "fuzzy match" or dedupe columns in Redshift?
If you are open to using a framework for this, check Zingg at https://github.com/zinggAI/zingg. It connects to Redshift, snowflake and other warehouses and can handle multiple columns
-
Show HN: Zingg – open-source entity resolution for single source of truth
Thanks for your support. Yes we do ship with some examples and their models which can be run out of the box. We have 3 customer demographic datasets and an ecommerce items matching across Google and Amazon. You can check them here https://github.com/zinggAI/zingg/tree/main/examples
-
Question about Github Referring Sites
I have an open source project hosted at https://github.com/zinggAI/zingg/.
- How do I promote the project appropriately?
-
GitHub Java Projects to Contribute
Check Zingg out at https://github.com/zinggAI/zingg and let me know if you would like to contribute
-
Match over 1 GB of data with inconsistent names
This is interesting, would love to get your feedback on Zingg(https://github.com/zinggAI/zingg) if you are upto it. Thanks!
-
Open source entity resolution - need your feedback!
I have released an open source entity resolution tool Zingg(https://github.com/zinggAI/zingg). Zingg uses Spark and ML to build single source of truth directly in the warehouse or the datalake. Would love to hear from the Reddit folks here what they think about it - do you find it useful? what can I do to make it better? any advice on the problem or the solution?
Optimus
We haven't tracked posts mentioning Optimus yet.
Tracking mentions began in Dec 2020.
What are some alternatives?
splink - Fast, accurate and scalable probabilistic data linkage with support for multiple SQL backends
AWS Data Wrangler - pandas on AWS - Easy integration with Athena, Glue, Redshift, Timestream, Neptune, OpenSearch, QuickSight, Chime, CloudWatchLogs, DynamoDB, EMR, SecretManager, PostgreSQL, MySQL, SQLServer and S3 (Parquet, CSV, JSON and EXCEL).
clrs
sweetviz - Visualize and compare datasets, target values and associations, with one line of code.
rumble - ⛈️ RumbleDB 1.21.0 "Hawthorn blossom" 🌳 for Apache Spark | Run queries on your large-scale, messy JSON-like data (JSON, text, CSV, Parquet, ROOT, AVRO, SVM...) | No install required (just a jar to download) | Declarative Machine Learning and more
ga-extractor - Tool for extracting Google Analytics data suitable for migrating to other platforms/databases
CLRS - Algorithms implementation in C++ and solutions of questions (both code and math proof) from “Introduction to Algorithms” (3e) (CLRS) in LaTeX.
zef - Toolkit for graph-relational data across space and time
skipledger - Differential privacy solution for maintaining and exposing information from evolving, append-only journals / ledgers.
anovos - Anovos - An Open Source Library for Scalable feature engineering Using Apache-Spark
yt-channels-DS-AI-ML-CS - A comprehensive list of 180+ YouTube Channels for Data Science, Data Engineering, Machine Learning, Deep learning, Computer Science, programming, software engineering, etc.
flashtext - Extract Keywords from sentence or Replace keywords in sentences.