usearch
pgrx
usearch | pgrx | |
---|---|---|
21 | 13 | |
1,691 | 3,257 | |
8.9% | 3.7% | |
9.8 | 9.5 | |
5 days ago | 2 days ago | |
C++ | Rust | |
Apache License 2.0 | GNU General Public License v3.0 or later |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
usearch
-
I'm writing a new vector search SQLite Extension
Might have a look at this library:
https://github.com/unum-cloud/usearch
It does HNSW and there is a SQLite related project, though not quite the same thing.
- USearch SQLite Extensions for Vector and Text Search
-
Ask HN: What is the state of art approximate k-NN search algorithm today?
Another worth mentioning in this thread is usearch, though not a separate algorithm, based on HNSW with a bunch of optimizations https://github.com/unum-cloud/usearch
-
Vector Databases: A Technical Primer [pdf]
I've used usearch successfully for a small project: https://github.com/unum-cloud/usearch/
- 90x Faster Than Pgvector โ Lantern's HNSW Index Creation Time
-
Python, C, Assembly โ Faster Cosine Similarity
The hardest (still missing) part of efficient cosine computation distance computation is picking a good epsilon for the `sqrt` calculation and avoiding "division by zero" problems.
We have an open issue about it in USearch and a related one in SimSIMD itself, so if you have any suggestions, please share your insights - they would impact millions of devices using the library (directly on servers and mobile, and through projects like ClickHouse and some of the Google repos): https://github.com/unum-cloud/usearch/issues/320
-
Show HN: I scraped 25M Shopify products to build a search engine
As you scale, you may benefit from these two projects I maintain, and the Big Tech uses :)
https://github.com/unum-cloud/usearch - for faster search
https://github.com/unum-cloud/uform - for cheaper multi-lingual multi-modal embeddings
- [P] unum-cloud/usearch: Fastest Open-Source Similarity Search engine for Vectors in Python, JavaScript, C++, C, Rust, Java, Objective-C, Swift, C#, GoLang, and Wolfram ๐
- USearch: SIMD-accelerated Vector Search Structure for 10 Programming Languages
-
Stringzilla: Fastest string sort, search, split, and shuffle using SIMD
> It doesn't appear to query CPUID
Yes, I'm actually looking for a good way to do it for other projects as well. I've looked into a couple more libs, and here is the best I've come up with so far: https://github.com/unum-cloud/usearch/blob/f942b6f334b31716f...
> Your substring routines have multiplicative worst case
Yes, that is true. It's a very simple stupid trick, just happens to work well for me :)
> It seems quite likely that your confirmation step
We have a different library internally at Unum, that avoids this shortcoming. It has a few thousand lines of C++ templates with SIMD intrinsics... and it's definitely more efficient, but the margins aren't always high. So I kept the pure C version with inlined functions as minimal and simple as possible.
> It would actually be possible to hook Stringzilla up to `memchr`'s benchmark suite if you were interested. :-)
Yes, that would be a fun thing to do! I haven't had time to look into `memchr` yet, but would expect great perf from your lib as well. For me the State of the Art is Intel HyperScan. Probably the most advanced SIMD library overall, not just for strings. I was very impressed with their perf ~5 years ago. But the repo is 200 K LOC... So get ready to invest a weekend :)
That said, I'm a bit slammed with work right now, including open-source. Hoping to ship a new major release in UCall this week, and a minor one in USearch :)
pgrx
-
Building a Managed Postgres Service in Rust
Consider also the companies and work behind pgrx [0] and pgzx [1]:
[0] https://github.com/pgcentralfoundation/pgrx
[1] https://github.com/xataio/pgzx
-
UUIDv7 is coming in PostgreSQL 17
If you like this (I do very much), you might also like pg_idkit[0] which is a little extension with a bunch of other kinds of IDs that you can generate inside PG, thanks to the seriously awesome pgrx[1] and Rust.
[0]: https://github.com/VADOSWARE/pg_idkit
[1]: https://github.com/pgcentralfoundation/pgrx
-
90x Faster Than Pgvector โ Lantern's HNSW Index Creation Time
(disclosure, i work at supabase and have been developing TLEs with the RDS team)
Trusted Language Extensions refer to an extension written in any trusted language. In this case Rust, but it also includes: plpgsql, plv8, etc. See [0]
> PL/Rust is a more performant and more feature-rich alternative to PL/pgSQL
This is only partially true. plpgsql has bindings to low-level Postgres APIs, so in some cases it is just as fast (or faster) than Rust.
> Building a vector index (or any index for that matter) inside Postgres is a more involved process and can not be done via the UDF interface, be it Rust, C or PL/pgSQL
Most PG Rust extensions are written with the excellent pgrx framework [1]. While it doesn't have index bindings right now, I can certainly imagine a future where this is possible[2].
All that said - I think there are a lot of hoops to jump through right now and I doubt it's worth it for the Latern team. I think they are right to focus on developing a separate C extension
[0] TLE: https://supabase.com/blog/pg-tle
[1] pgrx: https://github.com/pgcentralfoundation/pgrx
[2] https://github.com/pgcentralfoundation/pgrx/issues/190#issue...
-
SQL as API
Iโm currently playing with PostgreSQL, foreign data wrappers, and pgrx rust extensions. My development experience has been surprisingly smooth and enjoyable.
My main issue is that joins will be processed locally, so all the foreign data will be fetched before the join happens. But otherwise basic CRUD is easy.
https://wiki.postgresql.org/wiki/Foreign_data_wrappers
https://github.com/pgcentralfoundation/pgrx
https://github.com/supabase/wrappers
-
Postgres: The Next Generation
I think maybe what youโre really looking for are the files here: https://github.com/pgcentralfoundation/pgrx/tree/c2eac033856...
Those are the internals we currently expose as unsafe โsysโ bindings.
As we/contributors identify more that are desired we add them.
pgrxโ focus is on providing safe wrappers and general interfaces to the Postgres internals, which is the bulk of our work and is what will take many years.
As unsafe bindings go, we could just expose everything, and likely eventually will. Thereโs just some practical management concerns around doing that without a better namespace organization โ- something weโve been working.
The Postgres sources are not small. They are very complex, inconsistent in places, and often follow patterns that are specific to Postgres and not easy to generalize.
If youโve never built an extension with pgrx, give it a shot one afternoon. Itโs very exciting to see your own code running in your database.
- Pgrx โ Build Postgres Extensions with Rust
-
Pg_bm25: Elastic-Quality Full Text Search Inside Postgres
pgrx is one of the greatest enabling innovations in the PG ecosystem in a long time.
Awesome to see so many high quality extensions come out of it.
https://github.com/pgcentralfoundation/pgrx
- PGRX v0.9.7
- Let's make PostgreSQL multi-threaded (pgsql-hackers)
-
Build high-performance functions in Rust on Amazon RDS for PostgreSQL
If you're interested in what my Threadripper 3970X does with it, there's some numbers in this PR: https://github.com/tcdi/pgrx/pull/1147
What are some alternatives?
StringZilla - Up to 10x faster strings for C, C++, Python, Rust, and Swift, leveraging SWAR and SIMD on Arm Neon and x86 AVX2 & AVX-512-capable chips to accelerate search, sort, edit distances, alignment scores, etc ๐ฆ
api - ๐ Core REST API & Gateway for Zaun
ustore - Multi-Modal Database replacing MongoDB, Neo4J, and Elastic with 1 faster ACID solution, with NetworkX and Pandas interfaces, and bindings for C 99, C++ 17, Python 3, Java, GoLang ๐๏ธ
plrust - A Rust procedural language handler for PostgreSQL
uform - Pocket-Sized Multimodal AI for content understanding and generation across multilingual texts, images, and ๐ video, up to 5x faster than OpenAI CLIP and LLaVA ๐ผ๏ธ & ๐๏ธ
readyset - Readyset is a MySQL and Postgres wire-compatible caching layer that sits in front of existing databases to speed up queries and horizontally scale read throughput. Under the hood, ReadySet caches the results of cached select statements and incrementally updates these results over time as the underlying data changes.
faiss - A library for efficient similarity search and clustering of dense vectors.
mimir - โก Supercharged Flutter/Dart Database
SimSIMD - Up to 200x Faster Inner Products and Vector Similarity โ for Python, JavaScript, Rust, and C, supporting f64, f32, f16 real & complex, i8, and binary vectors using SIMD for both x86 AVX2 & AVX-512 and Arm NEON & SVE ๐
paradedb - Postgres for Search and Analytics
kuzu - Embeddable property graph database management system built for query speed and scalability. Implements Cypher.
influxdb_iox - Pronounced (influxdb eye-ox), short for iron oxide. This is the new core of InfluxDB written in Rust on top of Apache Arrow.