RoaringBitmap
pyroscope
RoaringBitmap | pyroscope | |
---|---|---|
24 | 56 | |
3,388 | 7,382 | |
0.8% | - | |
8.5 | 9.6 | |
10 days ago | about 1 year ago | |
Java | Go | |
Apache License 2.0 | GNU Affero General Public License v3.0 |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
RoaringBitmap
-
Iterating over Bit Sets Quickly
I was recently reading about Roaring https://roaringbitmap.org/ which is a highly optimized compressed bitset implementation. I reccomend reading about it if you are interested in this sort of thing. The talk at https://roaringbitmap.org/talks/ is especially good.
- Roaring Bitmaps
- Roaring bitmaps are compressed bitmaps, can be 100x faster
-
What feature would you like to remove in C++26?
However, I would love compressed (not just packed) bitsets too, which is something different to me. I would make it another class with a similar interface, based on something like roaring. It doesn't need to be in the standard, but it would be nice if the API was a such that one could easily swap implementations.
-
Jaccard Index
As an aside if you find yourself having to compute them on the fly, know that the Roaring Bitmaps libraries is the way to go [1]. The bitmaps are compressed, and can be streamed directly into SIMD computations (batching XORs and popcnts 256 bits wide!). The Jaccard index is just intersection_len / union_len [2] away
[1] https://roaringbitmap.org/
[2] https://roaringbitmap.readthedocs.io/en/latest/#roaringbitma...
-
Looking for fast, space-efficient key-lookup
Use a two stage approach, with a bloom/cuckoo filter stored as a https://roaringbitmap.org/ in memory. Then a secondary key/value store on disk (bolt or anything else).
-
BitSet Vs BigInteger
As an aside, if you're dealing with large bit sets, you might also want to evaluate Roaring Bitmaps.
-
Negative Incentives in Academic Research
Sidetracking a bit the conversation. What a coincidence that the author (Lemire) is also represented on Today's #1 "Ask HN: What are some cool but obscure data structures you know about?" as he is the main contributor of RoaringBitmap https://github.com/RoaringBitmap/RoaringBitmap and one of the main authors of the data structure.
- Ask HN: What are some 'cool' but obscure data structures you know about?
- Roaring bitmaps: A better compressed bitset
pyroscope
- Grafana Phlare, open source database for continuous profiling at scale
-
The pros and cons of eBPF profiling in K8s
What do you mean? pyroscope.io was slow for you? or the blog?
- Go garbage collector doesn't release memory
- Pyroscope - Continuous profiling platform
-
Ask HN: What are some 'cool' but obscure data structures you know about?
Tries (or prefix trees).
We use them a lot at Pyroscope for compressing strings that have common prefixes. They are also used in databases (e.g indexes in Mongo) or file formats (e.g debug symbols in macOS/iOS Mach-O format are compressed using tries).
We have an article with some animations that go into details about tries in case anyone's interested [0].
[0] https://github.com/pyroscope-io/pyroscope/blob/main/docs/sto...
- How to add dynamic tags/labels to Java profiles (example)
-
Question: How do you handle oversized heap analysis?
You could use continuous profiling with Pyroscope which uses async-profiler under the hood, but with the added functionality that you can add relevant tags to your VMs (example).
-
JFR (Java Flight Recorder) Parser written in Go
Java Flight Recorder (JFR) is a format for collecting diagnostic and profiling data from Java applications. A while back someone created an issue for Pyroscope , an open source continuous profiler written in Go, to support ingesting profiles in JFR format, but there were no existing parsers that were also written in Go.
-
flamegraph.com - a new website for uploading, analyzing, and sharing pprof profiles
This cloud version is actually a slimmed-down version of Pyroscope which is open source and so you can run it locally.
-
We created flamegraph.com - A website for uploading, analyzing, and sharing flamegraphs
At Pyroscope (open source continuous profiling) we use flamegraphs extensively to visualize and analyze profiling data. However, one of the worst parts about using flamegraphs for analysis is that they are kind of annoying to share.
What are some alternatives?
HyperMinHash-java - Union, intersection, and set cardinality in loglog space
parca - Continuous profiling for analysis of CPU and memory usage, down to the line number and throughout time. Saving infrastructure cost, improving performance, and increasing reliability.
lucene - Apache Lucene open-source search software
profefe - Continuous profiling for long-term postmortem analysis
CQEngine - Ultra-fast SQL-like queries on Java collections
barrier - Open-source KVM software
Primes - Prime Number Projects in C#/C++/Python
Grafana - The open and composable observability and data visualization platform. Visualize metrics, logs, and traces from multiple sources like Prometheus, Loki, Elasticsearch, InfluxDB, Postgres and many more.
Feign - Feign makes writing java http clients easier
SheetJS js-xlsx - 📗 SheetJS Spreadsheet Data Toolkit -- New home https://git.sheetjs.com/SheetJS/sheetjs
maven-compiler-plugin - Apache Maven Compiler Plugin
Oat++ - 🌱Light and powerful C++ web framework for highly scalable and resource-efficient web application. It's zero-dependency and easy-portable.