rebar vs ripgrep

rebar

A biased barometer for gauging the relative speed of some regex engines on a curated set of tasks. (by BurntSushi)

ripgrep

ripgrep recursively searches directories for a regex pattern while respecting your gitignore (by BurntSushi)

Applications written in Rust Text processing Ripgrep recursively-search Search Regex Gitignore Grep Command Line Tool Command-line CLI Rust

Source Code

Suggest alternative

Edit details

Our great sponsors

InfluxDB - Power Real-Time Data Analytics at Scale

WorkOS - The modern identity platform for B2B SaaS

SaaSHub - Software Alternatives and Reviews

Our great sponsors

rebar		ripgrep
	Project
22	Mentions	348
197	Stars	44,901
-	Growth	-
8.5	Activity	9.3
about 1 month ago	Latest Commit	4 days ago
Python	Language	Rust
The Unlicense	License	The Unlicense

The number of mentions indicates the total number of mentions that we've tracked plus the number of user suggested alternatives.
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.

rebar

Posts with mentions or reviews of rebar. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2024-04-16.

Knuth–Morris–Pratt Illustrated
2 projects | news.ycombinator.com | 16 Apr 2024

https://github.com/BurntSushi/rebar
For regex, you can't really distill it down to one single fastest algorithm.
It's somewhat similar even for substring search. But certainly, the fastest algorithms are going to be the ones that make use of SIMD in some way.
Regex character "$" doesn't mean "end-of-string"
1 project | news.ycombinator.com | 20 Mar 2024

I'll add two notes to this:
* Finite automata based regex engines don't necessarily have to be slower than backtracking engines like PCRE. Go's regexp is in practice slower in a lot of cases, but this is more a property of its implementation than its concept. See: https://github.com/BurntSushi/rebar?tab=readme-ov-file#summa... --- Given "sufficient" implementation effort, backtrackers and finite automata engines can both perform very well, with one beating the other in some cases but not in others. It depends.
* Fun fact is that if you're iterating over all matches in a haystack (e.g., Go's `FindAll` routines), then you're susceptible to O(m * n^2) search time. This applies to all regex engines that implement some kind of leftmost match priority. See https://github.com/BurntSushi/rebar?tab=readme-ov-file#quadr... for a more detailed elaboration on this point.
Re2c
4 projects | news.ycombinator.com | 22 Feb 2024

They are extremely fast too: https://github.com/BurntSushi/rebar?tab=readme-ov-file#summa...
C# Regex engine is now 3rd fastest in the world
3 projects | news.ycombinator.com | 31 Dec 2023

I love the flourish of "in the world." I had never thought about it that way. Which makes me think if there are any regex engines that aren't in rebar that could conceivably by competitive with the top engines in rebar. I do maintained a WANTED list of engines[1], but none of them jump out to me except for maybe Nim's engine.
Of course, there's also the question of whether the benchmarks are representative enough to make such extrapolations. I don't have a good answer for that one. All models are wrong, but, some are useful.
[1]: https://github.com/BurntSushi/rebar/blob/96c6779b7e1cdd850b8...
Ugrep – a more powerful, ultra fast, user-friendly, compatible grep
27 projects | news.ycombinator.com | 30 Dec 2023

I'm the author of ripgrep and its regex engine.
Your claim is true to a first approximation. But greps are line oriented, and that means there are optimizations that can be done that are hard to do in a general regex library.
If you read my commentary in the ripgrep discussion above, you'll note that it isn't just about the benchmarks themselves being accurate, but the model they represent. Nevertheless, I linked the hypergrep benchmarks not because of Hyperscan, but because they were done by someone who isn't the author of either ripgrep or ugrep.
As for regex benchmarks, you'll want to check out rebar: https://github.com/BurntSushi/rebar
You can see my full thoughts around benchmark design and philosophy if you read the rebar documentation. Be warned though, you'll need some time.
There is a fork of ripgrep with Hyperscan support: https://sr.ht/~pierrenn/ripgrep/
Translations of Russ Cox's Thompson NFA C Program to Rust
3 projects | news.ycombinator.com | 2 Nov 2023

Before getting to your actual question, it might help to look at a regex benchmark that compares engines (perhaps JITs are not the fastest in all cases!): https://github.com/BurntSushi/rebar
In particular, the `regex-lite` engine is strictly just the PikeVM without any frills. No prefilters or literal optimizations. No other engines. Just the PikeVM.
As to your question, the PikeVM is, essentially, an NFA simulation. The PikeVM just refers to the layering of capture state on top of the NFA simulation. But you can peel back the capture state and you're still left with a slow NFA simulation. I mention this because you seem to compare the PikeVM with "big graph structures with NFAs/DFAs." But the PikeVM is using a big NFA graph structure.
At a very high level, the time complexity of a Thompson NFA simulation and a DFA hints strongly at the answer to your question: searching with a Thompson NFA has worst case O(m*n) time while a DFA has worst case O(n) time, where m is proportional to the size of the regex and n is proportional to the size of the haystack. That is, for each character of the haystack, the Thompson NFA is potentially doing up to `m` amount of work. And indeed, in practice, it really does need to do some work for each character.
A Thompson NFA simulation needs to keep track of every state it is simultaneously in at any given point. And in order to compute the transition function, you need to compute it for every state you're in. The epsilon transitions that are added as part of the Thompson NFA construction (and are, crucially, what make building a Thompson NFA so fast) exacerbate this. So what happens is that you wind up chasing epsilon transitions over and over for each character.
A DFA pre-computes these epsilon closures during powerset construction. Of course, that takes worst case O(2^m) time, which is why real DFAs aren't really used in general purpose engines. Instead, lazy DFAs are used.
As for things like V8, they are backtrackers. They don't need to keep track of every state they're simultaneously in because they don't mind taking a very long time to complete some searches. But in practice, this can make them much faster for some inputs.
Feel free to ask more questions. I'll stop here.
Compile time regular expression in C++
5 projects | news.ycombinator.com | 12 Sep 2023

I'd love for someone to add this to rebar[1] so that we can get a good sense of how well it does against other general purpose regex engines. It will be a little tricky to add (since the build step will require emitting a C++ program and compiling it), but it should be possible.
[1]: https://github.com/BurntSushi/rebar
Stringzilla: Fastest string sort, search, split, and shuffle using SIMD
9 projects | news.ycombinator.com | 29 Aug 2023
Rust vs. Go in 2023
9 projects | news.ycombinator.com | 13 Aug 2023

https://github.com/BurntSushi/rebar#summary-of-search-time-b...
Further, Go refusing to have macros means that many libraries use reflection instead, which often makes those parts of the Go program perform no better than Python and in some cases worse. Rust can just generate all of that at compile time with macros, and optimize them with LLVM like any other code. Some Go libraries go to enormous lengths to reduce reflection overhead, but that's hard to justify for most things, and hard to maintain even once done. The legendary https://github.com/segmentio/encoding seems to be abandoned now and progress on Go JSON in general seems to have died with https://github.com/go-json-experiment/json .
Many people claiming their projects are IO-bound are just assuming that's the case because most of the time is spent in their input reader. If they actually measured they'd see it's not even saturating a 100Mbps link, let alone 1-100Gbps, so by definition it is not IO-bound. Even if they didn't need more throughput than that, they still could have put those cycles to better use or at worst saved energy. Isn't that what people like to say about Go vs Python, that Go saves energy? Sure, but it still burns a lot more energy than it would if it had macros.
Rust can use state-of-the-art memory allocators like mimalloc, while Go is still stuck on an old fork of tcmalloc, and not just tcmalloc in its original C, but transpiled to Go so it optimizes much less than LLVM would optimize it. (Many people benchmarking them forget to even try substitute allocators in Rust, so they're actually underestimating just how much faster Rust is)
Finally, even Go Generics have failed to improve performance, and in many cases can make it unimaginably worse through -- I kid you not -- global lock contention hidden behind innocent type assertion syntax: https://planetscale.com/blog/generics-can-make-your-go-code-...
It's not even close. There are many reasons Go is a lot slower than Rust and many of them are likely to remain forever. Most of them have not seen meaningful progress in a decade or more. The GC has improved, which is great, but that's not even a factor on the Rust side.
A Regex Barometer
1 project | /r/hypeurls | 5 Jul 2023

ripgrep

Posts with mentions or reviews of ripgrep. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2024-04-17.

Ask HN: What software sparks joy when using?
10 projects | news.ycombinator.com | 17 Apr 2024

ripgrep - https://github.com/BurntSushi/ripgrep
Code Search Is Hard
13 projects | news.ycombinator.com | 10 Apr 2024

Basic code searching skills seems like something new developers are never explicitly taught, but which is an absolutely crucial skill to build early on.
I guess the knowledge progression I would recommend would look something kind this:
- Learning about Ctrl+F, which works basically everywhere.
- Transitioning to ripgrep https://github.com/BurntSushi/ripgrep - I wouldn't even call this optional, it's truly an incredible and very discoverable tool. Requires keeping a terminal open, but that's a good thing for a newbie!
- Optional, but highly recommended: Learning one of the powerhouse command line editors. Teenage me recommended Emacs; current me recommends vanilla vim, purely because some flavor of it is installed almost everywhere. This is so that you can grep around and edit in the same window.
- In the same vein, moving back from ripgrep and learning about good old fashioned grep, with a few flags rg uses by default: `grep -r` for recursive search, `grep -ri` for case insensitive recursive search, and `grep -ril` for case insensitive recursive "just show me which files this string is found in" search. Some others too, season to taste.
- Finally hitting the wall with what ripgrep can do for you and switching to an actual indexed, dedicated code search tool.
Level Up Your Dev Workflow: Conquer Web Development with a Blazing Fast Neovim Setup (Part 1)
12 projects | dev.to | 16 Mar 2024

live grep: ripgrep
Ripgrep
1 project | news.ycombinator.com | 25 Feb 2024
Modern Java/JVM Build Practices
9 projects | news.ycombinator.com | 4 Jan 2024

The world has moved on though to opinionated tools, and Rust isn't even the furthest in that direction (That would be Go). The equivalent of those two lines in Cargo.toml would be this example of a basic configuration from the jacoco-maven-plugin: https://www.jacoco.org/jacoco/trunk/doc/examples/build/pom.x... - That's 40 lines in the section to do the "defaults".
Yes, you could add a load of config for files to include/exclude from coverage and so on, but the idea that that's a norm is way more common in Java projects than other languages. Like here's some example Cargo.toml files from complicated Rust projects:
Servo: https://github.com/servo/servo/blob/main/Cargo.toml
rust-gdext: https://github.com/godot-rust/gdext/blob/master/godot-core/C...
ripgrep: https://github.com/BurntSushi/ripgrep/blob/master/Cargo.toml
socketio: https://github.com/1c3t3a/rust-socketio/blob/main/socketio/C...
Ugrep – a more powerful, ultra fast, user-friendly, compatible grep
27 projects | news.ycombinator.com | 30 Dec 2023

I'm not clear on why you're seeing the results you are. It could be because your haystack is so small that you're mostly just measuring noise. ripgrep 14 did introduce some optimizations in workloads like this by reducing match overhead, but I don't think it's anything huge in this case. (And I just tried ripgrep 13 on the same commands above and the timings are similar if a tiny bit slower.)
[1]: https://github.com/radare/ired
[2]: https://github.com/BurntSushi/ripgrep/discussions/2597
Tell HN: My Favorite Tools
14 projects | news.ycombinator.com | 24 Dec 2023
Potencializando Sua Experiência no Linux: Conheça as Ferramentas em Rust para um Desenvolvimento Eficiente
5 projects | dev.to | 12 Dec 2023

Explore o Ripgrep no repositório oficial: https://github.com/BurntSushi/ripgrep
Scrybble is the ReMarkable highlights to Obsidian exporter I have been looking for
9 projects | /r/RemarkableTablet | 7 Dec 2023

🔎🗃️ ripgrep or ugrep (search fast, use regex patterns or fuzzy search, pipe output to bash/zsh shell for further processing V coloring)
RFC: Add ngram indexing support to ripgrep (2020)
2 projects | news.ycombinator.com | 30 Nov 2023

What are some alternatives?

When comparing rebar and ripgrep you can also consider the following projects:

Rebar3 - Erlang build tool that makes it easy to compile and test Erlang applications and releases.

telescope-live-grep-args.nvim - Live grep with args

cl-ppcre - Common Lisp regular expression library

fd - A simple, fast and user-friendly alternative to 'find'

hypergrep - Recursively search directories for a regex pattern

ugrep - NEW ugrep 5.1: an ultra fast, user-friendly, compatible grep. Ugrep combines the best features of other grep, adds new features, and searches fast. Includes a TUI and adds Google-like search, fuzzy search, hexdumps, searches nested archives (zip, 7z, tar, pax, cpio), compressed files (gz, Z, bz2, lzma, xz, lz4, zstd, brotli), pdfs, docs, and more

StringZilla - Up to 10x faster strings for C, C++, Python, Rust, and Swift, leveraging SWAR and SIMD on Arm Neon and x86 AVX2 & AVX-512-capable chips to accelerate search, sort, edit distances, alignment scores, etc 🦖

the_silver_searcher - A code-searching tool similar to ack, but faster.

moar - Moar is a pager. It's designed to just do the right thing without any configuration.

fzf - :cherry_blossom: A command-line fuzzy finder

bat - A cat(1) clone with wings.

alacritty - A cross-platform, OpenGL terminal emulator.

rebar vs Rebar3 ripgrep vs telescope-live-grep-args.nvim rebar vs cl-ppcre ripgrep vs fd rebar vs hypergrep ripgrep vs ugrep rebar vs StringZilla ripgrep vs the_silver_searcher rebar vs moar ripgrep vs fzf rebar vs bat ripgrep vs alacritty

Compare rebar vs ripgrep and see what are their differences.

rebar

ripgrep

rebar

ripgrep

What are some alternatives?