ucd-generate
fst
ucd-generate | fst | |
---|---|---|
3 | 11 | |
90 | 1,714 | |
- | - | |
6.6 | 3.5 | |
4 months ago | 4 months ago | |
Rust | Rust | |
Apache License 2.0 | The Unlicense |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
ucd-generate
-
Using unwrap() in Rust is Okay
So you're saying that the 'expect()' message when a regex compilation error occurs should be a translation from a terse domain specific language to bloviating prose? :-)
What 'expect()' message would you write for this regex? https://github.com/BurntSushi/ucd-generate/blob/6d3aae3b8005...
I think 'unwrap()' there is perfectly appropriate.
> I think it'd be desirable to have a `.unwrap_with_context("Context: {}")`, and the you'd get `Context: Inner Panic Info`.
Why?
-
Debian discusses vendoring again
I've also embedded Unicode tables a number of times. It's very easy to do, and I do it enough that I even have a tool to do it. Having tooling and scripts to do it is important for reasons of provenance and also for when the tables need to be updated (every year or so).
-
Announcing chr 1.0.0: A command-line tool that gives information about Unicode characters
ucd-generate for generating Unicode tables. It is what the regex crate uses to generate all of its tables, and it supports many properties already. It also provides a way to represent Unicode character names in a compressed data structure.
fst
- fst: Represent large sets and maps compactly with finite state transducers
-
Creating a perfect HashMap from string keys known in advance
I'd point you towards BurntSushi's fst crate: https://github.com/BurntSushi/fst
-
How to use mmap safely in Rust?
The fst crate effectively relies on mmap for it to work right. The folks here suggesting you just use the heap might be right, but only if using the heap is actually plausible. If your dictionary is GBs big (an FST might be bigger than available memory), then copying it the heap first would be disastrous.
-
Official /r/rust "Who's Hiring" thread for job-seekers and job-offerers [Rust 1.64]
You'll love what we're working on if you're interested in the implementation of:- Tantivy- Meilisearch- Finite State Transducers
-
rustc is unacceptably slow compiling long lists of constant slices
Here's an example of longest prefix matching using a FST which I based my approach on: https://github.com/BurntSushi/fst/pull/104/files
-
Official /r/rust "Who's Hiring" thread for job-seekers and job-offerers [Rust 1.63]
Finite State Transducers
-
Wikit Desktop - A dictionary application using tauri GUI framework
As a result, I have a plan to implement a desktop version from then and I finished today with a beta version. The desktop is based on tauri, and the dictionary index algorithm is FST (it is an awesome index algorithm).
-
WordBueno.com online dictionary. Fast, no frills, mobile friendly.
WordBueno’s data is currently derived from Wiktionary. The backend is using Rust’s warp with fst for indexing.
- Show HN: WordBueno: sleek dictionary built with Rust and Svelte
-
Speed of Rust vs. C
No you don't. I've written multiple programs that load things instantly off the file system via memory maps. See the fst crate[1], for example, which is designed to work with memory maps.
Rust "works badly with memory mapped files" doesn't mean, "Rust can't use memory mapped files." It means, "it is difficult to reconcile Rust's safety story with memory maps." ripgrep for example uses memory maps because they are faster sometimes, and its safety contract[2] is a bit strained. But it works.
[1] - https://github.com/BurntSushi/fst/
[2] - https://docs.rs/grep-searcher/0.1.7/grep_searcher/struct.Mma...
What are some alternatives?
ripgrep-all - rga: ripgrep, but also search in PDFs, E-Books, Office documents, zip, tar.gz, etc.
smartstring - Compact inlined strings for Rust.
itoa - Fast function for printing integer primitives to a decimal string
libskry_r - Lucky imaging library
rust-fnv - Fowler–Noll–Vo hash function
unicode-xid
character - tool for character manipulations
redgrep - ♥ Janusz Brzozowski
perl5 - 🐪 The Perl programming language
tao - The TAO of cross-platform windowing. A library in Rust built for Tauri.