fluvio
datafusion
fluvio | datafusion | |
---|---|---|
26 | 55 | |
2,663 | 5,086 | |
3.1% | 5.2% | |
9.5 | 9.9 | |
3 days ago | 3 days ago | |
Rust | Rust | |
Apache License 2.0 | Apache License 2.0 |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
fluvio
- Ask HN: WebSocket Relay?
- XFaaS: Hyperscale and Low Cost Serverless Functions at Meta
-
Iggy.rs – building message streaming in Rust
I'm not quite sure how this compares to Kafka and fluvio [1], a Kafka competitor also written in Rust?
Is it more of a message queue like rabbitmq?
[1] https://www.fluvio.io/
- Fluvio: Open-source data streaming platform
- Show HN: Fluvio – Distributed stream processing system written in Rust and WASM
-
Thank you for checking out the Fluvio repo
Seeing signs of the hockey stick traffic on the Fluvio Open Source repo: https://github.com/infinyon/fluvio
Stopped the mobile notifications for the stargazer bot on discord and slack to stop the dopamine rush!!!
But thank you for checking us out.
- Opens Source Rust and WASM Alternative to Kafka and Flink
- Fluvio is a high-performance distributed data streaming platform in Rust
- RabbitMQ vs. Kafka – An Architect’s Dilemma (Part 1)
-
Advice: I am 36, and i want to transition from ETL developer to Rust
Look at open source project fluvio https://github.com/infinyon/fluvio and see how we’re building a modern bus that can do ETL streaming style :).
datafusion
-
Velox: Meta's Unified Execution Engine [pdf]
Python's Substrait seems like the biggest/most-used competitor-ish out there. I'd love some compare & contrast; my sense is that Substrait has a smaller ambition, and more wants to be a language for talking about execution rather than a full on execution engine. https://github.com/substrait-io/substrait
We can also see from the DataFusion discussion that they too see themselves as a bit of a Velox competitor. https://github.com/apache/arrow-datafusion/discussions/6441
-
What I Talk About When I Talk About Query Optimizer (Part 1): IR Design
Agree, substrait is a really cool project! Related: if you like substrait you might want to check out datafusion too. The project is a query execution engine built on top of Apache Arrow (with SQL parser, query planner & optimizer, execution engine, extensible user defined functions, among others) and it implements a substrait provider and consumer: https://github.com/apache/arrow-datafusion/tree/main/datafus...
-
DuckDB performance improvements with the latest release
The draft contains some preliminary benchmark results, comparing it to DuckDB.
https://github.com/apache/arrow-datafusion/issues/6782
- Apache Arrow DataFusion
-
GlareDB: An open source SQL database to query and analyze distributed data
Apache Arrow is a pretty common memory structure these days. Datafusion is an open query engine built in Rust started by Andy Grove.
-
DuckDB 0.8.0
DuckDB is a great piece of software if you are
If you are looking for a query engine implemented in a safe language (Rust) I definitely suggest checking out DataFusion. It is comparable to DuckDB in performance, has all the standard built in SQL functionality, and is extensible in pretty much all areas (query language, data formats, catalogs, user defined functions, etc)
https://arrow.apache.org/datafusion/
Disclaimer I am a maintainer of DataFusion
-
Data Engineering with Rust
https://github.com/jorgecarleitao/arrow2 https://github.com/apache/arrow-datafusion https://github.com/apache/arrow-ballista https://github.com/pola-rs/polars https://github.com/duckdb/duckdb
- Polars: Computing a new column from multiple columns - there must be a better way
-
Bridging Async and Sync Rust Code - A lesson learned while working with Tokio
Problem comes when you want to do this inside an async context since we couldn't block an async task. https://users.rust-lang.org/t/sync-function-invoking-async/43364/6 You might need to do it in another runtime/thread. It is not recommended to do this, but sometimes it is unavoidable while implementing a third-party trait. https://github.com/apache/arrow-datafusion/issues/3777 However, I believe this isn't a problem particular to tokio, or any specific runtime.
- Using Rust to write a Data Pipeline. Thoughts. Musings.
What are some alternatives?
nsq - A realtime distributed messaging platform
polars - Dataframes powered by a multithreaded, vectorized query engine, written in Rust
datafuse - An elastic and reliable Cloud Warehouse, offers Blazing Fast Query and combines Elasticity, Simplicity, Low cost of the Cloud, built to make the Data Cloud easy [Moved to: https://github.com/datafuselabs/databend]
ClickHouse - ClickHouse® is a free analytics DBMS for big data
roapi - Create full-fledged APIs for slowly moving datasets without writing a single line of code.
databend - 𝗗𝗮𝘁𝗮, 𝗔𝗻𝗮𝗹𝘆𝘁𝗶𝗰𝘀 & 𝗔𝗜. Modern alternative to Snowflake. Cost-effective and simple for massive-scale analytics. https://databend.com
sccache - Sccache is a ccache-like tool. It is used as a compiler wrapper and avoids compilation when possible. Sccache has the capability to utilize caching in remote storage environments, including various cloud storage options, or alternatively, in local storage.
db-benchmark - reproducible benchmark of database-like ops
rust-blog - Educational blog posts for Rust beginners
duckdb - DuckDB is an in-process SQL OLAP Database Management System
cache - Cache dependencies and build outputs in GitHub Actions
nushell - A new type of shell