db-benchmark vs db-benchmark

db-benchmark

reproducible benchmark of database-like ops (by duckdblabs)

Suggest topics

Source Code

duckdblabs.github.io

Suggest alternative

Edit details

db-benchmark

reproducible benchmark of database-like ops (by h2oai)

Suggest topics

Source Code

h2oai.github.io

Suggest alternative

Edit details

InfluxDB - Power Real-Time Data Analytics at Scale

Get real-time insights from all types of time series data with InfluxDB. Ingest, query, and analyze billions of data points in real-time with unbounded cardinality.

www.influxdata.com

featured

SaaSHub - Software Alternatives and Reviews

SaaSHub helps you find the best software and product alternatives

www.saashub.com

featured

db-benchmark		db-benchmark
	Project
12	Mentions	91
124	Stars	320
6.5%	Growth	0.0%
8.0	Activity	0.0
4 months ago	Latest Commit	11 months ago
R	Language	R
Mozilla Public License 2.0	License	Mozilla Public License 2.0

The number of mentions indicates the total number of mentions that we've tracked plus the number of user suggested alternatives.
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.

db-benchmark

Posts with mentions or reviews of db-benchmark. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2024-01-08.

Database-Like Ops Benchmark
1 project | news.ycombinator.com | 9 Mar 2024

1 project | news.ycombinator.com | 16 Jan 2024

1 project | news.ycombinator.com | 7 Dec 2023
Polars
11 projects | news.ycombinator.com | 8 Jan 2024

DuckDB maintains a benchmark of open source database-like tools, including Polars and Pandas
https://duckdblabs.github.io/db-benchmark/
Planning a New Benchmarking for Comparing Filter2Groupby for 3,000 Files (100,000 Rows/Files)
1 project | /r/Python | 6 Jun 2023

1 project | /r/datascience | 6 Jun 2023
Pandas vs. Julia – cheat sheet and comparison
7 projects | news.ycombinator.com | 17 May 2023
Polars supports SQL statement in Python Plus CLI Verion (Polars.exe 24.4MB)
1 project | /r/Python | 11 May 2023

DuckDB is also a SQL/Python app, refer to this benchmark, seem it run very fast https://duckdblabs.github.io/db-benchmark/
The Return of the H2o.ai Database-Like Ops Benchmark
1 project | news.ycombinator.com | 21 Apr 2023
I discovered that the fastest way to create a Pandas DataFrame from a CSV file is to actually use Polars
2 projects | /r/Python | 15 Apr 2023

db-benchmark

Posts with mentions or reviews of db-benchmark. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2024-01-08.

Database-Like Ops Benchmark
1 project | news.ycombinator.com | 28 Jan 2024
Polars
11 projects | news.ycombinator.com | 8 Jan 2024

Real-world performance is complicated since data science covers a lot of use cases.
If you're just reading a small CSV to do analysis on it, then there will be no human-perceptible difference between Polars and Pandas. If you're reading a larger CSV with 100k rows, there still won't be much of a perceptible difference.
Per this (old) benchmark, there are differences once you get into 500MB+ territory: https://h2oai.github.io/db-benchmark/
DuckDB performance improvements with the latest release
8 projects | news.ycombinator.com | 6 Nov 2023

I do think it was important for duckdb to put out a new version of the results as the earlier version of that benchmark [1] went dormant with a very old version of duckdb with very bad performance, especially against polars.
[1] https://h2oai.github.io/db-benchmark/
Show HN: SimSIMD vs. SciPy: How AVX-512 and SVE make SIMD cleaner and ML faster
16 projects | news.ycombinator.com | 7 Oct 2023

https://news.ycombinator.com/item?id=33270638 :
> Apache Ballista and Polars do Apache Arrow and SIMD.
> The Polars homepage links to the "Database-like ops benchmark" of {Polars, data.table, DataFrames.jl, ClickHouse, cuDF, spark, (py)datatable, dplyr, pandas, dask, Arrow, DuckDB, Modin,} but not yet PostgresML? https://h2oai.github.io/db-benchmark/ *
LLM -> Vector database: https://en.wikipedia.org/wiki/Vector_database
/? inurl:awesome site:github.com "vector database"
Pandas vs. Julia – cheat sheet and comparison
7 projects | news.ycombinator.com | 17 May 2023

I agree with your conclusion but want to add that switching from Julia may not make sense either.
According to these benchmarks: https://h2oai.github.io/db-benchmark/, DF.jl is the fastest library for some things, data.table for others, polars for others. Which is fastest depends on the query and whether it takes advantage of the features/properties of each.
For what it's worth, data.table is my favourite to use and I believe it has the nicest ergonomics of the three I spoke about.
Any faster Python alternatives?
6 projects | /r/learnprogramming | 12 Apr 2023

Same. Numba does wonders for me in most scenarios. Yesterday I've discovered pola-rs and looks like I will add it to the stack. It's API is similar to pandas. Have a look at the benchmarks of cuDF, spark, dask, pandas compared to it: Benchmarks
Pandas 2.0 (with pyarrow) vs Pandas 1.3 - Performance comparison
1 project | /r/datascience | 8 Apr 2023

The syntax has similarities with dplyr in terms of the way you chain operations, and it’s around an order of magnitude faster than pandas and dplyr (there’s a nice benchmark here). It’s also more memory-efficient and can handle larger-than-memory datasets via streaming if needed.
Pandas v2.0 Released
5 projects | news.ycombinator.com | 3 Apr 2023

If interested in benchmarks comparing different dataframe implementations, here is one:
https://h2oai.github.io/db-benchmark/
Database-like ops benchmark
1 project | /r/dataengineering | 16 Feb 2023
Python "programmers" when I show them how much faster their naive code runs when translated to C++ (this is a joke, I love python)
2 projects | /r/ProgrammerHumor | 17 Jan 2023

Bad examples. Both numpy and pandas are notoriously un-optimized packages, losing handily to pretty much all their competitors (R, Julia, kdb+, vaex, polars). See https://h2oai.github.io/db-benchmark/ for a partial comparison.

What are some alternatives?

When comparing db-benchmark and db-benchmark you can also consider the following projects:

Tidier.jl - Meta-package for data analysis in Julia, modeled after the R tidyverse.

polars - Dataframes powered by a multithreaded, vectorized query engine, written in Rust

DataFramesMeta.jl - Metaprogramming tools for DataFrames

datafusion - Apache DataFusion SQL Query Engine

Apache Arrow - Apache Arrow is a multi-language toolbox for accelerated data interchange and in-memory processing

databend - 𝗗𝗮𝘁𝗮, 𝗔𝗻𝗮𝗹𝘆𝘁𝗶𝗰𝘀 & 𝗔𝗜. Modern alternative to Snowflake. Cost-effective and simple for massive-scale analytics. https://databend.com

sktime - A unified framework for machine learning with time series

arrow2 - Transmute-free Rust library to work with the Arrow format

disk.frame - Fast Disk-Based Parallelized Data Manipulation Framework for Larger-than-RAM Data

datatable - A Python package for manipulating 2-dimensional tabular data structures

DataFrame - C++ DataFrame for statistical, Financial, and ML analysis -- in modern C++ using native types and contiguous memory storage

db-benchmark vs Tidier.jl db-benchmark vs polars db-benchmark vs DataFramesMeta.jl db-benchmark vs datafusion db-benchmark vs Apache Arrow db-benchmark vs databend db-benchmark vs sktime db-benchmark vs DataFramesMeta.jl db-benchmark vs arrow2 db-benchmark vs disk.frame db-benchmark vs datatable db-benchmark vs DataFrame

Compare db-benchmark vs db-benchmark and see what are their differences.

db-benchmark

db-benchmark

db-benchmark

db-benchmark

What are some alternatives?