APM, error tracking, and dashboards for modern web apps. Ten-minute setup, transparent flat pricing, and support from engineers who actually use the product. Learn more →
Top 17 Jupyter Notebook Benchmark Projects
-
-
AppSignal
Monitoring that respects your time & budget. APM, error tracking, and dashboards for modern web apps. Ten-minute setup, transparent flat pricing, and support from engineers who actually use the product.
-
llm-colosseum
Benchmark LLMs by fighting in Street Fighter 3! The new way to evaluate the quality of an LLM
-
KernelBench
KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)
-
-
indonlu
The first-ever vast natural language processing benchmark for Indonesian Language. We provide multiple downstream tasks, pre-trained IndoBERT models, and a starter code! (AACL-IJCNLP 2020)
-
SKAB
SKAB - Skoltech Anomaly Benchmark. Time-series data for evaluating Anomaly Detection algorithms.
-
Awesome_Satellite_Benchmark_Datasets
Supplementary material for our paper "THERE IS NO DATA LIKE MORE DATA" is provided.
-
Kargo
Stop Scripting Promotions. Start Shipping with Kargo. Kargo automates promotion across dev, staging, and prod with approval gates and verification. Open source, built by the team behind Argo CD. Download now.
-
-
Project mention: Spaced Repetition Algorithm: A Three‐Day Journey from Novice to Expert | news.ycombinator.com | 2026-03-13
Nice article but small note:
>You might be thinking "If spaced repetition is so effective, why isn't it more popular?"
I might have thought that 30 years ago - but I can't remember the last time I used a flashcard app (phone, Desktop, etc) that didn't include SRS as the default algorithm. These days the new hotness is FSRS [1] which Anki uses.
[1] - https://github.com/open-spaced-repetition/srs-benchmark?tab=...
-
-
food-recognition-benchmark-starter-kit
This repository is the main Food Recognition Benchmark template and Starter kit. Clone the repository to compete now!
-
hashtable-bench
A benchmark for hash tables and hash functions in C++, evaluate on different data as comprehensively as possible
-
H.E.I.M.D.A.L.L
H.E.I.M.D.A.L.L looks at fleet telemetry and gives you natural-language insights. GPU data loading (cuDF), local LLM inference (Gemma 2), and production NIM on GKE. Open the notebooks, run cells, get answers! Quick start should not take longer than 10 minutes and the T4 path is completely free!
Project mention: Show HN: H.e.i.m.d.a.l.l – Telemetry-to-insight pipeline for fleet telemetry | news.ycombinator.com | 2026-02-17 -
-
-
tax-retrieval-benchmark
An implementation of the TaxRetrievalBenchmark task for the 🤗 Massive Text Embedding Benchmark (MTEB) framework.
-
longctx-bench-honest
Honest measurement of 1M-token long-context benchmarks (RULER + LongBench v2 + NIAH) on Qwen2.5-7B-1M local vs GitHub Models cloud. All zero credit card, drift-checked, reproducible.
Project mention: Counterintuitive: WSL2 + vllm cannot fit Qwen2.5-7B-1M on 6GB VRAM where Windows transformers can | dev.to | 2026-05-11Full numbers + 11 JSON evidence cells + 3 ADRs at: https://github.com/leagames0221-sys/longctx-bench-honest
-
SaaSHub
SaaSHub - Software Alternatives and Reviews. SaaSHub helps you find the best software and product alternatives
Jupyter Notebook Benchmark discussion
Jupyter Notebook Benchmark related posts
-
Can LLMs Beat Classical Hyperparameter Optimization Algorithms?
-
Text Embedding Benchmark (2022)
-
Anthropic – Introducing Contextual Retrieval
-
LLM Colosseum
-
Evaluate LLMs in Real Time with Street Fighter III
-
SKAB: NEW Data - star count:238.0
-
SKAB: NEW Data - star count:238.0
-
A note from our sponsor - AppSignal
www.appsignal.com | 4 Sep 2026
Index
What are some of the best open-source Benchmark projects in Jupyter Notebook? This list will help you:
| # | Project | Stars |
|---|---|---|
| 1 | tapnet | 1,971 |
| 2 | llm-colosseum | 1,483 |
| 3 | KernelBench | 1,220 |
| 4 | human-learn | 833 |
| 5 | indonlu | 655 |
| 6 | SKAB | 403 |
| 7 | Awesome_Satellite_Benchmark_Datasets | 362 |
| 8 | tf-metal-experiments | 279 |
| 9 | srs-benchmark | 258 |
| 10 | benchmarks | 172 |
| 11 | food-recognition-benchmark-starter-kit | 69 |
| 12 | hashtable-bench | 26 |
| 13 | H.E.I.M.D.A.L.L | 17 |
| 14 | SciTS | 14 |
| 15 | file-format-benchmark | 3 |
| 16 | tax-retrieval-benchmark | 1 |
| 17 | longctx-bench-honest | 0 |