Promptbench Alternatives

Similar projects and alternatives to promptbench

k3s

291 26,483 9.6 Go promptbench VS k3s

Lightweight Kubernetes
FLiPStackWeekly

80 14 9.9 promptbench VS FLiPStackWeekly

FLaNK AI Weekly covering Apache NiFi, Apache Flink, Apache Kafka, Apache Spark, Apache Iceberg, Apache Ozone, Apache Pulsar, and more...
WorkOS

workos.com sponsored

The modern identity platform for B2B SaaS. The APIs are flexible and easy-to-use, supporting authentication, user identity, and complex enterprise features like SSO and SCIM provisioning.
yq

66 10,802 9.2 Go promptbench VS yq

yq is a portable command-line YAML, JSON, XML, CSV, TOML and properties processor
evals

49 13,878 9.4 Python promptbench VS evals

Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
ydata-profiling

43 12,053 8.5 Python promptbench VS ydata-profiling

1 Line of code data quality profiling & exploratory data analysis for Pandas and Spark DataFrames.
lm-evaluation-harness

34 4,957 9.9 Python promptbench VS lm-evaluation-harness

A framework for few-shot evaluation of language models.
seatunnel

31 7,223 9.8 Java promptbench VS seatunnel

SeaTunnel is a next-generation super high-performance, distributed, massive data integration tool.
InfluxDB

www.influxdata.com sponsored

Power Real-Time Data Analytics at Scale. Get real-time insights from all types of time series data with InfluxDB. Ingest, query, and analyze billions of data points in real-time with unbounded cardinality.
osgameclones

17 1,623 9.5 JavaScript promptbench VS osgameclones

Open Source Clones of Popular Games
rewrite

24 1,830 9.9 Java promptbench VS rewrite

Automated mass refactoring of source code.
Stirling-PDF

20 21,832 9.9 Java promptbench VS Stirling-PDF

#1 Locally hosted web application that allows you to perform various operations on PDF files
TornadoVM

22 1,105 9.9 Java promptbench VS TornadoVM

TornadoVM: A practical and efficient heterogeneous programming framework for managed languages
qsv

13 2,228 9.9 Rust promptbench VS qsv

CSVs sliced, diced & analyzed.
kubernetes-client

11 3,307 9.7 Java promptbench VS kubernetes-client

Java client for Kubernetes & OpenShift
JavaOnRaspberryPi

3 71 5.6 Java promptbench VS JavaOnRaspberryPi

Sources and scripts for the book "Getting started with Java on the Raspberry Pi"
awesome-gpt-prompt-engineering

18 784 7.1 Python promptbench VS awesome-gpt-prompt-engineering

A curated list of awesome resources, tools, and other shiny things for GPT prompt engineering.
tbls

6 3,068 8.8 Go promptbench VS tbls

tbls is a CI-Friendly tool for document a database, written in Go.
FLaNK-EveryTransitSystem

8 3 6.1 promptbench VS FLaNK-EveryTransitSystem

Every transit system
FLaNK-Ice

8 1 6.0 promptbench VS FLaNK-Ice

Apache Iceberg - Cloud Data Lakehouse
Zolver

2 143 0.0 Python promptbench VS Zolver

Automatic jigsaw puzzle solver
opencompass

1 2,481 9.7 Python promptbench VS opencompass

OpenCompass is an LLM evaluation platform, supporting a wide range of models (InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.
SaaSHub

www.saashub.com sponsored

SaaSHub - Software Alternatives and Reviews. SaaSHub helps you find the best software and product alternatives

NOTE: The number of mentions on this list indicates mentions on common posts plus user suggested alternatives. Hence, a higher number means a better promptbench alternative or higher similarity.

Suggest an alternative to promptbench

promptbench reviews and mentions

Posts with mentions or reviews of promptbench. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2024-02-13.

Show HN: Times faster LLM evaluation with Bayesian optimization
6 projects | news.ycombinator.com | 13 Feb 2024

Fair question.
Evaluate refers to the phase after training to check if the training is good.
Usually the flow goes training -> evaluation -> deployment (what you called inference). This project is aimed for evaluation. Evaluation can be slow (might even be slower than training if you're finetuning on a small domain specific subset)!
So there are [quite](https://github.com/microsoft/promptbench) [a](https://github.com/confident-ai/deepeval) [few](https://github.com/openai/evals) [frameworks](https://github.com/EleutherAI/lm-evaluation-harness) working on evaluation, however, all of them are quite slow, because LLM are slow if you don't have infinite money. [This](https://github.com/open-compass/opencompass) one tries to speed up by parallelizing on multiple computers, but none of them takes advantage of the fact that many evaluation queries might be similar and all try to evaluate on all given queries. And that's where this project might come in handy.
FLaNK Weekly 31 December 2023
25 projects | dev.to | 31 Dec 2023
FLaNK 25 December 2023
33 projects | dev.to | 26 Dec 2023
Promptbench: A Unified Library for Evaluating and Understanding LLMs
1 project | news.ycombinator.com | 25 Dec 2023
A note from our sponsor - SaaSHub
www.saashub.com | 29 Apr 2024

SaaSHub helps you find the best software and product alternatives Learn more →