-
hermes-workspace
Does your AI agent actually follow rules? 13 pre-registered experiments + 5-layer verification architecture. Paper, data, code — all public.
hermes-workspace: n=30 causal experiment (p=0.0092)
-
SaaSHub
SaaSHub - Software Alternatives and Reviews. SaaSHub helps you find the best software and product alternatives
-
training-gate: Model-layer quality gates
-
PR #778: Metric submitted to HF evaluate
NOTE:
The number of mentions on this list indicates mentions on common posts plus user suggested alternatives.
Hence, a higher number means a more popular project.
Related posts
-
Your Feedback Made This Better — Here's What Changed
-
I Told My AI "You're Safe to Say I Don't Know." Then I Measured What Changed — With Logprobs.
-
My Experiment Showed Zero Effect. A Statistician Told Me My Measurement Was Broken.
-
[D] The MMSegmentation library from OpenMMLab appears to return the wrong results when computing basic image segmentation metrics such as the Jaccard index (IoU - intersection-over-union). It appears to compute recall (sensitivity) instead of IoU, which artificially inflates the performance metrics.
-
[P] Releasing 🤗 Evaluate - an evaluation library for ML