APM, error tracking, and dashboards for modern web apps. Ten-minute setup, transparent flat pricing, and support from engineers who actually use the product. Learn more →
Top 23 ai-safety Open-Source Projects
-
iFixAi
Independent Auditing of AI Agents. Run by human or the agent itself, to answer the most crucial question in the AI Agent Economy. Is the agent doing what is supposed to do? With iFixAi you can have this answer in less than 120 seconds.
Project mention: iFixAi,open-source auditor that checks if your AI agent does its job | news.ycombinator.com | 2026-08-12 -
AppSignal
Monitoring that respects your time & budget. APM, error tracking, and dashboards for modern web apps. Ten-minute setup, transparent flat pricing, and support from engineers who actually use the product.
-
agent-governance-toolkit
AI Agent Governance Toolkit — Policy enforcement, zero-trust identity, execution sandboxing, and reliability engineering for autonomous AI agents. Covers 10/10 OWASP Agentic Top 10.
Microsoft, Agent Framework overview and the Agent Governance Toolkit
-
WFGY
Verification-first reasoning engine for LLMs, with reproducible demos and audit-oriented specifications. Includes WFGY 3.0 Singularity Demo (public spec) and engineering failure maps for real systems.
Project mention: Show HN: A text-only reasoning core for LLMs (MIT, system prompt and self-test) | news.ycombinator.com | 2026-02-13 -
safe-rlhf
Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback
-
Project mention: Show HN: UQLM – Closed-book hallucination detection with UQ | news.ycombinator.com | 2026-06-01
-
Internal-Safety-Collapse
We built an adversarial codespace setup. Place any AI agent into a normal workflow inside it, and the agent will fill in whatever is missing.
Project mention: A Big Alignment Loophole of Current Froniter LLMs | news.ycombinator.com | 2026-04-03I want to share this finding because I think both developers building on LLMs and normal users need to be aware. This is real — I've included live demos as proof so you can see it happening, not just take my word for it:
85 reproducible prompt if you want to try it yourself: https://github.com/wuyoscar/ISC-Bench
-
-
Kargo
Stop Scripting Promotions. Start Shipping with Kargo. Kargo automates promotion across dev, staging, and prod with approval gates and verification. Open source, built by the team behind Argo CD. Download now.
-
cordum
The action firewall for AI agents. Enforce policy and human approval before risky tool calls, shell commands, workflows, and production changes, with auditable evidence.
Project mention: A silent failure looks exactly like a feature you never built | dev.to | 2026-08-26And Cordum's ADR-010 contains the cleanest statement of the failure class I know, about a Claude Code PreToolUse deny hook: HTTP hooks can deny with a 2xx JSON response, but connection failures, non-2xx responses and timeouts are non-blocking. A guard whose failure mode is "allow" is not a guard. That is my unbound Mac shortcut with higher stakes.
-
-
tiger
Open Source LLM toolkit to build trustworthy LLM applications. TigerArmor (AI safety), TigerRAG (embedding, RAG), TigerTune (fine-tuning) (by tigerlab-ai)
-
orbit
Self-hosted AI gateway for private RAG, natural-language data access, and tool-calling agents. (by schmitech)
-
Aegis
Runtime policy enforcement for AI agents. Cryptographic audit trail, human-in-the-loop approvals, kill switch. Zero code changes. (by Justin0504)
Project mention: Show HN: Mcpsnoop – Wireshark for MCP (transparent proxy and live TUI) | news.ycombinator.com | 2026-07-03Remote debugging and post-mortem debugging support might be useful.
There are many AI auditability proxies;
awesome-auditable-ai: "A curated list of papers, tools, datasets, benchmarks, and standards for building, evaluating, and auditing reliable AI agents" https://github.com/yzhao062/awesome-auditable-ai
Aegis and LiteLLM, for example, are pre-execution firewalls that add a cryptographic audit trail. https://github.com/Justin0504/aegis
-
-
agent-control
Centralized agent control plane for governing runtime agent behavior at scale. Configurable, extensible, and production-ready.
Enforce runtime guardrails through a centralized control layer—configure once and apply across all agents. Agent Control evaluates inputs and outputs against configurable rules to block prompt injections, PII leakage, and other risks without changing your agent’s code. - link [tool] - ( Added: 2026-03-31 08:12:52 )
-
Thought-Cloning
[NeurIPS '23 Spotlight] Thought Cloning: Learning to Think while Acting by Imitating Human Thinking
-
-
node9-proxy
The Execution Security Layer for the Agentic Era. Providing deterministic "Sudo" governance and audit logs for autonomous AI agents.
Project mention: Running Hermes Agent in the Cloud Safely: A Reader's Guide to Their Trust Model | dev.to | 2026-06-10If you want the in-process gate to be sharper, you can layer one on. This is where Node9 fits in a Hermes deployment: an AST-based policy engine that parses shell commands the way the OS does (not the way regex does), so obfuscated payloads (echo "Y3VybCAuLi4="| base64 -d | bash) collapse into their actual execution graph before the approval decision is made. It also runs a per-call inspection layer that catches credentials in outbound arguments, anomalously large payloads, and force-push patterns that simple denylists miss. The AST-parsing approach is covered in detail in Why Regex Is Not Enough.
-
ToolEmu
[ICLR'24 Spotlight] A language model (LM)-based emulation framework for identifying the risks of LM agents with tool use
-
Doberman-Core
Your AI's guard dog. Doberman sits at runtime, gating every input, output and tool call to stop unsafe or unintended actions before they execute.
-
phantasm
Toolkits to create a human-in-the-loop approval layer to monitor and guide AI agents workflow in real-time.
-
-
Project mention: Show HN: Hegelion – Force your LLM to argue with itself before answering | news.ycombinator.com | 2025-11-24
-
diagnostic
iFixAi. The open-source diagnostic for AI misalignment. 32 tests across fabrication, manipulation, deception, unpredictability, and opacity. Provider-agnostic. Runs against OpenAI, Anthropic, Bedrock, Azure, Gemini, and more. Letter grade in under 5 minutes, content-addressed manifest for bit-identical replay. Built by iMe.
Project mention: Open-source diagnostic for Al misalignment. Model agnostic, industry agnostic | news.ycombinator.com | 2026-05-04 -
SaaSHub
SaaSHub - Software Alternatives and Reviews. SaaSHub helps you find the best software and product alternatives
ai-safety discussion
ai-safety related posts
-
Same curl, two verdicts: taint tracking for coding agents
-
Show HN: GuardRail, shell guards that stop Claude Code before it pushes to main
-
I let Claude Code run my company at night. These 13 free guards are why I can sleep.
-
I built an AI agent that forges its own tools mid-task — and asks first (open source, 60-second demo)
-
Four verdicts instead of "done": grading an AI agent's claims on an evidence ladder
-
How I benchmark an agent guardrail without lying to myself
-
How execution boundaries reduce the blast radius of AI agent mistakes
-
A note from our sponsor - AppSignal
www.appsignal.com | 13 Sep 2026
Index
What are some of the best open-source ai-safety projects? This list will help you:
| # | Project | Stars |
|---|---|---|
| 1 | iFixAi | 14,200 |
| 2 | agent-governance-toolkit | 6,229 |
| 3 | WFGY | 1,788 |
| 4 | safe-rlhf | 1,615 |
| 5 | uqlm | 1,198 |
| 6 | Internal-Safety-Collapse | 1,183 |
| 7 | langtest | 559 |
| 8 | cordum | 504 |
| 9 | reins | 408 |
| 10 | tiger | 404 |
| 11 | orbit | 344 |
| 12 | Aegis | 340 |
| 13 | ethics | 324 |
| 14 | agent-control | 305 |
| 15 | Thought-Cloning | 268 |
| 16 | awesome-ai-safety | 218 |
| 17 | node9-proxy | 210 |
| 18 | ToolEmu | 209 |
| 19 | Doberman-Core | 198 |
| 20 | phantasm | 195 |
| 21 | make-safe-ai | 172 |
| 22 | Hegelion | 172 |
| 23 | diagnostic | 160 |