ai-safety

Open-source projects categorized as ai-safety

Top 23 ai-safety Open-Source Projects

ai-safety
  1. iFixAi

    Independent Auditing of AI Agents. Run by human or the agent itself, to answer the most crucial question in the AI Agent Economy. Is the agent doing what is supposed to do? With iFixAi you can have this answer in less than 120 seconds.

    Project mention: iFixAi,open-source auditor that checks if your AI agent does its job | news.ycombinator.com | 2026-08-12
  2. AppSignal

    Monitoring that respects your time & budget. APM, error tracking, and dashboards for modern web apps. Ten-minute setup, transparent flat pricing, and support from engineers who actually use the product.

    AppSignal logo
  3. agent-governance-toolkit

    AI Agent Governance Toolkit — Policy enforcement, zero-trust identity, execution sandboxing, and reliability engineering for autonomous AI agents. Covers 10/10 OWASP Agentic Top 10.

    Project mention: The Two Layers That Decide Whether Your Agent Survives | dev.to | 2026-08-08

    Microsoft, Agent Framework overview and the Agent Governance Toolkit

  4. WFGY

    Verification-first reasoning engine for LLMs, with reproducible demos and audit-oriented specifications. Includes WFGY 3.0 Singularity Demo (public spec) and engineering failure maps for real systems.

    Project mention: Show HN: A text-only reasoning core for LLMs (MIT, system prompt and self-test) | news.ycombinator.com | 2026-02-13
  5. safe-rlhf

    Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback

  6. uqlm

    [JMLR 2026] "UQLM: A Python Package for Uncertainty Quantification in Large Language Models"

    Project mention: Show HN: UQLM – Closed-book hallucination detection with UQ | news.ycombinator.com | 2026-06-01
  7. Internal-Safety-Collapse

    We built an adversarial codespace setup. Place any AI agent into a normal workflow inside it, and the agent will fill in whatever is missing.

    Project mention: A Big Alignment Loophole of Current Froniter LLMs | news.ycombinator.com | 2026-04-03

    I want to share this finding because I think both developers building on LLMs and normal users need to be aware. This is real — I've included live demos as proof so you can see it happening, not just take my word for it:

    85 reproducible prompt if you want to try it yourself: https://github.com/wuyoscar/ISC-Bench

  8. langtest

    Deliver safe & effective language models

  9. Kargo

    Stop Scripting Promotions. Start Shipping with Kargo. Kargo automates promotion across dev, staging, and prod with approval gates and verification. Open source, built by the team behind Argo CD. Download now.

    Kargo logo
  10. cordum

    The action firewall for AI agents. Enforce policy and human approval before risky tool calls, shell commands, workflows, and production changes, with auditable evidence.

    Project mention: A silent failure looks exactly like a feature you never built | dev.to | 2026-08-26

    And Cordum's ADR-010 contains the cleanest statement of the failure class I know, about a Claude Code PreToolUse deny hook: HTTP hooks can deny with a 2xx JSON response, but connection failures, non-2xx responses and timeouts are non-blocking. A guard whose failure mode is "allow" is not a guard. That is my unbound Mac shortcut with higher stakes.

  11. reins

    Stop AI agents from doing things you didn't ask for.

  12. tiger

    Open Source LLM toolkit to build trustworthy LLM applications. TigerArmor (AI safety), TigerRAG (embedding, RAG), TigerTune (fine-tuning) (by tigerlab-ai)

  13. orbit

    Self-hosted AI gateway for private RAG, natural-language data access, and tool-calling agents. (by schmitech)

    Project mention: Open-Source AI Platform Orbit | news.ycombinator.com | 2026-08-01
  14. Aegis

    Runtime policy enforcement for AI agents. Cryptographic audit trail, human-in-the-loop approvals, kill switch. Zero code changes. (by Justin0504)

    Project mention: Show HN: Mcpsnoop – Wireshark for MCP (transparent proxy and live TUI) | news.ycombinator.com | 2026-07-03

    Remote debugging and post-mortem debugging support might be useful.

    There are many AI auditability proxies;

    awesome-auditable-ai: "A curated list of papers, tools, datasets, benchmarks, and standards for building, evaluating, and auditing reliable AI agents" https://github.com/yzhao062/awesome-auditable-ai

    Aegis and LiteLLM, for example, are pre-execution firewalls that add a cryptographic audit trail. https://github.com/Justin0504/aegis

  15. ethics

    Aligning AI With Shared Human Values (ICLR 2021)

  16. agent-control

    Centralized agent control plane for governing runtime agent behavior at scale. Configurable, extensible, and production-ready.

    Project mention: Reading list (29th March to April 20th) | dev.to | 2026-04-20

    Enforce runtime guardrails through a centralized control layer—configure once and apply across all agents. Agent Control evaluates inputs and outputs against configurable rules to block prompt injections, PII leakage, and other risks without changing your agent’s code. - link [tool] - ( Added: 2026-03-31 08:12:52 )

  17. Thought-Cloning

    [NeurIPS '23 Spotlight] Thought Cloning: Learning to Think while Acting by Imitating Human Thinking

  18. awesome-ai-safety

    📚 A curated list of papers & technical articles on AI Quality & Safety

  19. node9-proxy

    The Execution Security Layer for the Agentic Era. Providing deterministic "Sudo" governance and audit logs for autonomous AI agents.

    Project mention: Running Hermes Agent in the Cloud Safely: A Reader's Guide to Their Trust Model | dev.to | 2026-06-10

    If you want the in-process gate to be sharper, you can layer one on. This is where Node9 fits in a Hermes deployment: an AST-based policy engine that parses shell commands the way the OS does (not the way regex does), so obfuscated payloads (echo "Y3VybCAuLi4="| base64 -d | bash) collapse into their actual execution graph before the approval decision is made. It also runs a per-call inspection layer that catches credentials in outbound arguments, anomalously large payloads, and force-push patterns that simple denylists miss. The AST-parsing approach is covered in detail in Why Regex Is Not Enough.

  20. ToolEmu

    [ICLR'24 Spotlight] A language model (LM)-based emulation framework for identifying the risks of LM agents with tool use

  21. Doberman-Core

    Your AI's guard dog. Doberman sits at runtime, gating every input, output and tool call to stop unsafe or unintended actions before they execute.

    Project mention: Same curl, two verdicts: taint tracking for coding agents | dev.to | 2026-09-09
  22. phantasm

    Toolkits to create a human-in-the-loop approval layer to monitor and guide AI agents workflow in real-time.

  23. make-safe-ai

    How to Make Safe AI? Let's Discuss! 💡|💬|🙌|📚

  24. Hegelion

    Dialectical reasoning architecture for LLMs (Thesis → Antithesis → Synthesis)

    Project mention: Show HN: Hegelion – Force your LLM to argue with itself before answering | news.ycombinator.com | 2025-11-24
  25. diagnostic

    iFixAi. The open-source diagnostic for AI misalignment. 32 tests across fabrication, manipulation, deception, unpredictability, and opacity. Provider-agnostic. Runs against OpenAI, Anthropic, Bedrock, Azure, Gemini, and more. Letter grade in under 5 minutes, content-addressed manifest for bit-identical replay. Built by iMe.

    Project mention: Open-source diagnostic for Al misalignment. Model agnostic, industry agnostic | news.ycombinator.com | 2026-05-04
  26. SaaSHub

    SaaSHub - Software Alternatives and Reviews. SaaSHub helps you find the best software and product alternatives

    SaaSHub logo
NOTE: The open source projects on this list are ordered by number of github stars. The number of mentions indicates repo mentiontions in the last 12 Months or since we started tracking (Dec 2020).

ai-safety discussion

Log in or Post with

ai-safety related posts

  • Same curl, two verdicts: taint tracking for coding agents

    1 project | dev.to | 9 Sep 2026
  • Show HN: GuardRail, shell guards that stop Claude Code before it pushes to main

    1 project | news.ycombinator.com | 9 Sep 2026
  • I let Claude Code run my company at night. These 13 free guards are why I can sleep.

    2 projects | dev.to | 9 Sep 2026
  • I built an AI agent that forges its own tools mid-task — and asks first (open source, 60-second demo)

    1 project | dev.to | 8 Sep 2026
  • Four verdicts instead of "done": grading an AI agent's claims on an evidence ladder

    1 project | dev.to | 6 Sep 2026
  • How I benchmark an agent guardrail without lying to myself

    1 project | dev.to | 6 Sep 2026
  • How execution boundaries reduce the blast radius of AI agent mistakes

    1 project | dev.to | 5 Sep 2026
  • A note from our sponsor - AppSignal
    www.appsignal.com | 13 Sep 2026
    APM, error tracking, and dashboards for modern web apps. Ten-minute setup, transparent flat pricing, and support from engineers who actually use the product. Learn more →

Index

What are some of the best open-source ai-safety projects? This list will help you:

# Project Stars
1 iFixAi 14,200
2 agent-governance-toolkit 6,229
3 WFGY 1,788
4 safe-rlhf 1,615
5 uqlm 1,198
6 Internal-Safety-Collapse 1,183
7 langtest 559
8 cordum 504
9 reins 408
10 tiger 404
11 orbit 344
12 Aegis 340
13 ethics 324
14 agent-control 305
15 Thought-Cloning 268
16 awesome-ai-safety 218
17 node9-proxy 210
18 ToolEmu 209
19 Doberman-Core 198
20 phantasm 195
21 make-safe-ai 172
22 Hegelion 172
23 diagnostic 160

Sponsored
Monitoring that respects your time & budget
APM, error tracking, and dashboards for modern web apps. Ten-minute setup, transparent flat pricing, and support from engineers who actually use the product.
www.appsignal.com