Skip to content
#

llm-safety-benchmark

Here are 9 public repositories matching this topic...

Policy-conformance harness for money-touching AI agents — catches over-promises against refund policy, with mechanically verified evidence (Python-derived labels, span-verified judge citations, frozen agent under test).

  • Updated Sep 3, 2026
  • Python

Add this topic to your repo

To associate your repository with the llm-safety-benchmark topic, visit your repo's landing page and select "manage topics."

Learn more