🔬 Repository for Claude 5.0 adversarial testing, safety alignment research, and LLM vulnerability documentation. Educational evaluation of safety guardrails.
-
Updated
Aug 12, 2026
🔬 Repository for Claude 5.0 adversarial testing, safety alignment research, and LLM vulnerability documentation. Educational evaluation of safety guardrails.
LLM Attack Testing Toolkit is a structured methodology and mindset framework for testing Large Language Model (LLM) applications against logic abuse, prompt injection, jailbreaks, and workflow manipulation.
LLM Sentinel Red Teaming Platform is an enterprise-grade framework for automated security testing of Large Language Models, detecting vulnerabilities such as jailbreaks, prompt injection, and system prompt leakage across multiple providers, with structured attack orchestration, risk scoring, and security reporting to harden models before production
AI testing portfolio project - LLM output testing, prompt injection, and non-deterministic testing patterns using Python, pytest, DeepEval, and Claude API
Promptfoo is an open-source CLI and TypeScript/Node.js library for evaluating, red-teaming, and security-testing LLM applications, agents, and RAG pipelines. It runs deterministic prompt evals with model-graded and rule-based assertions, generates dynamic adversarial attack probes across 50+ vulnerability categories (prompt injection, jailbreaks…
Mindgard is a UK-based offensive AI security company (London/Lancaster, spun out of Lancaster University) that provides an automated AI red-teaming and security testing platform for large language models, AI agents, and generative AI systems.
Add a description, image, and links to the jailbreak-testing topic page so that developers can more easily learn about it.
To associate your repository with the jailbreak-testing topic, visit your repo's landing page and select "manage topics."