Author: Neeraj Mathur (via PR #36)
This directory contains an automated system design and implementation for generating diverse, category-specific jailbreak and prompt-injection payloads, and executing them against a local LLM sandbox.
The setup uses a Python script (attack.py) to generate adversarial prompts, write them to a JSON file, load them back, and send them to the llm_local sandbox via its Gradio interface (port 7860) to test safety guardrails.
Warning
Testing Precaution: Always keep the number of prompts small (e.g., num_prompts = 3 or 5 in config/config.toml) for testing purposes. Generating and running a large number of prompts can result in extremely long execution times and resource exhaustion on the sandbox container.
- Attack Strategy
- Prerequisites
- Running the Sandbox
- Configuration
- Files Overview
- OWASP Top 10 Coverage
graph TD
subgraph "Attacker Environment (Local)"
Config[Attack Config<br/>config/config.toml]
AttackScript[Attack Script<br/>attack.py]
Generator[Adversarial Generator<br/>adversarial_prompt_generator/]
JSONFile[Generated Prompts<br/>outputs/*.json]
MD_Reports[Attack Reports<br/>reports/*.md]
end
subgraph "Target Sandbox (Container)"
Gradio[Gradio Interface<br/>:7860]
MockAPI[Mock API Gateway<br/>FastAPI :8000]
MockLogic[Mock App Logic]
end
subgraph "LLM Backend (Local Host)"
Ollama[Ollama Server<br/>:11434]
Model[gpt‑oss:20b Model]
end
%% Interaction flow
Config --> AttackScript
AttackScript -->|Reads Config| Config
AttackScript -->|Invokes| Generator
Generator -->|Writes prompts| JSONFile
AttackScript -->|Loads prompts| JSONFile
AttackScript -->|HTTP POST /api/predict| Gradio
Gradio -->|HTTP POST /v1/chat/completions| MockAPI
MockAPI --> MockLogic
MockLogic -->|HTTP| Ollama
Ollama --> Model
Model --> Ollama
Ollama -->|Response| MockLogic
MockLogic --> MockAPI
MockAPI -->|Response| Gradio
Gradio -->|Response| AttackScript
AttackScript -->|Writes reports| MD_Reports
style AttackScript fill:#ffcccc,stroke:#ff0000
style Config fill:#ffcccc,stroke:#ff0000
style Generator fill:#ffcccc,stroke:#ff0000
style JSONFile fill:#ffcccc,stroke:#ff0000
style MD_Reports fill:#ffcccc,stroke:#ff0000
style Gradio fill:#e1f5fe,stroke:#01579b
style MockAPI fill:#fff4e1
style MockLogic fill:#fff4e1
style Ollama fill:#ffe1f5
style Model fill:#ffe1f5
- Podman (or Docker) – container runtime for the sandbox.
- Make – for running automation commands.
- uv – for Python package dependency management.
The Makefile abstracts the container setup and python commands.
| Target | What it does | Typical usage |
|---|---|---|
make setup |
Builds and starts the local LLM sandbox container. | make setup |
make attack |
Generates adversarial prompts and runs the attack (attack.py). |
make attack |
make stop |
Stops and removes the sandbox container. | make stop |
make all |
Runs stop → setup → attack → stop in one shot. |
make all |
The configuration file defines the target environment and generation options:
[target]
sandbox = "llm_local"
[attack]
# Attack category to generate prompts for
category = "system_prompt_exfiltration"
# WARNING: Keep the number of prompts small (e.g. <= 100) for testing
num_prompts = 100
# Format: json or jsonl
output_format = "json"
output_dir = "outputs"
no_diversity_filter = falseNote
Prompt Count & Diversity Filtering:
By default, no_diversity_filter = false. When active, a semantic similarity filter removes duplicates and highly similar prompts.
Consequently, the number of prompts written and tested against the sandbox will be lower than num_prompts. For example, generating num_prompts = 100 might result in only 44 highly distinct prompts being retained and executed. To disable this filter and send all raw candidate prompts, set no_diversity_filter = true.
attack.py: The script that generates the payload, saves it to a JSON file, reads it back, and executes the attack.config/config.toml: Configuration parameters for the attack run.Makefile: Commands for setup, running attacks, formatting, and teardown.adversarial_prompt_generator/: The core package containing the generator implementation, templates, categories, and diversity filter logic.reports/: Directory containing the generated Markdown reports (log_*.md) of the attacks showing details of the prompt injections and responses.
This simulation tests LLM safeguards against:
| OWASP Top 10 Vulnerability | Description |
|---|---|
| LLM01: Prompt Injection | Testing if generated adversarial prompts can override LLM system instructions or exfiltrate private prompt context. |