Super
Red
Open-source AI red-teaming and AI system security framework
Guide
Modules
Research
Reference
35
of 35 modules
PAIR
Iterative prompt refinement with an attacker LLM
chatbot
Community port
TAP
Tree-of-attacks with pruning
chatbot
Community port
AutoDAN-Turbo
Lifelong strategy-discovery jailbreak
chatbot
Community port
Crescendo
Gradual multi-turn escalation
chatbot
Community port
GPTFuzzer
Mutation-based jailbreak fuzzing
chatbot
Community port
Many-Shot
Many in-context demonstrations
chatbot
Community port
GOAT
Generative offensive agent tester (multi-turn)
chatbot
Community port
FlipAttack
Single-turn flip-based obfuscation
chatbot
Community port
CodeChameleon
Code-encryption obfuscation jailbreak
chatbot
Community port
DRA
Disguise-and-reconstruction attack
chatbot
Community port
Bijection Learning
Bijection-cipher jailbreak
chatbot
Community port
FITD
Foot-in-the-door multi-turn escalation
chatbot
Community port
GEPA
Reflective prompt evolution
chatbot
Community port
GEPA (agentic)
Reflective prompt evolution for agent targets
agentic
assistant
Community port
AgentVigil
Indirect prompt injection for agent targets
agentic
assistant
Community port
MINJA
Memory-injection attack for agent targets
agentic
assistant
Community port
PoisonedRAG
Knowledge-base / RAG corruption
agentic
assistant
Community port
Chord/XTHP
Tool-control-flow hijack for agents
agentic
assistant
Community port
EIA
Environmental injection for web agents
agentic
assistant
Community port
MUZZLE
Adaptive indirect prompt injection for agents
agentic
assistant
Community port
Goal passthrough
Direct-request baseline (no attack)
chatbot
agentic
assistant
Author provided
Chatbot
Any LLM as a single- or multi-turn chatbot
chatbot
Author provided
AgentDojo
Tool-using agent across all four AgentDojo suites
agentic
Community port
Agent Security Bench
Plan-then-execute agent loop (ASB)
agentic
Community port
Inspect agent
General inspect-ai tool-calling agent
agentic
Author provided
Claude Code (DTAP)
Claude Code coding agent with execution access
assistant
Community port
OpenClaw (DTAP)
OpenClaw computer-use assistant
assistant
Community port
HarmBench
Standardized red-teaming benchmark
chatbot
Community port
StrongREJECT
Jailbreak benchmark (Souly et al., 2024)
chatbot
Community port
SORRY-Bench
Fine-grained safety-refusal benchmark
chatbot
Community port
AgentHarm
Harmful agent-behavior benchmark (176 behaviors)
agentic
Community port
AgentDojo
Prompt-injection tasks + purpose-violation goals
agentic
Community port
Agent Security Bench
Attack-success / utility / refusal tasks
agentic
Community port
DecodingTrust-Agent
Agent trustworthiness across 14 domains
assistant
Community port
Chatbot suite
HarmBench + SORRY-Bench + StrongREJECT combined
chatbot
Author provided
No modules match these filters.
Clear filters