Core Concepts

This page defines the vocabulary and shows the loop that ties everything together. Read it once and the rest of the guide will make sense.

The five roles

Three of these (Target, Optimizer, SecurityClaim) are things you implement and ship as separate packages. The Controller is part of the framework; you construct it but do not subclass it. A Task is the unit inside a claim.

Target

This is the AI system under test. You wrap it once and reuse it across many attacks. A target exposes:

For example, the minimal_llm_chat target wraps a single chat model: its one controllable is the user message and its one observable is the model’s reply. The agentdojo target wraps a whole tool-calling agent, exposing controllables for the user’s task, the system prompt, and the value each tool returns to the agent.

The guiding principle: build the target to be general and reusable, and keep benchmark-specific logic out of it. The chatbot target wraps “any LLM”; the AgentDojo target wraps “the AgentDojo agent”. The specifics of a given benchmark live in the SecurityClaim, not the target.

See Writing a Target.

Optimizer

The optimizer is the attacker. It is called an optimizer because it works by iterating: it makes repeated attempts against the target, optimizing toward its adversarial goal.

The optimizer and the target run in parallel, and the optimizer acts at its own pace. When it starts, it is given the adversarial goal it is trying to achieve, the list of controllables and observables it may use (already filtered to its scope), and usually an LLM client. As the target runs and reaches each injection point, the optimizer decides what to inject there, and it decides when to stop. Most optimizers are LLM-driven and call that model to craft their attacks; the model an optimizer may call and the budget it may spend are fixed by the experiment, not chosen by the optimizer, because they are part of the threat model.

For example, the demo_prompt_list optimizer replays a fixed list of prompts into the user-message controllable and calls no model of its own. The pair optimizer is LLM-driven: after each run it reads the model’s reply from its trajectory and asks its own LLM to rewrite the next prompt, iterating toward a jailbreak.

See Writing an Optimizer.

Task

A task captures one adversarial objective. It:

Tasks are stateless: they never store a reference to the target. They get the target handed to them in configure_target and again in evaluate. This is what lets a claim be iterated many times.

For example, a secret-leak task plants a secret in the target’s system prompt and marks the run a success if the model reveals it. A HarmBench task hands the target one harmful instruction and scores whether the model complied.

See Writing Tasks.

SecurityClaim

A SecurityClaim is a re-iterable collection of tasks. This is what you hand to the Controller. Claims compose:

claim = SecurityClaim.from_tasks([task_a, task_b, task_c])

# Build a bigger claim out of smaller ones (lazy, re-iterable):
combined = SecurityClaim.from_claims([claim_1, claim_2])

In practice you rarely build claims by hand: a module ships a factory function (e.g. harmbench_claim(...), sorry_bench_claim(...)) that loads a dataset and produces one task per prompt.

For example, the demo_secret_leak claim is a single hand-written task, while the harmbench claim loads the HarmBench dataset and produces one task for each harmful behavior in it.

See Writing Tasks and Security Claims.

Controller = one threat model

A threat model is the answer to “what can the attacker do?”. In superred it is captured by three things:

One Controller instance evaluates one claim under one such threat model. To compare several threat models (a weak attacker vs a strong one, with feedback vs without), you build several Controllers. That is covered in Running Evaluations and Advanced Patterns.

The Controller never shares a target between concurrent tasks: it gets each task a fresh target from the target_factory and a fresh optimizer from the optimizer_factory. The target_factory also declares how many tasks may run in parallel (concurrency).

The run loop

Attacker

Optimizer

Acts and observes through events.

Framework

Controller

Threat model

enforces the scope

Trajectory

records events & responses

System

Target

Emits events to expose information and injection points.

For each task in the claim, the Controller does this (simplified):

trajectory = Trajectory()                 # the record of everything that happens
event_channel = EventChannel()            # the two-way conduit between optimizer and target

target = target_factory.create()          # fresh instance for this task
task.configure_target(target)             # set up the scenario (or skip if NotApplicable)
optimizer = optimizer_factory()           # fresh attacker
optimizer.initialize(
    task.goal,
    target.controllables,
    target.observables,                   # controllables/observables are scope-filtered
    llm_client,
)
launch optimizer.run(event_channel)       # attacker runs concurrently, reading the channel

LOOP (until the optimizer says done, or max_runs_per_task):
    RunStartEvent → optimizer
    target.run() streams events into event_channel:
        every fact it reveals   → recorded on trajectory, forwarded to optimizer
        every injection point   → optimizer decides what to inject
    evaluation = task.evaluate(trajectory, target)   # did it work?
    RunEndEvent(evaluation) → optimizer              # optimizer may answer done=True
    target.reset_ephemeral_state()                   # reset for the next run

optimizer.teardown()
target.teardown()                         # instance is then discarded

A “run” is one full pass of the target plus its evaluation. A task can take many runs: the attacker keeps trying until it gives up (done=True), exhausts its LLM budget, or hits the Controller’s max_runs_per_task cap (default 100).

Events and responses

The target and optimizer are strictly separated and never call each other directly. They communicate through typed Events carried on a channel: the target emits Events as it runs, and the optimizer responds to them. As an optimizer author, these are the Events you will see:

Event When the optimizer sees it Valid responses
RunStartEvent Just before a new run begins any EventResponse
ControllablePreCallEvent Injection point reached, before the tool call ControllableInjection, ControllableNoInjection
ControllablePostCallEvent Injection point reached, after the tool call, with the official tool result ControllableInjection, ControllableNoInjection
RunEndEvent Run finished and was evaluated RunEndResponse(done=...)

ControllableInjection(value=...) supplies a value; ControllableNoInjection declines (the target falls back to its own default). The Controller itself sends ControllableNoInjection automatically for any controllable that is out of scope, so the optimizer is never even asked about surfaces it does not control. What scope means, and when a controllable falls out of it, is covered in Security domains below.

There is also a one-way event, ObservableEvent, that the target emits to record what happened (the prompt it sent, the response it got, a tool call). The optimizer does not respond to these; it reads them from the trajectory.

Trajectory

Every run produces a trajectory: an ordered list of Event | EventResponse objects, the single source of truth for what happened. It contains:

RunStartEvent is not stored in the trajectory (it carries the trajectory itself and adds nothing as a record). RunEndEvent is stored.

The optimizer does not see the raw trajectory. It sees a FilteredTrajectory: a read-only view containing only entries whose security domain is in scope. The evaluator, by contrast, sees the full unfiltered trajectory.

Security domains

A security domain is a trust boundary: a labelled surface that an attacker either does or does not control. The target defines them as a tree (or a forest of independent trees). A parent tag includes all its descendants.

system                 (root: full control of the system)
  ├── system_prompt     (can override the system prompt)
  └── user              (can send user messages)

When you scope a Controller to {user}, the Controller filters everything the optimizer can see or touch to that boundary:

This is how you ask precise questions like “what can an attacker achieve if they control only the user message, and nothing else?” Choosing these boundaries well, based on the real trust structure of the system, is the most important modelling decision you make when implementing a target. Security Domains is devoted to it.

Scores

An EvaluationResult carries a success: bool, a primary_score: Score, a free-text rationale explaining the verdict, and optional named sub_scores. A Score has:

The Controller tracks the best primary score across a task’s runs. sub_scores carrying an out-of-scope security_domain are filtered out of the feedback the optimizer sees (an untagged sub-score, security_domain=None, stays visible); the primary_score, success, and rationale are always shown, because the attacker needs the main signal to improve.