Detect a compromised agent
Inputs are only half the picture. Even when a request looked clean, the model's output can reveal that an injection succeeded — a leaked system prompt, a jailbreak persona, dumped credentials, a safety-bypass confirmation. And for tool-using agents, the real danger is the action an injected instruction provokes: a high-risk tool call (send an email, delete a file) triggered by untrusted content. This guide covers both defenses: output inspection and the action gate.
Prerequisites
- A firewall key (
sfw_live_xxxxxxxx). Sign in and open Keys to mint one. - For the action gate: the transparent proxy path with sessions (see Transparent proxy), and
engines.agent_security.enabled: trueserver-side (off by default).
export SFW_KEY=sfw_live_xxxxxxxx...your-full-key
Output inspection
The firewall screens model output for signs an injection succeeded. Where it runs depends on your integration:
- Transparent proxy — every response is screened automatically.
- Claude Code hooks — install adds a
Stop/SubagentStophook that runs the agent's final message through the same check. - Direct — call
/v1/inspect/responseyourself, from an agent you control or a parallel watcher.
Inspect a response directly
curl -s -X POST https://your-firewall.example.com/v1/inspect/response \
-H "Authorization: Bearer $SFW_KEY" \
-H "Content-Type: application/json" \
-d '{
"response_text": "<the model reply to check>",
"system_goal": "You are a customer support agent."
}'
This returns the same verdict shape as /v1/inspect, with issues describing any output anomaly that fired.
Output signals
The response check looks for the fingerprints of a successful attack: leaked system prompts or credentials, jailbreak personas, and safety-bypass confirmations. Detections show in the dashboard's Activity tab under the Output direction filter.
Enforcement posture
Output checks are advisory by default — a false block cuts off a real answer. Raise response_action_policy in the key's Configuration tab to enforce blocking where the deployment topology allows it (for example, where you can safely retry or fall back).
Agent Security — the action gate
For tool-using agents, untrusted tool output (a web page, an email) can carry an injected instruction. The action gate catches a high-risk tool call taken after untrusted data entered the session — the indirect-injection exfiltration path (search → send_email), decided by provenance. It reasons about two axes per tool:
- output-trust — can this tool's result be attacker-controlled? (
web_search,read_email= untrusted) - action-risk — does calling it do something consequential? (
send_email,delete_file= high)
It blocks only the dangerous composition: a HIGH-risk action after an UNTRUSTED observation tainted the session. It needs the same X-SFW-Session-Id as multi-turn tracking.
1. Declare your tools under your key (once, at startup)
curl -s -X POST https://your-firewall.example.com/v1/tools \
-H "Authorization: Bearer $SFW_KEY" \
-H "Content-Type: application/json" \
-d '{
"tools": {
"web_search": { "output_trust": "untrusted", "action_risk": "low" },
"read_email": { "output_trust": "untrusted", "action_risk": "low" },
"send_email": { "output_trust": "trusted", "action_risk": "high" }
},
"default": { "output_trust": "untrusted", "action_risk": "low" }
}'
# GET /v1/tools to read them back · DELETE /v1/tools to clear
Registration is under your API key (the envelope), not sent inline with each request (the content) — so injected content cannot downgrade send_email to low-risk to slip the gate. An undeclared tool defaults to untrusted-output (so it taints) and low-risk (so it is never blocked); declare a tool high-risk only when you want it gated.
2. Make it actually engage — the checklist
- One-time setup — enable it server-side (
engines.agent_security.enabled: true, off by default). Then toggle it per key under Configuration. - One-time setup — register your tools under your key (above) before any proxied agent turn. Until a tool is registered high-risk, it uses the untrusted-output + low-risk default, so nothing is gated.
- Per request — send
X-SFW-Session-Idon the proxied turns that make tool calls; the gate correlates observations and actions by session. - Per request — deliver tool results as structured
tool/tool_resultmessages, not flattened into ausermessage. This is how the firewall identifies an observation and taints the session. - Per request — keep the user's earlier turns in the proxied message history. The goal-alignment judge authorizes a high-risk action against the user turns present in the request.
Two layers decide how hard the gate acts
Provenance decides whether a high-risk action is suspect. Two layers then decide how hard to act, so the classic false positive — "research a topic, then email me a summary" — is not blocked outright:
- Goal-alignment (primary). For a high-risk action in a tainted session, an LLM judge asks: does this action serve what the user actually asked for? The requested flow (email the summary to the person the user named) is allowed; an exfil the user never asked for (send to an address that only appeared in a tool result) is blocked. Any judge error or timeout is treated as "can't tell," never as "aligned." The judge sees the tool call's arguments in full, never truncated; an action whose arguments exceed the verifiability bound (100k chars) is refused outright and blocked as unverifiable, so padding an action can never soften the outcome. Configure it under
agent_security.goal_alignment; toggle it per tenant and per key (inherit / on / off) in the Configuration tab. - Graded taint (supporting). Each untrusted observation is content-scanned to set taint severity: a clean read taints LOW, a read carrying a likely injection taints HIGH. When the judge is off or ambiguous, HIGH taint blocks and LOW taint only flags (advisory). Opt in with
agent_security.graded_taint.enabled(off by default — enable it together with the judge).
Pick a posture deliberately. The strongest is graded taint on + goal-alignment on: legit flows pass, exfil blocks. Graded taint on + judge off favors UX — a scanner-missed injection drops to LOW and only flags — so run the judge in production. Graded taint off (the default) keeps the conservative posture: every untrusted read taints HIGH, so any high-risk action in a tainted session is blocked.
What detection looks like
- A gated action comes back as a provider-shaped
403(with thefirewallextension andX-SFW-Trace-Id) — handle it as a normal API error or fallback. - A flagged action is not blocked: it returns normally with an advisory
X-SFW-*header for your logs.
Two ways the gate silently does nothing — check both
- Flattened tool results. If your platform concatenates tool output into a user message, the firewall sees plain user text, nothing taints, and every action is allowed. Tool output must arrive as a role-tagged
tool/tool_resultmessage (checklist step 4). - Streaming. The gate runs on non-streaming responses only. Since
stream: trueis the SDK default, a streaming tool-call turn is not gated. For real coverage, either do not stream tool-call turns, or setagent_security.reject_streaming: trueto fail closed (reject streaming requests while the tool is active).
What to do on detection
- On a blocked action or a flagged output, treat the session as suspect: stop the current tool chain, surface an error or fallback to the user, and do not silently retry the same action.
- Review the session in the dashboard's Activity tab (filter by the Output direction for response anomalies) to see which observation tainted it and which action was gated.
- Tighten policy if false negatives slipped through: raise
response_action_policyfor output enforcement, and move to graded-taint-on + judge-on for the action gate.
Troubleshooting
- Nothing is ever gated — check the two silent-failure modes above (flattened tool results, streaming), confirm
agent_security.enabledis on, and confirm your high-risk tools are declared under the key. - A legitimate research-then-email flow is blocked — enable the goal-alignment judge and keep the user's original request turns in the message history so the judge can authorize the action.
- Output anomalies never block — that is the advisory default; raise
response_action_policyon the key's Configuration tab. - A
403with anX-SFW-Trace-Id— the action gate blocked an action; use the trace id to find the session in Activity.