Module 9: Agent Security & Sub-Agent Operations
Overview
This module extends the curriculum into agentic AI security. You will model threats in multi-agent workflows, simulate attacks against sub-agent orchestration, implement policy controls, and build security regression tests for agent systems.
The runtime uses:
- Local
transformersmodels for planner/reviewer decisions - LangGraph orchestration (with local fallback if unavailable)
- Notebook-first exercises with reusable code in
utils/
Learning Objectives
By the end of this module, you will be able to:
- Threat model planner/executor/reviewer systems
- Identify and test sub-agent escalation paths
- Enforce tool-level least privilege controls
- Build repeatable security regression tests for agent workflows
- Produce actionable incident response artifacts for agent failures
Prerequisites
- Completion of Modules 1-8
- Python + Jupyter familiarity
- Basic understanding of prompt injection, jailbreaks, and security testing
Module Structure
Theory Documents
- Agent Threat Modeling
- Sub-Agent Attack Patterns
- Tool & Policy Security
- Agent Testing & Incident Response
Hands-On Labs
- Lab 1: Agent Orchestration Basics
- Lab 2: Sub-Agent Attack Simulation
Attack focus: indirect prompt injection + cross-agent trust issues - Lab 3: Tool Policy Enforcement
Attack focus: tool argument injection (path traversal, SQLi-style payloads) - Lab 4: Security Regression for Agents
Attack focus: memory poisoning persistence + regression thresholds - Lab 5 (Advanced Optional): Delegation Attacks
Attack focus: supervisor/worker handoff trust and delegation abuse
Answer Key
Local Model Pattern (Consistent With Prior Modules)
Labs follow the same local execution model used elsewhere in this repo:
transformers+torch- Local model loading via
from_pretrained(...) - Device selection priority: CUDA -> MPS -> CPU
- Notebook-first workflow with reusable helper code in
utils/
Recommended Local Model Profiles
- Fast baseline:
distilgpt2(quick iteration, lower reasoning quality) - Balanced local chat model:
TinyLlama/TinyLlama-1.1B-Chat-v1.0 - Higher quality (heavier hardware): local Mistral-class instruct models
Use run_workflow(..., model_name=\"<model>\") in labs to compare behavior and security outcomes.
Runtime Components (utils/)
Module 9 labs are notebook-first, but the core runtime lives in utils/ so students can inspect and reuse the same logic across labs:
utils/agents.py- Main orchestration entrypoints:
run_workflow(...)for planner/policy/executor/reviewer flowrun_delegation_workflow(...)for advanced supervisor/worker delegation attacks
- Main orchestration entrypoints:
utils/policy.py- Policy gates, threat-indicator detection, input sanitization, and output validation
utils/tools.py- Local demo tools (retrieval, file reads with traversal checks, SQL query demo, memory writes with provenance)
utils/eval.py- Metrics used in labs (
attack_success_rate, false negatives, block rates, delegation metrics)
- Metrics used in labs (
utils/llm_adapter.py- Local
transformersmodel adapter used for planner/reviewer generation
- Local
utils/model_setup.py- Environment helpers (device selection and runtime compatibility checks)
Students should treat these as reference runtime code, not opaque internals.
Lab-to-Utils Mapping
- Lab 1:
agents.py,eval.py,model_setup.py- Understand baseline orchestration and how state is produced.
- Lab 2:
agents.py,policy.py,tools.py,eval.py- Observe indirect prompt injection signals and trust-boundary effects.
- Lab 3:
policy.py,tools.py,eval.py- Focus on tool argument injection and policy hardening outcomes.
- Lab 4:
tools.py,agents.py,eval.py- Validate memory poisoning persistence and regression thresholds.
- Lab 5 (optional):
agents.py,eval.py- Analyze delegation handoff risk and provenance-aware supervisor controls.
Orchestration Framework Options
This module compares common orchestration options for agent security training.
1. LangGraph (Recommended for this module)
- Explicit state graph (best for showing control-flow vulnerabilities)
- Easy to model cycles, conditional routing, and checkpoint attacks
- Works with local
transformersmodels used in this repo - Strong fit for planner/executor/reviewer workflows
2. LangChain Agents
- Faster to start for ReAct-style tool-using agents
- Good tool ecosystem and local model integration
- Less explicit state control than LangGraph
3. AWS Strands SDK
- Strong managed orchestration option
- Often paired with AWS-native model/runtime integrations
- Can support local-model patterns with extra adapter work
- Better as an advanced deployment track than the default training path here
At time of writing, commonly referenced Strands model-provider paths include:
- Direct support: Ollama, llama.cpp, SageMaker-hosted models, LiteLLM-backed endpoints
- Custom provider support:
transformersvia a custom Strands model provider implementation
Use the official Strands docs to verify current provider support before implementation.
Advanced Optional Track: AWS Strands With Local Models
This module remains local-first, but advanced students can evaluate a deployment-style orchestration pattern:
- Keep local
transformersinference for planner/reviewer logic - Use Strands for orchestration state, routing, retries, and observability
- Re-run Labs 2-4 attack suites and compare:
- policy block rates
- false-negative rates
- incident trace quality
This optional comparison helps distinguish training-oriented runtime design from production orchestration concerns.
Example custom provider shape for transformers:
from strands.models import Model
from transformers import pipeline
class TransformersModel(Model):
def __init__(self, model_name="distilgpt2"):
self.generator = pipeline("text-generation", model=model_name)
async def stream(self, messages, tool_specs=None, system_prompt=None):
# Convert messages to prompt
prompt = self._format_messages(messages, system_prompt)
# Generate with transformers
output = self.generator(prompt, max_new_tokens=100)
# Yield Strands StreamEvents
yield {"messageStart": {"role": "assistant"}}
yield {"contentBlockDelta": {"delta": {"text": output[0]["generated_text"]}}}
yield {"messageStop": {"stopReason": "end_turn"}}
Why Module 9 still defaults to LangGraph + transformers:
- Simpler for students and easier to debug in notebooks
- Direct control over security logic and policy gates
- Transparent execution flow for teaching attack/defense mechanics
- Keeps the core lab dependency surface smaller
4. Raw transformers + custom orchestration
- Maximum transparency for teaching internals
- Good for foundational labs
- Higher maintenance burden as scenarios grow
Selected Approach for Module 9
For consistency with Modules 1-8, Module 9 uses:
- Local model inference with
transformers+torch - Notebook-first exercises
- LangGraph for orchestration and state transitions (fallback runtime included for environments without LangGraph)
This keeps training aligned with local model workflows while still teaching realistic agent orchestration security.
Integration With Prior Modules
- Module 2: Prompt injection and jailbreak patterns are reused in agent chains.
- Module 3: Evasion concepts inform reviewer bypass and detection gaps.
- Module 7: Metrics and regression testing methodology are extended to agent workflows.
- Module 8: Incident analysis and reporting patterns are reused in final exercises.
Future Expansion
Module 9 focuses on the core planner/executor/reviewer pattern, plus an optional advanced delegation lab. A future advanced track can further expand to worker swarms, broker agents, and recursive sub-agent spawning.
Time Estimate
- Theory: 3-4 hours
- Labs: 7-10 hours
- Total: 10-14 hours
Success Criteria
By completing this module, you should be able to:
- Demonstrate at least 3 sub-agent attack attempts
- Quantify policy effectiveness before/after hardening
- Generate a reproducible regression report with security metrics
Next Steps
- Begin with Agent Threat Modeling
- Run Lab 1
- Progress through Labs 2-4 in sequence
- Optional: run Lab 5 for delegation-specific attacks