[Speakers]
Adversary Village at
DEF CON 34

Dominika Pietrzak

Offensive AI Engineer

Building AI-led tooling to scale offensive security operations.

Field Notes on Offensive Agents: Reusability, Reliability, and What Breaks

13:30-14:00 PDT | Saturday, Aug 8th 2026 | DEF CON Creator Stage 3, Las Vegas Convention Center
Talk

Co-presented with: Ibai Castells

Abstract

This session shares field notes from building and running an agentic web-testing system in real offensive security engagements. Rather than presenting a product pitch or claiming that offensive agents are “solved,” we focus on the practical middle ground: what we found to be relevant in making agentic workflows reusable, reliable, and useful once they leave controlled environments. We'll cover the design choices that helped our framework survive contact with new models, unfamiliar targets and messy real-world constraints. We will also discuss where complexity exposed weak assumptions, how we approached evaluation, and how we kept agent behavior scoped, non-destructive, and under human control despite limited engineering time and fast-moving model releases.
Offensive agents are best understood as force multipliers for skilled operators, not replacements for expertise. Used well, they can help teams run longer, cover more ground, and preserve tradecraft across many targets. Used poorly, they simply automate fragility at scale. This talk shares what we learned about the engineering and operational choices that can make the difference between the two.

Talk outline

0–4 minutes: Framing the problem
Why offensive agents for adversary emulation are easy to demo, but hard to make reusable, reliable, and safe in practice.


4–10 minutes: Designing reusable agents
How we built agents that can be applied across tech stacks and plug into different orchestration layers without being tied to one workflow or framework.


10–15 minutes: Building portable workflows
How we structured agents so they can support different assessment types, attack paths, and toolchains.


15–22 minutes: Evaluating performance over time
What we have found to be most impactful on agent performance and how to measure whether agents improve, regress, or behave inconsistently as models, prompts, scoring methods, and tools change.


22–27 minutes: Lessons from testing
Common failure modes, what broke in practice, and how those lessons informed better design and evaluation.


27–30 minutes: Safety, takeaways, and Q&A
Key guardrails, responsible-use considerations, and practical takeaways for framework-agnostic offensive agents.

Agency.


Join Adversary Village Discord Server.

Join Adversary Village official Discord server to connect with our amazing community of adversary simulation experts and offensive security researchers!