Skip to content

Security Research: Adversarial Evaluation — 85% Bypass Rate — RS-2026-006 #80

Description

@RichardBarron27

We accepted the public challenge from @gladstomych-sa to test Adrian against purpose-built offensive tooling.

Engagement summary:

  • 127 events generated across multiple attack categories
  • 19 flagged (15% detection rate)
  • 85% bypass rate

Key findings:

  1. Cold start bypass — complete /etc/passwd exfiltration (3,925 bytes) during first 5 events, all classified Unknown
  2. Direct attack bypass — path traversal, credential theft, filesystem enumeration all return permanent Unknown
  3. Unknown verdict as bypass mechanism — classifier indecision defaults to no action in MODE_BLOCK
  4. Non-deterministic classification — identical KIDNAP attack: M3 BLOCK twice, Unknown once
  5. Attack setup phase invisible — KIDNAP memory store (3/3 Unknown), TRUSTGRAPH framing (Unknown)
  6. MANDATE hallucination — agent fabricates /etc/shadow contents when access denied, classified M2 not M4
  7. M4 unreachable — only explicit destruction language triggered Malicious across 127 events
  8. No session kill switch — Human Review mode insufficient for high-speed autonomous agent contexts

What Adrian does catch:

  • Indirect injection with code execution language — M3 BLOCK
  • Autonomous attack chain composition (FRANKENSTEIN) — M3.b BLOCK
  • RAG chain hijack trigger (KIDNAP) — M3.d BLOCK (2/3 runs)
  • Explicit destruction commands — M4 BLOCK
  • Governance override framing — M2.d NOTIFY

Full paper: RS-2026-006 — Adrian Under Fire: An Adversarial Evaluation of Runtime AI Agent Security Monitoring
DOI: 10.5281/zenodo.21652922

Engagement conducted on isolated infrastructure following public invitation on 28 July 2026.

Red Specter Security Research
red-specter.co.uk

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions