Open-source, 100% reproducible AI Agent Runtime Security Benchmark & Sandbox Environment (RFC-010 Draft Protocol).
-
Updated
Aug 6, 2026 - HTML
Open-source, 100% reproducible AI Agent Runtime Security Benchmark & Sandbox Environment (RFC-010 Draft Protocol).
Adversarial security benchmark for agent authorization: does a compromised agent's policy-violating proposal become an unauthorized external effect? 73 trials, nine families, an independent oracle, per-mechanism ablation, confidence intervals. 0 unauthorized effects in 61 attack trials (95% CI [0.0%, 5.9%]). Reproduction is partial.
Open deterministic security tests for unsafe multi-agent handoffs and authority escalation.
Open-source benchmark for adversarial evidence attacks on LLM-based cybersecurity auditors, targeting ACM AsiaCCS 2027.
Vendor-neutral benchmark measuring how MCP security proxies/gateways DEFEND against 22+ attack vectors — crosswalked to NIST AI RMF & OWASP LLM/Agentic Top 10. CI-gated, reproducible, DOI-cited. Submit your tool to the leaderboard.
Deterministic security benchmark for tool-using AI agents
Open AI-for-security validation benchmark: non-LLM scorer + a SOTA-validation loop. Labeled positive corpus withheld pending coordinated disclosure.
Cross-framework AI agent red-teaming benchmark. Tests LangChain, CrewAI, AutoGen, LlamaIndex & OpenAI Agents SDK for prompt injection, scope violations, and jailbreaks against a shared, OWASP ASI-aligned attack payload set.
ReplayBench-IoT: reproducible IoT replay-defense benchmark with Monte Carlo sweeps, CI, static demo, and hardware-validation artifacts.
GitHub action for Maester
FreightSkillBench is a reproducible benchmark for evaluating document-to-transaction integrity, prompt-injection risk, and security controls in AI-enabled shipping and logistics workflows.
Reproducible benchmark for smart-contract security tools, measuring precision, recall, and false positives against executable PoCs and versioned ground truth.
Product-security LLM benchmark harness for realistic AppSec, supply-chain, and LLM application security evaluations.
The core repository for the Maester module with helper cmdlets that will be called from the Pester tests.
Add a description, image, and links to the security-benchmark topic page so that developers can more easily learn about it.
To associate your repository with the security-benchmark topic, visit your repo's landing page and select "manage topics."