Skip to content
#

security-benchmark

Here are 14 public repositories matching this topic...

Adversarial security benchmark for agent authorization: does a compromised agent's policy-violating proposal become an unauthorized external effect? 73 trials, nine families, an independent oracle, per-mechanism ablation, confidence intervals. 0 unauthorized effects in 61 attack trials (95% CI [0.0%, 5.9%]). Reproduction is partial.

  • Updated Aug 4, 2026
  • Elixir

Vendor-neutral benchmark measuring how MCP security proxies/gateways DEFEND against 22+ attack vectors — crosswalked to NIST AI RMF & OWASP LLM/Agentic Top 10. CI-gated, reproducible, DOI-cited. Submit your tool to the leaderboard.

  • Updated Jul 28, 2026
  • JavaScript

Improve this page

Add a description, image, and links to the security-benchmark topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the security-benchmark topic, visit your repo's landing page and select "manage topics."

Learn more