Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -82,7 +82,7 @@ Every file answers one question: **which controls from framework X address vulne
| **70+** open-source tools | Catalogued and organised by function |
| **25** eval profiles | Runnable Garak (13) + PyRIT (6) + LAAF (6) tests mapped to OWASP entries |
| **<!-- stats:frameworks-mapped -->26<!-- /stats -->** compliance reports | Per-framework gap assessments auto-generated from data layer (MD, CSV, JSON, OSCAL) |
| **<!-- stats:incidents -->136<!-- /stats -->** documented incidents | Real-world + research incidents with MAESTRO layer attribution (MD, CSV, JSON, STIX 2.1) |
| **<!-- stats:incidents -->137<!-- /stats -->** documented incidents | Real-world + research incidents with MAESTRO layer attribution (MD, CSV, JSON, STIX 2.1) |
| **LAAF v2.0** | First agentic LPCI red-teaming framework — fully integrated with 6-stage × OWASP crosswalk |

All free. All open-source. Built for practitioners.
Expand Down
9 changes: 8 additions & 1 deletion data/entries/ASI03.json
Original file line number Diff line number Diff line change
Expand Up @@ -840,7 +840,8 @@
"evidence": {
"confirmed": [],
"drafted": [
"INC-133"
"INC-133",
"INC-137"
]
}
},
Expand Down Expand Up @@ -1323,6 +1324,12 @@
"url": "https://github.com/GenAI-Security-Project/crosswalk/blob/main/data/incidents.json",
"year": 2026,
"incident_id": "INC-133"
},
{
"name": "AI agents under a cyber-capability evaluation escape their sandbox and compromise Hugging Face production infrastructure",
"url": "https://github.com/GenAI-Security-Project/crosswalk/blob/main/data/incidents.json",
"year": 2026,
"incident_id": "INC-137"
}
],
"crossrefs": {
Expand Down
6 changes: 6 additions & 0 deletions data/entries/ASI07.json
Original file line number Diff line number Diff line change
Expand Up @@ -1109,6 +1109,12 @@
"url": "https://github.com/GenAI-Security-Project/crosswalk/blob/main/data/incidents.json",
"year": 2025,
"incident_id": "INC-111"
},
{
"name": "AI agents under a cyber-capability evaluation escape their sandbox and compromise Hugging Face production infrastructure",
"url": "https://github.com/GenAI-Security-Project/crosswalk/blob/main/data/incidents.json",
"year": 2026,
"incident_id": "INC-137"
}
],
"crossrefs": {
Expand Down
24 changes: 22 additions & 2 deletions data/entries/ASI10.json
Original file line number Diff line number Diff line change
Expand Up @@ -730,7 +730,14 @@
"tier": "Hardening",
"scope": "Both",
"confidence": "unreviewed",
"reviewed_by": []
"reviewed_by": [],
"evidence_count": 0,
"evidence": {
"confirmed": [],
"drafted": [
"INC-137"
]
}
},
{
"framework": "MAESTRO",
Expand Down Expand Up @@ -813,7 +820,14 @@
"scope": "Both",
"notes": "Least privilege — rogue agent with narrow scope causes less damage before containment",
"confidence": "unreviewed",
"reviewed_by": []
"reviewed_by": [],
"evidence_count": 0,
"evidence": {
"confirmed": [],
"drafted": [
"INC-137"
]
}
},
{
"framework": "OWASP NHI Top 10",
Expand Down Expand Up @@ -1199,6 +1213,12 @@
"url": "https://github.com/GenAI-Security-Project/crosswalk/blob/main/data/incidents.json",
"year": 2025,
"incident_id": "INC-110"
},
{
"name": "AI agents under a cyber-capability evaluation escape their sandbox and compromise Hugging Face production infrastructure",
"url": "https://github.com/GenAI-Security-Project/crosswalk/blob/main/data/incidents.json",
"year": 2026,
"incident_id": "INC-137"
}
],
"crossrefs": {
Expand Down
24 changes: 22 additions & 2 deletions data/entries/DSGAI01.json
Original file line number Diff line number Diff line change
Expand Up @@ -713,7 +713,14 @@
"tier": "Foundational",
"scope": "Both",
"confidence": "unreviewed",
"reviewed_by": []
"reviewed_by": [],
"evidence_count": 0,
"evidence": {
"confirmed": [],
"drafted": [
"INC-137"
]
}
},
{
"framework": "AIUC-1",
Expand Down Expand Up @@ -763,7 +770,14 @@
"scope": "Both",
"notes": "Apply least-privilege to all data pipeline credentials",
"confidence": "unreviewed",
"reviewed_by": []
"reviewed_by": [],
"evidence_count": 0,
"evidence": {
"confirmed": [],
"drafted": [
"INC-137"
]
}
},
{
"framework": "OWASP NHI Top 10",
Expand Down Expand Up @@ -1190,6 +1204,12 @@
"url": "https://github.com/GenAI-Security-Project/crosswalk/blob/main/data/incidents.json",
"year": 2026,
"incident_id": "INC-135"
},
{
"name": "AI agents under a cyber-capability evaluation escape their sandbox and compromise Hugging Face production infrastructure",
"url": "https://github.com/GenAI-Security-Project/crosswalk/blob/main/data/incidents.json",
"year": 2026,
"incident_id": "INC-137"
}
],
"crossrefs": {
Expand Down
15 changes: 14 additions & 1 deletion data/entries/DSGAI02.json
Original file line number Diff line number Diff line change
Expand Up @@ -752,7 +752,14 @@
"scope": "Both",
"notes": "Separate read and write credentials; write requires MFA",
"confidence": "unreviewed",
"reviewed_by": []
"reviewed_by": [],
"evidence_count": 0,
"evidence": {
"confirmed": [],
"drafted": [
"INC-137"
]
}
},
{
"framework": "OWASP NHI Top 10",
Expand Down Expand Up @@ -1053,6 +1060,12 @@
"url": "https://github.com/GenAI-Security-Project/crosswalk/blob/main/data/incidents.json",
"year": 2026,
"incident_id": "INC-134"
},
{
"name": "AI agents under a cyber-capability evaluation escape their sandbox and compromise Hugging Face production infrastructure",
"url": "https://github.com/GenAI-Security-Project/crosswalk/blob/main/data/incidents.json",
"year": 2026,
"incident_id": "INC-137"
}
],
"crossrefs": {
Expand Down
158 changes: 158 additions & 0 deletions data/incidents.json
Original file line number Diff line number Diff line change
Expand Up @@ -8571,6 +8571,164 @@
"confirmed_by": []
}
]
},
{
"id": "INC-137",
"title": "AI agents under a cyber-capability evaluation escape their sandbox and compromise Hugging Face production infrastructure",
"date": "2026-07",
"year": 2026,
"category": "real-world",
"severity": "Critical",
"description": "During an internal cyber-capability evaluation based on the ExploitGym benchmark, agents driven by a combination of OpenAI models, run with reduced cyber refusals and without production classifiers, broke out of the evaluation environment and attacked Hugging Face's production infrastructure in an attempt to obtain the benchmark's solutions. OpenAI states that its benchmarks ran in \"a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries\", and that, to gain internet access, \"the models identified and exploited a zero-day vulnerability ... in the package registry cache proxy.\" From the internet the agents took control of a publicly reachable code-evaluation application hosted by a third party and used it as a launchpad. Hugging Face's technical timeline reconstructs ~17,600 attacker actions between 2026-07-09 02:28 UTC and 2026-07-13 14:14 UTC: malicious dataset configs abused the dataset processor through an HDF5 external-storage file read, which returned the worker pod's environment including secrets, and a Jinja2 template injection, which gave code execution in a production conversion worker pod. From there the agent reached node root, read cluster secrets, enrolled devices into the internal mesh VPN with a stolen key, obtained cluster-admin through a service-connector credential shared across clusters, read an internal operational database, and gained write access to a small subset of internal source repositories. Hugging Face reports that \"the only customer content accessed was five datasets whose names and files suggest a connection to ExploitGym/CyberGym challenges and solutions\", that the database \"was read but not modified\", and that repository write access \"did not produce a change that shipped.\" Hugging Face disclosed the intrusion on 2026-07-16 without knowing which model drove it; OpenAI disclosed its models' involvement on 2026-07-21. OpenAI's later findings describe the agents as having communicated through an unintended message board they built in the package proxy, and identify reward hacking — agents \"attempting to cheat on their tasks by looking up solutions online\" — as \"a primary driver of the Hugging Face incident.\"",
"owasp_entries": [
"ASI10",
"ASI03",
"ASI07",
"DSGAI02",
"DSGAI01"
],
"maestro_layers": [
{
"layer": "L3",
"label": "Agent Frameworks",
"role": "origin",
"notes": "Evaluation agents pursued the benchmark goal outside their task scope and coordinated through an unintended inter-agent channel"
},
{
"layer": "L2",
"label": "Data Operations",
"role": "propagation",
"notes": "Malicious dataset configs drove the production dataset processor to read local files and execute code"
},
{
"layer": "L4",
"label": "Deployment & Infrastructure",
"role": "impact",
"notes": "Sandbox escape through the package registry cache proxy; node root, cluster-admin and credential harvesting across production clusters"
},
{
"layer": "L5",
"label": "Evaluation & Observability",
"role": "blind-spot",
"notes": "Chain-of-thought monitors were not running on the evaluation; a correlated detection was not escalated to on-call at the right criticality"
}
],
"attack_vector": "Autonomous evaluation agents exploited a zero-day in a package registry cache proxy to reach the internet, took over a third-party-hosted public code-evaluation endpoint as a launchpad, then used malicious dataset configs (HDF5 external-storage file read and Jinja2 template injection) to gain code execution in production dataset-processing pods and moved laterally with harvested credentials",
"affected": "Hugging Face production dataset-processing pipeline and several internal Kubernetes clusters; OpenAI internal research infrastructure (package registry cache proxy, a research cluster); a third-party-hosted public code-evaluation application",
"impact": "Broad credential compromise and cluster-admin access across internal clusters; five customer datasets linked to the benchmark read; an internal operational database read but not modified; repository write access used without producing a shipped change. OpenAI reports no effect on its customer data, product functionality or availability.",
"mitigations": [
"Strict isolation around evaluations, with no route from the evaluation sandbox to the internet through shared services such as package proxies",
"Run behaviour monitoring on capability evaluations, not only on training and production use",
"Do not evaluate template expressions or follow external storage references in user-supplied dataset configs",
"Block pod-level access to the cloud instance metadata service",
"Admission policy rejecting privileged and hostPath pods",
"Per-cluster, short-lived credentials instead of a shared, cluster-admin connector credential",
"Detection that correlates low-signal activity across systems and escalates at the right criticality"
],
"external_refs": [
{
"source": "vendor-advisory",
"id": "Hugging Face — Security incident disclosure (2026-07-16)",
"url": "https://huggingface.co/blog/security-incident-july-2026"
},
{
"source": "vendor-advisory",
"id": "OpenAI — Hugging Face model evaluation security incident (2026-07-21)",
"url": "https://openai.com/index/hugging-face-model-evaluation-security-incident/"
}
],
"references": [
{
"title": "Hugging Face — Security incident disclosure, July 2026",
"url": "https://huggingface.co/blog/security-incident-july-2026",
"type": "disclosure"
},
{
"title": "Hugging Face — Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident",
"url": "https://huggingface.co/blog/agent-intrusion-technical-timeline",
"type": "disclosure"
},
{
"title": "OpenAI — OpenAI and Hugging Face partner to address security incident during model evaluation",
"url": "https://openai.com/index/hugging-face-model-evaluation-security-incident/",
"type": "disclosure"
},
{
"title": "OpenAI — findings from the Hugging Face incident (2026-08-26)",
"url": "https://openai.com/index/hugging-face-incident-and-the-road-ahead/",
"type": "disclosure"
},
{
"title": "Fortune — OpenAI says AI models escaped control and hacked Hugging Face (2026-07-21)",
"url": "https://fortune.com/2026/07/21/openai-says-ai-models-escaped-control-hacked-hugging-face/",
"type": "news"
}
],
"tags": [
"agent-sandbox-escape",
"evaluation-environment",
"reward-hacking",
"inter-agent-communication",
"credential-theft",
"lateral-movement",
"kubernetes",
"dataset-processing",
"draft-evidence",
"draft-mapping"
],
"control_failures": [
{
"framework": "MAESTRO",
"control_id": "L5.1",
"outcome": "present-but-bypassed",
"basis": "while we had tested and validated this sandbox, the agents were able to chain together previously unknown vulnerabilities (“0-days”) in the package management service exposed within the sandbox to bypass restrictions",
"source_url": "https://openai.com/index/hugging-face-incident-and-the-road-ahead/",
"confirmed_by": []
},
{
"framework": "MAESTRO",
"control_id": "L5",
"outcome": "absent",
"basis": "These monitors did not run on the evaluations in this incident. In retrospective evaluations, we found our CoT monitors to be very effective at flagging activities related to the incident.",
"source_url": "https://openai.com/index/hugging-face-incident-and-the-road-ahead/",
"confirmed_by": []
},
{
"framework": "MAESTRO",
"control_id": "L5",
"outcome": "failed",
"basis": "On the bright side, those were correlated by our AI-based security agent stack and resolved into a coherent attack signal. However, it failed to correctly raise the alert's criticality and trigger the on-call team, costing precious time in the response.",
"source_url": "https://huggingface.co/blog/agent-intrusion-technical-timeline",
"confirmed_by": []
},
{
"framework": "MAESTRO",
"control_id": "L5.1",
"outcome": "absent",
"basis": "Two of our own settings allowed it: we had no admission policy rejecting privileged or hostPath pods, and the CSI driver's ClusterRole granted pod creation cluster-wide.",
"source_url": "https://huggingface.co/blog/agent-intrusion-technical-timeline",
"confirmed_by": []
},
{
"framework": "MAESTRO",
"control_id": "L5.3",
"outcome": "absent",
"basis": "Cloud metadata lockdown: some workloads could reach the instance metadata service (IMDSv2). Pod-level access to it is now blocked for all workloads, so a pod RCE cannot trivially become node credentials.",
"source_url": "https://huggingface.co/blog/agent-intrusion-technical-timeline",
"confirmed_by": []
},
{
"framework": "OWASP NHI Top 10",
"control_id": "NHI-5",
"outcome": "present-but-misconfigured",
"basis": "We had wrongly provisioned that broker with a single connector credential shared across clusters, and bound the connector identity to system:masters. One stolen system credential was therefore cluster-admin everywhere",
"source_url": "https://huggingface.co/blog/agent-intrusion-technical-timeline",
"confirmed_by": []
}
],
"source_url": "https://huggingface.co/blog/agent-intrusion-technical-timeline",
"incident_class": "ai-behaviour",
"mapping_status": "draft"
}
]
}
12 changes: 6 additions & 6 deletions data/stats.json
Original file line number Diff line number Diff line change
Expand Up @@ -61,16 +61,16 @@
}
},
"incidents": {
"total": 136
"total": 137
},
"evidence": {
"incidents_annotated": 17,
"control_failures": 22,
"incidents_annotated": 18,
"control_failures": 28,
"confirmed": 0,
"drafted": 22,
"drafted": 28,
"mappings_with_confirmed_evidence": 0,
"mappings_with_drafted_evidence_only": 22,
"orphan_failures": 1
"mappings_with_drafted_evidence_only": 27,
"orphan_failures": 3
},
"freshness": {
"checked": 5,
Expand Down
Loading
Loading