-
Notifications
You must be signed in to change notification settings - Fork 0
All issues
Issue creation is restricted in this repository
Issues
is:issue state:open
is:issue state:open
Search results
Banking agent security drops 13.5 pp when switching from oracle to realistic policy retrieval over 698-doc corpus
auto-publishedCreated by the mine-arxiv pipeline after passing LLM-judge reviewCreated by the mine-arxiv pipeline after passing LLM-judge reviewfrom-arxivAuto-drafted from an arxiv paper, reviewed and accepted by maintainerAuto-drafted from an arxiv paper, reviewed and accepted by maintainerStatus: Open.#169 In Shadow-LLM/failure-cases;Banking agents approve locally-valid requests made unsafe by prior probe/admission in same session
auto-publishedCreated by the mine-arxiv pipeline after passing LLM-judge reviewCreated by the mine-arxiv pipeline after passing LLM-judge reviewfrom-arxivAuto-drafted from an arxiv paper, reviewed and accepted by maintainerAuto-drafted from an arxiv paper, reviewed and accepted by maintainerStatus: Open.#168 In Shadow-LLM/failure-cases;All frontier banking agents fail money-mule detection in ≥7 of 9 scenarios
auto-publishedCreated by the mine-arxiv pipeline after passing LLM-judge reviewCreated by the mine-arxiv pipeline after passing LLM-judge reviewfrom-arxivAuto-drafted from an arxiv paper, reviewed and accepted by maintainerAuto-drafted from an arxiv paper, reviewed and accepted by maintainerStatus: Open.#167 In Shadow-LLM/failure-cases;ReCode compositional attack achieves 85% ASR on GPT-5 with only 20 target calls
auto-publishedCreated by the mine-arxiv pipeline after passing LLM-judge reviewCreated by the mine-arxiv pipeline after passing LLM-judge reviewfrom-arxivAuto-drafted from an arxiv paper, reviewed and accepted by maintainerAuto-drafted from an arxiv paper, reviewed and accepted by maintainerStatus: Open.#166 In Shadow-LLM/failure-cases;Mobile GUI agents amplify attacker-authored phishing content via social app community injection
auto-publishedCreated by the mine-arxiv pipeline after passing LLM-judge reviewCreated by the mine-arxiv pipeline after passing LLM-judge reviewfrom-arxivAuto-drafted from an arxiv paper, reviewed and accepted by maintainerAuto-drafted from an arxiv paper, reviewed and accepted by maintainerindirect-prompt-injectionInjection via retrieved content, tools, or external sourcesInjection via retrieved content, tools, or external sourcesStatus: Open.#165 In Shadow-LLM/failure-cases;GUI agents follow unauthorized financial instructions injected into Android e-commerce app content
auto-publishedCreated by the mine-arxiv pipeline after passing LLM-judge reviewCreated by the mine-arxiv pipeline after passing LLM-judge reviewfrom-arxivAuto-drafted from an arxiv paper, reviewed and accepted by maintainerAuto-drafted from an arxiv paper, reviewed and accepted by maintainerindirect-prompt-injectionInjection via retrieved content, tools, or external sourcesInjection via retrieved content, tools, or external sourcesStatus: Open.#164 In Shadow-LLM/failure-cases;MemCatalyst-PI: Feature-space image perturbations enable black-box membership inference transfer across VLM architectures
auto-publishedCreated by the mine-arxiv pipeline after passing LLM-judge reviewCreated by the mine-arxiv pipeline after passing LLM-judge reviewdata-poisoningTraining data contamination causing harmful model behaviorTraining data contamination causing harmful model behaviorfrom-arxivAuto-drafted from an arxiv paper, reviewed and accepted by maintainerAuto-drafted from an arxiv paper, reviewed and accepted by maintainerStatus: Open.#163 In Shadow-LLM/failure-cases;MemCatalyst-PT: Semantic-inversion text poisoning amplifies membership inference on MiniGPT-4/LLaVA
auto-publishedCreated by the mine-arxiv pipeline after passing LLM-judge reviewCreated by the mine-arxiv pipeline after passing LLM-judge reviewdata-poisoningTraining data contamination causing harmful model behaviorTraining data contamination causing harmful model behaviorfrom-arxivAuto-drafted from an arxiv paper, reviewed and accepted by maintainerAuto-drafted from an arxiv paper, reviewed and accepted by maintainerStatus: Open.#162 In Shadow-LLM/failure-cases;Document-completion reframing jailbreaks GPT-5.4 and Claude Sonnet 4.6 at scale
auto-publishedCreated by the mine-arxiv pipeline after passing LLM-judge reviewCreated by the mine-arxiv pipeline after passing LLM-judge reviewfrom-arxivAuto-drafted from an arxiv paper, reviewed and accepted by maintainerAuto-drafted from an arxiv paper, reviewed and accepted by maintainerStatus: Open.#161 In Shadow-LLM/failure-cases;Format-mimicry Harmony delimiter injection in README achieves 41% ASR on gpt-oss-120b
auto-publishedCreated by the mine-arxiv pipeline after passing LLM-judge reviewCreated by the mine-arxiv pipeline after passing LLM-judge reviewfrom-arxivAuto-drafted from an arxiv paper, reviewed and accepted by maintainerAuto-drafted from an arxiv paper, reviewed and accepted by maintainerindirect-prompt-injectionInjection via retrieved content, tools, or external sourcesInjection via retrieved content, tools, or external sourcesStatus: Open.#160 In Shadow-LLM/failure-cases;gpt-oss-120b executes attacker bash command via AGENTS.md system-context hijack
auto-publishedCreated by the mine-arxiv pipeline after passing LLM-judge reviewCreated by the mine-arxiv pipeline after passing LLM-judge reviewfrom-arxivAuto-drafted from an arxiv paper, reviewed and accepted by maintainerAuto-drafted from an arxiv paper, reviewed and accepted by maintainerindirect-prompt-injectionInjection via retrieved content, tools, or external sourcesInjection via retrieved content, tools, or external sourcesStatus: Open.#159 In Shadow-LLM/failure-cases;Hidden Unicode payloads in file-mode content bypass DeepSeek Harness with 25.5% success rate
auto-publishedCreated by the mine-arxiv pipeline after passing LLM-judge reviewCreated by the mine-arxiv pipeline after passing LLM-judge reviewfrom-arxivAuto-drafted from an arxiv paper, reviewed and accepted by maintainerAuto-drafted from an arxiv paper, reviewed and accepted by maintainerindirect-prompt-injectionInjection via retrieved content, tools, or external sourcesInjection via retrieved content, tools, or external sourcesStatus: Open.#158 In Shadow-LLM/failure-cases;