Presentation overview
This presentation introduces a full-chain capability-laundering threat in self-evolving multi-agent systems. It examines how instructions embedded in untrusted external content can be promoted into persistent memory, retrieved in a later session, delegated across agents, and ultimately executed through high-privilege tools.
Research focus
- Model the attack as a sequence of memory write, retrieval, delegation, and execution stages.
- Compare the complete indirect attack chain with direct and single-agent baselines.
- Measure whether composition amplifies risk beyond the probability of failure at each individual stage.
- Evaluate provenance labels, authorization checks, and executor-side validation as defenses.
Experimental direction
The proposed OpenClaw-based harness isolates memory, agent roles, access controls, and tool execution so each stage can be measured independently. The deck also positions the work alongside SudoBench, OEP, and Skill-Inject, highlighting persistent memory as the mechanism that transfers attacker-controlled policy across sessions and privilege boundaries.