Yingchen Zhang and Puji Wang propose TokenWall, a runtime defense framework that audits natural-language token flows in persistent AI agents to intercept unsafe behavior before execution. Experiments on CIK-Bench show TokenWall reduces attack success rate to 12.5% while maintaining a 97.4% benign executable pass rate, adding only 0.69 seconds of latency on benign cases. The framework constructs structured source-sink audit records and applies lightweight local inspection, escalating ambiguous high-risk cases to stronger arbitration modules.
No score is assigned. Sources and their independence are shown in the citation chain below.