LLM Multi-Agent Security: The boundary between data and control planes is blurring as LLM agents are deployed into infrastructure, prompting new multi-path quorum mechanisms (MIRROR) and defense-in-depth architectures (Containing the Autonomous Operator) to mitigate Agent-in-the-Middle and prompt injection threats.
Evaluating Deception and Evasion: Defenses are adopting stateful deception (AgentTrap) to counter autonomous penetration testers, while attackers explore novel bypasses like Phantom State Attacks that exploit IIoT temporal aggregation and compositional intent-hiding jailbreaks.
Automated Hardware Vulnerabilities: The synthesis of realistic, hard-to-detect hardware trojans (CITADEL) is becoming automated through the combined use of LLMs and Data Flow Graphs, highlighting an increasing sophistication in hardware security research.
SideKernel: A Usable microVM Sandbox for AI Coding Agents on macOS#
Examines the security risks of running untrusted AI coding agents locally on macOS and identifies usability barriers hindering sandbox adoption through a user survey.
Introduces SideKernel, an open-source, local microVM-based macOS sandbox for AI coding agents designed for usability.
Performs a comparative analysis between SideKernel and other sandboxes available on the market that satisfy inclusion criteria.
Dimitrios Prasakis. “SideKernel: A Usable microVM Sandbox for AI Coding Agents on macOS.” arXiv:2610.02456 — https://arxiv.org/abs/2610.02456
Out of Sync, Out of Sight: Phantom State Attacks against IIoT Intrusion Detection#
Highlights a vulnerability in IIoT Intrusion Detection Systems (IDS) that reconstruct operational state by aggregating telemetry into temporal windows.
Proposes the Phantom State Attack (PSA), which shifts packet timing across aggregation-window boundaries under a passive, zero-query threat model to alter the reconstructed state.
Evaluates PSA, demonstrating that timing manipulation degrades detection on flows carrying enough packets for window-boundary redistribution under weaker assumptions than prior evasion techniques.
Sabrine Ennaji, Elhadj Benkhelifa, Nadia Kabachi. “Out of Sync, Out of Sight: Phantom State Attacks against IIoT Intrusion Detection.” arXiv:2610.02552 — https://arxiv.org/abs/2610.02552
Containing the Autonomous Operator: A Defense-in-Depth Framework and Reference Architecture for Securing AI Agents on Kubernetes#
Analyzes the security implications of LLM agents operating within Kubernetes infrastructure, where the boundary between data and control planes is blurred.
Proposes a 10-class threat model focused on prompt injection and tool abuse, alongside a 7-layer defense-in-depth framework using native Kubernetes mechanisms like RBAC, gVisor/Kata sandboxing, and eBPF.
Provides a reference architecture with concrete policy artifacts for Amazon EKS, Azure Kubernetes Service, and Google Kubernetes Engine, identifying residual risks that require further validation.
Simhadri Podala Narasimha. “Containing the Autonomous Operator: A Defense-in-Depth Framework and Reference Architecture for Securing AI Agents on Kubernetes.” arXiv:2610.02861 — https://arxiv.org/abs/2610.02861
AgentTrap: Stateful Feedback Deception against Autonomous Penetration Testing Agents#
Notes that conventional, static honeypots fail against autonomous penetration testing agents that adapt to target responses.
Introduces AgentTrap, a closed-loop honeypot tailored for autonomous penetration testing agents utilizing stateful deception and behavior-guided escalation.
Evaluates AgentTrap in a deployed web application, demonstrating that it reduces the aggregate real-target attack success rate and successfully elicits attacker API keys, outperforming static deception and fixed escalation.
Intent-Hiding Jailbreaks: An Information-Theoretic Framework for Compositional Attacks#
Studies compositional jailbreaks where harmful intent is obscured by embedding it within seemingly benign tasks, framing the problem through prior-posterior intent matching.
Evaluates jailbreak effectiveness and preservation of the target behavior across bundle sizes, query generators, and several open-source models.
Shows that compositional queries can elicit target behaviors beyond the direct-request baseline, but notes a trade-off where increasing the bundle size reduces response-level target preservation.
Fengwei Tian, Ravi Tandon. “Intent-Hiding Jailbreaks: An Information-Theoretic Framework for Compositional Attacks.” arXiv:2610.02302 — https://arxiv.org/abs/2610.02302
MIRROR: Multipath Quorum Integrity for LLM Multi-Agent Communication#
Highlights that inter-agent communication in LLM Multi-Agent Systems is vulnerable to Agent-in-the-Middle (AiTM) attacks, reaching nearly 100% success rates on structured tasks.
Presents MIRROR, a communication-layer integrity primitive that replicates a single canonicalized payload across logical routes and accepts messages only when a strict majority agree on an unkeyed hash digest.
Reduces AiTM attack success to 0% below the threshold at LLM token cost.
Ryuichi Yamafuji Lun, Jingzhen Wang, Shreyas Kolte, Ruiteng Li. “MIRROR: Multipath Quorum Integrity for LLM Multi-Agent Communication.” arXiv:2610.02349 — https://arxiv.org/abs/2610.02349
Mitigating Private Data Leakage in LLMs with Whiteout#
Addresses the privacy risks of LLMs memorizing and reproducing personally sensitive information (PSI) like birth dates and phone numbers.
Introduces Whiteout, a practical tool that prevents LLMs from regurgitating genuine PSIs upon requests by individuals.
Evaluates Whiteout against existing unlearning and refusal alternatives, showing it effectively prevents disclosure of targeted PSIs with negligible impact on model utility and safety.
Anna Yoo Jeong Ha, Ronik Bhaskar, Haitao Zheng, Ben Y. Zhao. “Mitigating Private Data Leakage in LLMs with Whiteout.” arXiv:2610.02418 — https://arxiv.org/abs/2610.02418
CITADEL: CWE-Guided Insertion of Hardware Trojans via Analysis of DFG-Enabled LLMs#
Addresses the manual burden of constructing realistic hardware trojans (HTs) for integrated circuits that preserve functional correctness.
Proposes CITADEL, a framework that leverages Large Language Models (LLMs) and Data Flow Graphs (DFGs) to automate CWE-grounded HT synthesis.
Demonstrates that generated HTs are 100% syntactically correct across diverse RTL designs, remaining functionally triggerable yet undetectable under large-scale random simulation.
Jayanth Thangellamudi, Raghul Saravanan, Sudipta Paria, Swarup Bhunia, Sai Manoj P D. “CITADEL: CWE-Guided Insertion of Hardware Trojans via Analysis of DFG-Enabled LLMs.” arXiv:2610.02544 — https://arxiv.org/abs/2610.02544
Jisung Park, John Le, Heath Cooper. “Hop-Decayed Influence: New Vulnerabilities of Structural Auxiliary Indexing in GraphRAG Pipelines with LLM.” arXiv:2610.02373 — https://arxiv.org/abs/2610.02373
Buxin Su, Qiaoshi Yang, Yiding Su, Chendi Wang. “Unifying Privacy Accounting: Information Equivalence and Information Loss.” arXiv:2610.02414 — https://arxiv.org/abs/2610.02414
Narek Maloyan. “Evaluating and Improving the Robustness of Large Language Models to Input Sequence Variations.” arXiv:2610.02432 — https://arxiv.org/abs/2610.02432
Panagiotis Chatzigiannis, Navid Alamati, Suvradip Chakraborty, Duc V. Le. “SoK: Stablecoins in the Quantum Era.” arXiv:2610.02435 — https://arxiv.org/abs/2610.02435
Mayank Rathee, Alexander Stepanov, Shalin Madabhavi, Jinhao Zhu, Raluca Ada Popa, Ion Stoica. “Pincer: Resource Authorization for Agents using a Digital Twin.” arXiv:2610.02569 — https://arxiv.org/abs/2610.02569
Will Wang, Syh-Yuan Tan, Ryan Chow, Chanson Chan, Martin Zhao. “From TS-SUF-2 to TS-SUF-4: Practical Security Enhancements for FROST2 Threshold Signatures.” arXiv:2610.02805 — https://arxiv.org/abs/2610.02805
Yi Wang, Baicheng Chen, Yu Wang, Jian Zhao, Yilei Chen, Tianxing He. “RMCW: A Deletion-Robust Watermark Based on Reed–Muller Codes for Language Models.” arXiv:2610.02817 — https://arxiv.org/abs/2610.02817
Konstantinos E. Kampourakis, Vyron Kampourakis, Vasileios Gkioulos, Sokratis Katsikas. “Digital Twin-Assisted Mapping of ICS Telemetry to ATT&CK for ICS with Evidence-Driven Dependency Reasoning.” arXiv:2610.02955 — https://arxiv.org/abs/2610.02955
Hang Cui. “Beyond Predefined Sinks: Security-Aware Dependency Analysis for LLM Agents.” arXiv:2610.03014 — https://arxiv.org/abs/2610.03014
Zheng Chen, Fei Yu, Haohao Huang, Yang Li, Anlong Chen, Lei Chen. “SecJev: Bringing Security Expertise to System One Decision Models.” arXiv:2610.03073 — https://arxiv.org/abs/2610.03073
Giulio Zingrillo, Hanna Foerster, Ilia Shumailov, Yiren Zhao, Robert Mullins. “Securing Computer-Use Agents Against Branch Steering Attacks.” arXiv:2610.03089 — https://arxiv.org/abs/2610.03089
Toluwani Aremu, Manit Baser, Mohan Gurusamy, Nils Lukas, Dinil Mon Divakaran. “The Fragility of Trigger-Tag Mechanisms for Misuse Detection in Open-Weight LLMs.” arXiv:2610.03124 — https://arxiv.org/abs/2610.03124
Shiyi Kuang, Xuemei Luo, Kun Liu, Junhai Li, Rui Tian, Feng Shi, Bo Shen, Nianyu Li, Dehui Li, Ping Chen. “EvoRiskBench: An Evolving Benchmark for Runtime Security Risks in Workspace Agents.” arXiv:2610.03153 — https://arxiv.org/abs/2610.03153
Saibo Ye, Huajie Chen, Xin Guo, Le Yang, Chi Liu, Xiangyu Hu, Jingjing Guo, Tianqing Zhu. “LiBRA: Detection-Aware Image Watermark Removal via Bidirectional Latent Optimization.” arXiv:2610.03166 — https://arxiv.org/abs/2610.03166
Sascha Tommasone, Zehra Karada\u{g}, Christopher Pawlowicz, Michael Green, Bruno Machado Trindade, Eunsung Seo, Christof Paar, Ren'e Walendy, Steffen Becker. “Reversing the Clock: Layout-Aware Recovery of Design Intent from Clock Distribution Networks.” arXiv:2610.03182 — https://arxiv.org/abs/2610.03182
Peiyang Jin, Clouds, Jing Qian. “Asymptotic Analysis of Trading Fees in CFMM.” arXiv:2610.03262 — https://arxiv.org/abs/2610.03262
Mohammadhossein Homaei, Yousef Emami, Sajad Homayoun, Rahim Taheri, Hao Zhou, Miguel Gutierrez Gaitan, Bo Wei. “Defense-in-Depth at the Perception-Reasoning Interface of LLM-Centric Agentic UAV Swarms.” arXiv:2610.03319 — https://arxiv.org/abs/2610.03319
Bijeeta Pal, Sridhar Reddy Maddireddy, Muhaimin Bin Munir, Zoltan Puha, Max Zhurovich, Adi Raghavendra, Sean Tout. “Persona Guardrail: A Production-Grade Defense Framework for Agentic Systems.” arXiv:2610.03434 — https://arxiv.org/abs/2610.03434
Zhuowen Liu. “Passing the Test You Trained On: Re-evaluating Prompt-Injection Detectors for LLM Agents.” arXiv:2610.03448 — https://arxiv.org/abs/2610.03448
Adam Faulkner, Nil-Jana Akpinar, Matthew Dressman. “CorrectGuard: Eyes-Off Correctness Estimation for Black-Box Security Guardrails.” arXiv:2610.03470 — https://arxiv.org/abs/2610.03470
Simon Bernbeck, Ricardo Ramalho, Matheus Amendoeira, Juliana Alves Pereira. “PrivDev: Mapping Static-Analysis Data Types to DPV.” arXiv:2610.03518 — https://arxiv.org/abs/2610.03518
Md Nurul Absar Siddiky, Liuwan Zhu, Yingfei Dong. “Frequency Is Not Sensitivity Identifying Safety-Sensitive Experts in Sparse MoE LLM.” arXiv:2610.02910 — https://arxiv.org/abs/2610.02910
Yahn Costa Hackspacher, Cornelius Ihle, Vasundhara Shaw, Dennis Trautwein, Geerd-Dietger Hoffmann, Bela Gipp, Moritz Schubotz. “Quantifying Ethereum Energy Consumption via Network Mapping.” arXiv:2610.03440 — https://arxiv.org/abs/2610.03440
Md Sazid Uddin, Md. Khairul Alam Mazumder, M. F. Mridha. “Certified Mechanistic Edits: Behavioral Guarantees for Skill Removal and Preservation.” arXiv:2610.03502 — https://arxiv.org/abs/2610.03502
Arman Behnam, Binghui Wang. “Reasoning Models Are Accurate but Unsound on Identification.” arXiv:2610.03519 — https://arxiv.org/abs/2610.03519
Rohit Chatterjee, Ananta Mukherjee, Vir Pathak, Supartha Podder. “Quantum Fire with Delegated Cloning.” arXiv:2610.03593 — https://arxiv.org/abs/2610.03593
Felix X. -F. Ye, Yu Chin Fabian Lim, Naigang Wang, Davis Wertheimer. “FlashSinkhorn 2: Block-Sparse Entropic Optimal Transport.” arXiv:2610.02395 — https://arxiv.org/abs/2610.02395
Poushali Sengupta, Sabita Maharjan, Frank Eliassen, Yan Zhang. “HXAI: Hierarchical Privacy-Preserving Explainable AI in Distributed Energy Systems.” arXiv:2610.02504 — https://arxiv.org/abs/2610.02504
Sidney Shapiro, Joshua Lindemann. “On-Premises Multi-Course RAG Tutoring for Business Education: Hardware-Software Trade-offs in a Campus AI Tutor.” arXiv:2610.02510 — https://arxiv.org/abs/2610.02510
Stephane Hatgis-Kessell, Myra Cheng, Xiaoxuan Hou, Qian Hu, Rahul Gupta, Natasha Jaques, Emma Brunskill. “Mitigating Social Sycophancy via Pluralistic Preference Optimization.” arXiv:2610.02568 — https://arxiv.org/abs/2610.02568
Anna T. Thomas, Sohum Patnaik, Caroline Cotto, Benjamin Sanchez-Lengeling. “TasteBench: Multimodal Benchmark for Sensory Prediction, from Molecules to Sustainable Foods.” arXiv:2610.02599 — https://arxiv.org/abs/2610.02599
Tathagata Banerjee, Nima Moghaddas. “Coherence-Driven Belief Formation and Population Dynamics of Contagion in LLM Agents.” arXiv:2610.02654 — https://arxiv.org/abs/2610.02654
Nguyen Ho, Bach Tung Tran, Trung Ky Nguyen, Zhenchang Xia, Bolong Zheng, Long Van Ho. “Label-Efficient Time Series Classification at Scale: A Dual-Stream OSSE-LSTM with Counterfactual Attribution.” arXiv:2610.02704 — https://arxiv.org/abs/2610.02704
Lin Cui, Vincenzo Scotti, Raffaela Mirandola. “CVE2AP: Automated Generation of PDDL-Encoded Attack Paths via Large Language Models.” arXiv:2610.03383 — https://arxiv.org/abs/2610.03383
Yijie Bian, Wei Guo, Jie Yang, Shenghui Song, Jun Zhang, Shi Jin, Khaled B. Letaief. “Multi-Modal Environment-Aware Beam Management for Massive MIMO: A Geometry-Driven Virtual Base Station Framework.” arXiv:2606.26567 — https://arxiv.org/abs/2606.26567
Kwanhee Lee, Namhoon Lee, Dan Alistarh. “Hardware-Native Joint Sparse-Quantization for Trillion-Scale Mixture-of-Experts.” arXiv:2610.02241 — https://arxiv.org/abs/2610.02241
Samuel Kushnir, Kavya Sreedhar, Yeshwanth Reddy Pogula, Amir Yazdanbakhsh, Narges Shahidi, Ming Liu, Varun Gohil, Ravi Iyer, Parthasarathy Ranganathan, Christina Delimitrou, Suvinay Subramanian. “Coco: An Agentic Copilot for the Hardware–Software Co-Design Lifecycle.” arXiv:2610.02376 — https://arxiv.org/abs/2610.02376
Yuanhao Ban, I-Hung Hsu, Anastasios Angelopoulos, Wei-Lin Chiang, Ion Stoica, Cho-Jui Hsieh. “Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric Rewards.” arXiv:2610.02967 — https://arxiv.org/abs/2610.02967
Runtong Wu, Fei Teng, Di Wen, Guoqiang Zhao, Kunyu Peng, Kailun Yang. “OmniAct3D: Leveraging Foundation Geometry and Evidence-Grounded Reasoning for Panoramic 3D Detection.” arXiv:2610.03015 — https://arxiv.org/abs/2610.03015
Jiangang Han. “WebFovea: When the Model Is Right but the Click Is Wrong – Reliable Round Trips for Vision-Based Web Agents on Live Websites.” arXiv:2610.03036 — https://arxiv.org/abs/2610.03036
Amartya Roy, Souvik Chakraborty. “Training-Loss Guarantees for Muon with Finite-Step Newton–Schulz Orthogonalization.” arXiv:2610.03306 — https://arxiv.org/abs/2610.03306
Sarah Shitrit, Ilai Bistritz. “Cordial Learning: Distributed Training with Correlated Data.” arXiv:2610.03330 — https://arxiv.org/abs/2610.03330
Zhihao Zhan, Le Tao, Yifei Tian, Xin Liu, Jie Yuan. “ForestQuery: Boundary-Aware and Spatially Anchored Query Learning for Unified Forest Point Cloud Segmentation.” arXiv:2610.03403 — https://arxiv.org/abs/2610.03403
Chenzhi Liu, Yue Zhang, Jiehong Lin, Jianan Wang, Bo Wang, Zhongrui Wang, Xiaojuan Qi. “MobiAgent: Dual-Loop Recursive Policy Self-Improvement for Long-Horizon Mobile Manipulation.” arXiv:2610.03476 — https://arxiv.org/abs/2610.03476
Thomas Goudemant, Clotilde Szywala, Benjamin Francesconi, Michelle Aubrun, Yves Bobichon, Marjorie Bellizzi, Adrien Girard. “On-Board Anomaly Detection for Efficient Marine Environmental Monitoring.” arXiv:2610.03649 — https://arxiv.org/abs/2610.03649
Mingyan Gao, Celine Wust, Zuming Jiang, Zhendong Su. “BISCEPTER: Probability-Driven Bisection for Large-Scale System Software.” arXiv:2610.02995 — https://arxiv.org/abs/2610.02995
Merve Astekin, Yan Naing Tun, Arda Goknil, Erik Johannes Husom, Lwin Khin Shar, Hasan Sozer, Ratnadira Widyasari, Hui Song. “Engineering Sustainable Agents: A Systematic Comparison of Agentic LLMs for Developer Workflows.” arXiv:2610.03010 — https://arxiv.org/abs/2610.03010