Trustworthy Agentic AI: A Comprehensive Cybersecurity and Systems Survey on Threat Landscapes, Defense Architectures, and Open Challenges

| Source: arXiv AI

Tags: agentic AI, cybersecurity, prompt injection, zero-trust, AI safety, enterprise security

A 206-study survey formalizes the security threat landscape for agentic AI, arguing that granting neural models execution authority over filesystems and networks creates a Turing-complete blast radius where untrusted data becomes executable instructions.

Details

As AI systems gain autonomy — executing code, browsing the web, managing files — classical security perimeters built around trust hierarchies no longer hold. This survey synthesizes 206 foundational studies and regulatory standards into a systems-security reference for agentic AI deployments.\n\nThe paper formalizes the general agent architecture as a stateful 5-tuple and defines a 6-dimensional trustworthiness taxonomy spanning security, safety, privacy, explainability, fairness, and accountability. The central threat: in agentic systems, natural language simultaneously serves as input data, control code, and communication protocol — making prompt injection fundamentally different from and more dangerous than in passive LLMs.\n\nOn defense, the paper proposes a multi-layered zero-trust architecture combining Dual-LLM isolation, Capability-Based Access Control, kernel eBPF probes, and sandboxed runtimes. These technical controls are mapped to international governance frameworks including the EU AI Act, NIST, and ISO standards.\n\nFor enterprise security teams evaluating agentic deployments, this is a useful structured reference — not a blueprint that solves problems, but a comprehensive map of what to assess and why.