ATTACKS AND DEFENSES FOR LLM-BASED AGENTS

PhD Qualifying Examination


Title: "ATTACKS AND DEFENSES FOR LLM-BASED AGENTS"

by

Mr. Liwen WANG


Abstract:

Large language models (LLMs) increasingly serve as planners for agents that 
retrieve information, invoke tools, retain memory, and communicate with 
peers. These components convert model outputs into persistent state and 
real-world actions, while security boundaries remain expressed in natural 
language and interpreted by probabilistic models. Adversaries exploit this 
design through injected instructions, poisoned memories, malicious tools, 
planner backdoors, and compromised peer agents. This survey develops a 
lifecycle-centered framework for attacks and defenses in LLM-based agents. It 
defines an agent as a composition of a language-model planner, memory, tools, 
an execution environment, and policy enforcement. It classifies attacks by 
entry point, targeted component, attacker capability, persistence, and 
violated security property. Defenses are organized by the invariants they 
enforce: instruction provenance, task alignment, information-flow control, 
least privilege, state integrity, and accountable delegation. The survey also 
compares major benchmarks and relates attack success to benign utility, 
action validity, attacker feasibility, and adaptive robustness. The 
literature increasingly supports secure-by-design architectures that place 
deterministic enforcement around an untrusted language-model core. Persistent 
state, dynamic tool ecosystems, compositional workflows, and adaptive 
evaluation remain central research challenges.


Date:                   Wednesday, 26 August 2026

Time:                   3:00pm - 5:00pm

Zoom Meeting:
https://hkust.zoom.us/j/94168110792?pwd=gblnLegPQYO9dZVqzOHOHyJvgyER9v.1

Committee Members:      Dr. Shuai Wang (Supervisor)
                        Dr. Dan Xu (Chairperson)
                        Dr. Chaojian Li