AI agents can hide memory attacks until it’s too late

New research finds AI agents can carry hidden threats that surface only later

|
Published September 03, 2026
AI agents can hide memory attacks until it’s too late

Researchers at the University of Calgary ran 2,614 simulated attack sequences against memory-enabled AI agents and found something unsettling: some attacks left no visible trace until well after the damage was already set in motion.

The study, led by Hadis Karimipour alongside a colleague, examined four distinct attack types, chain poisoning, policy rewriting, backdoor triggering, and slow drift.

The problem exists in one of the characteristics that makes today’s AI agents effective in the first place: memory. These technologies can now store their data over multiple sessions, can perform multi-phase tasks, and employ digital tools, including false information, until such information is extracted and treated as something that the agent learned itself and believes to be true.

Traditional cybersecurity threats become visible almost immediately: a malicious link is clicked, malware is executed, and a password is stolen. Memory poison does not follow this pattern at all. An agent may continue operating in the normal manner for days or even for multiple sessions until it starts using the poisoned piece of memory.

As discovered by the Calgary-based group of researchers, slow-drift and backdoor-trigger attacks, in particular, may remain undetected for quite some time until the moment when the agent uses the poisoned memory.

Risk didn't always climb steadily either; the researchers describe "non-monotonic" patterns, where an attack looked more concerning at one point and less concerning at another before fully developing.

That timing problem undermines a common security practice: testing an AI agent right after it encounters suspicious input, then declaring it safe if nothing goes wrong immediately. The researchers compare this to inspecting a notebook the moment after a false entry is written but before anyone has acted on it; the absence of immediate harm proves nothing.

Pareesa Afreen
Pareesa Afreen is a reporter and sub editor specialising in technology coverage, with 3 years of experience. She reports on digital innovation, gadgets, and emerging tech trends while ensuring clarity and accuracy through her editorial role, delivering accessible and engaging stories for a fast-evolving digital audience.
Share this story:
Advertisement