Focus

How does your AI behave when someone tries to break it?

Normal performance says little about how a system responds to manipulation, hostile inputs, unexpected context or failing dependencies.

2 min read Author: KeynesMoore

How does your AI behave when someone tries to break it?

Conventional evaluation asks whether an AI system succeeds with representative inputs. Adversarial evaluation asks what happens when an actor deliberately manipulates the data, instructions, context or tools on which that success depends. Both are necessary: normal accuracy says little about the security of an exposed workflow.

Build the threat model around assets and actions, not only the model. Map training and fine-tuning data, retrieval sources, prompts, identities, tool permissions, outputs and external components. NIST�s 2025 taxonomy distinguishes evasion, poisoning, privacy and misuse attacks across the AI lifecycle, making clear that risk extends far beyond a malicious chat message.

Exercise realistic attack paths: instructions hidden in a document retrieved by the system; an email that attempts to redirect an agent; sensitive-data extraction through repeated queries; unsafe tool chaining; output that exploits downstream software; resource exhaustion; and a compromised supplier. Test whether controls limit the blast radius when the model follows the attacker.

Red teaming should be continuous and independent enough to challenge design assumptions. Preserve attack cases as regression tests, seed canary data, log tool calls and privilege changes, and rehearse containment with security, engineering and business owners. OWASP identifies excessive functionality, permissions and autonomy as root causes of damaging agent behaviour; prompt tuning alone cannot remove them.

Track attack success rate, unauthorised actions, data exposure, detection time, containment time and recovery quality. No defence is foolproof, so resilience matters as much as prevention. A trustworthy system is not one that never encounters hostile input; it is one whose boundaries remain enforceable when that input arrives.

Registered access

Access exclusive content and member services

Register or log in to read the full content and access exclusive insights and services reserved for registered users.

Related macro

AI and autonomous systems

Shape AI and autonomous systems around business priorities, operating models, governance and measurable outcomes.

Discover the macro

Editorial overview

Articles

Focus

Strategic challenges

POV

Strategic impact

What we observe

Get in touch

Get in touch with our experts to discuss your priorities, explore potential opportunities, and understand how our capabilities can support your organization.

Contact us
The content on this website is provided for general information only and does not constitute financial, legal, tax, or professional advice. KeynesMoore makes no representations regarding the accuracy or completeness of the information provided. Users are solely responsible for any decisions made based on this material. For comprehensive analysis and tailored strategic guidance, please schedule a consultation with our expert team. All content is proprietary to KeynesMoore and protected by copyright. Any unauthorized reproduction, distribution, or use is strictly prohibited.
®2026 KeynesMoore. All Rights Reserved.