From AI pilots to enterprise performance
What separates companies that scale AI from those that accumulate experiments�and how operating models, economics and governance determine whether adoption creates measurable value.
Read articleHow does your AI behave when someone tries to break it?
Conventional evaluation asks whether an AI system succeeds with representative inputs. Adversarial evaluation asks what happens when an actor deliberately manipulates the data, instructions, context or tools on which that success depends. Both are necessary: normal accuracy says little about the security of an exposed workflow.
Build the threat model around assets and actions, not only the model. Map training and fine-tuning data, retrieval sources, prompts, identities, tool permissions, outputs and external components. NIST�s 2025 taxonomy distinguishes evasion, poisoning, privacy and misuse attacks across the AI lifecycle, making clear that risk extends far beyond a malicious chat message.
Exercise realistic attack paths: instructions hidden in a document retrieved by the system; an email that attempts to redirect an agent; sensitive-data extraction through repeated queries; unsafe tool chaining; output that exploits downstream software; resource exhaustion; and a compromised supplier. Test whether controls limit the blast radius when the model follows the attacker.
Red teaming should be continuous and independent enough to challenge design assumptions. Preserve attack cases as regression tests, seed canary data, log tool calls and privilege changes, and rehearse containment with security, engineering and business owners. OWASP identifies excessive functionality, permissions and autonomy as root causes of damaging agent behaviour; prompt tuning alone cannot remove them.
Track attack success rate, unauthorised actions, data exposure, detection time, containment time and recovery quality. No defence is foolproof, so resilience matters as much as prevention. A trustworthy system is not one that never encounters hostile input; it is one whose boundaries remain enforceable when that input arrives.
Related macro
Articles
What separates companies that scale AI from those that accumulate experiments�and how operating models, economics and governance determine whether adoption creates measurable value.
Read articleHow autonomous workflows could reshape decisions, coordination and productivity�and where human oversight remains essential as AI moves from assistance to execution.
Read articleFocus
Enterprise knowledge becomes useful to AI when evidence can be retrieved, contextualised and traced rather than merely placed inside a prompt.
The strategic question is not where AI can be used, but where it materially changes competitive position, economics or customer value.
Strategic challenges
Models, prompts, data and providers can change independently, creating operational dependencies conventional software practices may miss.
Independent models, platforms and integrations can create duplicated infrastructure and technical dependencies that compound over time.
POV
A strong AI architecture standardises what should be shared while preserving choice where technologies and requirements will continue to change.
A company can be highly capable overall and still be unready for the specific use cases it considers strategically important.
Strategic impact
Combining AI, automation and human judgement around the complete process can remove friction that task-level automation leaves untouched.
Selective sovereignty can protect critical workloads without forcing organisations to own infrastructure that offers little strategic advantage.
What we observe
We often see AI economics assessed after technology choices are made, leaving benefits estimated around investment rather than the reverse.
We frequently see collections of use cases and technology initiatives without explicit choices about competitive or business priorities.