Deep dives on the OWASP Top 10 for LLM applications, and a running record of the real-world AI security incidents of the last four years — what happened, why, and the lesson every security team should take from it.
The canonical risk framework for LLM apps, one clear explainer per category.
Unsanctioned LLMs, agents, and MCP servers are already in your fleet — here's why discovery has to come first.
Direct vs. indirect injection, real-world examples, and why OWASP ranks it as the number-one LLM risk.
The adversarial-testing duty, the evidence you'll be asked for, and the timelines that are already live.
Over-broad tool scope, typosquatting, and tool poisoning — plus a hardening checklist you can run today.
The second-ranked LLM risk is also the quietest: models that leak the very data they were trusted to handle.
Retrieval-augmented generation, the vector database it runs on, and the security seams that open up the moment you bolt your data onto a model.
A decades-old security flaw is having a renaissance — because an AI agent is the most powerful confused deputy we have ever built.
The Model Context Protocol exists to kill the N×M integration problem — and in doing so it turns every tool connection into a shared trust surface.
MCP gives an agent hands. The orchestrator is the nervous system that decides when to use them — which makes it the component that holds the credentials and the blast radius.
An LLM application is not just a model — it is a prompt-assembly pipeline with one critical trust boundary most teams draw in the wrong place.
Perceive, plan, act, observe — the loop that makes an agent an agent is also the mechanism that lets one poisoned observation hijack everything that follows.
Three agent capabilities that are each harmless alone become a data-theft machine the moment one system has all three at once.
An MCP tool description is prose the model obeys — which makes the server that writes it a place to hide instructions your agent will follow.
Vectors are not anonymised numbers — they are a lossy copy of your text, stored in a database that usually has none of your usual access controls.
Everything the model reads — your rules, the user's text, retrieved docs, tool output — lands in one flat buffer with no privilege separation. That is the whole problem.
Every model you ship is assembled from parts you did not build — scraped data, downloaded weights, community adapters and third-party plugins. Each is an injection point.
An agent acting for a user needs authority without becoming the user. That is the problem OAuth solved for apps — and the one we keep re-solving badly for agents.
Memory is what makes an agent feel personal — and what lets a single poisoned input follow a user across every future session.
Guardrails are useful and oversold in equal measure. Knowing precisely what they can enforce — and what they structurally cannot — is the difference between defence and theatre.
Traditional testing assumes the same input yields the same output. AI breaks that assumption — and with it, the meaning of a passing test.
If a passing test is only a sample, you need two disciplines to know an AI system is safe: repeatable evaluations for the known, red-teaming for the unknown.
The Model Context Protocol standardised how agents reach tools and data. Standardising the plumbing also standardised the attack surface — here is where to secure it.
The model you shipped is something you assembled, not something you built — and every borrowed piece is a link an attacker can pull.
Corrupt the data a model learns from and you corrupt the model — quietly, durably, and in ways that survive every later test you didn't think to run.
A model's output is untrusted input to whatever comes next. Treat it as trusted and you hand classic injection bugs a brand-new source.
Give a probabilistic system tools, permissions, and autonomy, and its worst mistake is bounded only by what you let it reach.
The danger isn't that someone reads your system prompt — it's what teams keep hiding inside it, assuming no one ever will.
Retrieval-augmented generation moved your sensitive data into a vector store — and most of the access controls stayed behind.
A confidently wrong model isn't just a quality problem — it's a liability, an attack surface, and, in code, a fresh supply-chain hole.
Metered, expensive inference turns old-fashioned resource abuse into a bill — and gives attackers a way to steal the model one query at a time.
Documented, publicly-reported incidents — reported factually, with the takeaway that matters.
A tribunal ruled an airline liable for its chatbot's bad advice — and rejected the idea that the bot was a separate legal entity.
Weeks after Samsung let engineers use ChatGPT, staff pasted semiconductor source code and meeting notes into it three separate times.
For a few hours in March 2023, a caching race condition let ChatGPT users see strangers' conversation titles and, for some, partial payment data.
A single misconfigured Azure token in a public AI repo exposed 38 terabytes of internal data — including workstation backups and secrets — for years.
Researchers surgically edited an open model to lie about specific facts, uploaded it under a look-alike name, and showed it passed standard benchmarks.
Lasso Security found over 1,600 valid Hugging Face tokens exposed in public code, many with write access to models from Meta, Google, and Microsoft.
Days after launch, a student coaxed Microsoft's new Bing Chat into reciting the confidential instructions it had been told never to reveal.
In March 2023, Italy's data regulator became the first in the West to block ChatGPT, turning AI privacy from an abstract worry into an operational one.
A dealership bolted ChatGPT onto its website. Within a viral afternoon, users had it agreeing to $1 SUVs and answering questions about rival brands.
Researcher Johann Rehberger showed how a rendered Markdown image could quietly smuggle a user's private data out to an attacker's server.
Vulcan Cyber showed that AI coding assistants confidently recommend software packages that don't exist — and that an attacker can register the name and wait.
In mid-2023, dark-web sellers began advertising ChatGPT clones with the safety filters stripped out, purpose-built for phishing and fraud.
A crowdsourced roleplay prompt called DAN spent 2023 trying to talk ChatGPT out of its own safety rules — and kept evolving as OpenAI patched it.
Group-IB found over 100,000 ChatGPT credentials in infostealer logs — not because ChatGPT was breached, but because of what users had typed into it.
A public AI app that turned math questions into Python was talked into running the attacker's Python instead — leaking its own API key.
Researchers showed Bing's chat could be hijacked not by what the user typed, but by invisible text on a web page it happened to read.
A UK delivery firm's support bot was talked into cursing and mocking its employer — a lesson in what happens when an update quietly removes guardrails.
NYC's official MyCity chatbot advised employers and landlords to do things that are plainly illegal — and stayed online after it was exposed.
PromptArmor showed how a message in a public Slack channel could coax Slack AI into leaking data from a private one via indirect prompt injection.
At Black Hat 2024, Zenity's Michael Bargury showed how prompt injection could bend Microsoft 365 Copilot into a phishing and data-extraction tool.
Oligo found thousands of internet-exposed Ray clusters being exploited — amid a dispute over whether it's a vulnerability or the framework working as designed.
Wiz uploaded malicious models to Replicate and SAP AI Core to cross tenant boundaries — showing a model file is executable code, not just data.
Malicious versions of Ultralytics YOLO shipped a cryptominer to PyPI — not by stealing a password, but by poisoning the GitHub Actions build cache.
As DeepSeek's models went viral, Wiz found one of its databases open to the internet with no authentication — plaintext chat logs and secret keys included.
A developer found OpenAI's ChatGPT Mac app saving every conversation in unencrypted local files, readable by any other app on the machine.
JFrog showed how a prompt to the Vanna.AI library could jump the gap from natural-language question to arbitrary code execution — CVE-2024-5565.
Microsoft's Recall feature captured everything on your screen into a local, unencrypted database — until a security backlash forced a redesign.
Aim Security disclosed a zero-click vulnerability in Microsoft 365 Copilot that could exfiltrate a user's data from a single unopened email — CVE-2025-32711.
Gemini's image generator produced historically inaccurate depictions of people and Google paused it — a case study in AI governance, not a breach.
A hacker breached OpenAI's internal employee messaging system in early 2023 — a fact the public did not learn until The New York Times reported it in July 2024.
ReversingLabs found malicious models on Hugging Face that hid a payload inside a deliberately broken pickle file to evade the platform's security scanner.