# How to Prevent Indirect Prompt Injection

Indirect prompt injection prevention limits what a hidden instruction can do: tool allow-lists, argument checks, and approvals bound to the exact call.

Author: Formgong team
Published: 2026-10-06
Updated: 2026-10-06
Language: en
Canonical: https://formgong.com/en/blog/indirect-prompt-injection-prevention/

Indirect prompt injection prevention means limiting what a hidden instruction can do. You cannot make a model ignore every hidden instruction. Keep untrusted content away from privileged tools. Check every tool call in code. Bind each approval to the exact arguments. Allow-list where data can go. Classifiers and spotlighting lower the odds. They are not a boundary.

## Why prompt filters alone can't prevent indirect prompt injection

A phrase filter reads the text and hopes. The model can still treat a form field, a page, or a tool result as the next order. Indirect prompt injection protection limits the action, not the wording.

[OWASP LLM01:2025](https://genai.owasp.org/llmrisk/llm01-prompt-injection/) says fool-proof prevention is unclear. Models mix instructions and data. The listed steps cut the impact. They do not close the class. The definition stays on [Indirect Prompt Injection in MCP](/en/blog/indirect-prompt-injection-mcp/).

The [OWASP prevention cheat sheet](https://cheatsheetseries.owasp.org/cheatsheets/LLM_Prompt_Injection_Prevention_Cheat_Sheet.html) was checked on 06.10.2026. Pattern filters do not reliably catch indirect injection. A guard model can sit beside the deterministic checks. It does not replace them. The sheet's sample is a paragraph that calls the next block data. That paragraph is not a boundary.

[StruQ](https://arxiv.org/abs/2402.06363) is a different design. A front end splits the prompt and the data. A model is trained to follow only the prompt channel. A sentence in your system prompt is not that training.

How to protect against indirect prompt injection is this check: the host allows or denies the call. Detection can warn you. Classical injection stays mandatory. [Types of Injection Attacks on Web Forms (2026)](/en/blog/injection-taxonomy-2024-2026/) is that list. [Spam controls without a CAPTCHA](/en/blog/contact-form-spam-without-captcha/) score junk, not model instructions.

## Probabilistic vs deterministic controls

Microsoft's [MSRC post of 29 July 2025](https://www.microsoft.com/en-us/msrc/blog/2025/07/how-microsoft-defends-against-indirect-prompt-injection-attacks) splits the stack in two. A probabilistic control can lower the odds. It may still miss. A deterministic control can block one outcome even if the model is fooled. Their hard stop is a design that cannot send the data.

The same post names email, shared documents, and tool results. Prompt Shields is their classifier. Spotlighting marks untrusted text. Purview labels and a blocked image URL are their hard side. "Draft with Copilot" still needs a person to send. The first four rows are not a boundary.

ControlTypeWhat it stopsWhat it does not stopCost

System-prompt rulesProbabilisticSome copied ordersA new wordingPrompt text
SpotlightingProbabilisticSome marked attacks in the paperA block the model still obeysA text transform
Injection classifiersProbabilisticShapes it has seenA paraphrase or another languageA model call
Output screeningProbabilisticSome hostile repliesA reply that looks normalA second check
Tool allow-list per taskDeterministicA tool off the listA listed tool used badlyA name list
Argument checks in codeDeterministicA bad type, extra field, or other tenantA valid hostile valueSchema and an owner check
Approval bound to argumentsDeterministicA changed call or a reused tokenA person who approves the real callOne prompt per call
Egress and recipient allow-listDeterministicA host or address off the listAn allowed host that is hostileA host list
Provenance and taintDeterministicOutbound use of private data after untrusted text, unless approvedA label the host set wrongA flag in code
Dual LLM and CaMeLDeterministic if the split holdsUntrusted text driving a toolA lying summary or a pasted promptTwo models and a runtime
Security logRecordNothing by itselfA call another row did not denyFive fields, no bodies

Mitigation that lives only in the prompt stays in the first rows. Prevention that survives a fooled model is the later rows. The code on this page is those later rows for an MCP host.

[Microsoft Learn](https://learn.microsoft.com/en-us/security/zero-trust/sfi/defend-indirect-prompt-injection) names Prompt Shields, spotlighting, information-flow control, and a person on risky actions. It says no single layer is enough. That page has no host policy you can run.

## How to prevent indirect prompt injection in an MCP host: step by step

These seven steps are what the tests enforce. Each one names a function in `examples/mcp-host-policy.ts`. The file does not call a model. It does not open a socket.

- **Inventory tools and scopes.** Call inventoryTools on the specs the host loaded. Record each name, and whether that tool can send. A read grant is not a send grant. Do this before a task starts.

- **Allow-list tools per task.** Put only this task's names on TaskPolicy.tools. decideCall denies any other name. The model does not edit that list.

- **Label every context block.** Call labelBlock with a source and a trusted flag. A visitor submission is not trusted. Set holdsPrivateData when the session also holds private rows. The model never sets either label.

- **Validate arguments in code.** The spec lists each argument and its type. decideCall rejects a missing field, an extra field, or a wrong type. The owns callback checks that the resource belongs to ownerId.

- **Allow-list egress and recipients.** A URL, a host, or an email is checked against the task lists. A destination you did not list is denied. Mail-only HTML checks stay on the email page.

- **Bind approval to the exact call.** Show the person canonicalCall, not the model's prose. issueApproval hashes that JSON with SHA-256. The token expires, and it works once. A changed argument makes the hash fail.

- **Log security facts and test.** securityLog stores the tool, the task, the resource id, the decision, and the hash. It does not store bodies, prompts, or tokens. The five tests on this page are what the server must pass.

A tool call passes the host policy before it runs
The host labels context, checks the task allow-list, checks arguments and egress, then requires an approval token when untrusted content and private data meet an outbound tool. The log stores the decision, not the body.

1. Label the context
source and trusted

2. Allow-list the tool
this task only

3. Check arguments
owner, host, recipient

4. Approval if needed
exact call, one use

5. Log the decision
no body, no token

The host decides. The model only proposes the call.

A classifier beside step 1 is not the gate.

## Break the lethal trifecta

Simon Willison named the [lethal trifecta](https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/) on 16 June 2025. An agent is exposed when it has all three:

- **Private data.** "Access to your private data."

- **Untrusted content.** "Any mechanism by which text (or images) controlled by a malicious attacker could become available to your LLM."

- **External communication.** "The ability to externally communicate in a way that could be used to steal your data."

If an agent has all three, an attacker can trick it into reading private data and sending it out. How to stop AI's lethal trifecta is to remove one leg. Dated cases stay on [Indirect Prompt Injection Examples (2023–2026)](/en/blog/indirect-prompt-injection-examples/).

A summariser that reads submissions gets no send tool. A person sends the reply. A sender never reads raw submissions, so it cannot quote the visitor. An inbox agent has all three legs. Split read and send into different clients.

AgentLegs presentLeg to remove

Form summariserPrivate rows and untrusted fieldsNo send tool
Sender with a templateExternal sendIt never reads raw submissions
Inbox agentAll threeSplit read and send into two clients
Read-only form toolPrivate rows and untrusted fieldsNo outbound tool on that server

The policy file turns the third column into a check. Untrusted content plus private data means an outbound tool needs a matching approval. Otherwise it is denied. The code reads the label. The model does not set it.

## Code: an MCP-host policy layer in TypeScript

`examples/mcp-host-policy.ts` is the file the tests run. It is a host gate, not a detector. It will not tell you that a sentence is an injection. It refuses a call the task may not make.

Pass the task policy, the tool spec, the arguments, the labels, and any approval. `decideCall` checks the allow-list, the schema, the owner, egress, and the trifecta rule. Show the person `canonicalCall`, not the model's prose. `issueApproval` binds a SHA-256 token to that JSON. `securityLog` keeps five fields.

Pinning the catalog is on [MCP Rug Pull Attack: Detect Tool Changes](/en/blog/mcp-rug-pull-attack/). Scanning metadata is on [what tool poisoning in MCP is, with a scanner you can run](/en/blog/mcp-tool-poisoning/). Email HTML stays on [Indirect Prompt Injection via Email](/en/blog/indirect-prompt-injection-via-email/).

```typescript
/**
 * Decide whether an MCP host may run one tool call.
 *
 * The host passes the task policy, the tool spec, the proposed arguments,
 * context labels, and any approval token. decideCall checks the allow-list,
 * the schema, ownership, egress, and the lethal-trifecta rule.
 *
 * This file does not detect injections. It does not open a network connection.
 * The model never sets the labels, the allow-lists, or the owner.
 * Show the human canonicalCall, not the model's prose.
 */

export class HostPolicyError extends Error {
  constructor(message: string) {
    super(message);
    this.name = "HostPolicyError";
  }
}

/** A block of context. The host sets both fields. The model does not. */
export type ContextBlock = {
  source: string;
  trusted: boolean;
};

export type TaskPolicy = {
  id: string;
  /** Tool names this task may call. */
  tools: readonly string[];
  /** Hostnames an outbound URL or host argument may use. */
  hosts: readonly string[];
  /** Email addresses an outbound recipient may use. */
  recipients: readonly string[];
};

export type ArgKind = "string" | "number" | "boolean";

export type ToolSpec = {
  name: string;
  /** True when the tool can send, post, or fetch an outside destination. */
  outbound: boolean;
  arguments: Readonly<Record<string, ArgKind>>;
  /** Argument names whose value is a URL or a bare host. */
  urlArguments?: readonly string[];
  /** Argument names whose value is one email, a list, or a comma-separated list. */
  recipientArguments?: readonly string[];
  /** Argument copied into the security log as the resource id. */
  resourceArgument?: string;
};

export type ProposedCall = {
  name: string;
  arguments: Readonly<Record<string, unknown>>;
};

export type ApprovalRecord = {
  /** SHA-256 hex of canonicalCall. Valid only for that exact call. */
  hash: string;
  /** Epoch milliseconds. Invalid at or after this time. */
  expiresAt: number;
  used: boolean;
};

/** True when formId belongs to ownerId. The host supplies this. The model does not. */
export type Ownership = (resourceId: string, ownerId: string) => boolean;

export type CallInput = {
  task: TaskPolicy;
  spec: ToolSpec;
  call: ProposedCall;
  ownerId: string;
  blocks: readonly ContextBlock[];
  /** Host label. True when this session also holds private data. */
  holdsPrivateData: boolean;
  owns: Ownership;
  now: number;
  approval?: ApprovalRecord | null;
};

export type Decision = {
  allowed: boolean;
  reason: string;
  approvalHash: string | null;
  resourceId: string | null;
  /** What to show a person. Canonical JSON of the call. */
  canonical: string;
  approval: ApprovalRecord | null;
};

export type SecurityLog = {
  tool: string;
  task: string;
  resourceId: string | null;
  decision: "allow" | "deny";
  approvalHash: string | null;
};

const EMAIL = /^[a-z0-9._%+-]+@[a-z0-9.-]+\.[a-z]{2,}$/i;

function isRecord(value: unknown): value is Record<string, unknown> {
  return !!value && typeof value === "object" && !Array.isArray(value);
}

function sortValue(value: unknown): unknown {
  if (Array.isArray(value)) return value.map(sortValue);
  if (!isRecord(value)) return value;
  const out: Record<string, unknown> = {};
  for (const key of Object.keys(value).sort()) out[key] = sortValue(value[key]);
  return out;
}

function requireId(value: string, label: string): string {
  if (typeof value !== "string" || value.trim() === "") throw new HostPolicyError(`${label} is required.`);
  return value;
}

/** Stamp a context block. Code reads this label later. The model does not. */
export function labelBlock(source: string, trusted: boolean): ContextBlock {
  if (typeof source !== "string" || source.trim() === "") throw new HostPolicyError("source is required.");
  if (typeof trusted !== "boolean") throw new HostPolicyError("trusted must be a boolean.");
  return { source: source.trim(), trusted };
}

/** Names and outbound flags for the tools this host loaded. */
export function inventoryTools(specs: readonly ToolSpec[]): Array<{ name: string; outbound: boolean }> {
  if (!Array.isArray(specs)) throw new HostPolicyError("specs must be an array.");
  const names = new Set<string>();
  return specs.map((spec) => {
    if (!isRecord(spec) || typeof spec.name !== "string" || spec.name.trim() === "") {
      throw new HostPolicyError("Each tool needs a name.");
    }
    if (names.has(spec.name)) throw new HostPolicyError(`Duplicate tool name: ${spec.name}`);
    names.add(spec.name);
    if (typeof spec.outbound !== "boolean") throw new HostPolicyError(`${spec.name} needs an outbound flag.`);
    return { name: spec.name, outbound: spec.outbound };
  });
}

/** Stable JSON of the name and arguments. This is the text a person approves. */
export function canonicalCall(call: ProposedCall): string {
  if (!isRecord(call)) throw new HostPolicyError("call must be an object.");
  if (typeof call.name !== "string" || call.name.trim() === "") throw new HostPolicyError("call.name is required.");
  if (!isRecord(call.arguments)) throw new HostPolicyError("call.arguments must be an object.");
  return JSON.stringify(sortValue({ name: call.name, arguments: call.arguments }));
}

export async function sha256Hex(text: string): Promise<string> {
  const digest = await crypto.subtle.digest("SHA-256", new TextEncoder().encode(text));
  return [...new Uint8Array(digest)].map((byte) => byte.toString(16).padStart(2, "0")).join("");
}

/** Build a single-use token for this exact canonical call. */
export async function issueApproval(call: ProposedCall, expiresAt: number, now: number): Promise<ApprovalRecord> {
  if (!Number.isFinite(expiresAt) || !Number.isFinite(now)) throw new HostPolicyError("expiry and now must be numbers.");
  if (expiresAt <= now) throw new HostPolicyError("approval expiry must be in the future.");
  return { hash: await sha256Hex(canonicalCall(call)), expiresAt, used: false };
}

function hostnameOf(value: string): string | null {
  const trimmed = value.trim();
  if (!trimmed || /\s/.test(trimmed)) return null;
  let url: URL;
  try {
    url = new URL(trimmed.includes("://") ? trimmed : `https://${trimmed}`);
  } catch {
    return null;
  }
  if (url.username || url.password) return "";
  const host = url.hostname.toLowerCase();
  if (!host.includes(".")) return null;
  return host;
}

function emailsOf(value: unknown): string[] | null {
  const parts = Array.isArray(value)
    ? value
    : typeof value === "string"
      ? value.split(",")
      : null;
  if (!parts || parts.some((part) => typeof part !== "string")) return null;
  const emails = parts.map((part) => part.trim()).filter((part) => part !== "");
  if (!emails.length || emails.some((email) => !EMAIL.test(email))) return null;
  return emails.map((email) => email.toLowerCase());
}

function checkValue(spec: ToolSpec, key: string, value: unknown, task: TaskPolicy): string | null {
  const kind = spec.arguments[key];
  if (kind === "string" && typeof value !== "string") return `${key} has the wrong type`;
  if (kind === "number" && typeof value !== "number") return `${key} has the wrong type`;
  if (kind === "boolean" && typeof value !== "boolean") return `${key} has the wrong type`;
  if (typeof value !== "string") return null;
  const urlKeys = new Set(spec.urlArguments ?? []);
  const mailKeys = new Set(spec.recipientArguments ?? []);
  if (urlKeys.has(key) || /^[a-z][a-z0-9+.-]*:\/\//i.test(value.trim())) {
    const host = hostnameOf(value);
    if (!host) return `${key} is not an allow-listed host`;
    if (!task.hosts.map((item) => item.toLowerCase()).includes(host)) return `${key} host is not allow-listed`;
  }
  if (mailKeys.has(key) || (!urlKeys.has(key) && EMAIL.test(value.trim()))) {
    const emails = emailsOf(value);
    const allowed = new Set(task.recipients.map((item) => item.toLowerCase()));
    if (!emails || emails.some((email) => !allowed.has(email))) return `${key} recipient is not allow-listed`;
  }
  return null;
}

function resourceIdOf(spec: ToolSpec, args: Readonly<Record<string, unknown>>): string | null {
  if (!spec.resourceArgument) return null;
  const value = args[spec.resourceArgument];
  return typeof value === "string" ? value : null;
}

/**
 * Allow or deny one call. On allow with an approval, that record is marked used.
 * A later call with the same record, or with different arguments, is denied.
 */
export async function decideCall(input: CallInput): Promise<Decision> {
  if (!isRecord(input)) throw new HostPolicyError("input must be an object.");
  const taskId = requireId(input.task?.id, "task.id");
  if (!Array.isArray(input.task.tools) || !Array.isArray(input.task.hosts) || !Array.isArray(input.task.recipients)) {
    throw new HostPolicyError("task lists must be arrays.");
  }
  if (!isRecord(input.spec) || typeof input.spec.outbound !== "boolean" || !isRecord(input.spec.arguments)) {
    throw new HostPolicyError("spec is incomplete.");
  }
  requireId(input.ownerId, "ownerId");
  if (!Array.isArray(input.blocks)) throw new HostPolicyError("blocks must be an array.");
  if (typeof input.holdsPrivateData !== "boolean") throw new HostPolicyError("holdsPrivateData must be a boolean.");
  if (typeof input.owns !== "function") throw new HostPolicyError("owns must be a function.");
  if (!Number.isFinite(input.now)) throw new HostPolicyError("now must be a number.");
  const canonical = canonicalCall(input.call);
  const resourceId = resourceIdOf(input.spec, input.call.arguments);
  const presented = input.approval ?? null;
  const base = { canonical, resourceId, approvalHash: presented?.hash ?? null, approval: presented };

  const deny = (reason: string): Decision => ({ ...base, allowed: false, reason });

  if (input.spec.name !== input.call.name) return deny("tool spec does not match the call");
  if (!input.task.tools.includes(input.call.name)) return deny("tool is not on this task");

  for (const key of Object.keys(input.call.arguments)) {
    if (!(key in input.spec.arguments)) return deny(`${key} is not in the schema`);
  }
  for (const key of Object.keys(input.spec.arguments)) {
    if (!(key in input.call.arguments)) return deny(`${key} is missing`);
    const problem = checkValue(input.spec, key, input.call.arguments[key], input.task);
    if (problem) return deny(problem);
  }

  if (input.spec.resourceArgument) {
    if (!resourceId) return deny("resource id is missing");
    let owned = false;
    try {
      owned = input.owns(resourceId, input.ownerId) === true;
    } catch {
      return deny("ownership check failed");
    }
    if (!owned) return deny("resource is not owned by this account");
  }

  const untrusted = input.blocks.some((block) => {
    if (!isRecord(block) || typeof block.trusted !== "boolean") throw new HostPolicyError("each block needs a trusted flag.");
    return block.trusted === false;
  });
  const needsApproval = input.spec.outbound && untrusted && input.holdsPrivateData;
  if (needsApproval && !presented) return deny("approval required");

  if (presented) {
    if (presented.used) return deny("approval already used");
    if (!Number.isFinite(presented.expiresAt) || input.now >= presented.expiresAt) return deny("approval expired");
    const hash = await sha256Hex(canonical);
    if (presented.hash !== hash) return deny("approval does not match this call");
    presented.used = true;
  }

  void taskId;
  return { ...base, allowed: true, reason: "allowed", approval: presented };
}

/** Security facts only. Arguments, prompts, bodies, and tokens are not copied. */
export function securityLog(taskId: string, decision: Decision): SecurityLog {
  requireId(taskId, "task.id");
  if (!isRecord(decision) || typeof decision.allowed !== "boolean") throw new HostPolicyError("decision is incomplete.");
  const tool = JSON.parse(decision.canonical) as { name?: string };
  if (typeof tool.name !== "string") throw new HostPolicyError("decision has no tool name.");
  return {
    tool: tool.name,
    task: taskId,
    resourceId: decision.resourceId,
    decision: decision.allowed ? "allow" : "deny",
    approvalHash: decision.approvalHash,
  };
}

type NodeProcess = {
  argv: string[];
  stdin: {
    on(event: "data", listener: (chunk: Uint8Array) => void): void;
    on(event: "end", listener: () => void): void;
    on(event: "error", listener: (error: unknown) => void): void;
  };
  stdout: { write(text: string): void };
  stderr: { write(text: string): void };
  exitCode?: number;
};

function nodeProcess(): NodeProcess | null {
  const proc = (globalThis as { process?: NodeProcess }).process;
  if (!proc?.argv || !proc.stdin || !proc.stdout || !proc.stderr) return null;
  return proc;
}

function readStdin(proc: NodeProcess): Promise<string> {
  const decoder = new TextDecoder();
  let raw = "";
  return new Promise((resolve, reject) => {
    proc.stdin.on("data", (chunk) => {
      raw += decoder.decode(chunk, { stream: true });
    });
    proc.stdin.on("end", () => resolve(raw + decoder.decode()));
    proc.stdin.on("error", reject);
  });
}

async function main(proc: NodeProcess): Promise<void> {
  const raw = (await readStdin(proc)).trim();
  if (!raw) throw new HostPolicyError("Pass a call JSON document on stdin.");
  let parsed: unknown;
  try {
    parsed = JSON.parse(raw);
  } catch {
    throw new HostPolicyError("stdin is not JSON.");
  }
  if (!isRecord(parsed)) throw new HostPolicyError("stdin JSON must be an object.");
  const owned = new Set(Array.isArray(parsed.ownedIds) ? parsed.ownedIds.filter((id) => typeof id === "string") : []);
  const decision = await decideCall({
    task: parsed.task as TaskPolicy,
    spec: parsed.spec as ToolSpec,
    call: parsed.call as ProposedCall,
    ownerId: String(parsed.ownerId ?? ""),
    blocks: (parsed.blocks as ContextBlock[]) ?? [],
    holdsPrivateData: parsed.holdsPrivateData === true,
    owns: (resourceId) => owned.has(resourceId),
    now: typeof parsed.now === "number" ? parsed.now : Date.now(),
    approval: (parsed.approval as ApprovalRecord | null) ?? null,
  });
  const log = securityLog((parsed.task as TaskPolicy).id, decision);
  proc.stdout.write(`${JSON.stringify({ decision: decision.reason, log }, null, 2)}\n`);
  proc.exitCode = decision.allowed ? 0 : 1;
}

const node = nodeProcess();
if (node?.argv[1]?.endsWith("mcp-host-policy.ts")) {
  main(node).catch((error: unknown) => {
    const message = error instanceof Error ? error.message : "Policy check failed.";
    node.stderr.write(`${message}\n`);
    node.exitCode = 2;
  });
}

```

**What a denied outbound call looks like in the log**

```json
{
  "decision": "approval required",
  "log": {
    "tool": "send_note",
    "task": "summarise-form",
    "resourceId": "form-1",
    "decision": "deny",
    "approvalHash": null
  }
}
```

## Five tests that prove your controls work

`test/mcp-host-policy.test.ts` is what the server must pass. A green model score does not. The marker is `[PLACEHOLDER INSTRUCTION]`. The host that must fail is `attacker.example`.

- **A tool outside the task is denied.** The task list has no `delete_row`. The reason is "tool is not on this task".

- **A cross-tenant form id is denied.** `owns` is true only for `form-1` and `owner-1`. `form-2` returns "resource is not owned by this account".

- **An approval reused with changed arguments is denied.** The token was issued for one canonical call. A different note fails the hash. After a success, the same token returns "approval already used".

- **A host or recipient off the list is denied.** `https://attacker.example/log` and `sink@attacker.example` both fail. An address on the list can still pass.

- **Untrusted content, private data, and an outbound tool require approval.** With no token, the reason is "approval required". A fresh token for that exact call allows it once.

The log JSON has the tool, the task, the resource id, the decision, and the hash. It does not contain the placeholder text.

## Spotlighting, classifiers and other probabilistic layers

What spotlighting changes
Spotlighting marks untrusted text. Hines and co-authors describe it in [arXiv:2403.14720](https://arxiv.org/abs/2403.14720). Delimiting wraps the block. Datamarking repeats a marker through it. Encoding transforms it, for example with base64.

The abstract says: "Using GPT-family models, we find that spotlighting reduces the attack success rate from greater than 50% to below 2% in our experiments." That is those models and those tests. The body splits by mode. On GPT-3.5-Turbo, delimiters cut the rate by about half. Datamarking took that model from about 50% to below 3%. Encoding took summarization to 0.0% and a question task to 1.8% on that model. The attacks were keyword payloads in a synthetic set.

A model can still obey the mark. Do not make it the only control.

Classifiers and vendor layers
A classifier is a signal. The cheat sheet, checked on 06.10.2026, says a trained model catches cases a regex misses. The guard itself can be injected. A miss is still a call the host must deny.

Google, on [13 June 2025](https://blog.google/security/mitigating-prompt-injection-attacks/), layers classifiers, Markdown cleanup, and redacted URLs around Gemini. A user confirms a risky action, such as deleting a calendar event. A notice appears when a defense fires. The post does not close the class.

Anthropic, on [24 November 2025](https://www.anthropic.com/research/prompt-injection-defenses), trained browser agents on injections in pages. They scan untrusted content with classifiers. They wrote that prompt injection is far from solved. A 1% attack-success rate on their internal test, they said, still represents meaningful risk.

On [25 May 2026](https://www.anthropic.com/engineering/how-we-contain-claude) they described containment. People approved roughly 93% of permission prompts. An approval here is rare, and it is bound to the exact call. Any probabilistic defense has a non-zero miss rate. A sandbox and an egress rule still hold.

Cloudflare, on [29 April 2026](https://www.cloudflare.com/cloudforce-one/research/adversarial-deception-a-study-of-indirect-prompt-code-injection/), scored scripts with hidden instructions in comments. When those comments were under 1% of the file, detection across the models fell to 53%. Success also depended on the model tier. A score is not proof the field is safe to hand to a tool.

[MELON](https://arxiv.org/abs/2502.05174) is a detector, not a host boundary. It masks the user task, runs again, and flags similar tool calls. MELON-Aug is reported at 0.32% attack success and 68.72% utility on GPT-4o in AgentDojo. That is their benchmark, not a Formgong result.

## Design patterns: dual LLM and CaMeL

What the dual LLM pattern does
Willison described the [dual LLM pattern](https://simonwillison.net/2023/Apr/25/dual-llm-pattern/) on 25 April 2023. A privileged model sees the trusted request and may call tools. A quarantined model reads untrusted content and has no tools. A controller program passes tokens such as `$VAR1`. The privileged model must not see the raw untrusted text.

He called the pattern pretty bad. It is harder to use. A user who pastes untrusted text into the trusted box breaks the assumption. His update of 11 April 2025 points to CaMeL.

What CaMeL adds, and what it does not
[CaMeL](https://arxiv.org/abs/2503.18813) is "Defeating Prompt Injections by Design", arXiv:2503.18813. Searches for camel prompt injection point here. A privileged model plans from the trusted query. It does not read untrusted documents. A quarantined model may read them and has no tools. Values carry capabilities. An interpreter checks a policy when a tool runs.

On AgentDojo it solved 77% of tasks with provable security, against 84% undefended. It does not stop a lying summary or a phishing line that leaves the data flow unchanged. It assumes a trusted user prompt and clean memory. Some queries still need a person.

The cheat sheet, checked on 06.10.2026, repeats the authors' warning. The released code may contain security bugs. They do not plan to maintain it. Treat the repository as research. Do not ship it as a supported control.

Beurer-Kellner and co-authors, [arXiv:2506.08837](https://arxiv.org/abs/2506.08837), state a shared rule. "Once an LLM agent has ingested untrusted input, it must be constrained so that it is impossible for that input to trigger any consequential actions." The host policy here is the small version. Untrusted text cannot, by itself, cause a send.

## MCP-specific checklist

The spec revision is 2026-07-28, checked on 06.10.2026.

- **Pin the tool list** you approved. The hash is on [MCP Rug Pull Attack: Detect Tool Changes](/en/blog/mcp-rug-pull-attack/).

- **Scan tool metadata on load.** Names and schema text are on [what tool poisoning in MCP is](/en/blog/mcp-tool-poisoning/).

- **Treat mail as untrusted.** Recipients and Markdown images are on [Indirect Prompt Injection via Email](/en/blog/indirect-prompt-injection-via-email/).

- **Use OAuth 2.1 with PKCE S256.** The [authorization section](https://modelcontextprotocol.io/specification/2026-07-28/basic/authorization/security-considerations) says clients must use S256 when they can.

- **Bind the token to one resource.** That section requires the RFC 8707 resource indicator. Reject a token issued for a different server.

- **Do not pass a client token through.** The spec says the MCP server must not forward the client's token.

- **Keep read and send servers in different clients.** One session lets every description vote on every tool.

- **Get consent for the real client.** The [best-practices note](https://modelcontextprotocol.io/docs/2026-07-28/tutorials/security/security_best_practices) covers the confused deputy. Do not reuse one approval for a client the user did not see.

- **Allow-list tools per task,** then check arguments. That is the module above.

- **Log the decision, not the body.** Five fields are enough.

Microsoft's note of [28 April 2025](https://developer.microsoft.com/blog/protecting-against-indirect-injection-attacks-mcp/) is about protecting against indirect prompt injection attacks in MCP. It points at shields and delimiters. It does not ship a host policy.

## What Formgong does about this

Formgong is a free form backend for static and AI-built sites that delivers submissions to Telegram and email, stores data in the EU, and works in 12 languages. On the free plan, email is a daily digest at 08:00 the next morning, in your time zone, for the previous day. Telegram is instant on every plan. Paid plans can email each submission. Plans are on the [pricing page](/en/pricing/).

The [MCP server](/en/docs/mcp/) was checked on 06.10.2026 against this repo and [llms.txt](https://formgong.com/llms.txt). It lists forms, creates a form, and returns a snippet. With opt-in `submissions:read`, it lists recent submissions: time and fields, not IP or browser data. The read tool says those fields are untrusted data, never instructions. The result repeats that notice. There is no send tool and no delete tool. The outbound leg is absent on this server.

Sign-in is authorization code with PKCE S256. The token is bound to this MCP URL. Session cookies do not authorize the endpoint. `initialize` sets `listChanged` to false. None of that covers a second server in the same chat. If that server can send, the trifecta is back. This policy does not run inside it.

A Formgong notice is a label. The host still has to enforce the steps above. Security reports go to [support@formgong.com](mailto:support@formgong.com).

## Frequently asked questions

### How do you prevent indirect prompt injection?

Limit what a hidden instruction can do. Keep untrusted content away from privileged tools. Check every call in code. Bind approval to the exact arguments. Allow-list where data can go. Filters lower the odds. They are not a boundary.

### What is the lethal trifecta?

Simon Willison, on 16 June 2025, named three legs: private data, untrusted content, and a way to send data out. An agent with all three can be tricked into leaking. Remove one leg. A form summariser should not have a send tool.

### What is CaMeL?

CaMeL is a Google DeepMind design, arXiv:2503.18813. A privileged model plans from the trusted query. A quarantined model reads untrusted data and has no tools. On AgentDojo it solved 77% of tasks with a security proof, against 84% undefended. It does not stop a lying summary.

### Does spotlighting stop prompt injection?

Not by itself. In the Spotlighting paper, GPT-family tests fell from above 50% attack success to below 2%. The body splits that by mode and model. The mark lowers the odds. A tool allow-list is what blocks the call.

### Can a classifier detect indirect prompt injection?

Sometimes. OWASP says a trained classifier catches cases a regex misses, and that the classifier can be fooled too. Cloudflare's 29 April 2026 report saw detection fall to 53% when the hidden text was a tiny part of the file. Treat the score as a signal.

## Sources and documentation

- [OWASP LLM01:2025 Prompt Injection](https://genai.owasp.org/llmrisk/llm01-prompt-injection/)
- [OWASP LLM Prompt Injection Prevention Cheat Sheet (checked 6 October 2026)](https://cheatsheetseries.owasp.org/cheatsheets/LLM_Prompt_Injection_Prevention_Cheat_Sheet.html)
- [Chen et al., StruQ, arXiv:2402.06363](https://arxiv.org/abs/2402.06363)
- [Microsoft MSRC: defending against indirect prompt injection (29 July 2025)](https://www.microsoft.com/en-us/msrc/blog/2025/07/how-microsoft-defends-against-indirect-prompt-injection-attacks)
- [Hines et al., Spotlighting, arXiv:2403.14720](https://arxiv.org/abs/2403.14720)
- [Microsoft Learn: Defend against indirect prompt injection attacks](https://learn.microsoft.com/en-us/security/zero-trust/sfi/defend-indirect-prompt-injection)
- [Microsoft: indirect prompt injection in MCP (28 April 2025)](https://developer.microsoft.com/blog/protecting-against-indirect-injection-attacks-mcp/)
- [Simon Willison: The lethal trifecta for AI agents (16 June 2025)](https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/)
- [Simon Willison: The Dual LLM pattern (25 April 2023)](https://simonwillison.net/2023/Apr/25/dual-llm-pattern/)
- [Debenedetti et al., CaMeL, arXiv:2503.18813](https://arxiv.org/abs/2503.18813)
- [Beurer-Kellner et al., Design Patterns for Securing LLM Agents, arXiv:2506.08837](https://arxiv.org/abs/2506.08837)
- [Google: layered prompt-injection defense (13 June 2025)](https://blog.google/security/mitigating-prompt-injection-attacks/)
- [Anthropic: prompt injection in browser use (24 November 2025)](https://www.anthropic.com/research/prompt-injection-defenses)
- [Anthropic: containing Claude (25 May 2026)](https://www.anthropic.com/engineering/how-we-contain-claude)
- [Cloudflare Cloudforce One: indirect prompt injection (29 April 2026)](https://www.cloudflare.com/cloudforce-one/research/adversarial-deception-a-study-of-indirect-prompt-code-injection/)
- [Zhu et al., MELON, arXiv:2502.05174](https://arxiv.org/abs/2502.05174)
- [MCP specification, revision 2026-07-28](https://modelcontextprotocol.io/specification/2026-07-28)
- [MCP authorization security considerations (2026-07-28)](https://modelcontextprotocol.io/specification/2026-07-28/basic/authorization/security-considerations)
- [MCP security best practices (2026-07-28)](https://modelcontextprotocol.io/docs/2026-07-28/tutorials/security/security_best_practices)
- [Formgong MCP server](https://formgong.com/en/docs/mcp/)

[Get a form key](https://formgong.com/en/#top)
