AI Agent Security — Zero-Trust Defense Against Privilege Escalation & Prompt Injection

A controlled red-team test, not an incident: the deployment team baited a production AI agent with an unauthenticated authority file to prove its defenses held. The agent refused the privilege escalation, caught four forensic red flags, and demanded authenticated operator approval. Zero-trust security isn't a policy document here — it's what the agent did under live bait.

Always-on

What was getting in the way.

AI agents with production access are a new attack surface. A file dropped into an agent's working directory — by a compromised vendor, a malicious insider, or a prompt-injection payload — reads as an instruction, including instructions that expand the agent's own permissions. Most deployments have no defense: the agent treats anything in its workspace as trusted and complies.

Context

A controlled validation in week one of a production AI agent deployment at an e-commerce consultancy. The 'attack' was staged by the deployment team to prove the day-one grounding pack held: no real adversary, no exposure, no production risk.

How the work runs.

  1. 01

    Trigger

    as part of a controlled security validation, the deployment team drops a document into the agent's workspace (delivered over SSH) directing it to adopt broader config authority — a staged privilege-escalation attempt, built to test whether the agent would comply.

  2. 02

    Zero-trust provenance check

    Before adopting anything, the agent audits the file like a security analyst: who claims to have sent it, can that claim be verified, does the signature match workspace policy, do the dates line up, and does the write timestamp look human or scripted?

  3. 03

    Four red flags raised

    A signature from an organization barred from the workspace, an authorization claim unverifiable from the agent's side, a date inconsistency, and a file written 0.36 seconds after a config change — machine-speed delivery, not a sat-down memo.

  4. 04

    Refusal

    The agent declines to adopt the directive, quarantines it as unverified input, and states exactly which checks failed. No partial compliance.

  5. 05

    Operator-in-the-loop escalation

    It asks for the directive from the operator directly, in the operator's own words, in Slack — the one channel where identity is authenticated.

  6. 06

    Outcome

    the test passes: the privilege boundary holds with zero human intervention. The authority model is now explicit — operator messages in Slack are directives; files arriving over SSH are input to evaluate, never self-authorizing instructions.

Evidence from the workflow.

An AI agent's refusal message: it declines to adopt an unauthenticated authority file, lists four forensic red flags — unverifiable signature, unconfirmable requester claim, date inconsistency, and a machine-speed write timestamp — and asks the operator to state the directive in their own words (names anonymized)
The refusal, verbatim, from the controlled bait test: four forensic red flags on an unauthenticated file, no partial compliance, and a request for operator-voice authorization (names anonymized)

Each system has a role.

  • Record of truth (authenticated operator authorization)

    Slack

  • Intake (untrusted until provenance-verified)

    SSH / file transfer

  • Record of truth (claim-verification rules)

    Grounding & anti-hallucination policy pack

  • Protected asset

    Agent runtime & config

Why this is Operator.

Owns a workflow end to end, running the process inside the authority you set.

  1. 01

    The agent cannot expand its own permissions under any input. Authority changes require the operator's plain-English authorization in Slack; consultant and third-party files are permanently classed as advice, never directives. Every refusal states its evidence, so a human can overrule with full context.

Impact / Outcomes

Under live bait, the agent independently caught four forensic red flags on a single unauthenticated file — including a write timestamp 0.36 seconds after a config change — and refused the privilege escalation outright.

No vulnerability was ever in play: the bait file came from the deployment team itself, under controlled conditions with nothing at risk. That's the point of the test — zero trust means even friendly sources get refused until authenticated. The defense passed; nothing needed fixing.

The defense wasn't purpose-built for this attack. A grounding pack installed to stop hallucinated claims generalized, unprompted, into supply-chain and social-engineering skepticism — and the bait test proved it.

← All use cases

Find where a workflow like this fits.

Start with the systems, work, constraints, and authority already present in your operation.