What Exploit Research Taught Me About AI Agent Security

REDSAND | THE CTO FIELD NOTES

Same Principles. A New Frontier.

By Tim Shelton | Founder & CTO, HAWK Network Defense

I have spent a significant part of my career looking at systems from a perspective their designers did not necessarily intend.

That is what exploit research teaches you to do.

You do not begin by asking:

What was this system designed to do?

You ask:

What else is possible?

What happens if inputs arrive in an unexpected sequence?

What if two individually legitimate capabilities are combined?

What if an assumption made by one component is not shared by another?

What if an access path exists that nobody considered important because it was never part of the expected workflow?

The interesting failures usually live in those gaps.

And that is why I think many of the lessons cybersecurity learned through exploit research are becoming newly important as organizations deploy AI agents.

AI changes the speed and scale of the problem.

It does not change the underlying security principle.

Intent is not a security boundary.

If a system is technically capable of doing something, we should assume that capability may eventually be discovered.

That was true for applications.

It was true for networks.

It was true for identity systems.

And it will be true for autonomous AI agents.

The difference is that an AI agent may actively reason about the environment in which it operates.

It may discover another tool. Another API. Another data source. Another credential. Another workflow.

Another path toward its objective.

That makes security architecture matter more, not less.

1. Never Confuse Intended Behavior With Possible Behavior

Security testing has always challenged assumptions.

A developer may say:

"The application should never receive that input."

The security researcher asks:

"What happens if it does?"

An AI agent deserves the same treatment.

Production environments are messy.

Permissions change. APIs evolve. Systems fail. Dependencies appear. Data conflicts. Unexpected capabilities become available.

Agent security therefore cannot be evaluated only against the expected workflow.

We need to understand what happens when:

  • Evidence is incomplete.

  • Two tools return conflicting results.

  • A connected system exposes more capability than expected.

  • Credentials provide broader access than the workflow requires.

  • The agent discovers an alternative route to its objective.

  • A dependency fails halfway through execution.

  • The environment no longer matches assumptions made at deployment.

A mature AI security architecture assumes the unexpected path exists before the agent finds it.

2. Capability and Authority Are Not the Same Thing

An AI agent may be technically capable of performing an action.

That does not mean it should be authorized to perform it.

Security has dealt with this distinction for decades.

A service account may technically have write permission.

A process may technically be able to reach another network segment.

A user may technically be capable of accessing information beyond what is necessary for the task.

Good architecture doesn't merely ask what is possible.

It determines what is permitted.

Agents make this problem more interesting because capabilities can combine.

An agent that can read customer information, call an external API, create content, and communicate externally may have a substantially different effective capability than any of those permissions suggests individually.

AI agent security architecture therefore needs to enforce:

  • Identity

  • Permissions

  • Purpose

  • Scope

  • Action-level authorization

  • Runtime boundaries

  • Revocation

  • Evidence preservation

  • Outcome verification

The architecture should determine what the agent is allowed to do even when the agent discovers that it could do more.

3. Assume AI Agent Capabilities Will Combine

Exploit chains are rarely interesting because of one capability in isolation.

They become interesting because several small conditions combine.

A weak permission.

A trusted relationship.

An overlooked endpoint.

A credential.

An unexpected sequence.

Each may appear relatively harmless independently.

Together, they create a path nobody intended.

AI agents introduce a similar problem.

We should therefore test agents not simply for individual permissions but for the authority that emerges when those permissions are combined.

Consider an agent that can retrieve sensitive information, transform the content, invoke external services, and communicate results.

Each capability may be legitimate within its intended workflow.

But what happens when the agent combines them in a sequence the system designer did not anticipate?

The architecture has to account for those combinations rather than assuming that individually approved permissions always produce an acceptable overall outcome.

What can the system accomplish when all of its legitimate capabilities are combined?

That is a much harder question than:

What tools does the agent have?

But it is the question security teams need to answer.

4. The Prompt Is Not the Security Boundary

Prompts matter.

System instructions matter.

Behavioral constraints matter.

But none should substitute for enforceable security controls.

"Do not access that information" is guidance.

Access control is enforcement.

"Do not send information outside this environment" is guidance.

Egress control is enforcement.

"Do not modify production" is guidance.

Permission boundaries are enforcement.

Prompts can influence behavior.

Architecture determines what behavior is actually possible.

This distinction is critical when AI agents interact with enterprise systems, privileged credentials, sensitive information, or production workflows.

A prompt should not be responsible for enforcing a boundary that the underlying system has the technical means to enforce directly.

That means treating authorization, connector access, tool permissions, network restrictions, and runtime controls as security architecture—not merely instructions provided to the model.

5. Human-in-the-Loop Should Not Mean Human-in-the-Way

One response to autonomous risk is to insert manual approval before every action.

That can feel safer.

It can also recreate the operational bottleneck automation was supposed to eliminate.

Human control does not require someone to click APPROVE every time an agent encounters a routine, well-understood condition.

Human judgment can be incorporated into the operating model before the event occurs.

People can define:

  • Required evidence

  • Permitted actions

  • Protected systems

  • Decision thresholds

  • Human-only decisions

  • Escalation conditions

  • Conditions that reduce or revoke authority

That allows an agent to operate within bounded authority while preserving human accountability.

I increasingly like the phrase human at the helm for this reason.

The human does not need to manipulate every control to remain in control of the system.

People establish the destination, boundaries, delegated authority, conditions for intervention, and definition of success.

Automation can operate the controls.

Accountability still belongs at the helm.

6. AI Agent Authority Should Be Earned and Reversible

We would not give a new employee unrestricted privileged access simply because that person appeared capable.

AI agents shouldn't receive maximum production authority by default either.

Production autonomy can be progressive:

OBSERVE → RECOMMEND → ACT → VERIFY → EARN GREATER AUTHORITY

An agent may begin by observing activity and assembling evidence.

As confidence in its performance and reliability increases, it may be permitted to recommend actions or execute specific actions within defined boundaries.

Greater authority should be tied to validated evidence, demonstrated reliability, and the risk associated with the action.

And that progression cannot operate in only one direction.

If evidence quality deteriorates, exceptions increase, permissions drift, outcomes become unreliable, or the operating environment materially changes, authority should be capable of contracting.

That may require additional verification, narrower permissions, human escalation, or revocation of previously delegated authority.

That is not failure.

That is governance working.

The mature system isn't merely capable of gaining autonomy.

It can lose autonomy when the evidence no longer justifies it.

7. Verification Is Part of the Action

Successful execution is not necessarily a successful outcome.

An API can return success.

A command can complete.

A workflow can close.

And the underlying security condition can remain unchanged.

If an agent isolates an endpoint, did isolation actually occur?

If it blocks malicious infrastructure, did the relevant control enforce that block?

If it terminates a session, is the session actually gone?

If it changes a configuration, did the intended condition improve without producing unacceptable downstream effects?

The stronger operating sequence is:

EVIDENCE → DECISION → AUTHORIZED ACTION → VERIFICATION

Verification should establish three things:

  • The action actually occurred.

  • The intended security or business condition improved.

  • The action did not create unacceptable downstream impact.

This requires more than an API response or a successful workflow status.

The system needs evidence that the intended result was achieved.

For security operations, that may mean validating the enforcement state of a control, checking whether an identity remains active, or determining whether a compromised system retains an unintended access path.

For other enterprise workflows, verification may involve validating data integrity, transaction state, or changes to authorized resources.

The principle is the same.

Without verification, automation can produce confidence without proof.

The more consequential the action, the more important that distinction becomes.

8. Assume the AI Agent Will Get Creative

This may be the exploit-research lesson I think about most.

A security researcher does not assume a system will remain inside the workflow somebody documented.

Neither should an architect designing autonomous systems.

If an agent is capable of discovering new paths, security should assume those paths eventually will be explored.

The answer isn't necessarily to make the agent less intelligent.

The answer is to make the architecture more disciplined.

Give the system enough capability to accomplish its objective.

Give it the validated evidence required to make the decision.

Give it enough authority to act when the conditions justify action.

But don't give it unnecessary authority simply because exposing that authority is convenient.

The objective is to enable useful autonomy without allowing that autonomy to exceed the boundaries the organization has established.

What Exploit Research Teaches Us About AI Agent Security Architecture

AI agents are creating a new technical frontier.

But some of the hardest security questions aren't new.

What can this system actually do?

What assumptions are we making about its environment?

What happens when legitimate capabilities combine?

Where does trust begin and end?

What is the potential blast radius?

How quickly can authority be reduced or revoked?

Can we reconstruct why the system acted?

Can we verify what happened afterward?

These are cybersecurity questions.

They require security engineering, architecture, runtime enforcement, and evidence-based validation.

They also require testing the system beyond the workflows its designers expect.

An AI agent's security posture cannot be established solely by observing that it behaves correctly under ordinary conditions.

We need to understand what happens when permissions drift, tools return unexpected information, assumptions fail, and an agent discovers a capability that its designers did not anticipate.

The Real Security Question: What Can the Agent Actually Do?

Exploit research taught us something important a long time ago.

The system does not care what we intended.

It responds to what the architecture permits.

As autonomous systems become more capable, that principle becomes more consequential.

The objective shouldn't be to prevent AI from discovering unexpected possibilities.

That capability may be part of what makes these systems valuable.

The objective is to make sure discovery never creates authority that the architecture did not intend to grant.

That requires a disciplined distinction between capability and authorization.

It requires enforceable boundaries rather than relying exclusively on behavioral instructions.

It requires human-defined accountability without making manual approvals an unavoidable bottleneck.

It requires progressive and reversible authority.

And it requires proof that actions produce the intended outcome.

These are not entirely new security ideas.

They are established principles applied to systems that can increasingly reason, discover, and act.

REDSAND | The CTO Field Notes

Assume the agent will eventually discover what it is able to do.

Then make sure what it is able to do never exceeds what it should be authorized to do.

That is the architectural challenge.

Same principles.

A new frontier.

Tim Shelton
Founder & CTO | HAWK Network Defense

REDSAND | THE CTO FIELD NOTES — Security engineering, AI architecture, exploit research, and the operating principles behind resilient systems.

Tim Shelton

Tim Shelton is the Founder and Chief Technology Officer of HAWK Network Defense, bringing more than two decades of experience in cybersecurity, exploit research, security engineering, and large-scale security analytics. His work focuses on advancing machine-speed security operations, reducing decision latency, and connecting threat detection to authorized action and verified outcomes.

Through REDSAND | THE CTO FIELD NOTES, Tim explores AI security architecture, autonomous agents, runtime controls, and the engineering principles required to build secure, resilient systems.

Previous
Previous

AI Agent Security Testing: What Happens When Agents Get Creative?

Next
Next

Incident Response Metrics: Measure Time to Verified Outcome