AI Agent Security Testing: What Happens When Agents Get Creative?

REDSAND | THE CTO FIELD NOTES

By Tim Shelton | Founder & CTO, HAWK Network Defense

The most important AI security test may not be whether an agent follows its instructions.

It may be what happens when the agent discovers something its designers never anticipated.

I learned something early in exploit research:

Systems rarely fail only in the ways their designers anticipated.

You don't find the most interesting vulnerabilities by asking what an application was intended to do.

You ask what it can do.

What happens if I use this input differently?

What if I call something in the wrong sequence?

What if two individually legitimate capabilities are combined in a way nobody expected?

What happens when a system discovers a path its developers never considered?

Those questions have shaped security research for decades.

AI agents make them newly relevant.

Why AI Agent Security Testing Must Go Beyond Expected Workflows

An AI agent may have legitimate access to several enterprise systems.

It may be able to retrieve information, invoke APIs, interact with tools, access credentials, and execute workflows.

Every permission may appear reasonable when evaluated independently.

But individual permissions do not necessarily describe the agent's effective capabilities.

Consider an AI agent with permission to:

  • Retrieve information from an internal knowledge repository.

  • Access a customer database.

  • Call an external API.

  • Generate reports using retrieved information.

  • Send messages through an approved communication service.

Each capability may support a legitimate business process.

But what can the agent accomplish when those capabilities are combined?

Could it retrieve information from one system, transform it using another, and transmit the result through a communication channel that was never intended for that data?

Could it use a legitimate tool to reach a resource that another component was expected to protect?

Could conflicting information or unexpected tool responses cause it to take a different path toward its objective?

Those are not necessarily failures of individual tools.

They may be failures of the boundaries between them.

This is precisely the kind of problem exploit research has taught us to examine.

The interesting question is not simply what each tool allows. It is what the agent can accomplish with all of them together.

Intended Behavior Versus Possible Behavior

Most systems are designed around expected workflows.

An administrator configures permissions.

A developer defines how an integration should operate.

A security team validates normal behavior.

The system passes its acceptance tests.

But successful operation under expected conditions does not establish that the architecture is secure under unexpected conditions.

In conventional security research, we examine malformed inputs, unusual execution sequences, trust relationships, and combinations of capabilities.

We look for the gap between what the designer expected and what the system actually permits.

AI agents require the same adversarial mindset.

And potentially greater attention to those gaps.

An agent capable of reasoning about available tools and selecting alternative approaches may discover combinations that were not explicitly programmed as a single workflow.

That does not mean every agent will discover or exploit every possible path.

It means security architecture should not depend on the assumption that the agent will never try.

Security testing must examine what is technically possible—not merely what was intended.

Four Areas AI Agent Security Testing Should Examine

There are four areas I would prioritize when evaluating the security of an agent operating inside an enterprise environment.

1. Test Beyond Expected Behavior

Don't stop testing when the agent completes its assigned task correctly.

Challenge the assumptions that make the expected workflow appear safe.

What happens if an API returns unexpected information?

What if the requested resource is unavailable?

What if the agent receives conflicting evidence?

What if a connected system exposes additional functionality?

What if a dependency fails halfway through execution?

The goal is not simply to confirm that the agent behaves correctly when everything works.

It is to understand how the system behaves when the environment stops matching the assumptions under which it was designed.

That includes whether the agent fails safely, preserves evidence, and respects its authority boundaries.

2. Verify That Architectural Boundaries Are Enforceable

Prompts and behavioral instructions can influence an agent's decisions.

But they are not substitutes for technical enforcement.

An instruction telling an agent not to access sensitive information is not equivalent to preventing access through authorization controls.

A prompt telling an agent not to transmit data externally is not equivalent to enforcing network and egress restrictions.

A behavioral rule telling an agent not to modify a production environment is not equivalent to denying production write privileges.

The distinction is fundamental.

Security instructions describe desired behavior. Architectural controls enforce permitted behavior.

Testing should therefore establish whether restrictions remain effective even when the agent attempts an action outside its authorized scope.

The important question is not whether the agent usually follows the rule.

It is whether the system prevents the action when the rule is not followed.

3. Evaluate Combined Capabilities and Emergent Behavior

This is where the lessons from exploit research become especially valuable.

Individual permissions can appear harmless while producing unexpected authority when combined.

An agent may legitimately read from one system, process information through another, and communicate through a third.

The combination may create a path that none of those system owners intended.

Security testing should investigate these interactions.

Can the agent move information between systems that should remain separated?

Can it invoke tools in an unexpected sequence?

Can a lower-trust source influence an action involving higher-trust resources?

Can the agent indirectly achieve an outcome that it is not permitted to perform directly?

Does the combination of legitimate permissions create an unintended access or execution path?

These questions reveal the agent's effective authority rather than merely its documented permissions.

That is where architectural risk can emerge.

4. Keep Human Judgment in the Operating Model

Autonomous systems need human accountability.

But human control should not automatically mean inserting manual approval before every routine action.

Security teams can establish boundaries before an agent operates.

They can define permitted actions, required evidence, decision thresholds, escalation conditions, and protected resources.

They can determine which decisions require human authorization and which can be safely delegated within defined limits.

They can also define when previously granted authority must be reduced or revoked.

Testing should validate those controls.

Does the agent escalate when conditions exceed its authority?

Does it stop when required evidence is unavailable?

Can privileged actions be denied or revoked?

Can investigators reconstruct why a decision was made?

Can the organization verify the result?

Human judgment should shape the operating model, while technical controls enforce its boundaries.

The objective is not unrestricted autonomy.

It is useful autonomy within accountable, verifiable limits.

The Security Test That Matters Most: Can the Agent Exceed Its Intended Authority?

Traditional security testing often examines access controls, execution paths, trust boundaries, and unintended interactions.

AI agent security testing needs to incorporate those same principles.

But agents introduce a particularly important question:

Can the agent combine legitimate capabilities to achieve an outcome that exceeds its intended authority?

Answering that question requires more than evaluating prompts, inspecting tool lists, or confirming that ordinary workflows complete successfully.

It requires adversarial testing across the environment in which the agent actually operates.

That includes examining identity, permissions, API access, credentials, data flows, execution sequences, and interactions between connected systems.

It also requires understanding the potential consequences when an unexpected path becomes available.

What is the blast radius?

Can authority be restricted?

Will the system generate sufficient evidence to reconstruct what happened?

Can security teams determine whether an unauthorized action succeeded?

Can they verify that corrective action restored the intended security condition?

These are not theoretical architecture questions.

They are practical questions about what an enterprise system can permit when its components interact.

Why AI Security Architecture Matters More as Agents Become Capable

The more useful an AI agent becomes, the more systems and capabilities it may need to interact with.

That can expand the number of possible execution paths.

The answer is not necessarily to make agents less capable.

It is to ensure their capabilities remain bounded by enforceable authority.

That means identifying the resources an agent needs, granting appropriate permissions, restricting unnecessary access, and evaluating how legitimate capabilities interact.

It also means continuing to test those assumptions as environments change.

APIs evolve.

Permissions drift.

New integrations appear.

Business workflows change.

The effective capability of an agent can change as its environment changes, even when its original instructions remain the same.

Security validation therefore cannot be limited to the moment an agent is first deployed.

It must remain part of the operating model.

REDSAND | The CTO Field Notes

I learned through exploit research that the most useful questions often begin where the expected workflow ends.

The designer asks what the system was built to do.

The security researcher asks what else is possible.

As organizations deploy increasingly capable AI agents, that distinction becomes more important.

The objective should not be to prevent an agent from reasoning, discovering, or using legitimate capabilities to accomplish its task.

Those capabilities are part of what makes the technology valuable.

The objective is to prevent that capability from becoming unintended authority.

Assume the agent will eventually discover what it is able to do.

Then make sure what it is able to do never exceeds what it should be authorized to do.

That is the security test I would want every AI agent architecture to pass.

Tim Shelton
Founder & CTO | HAWK Network Defense

REDSAND | THE CTO FIELD NOTES — Exploit research, security engineering, and the architectural principles behind resilient AI systems.

Tim Shelton

Tim Shelton is the Founder and Chief Technology Officer of HAWK Network Defense, bringing more than two decades of experience in cybersecurity, exploit research, security engineering, and large-scale security analytics. His work focuses on advancing machine-speed security operations, reducing decision latency, and connecting threat detection to authorized action and verified outcomes.

Through REDSAND | THE CTO FIELD NOTES, Tim explores AI security architecture, autonomous agents, runtime controls, and the engineering principles required to build secure, resilient systems.

Previous
Previous

AI Authority Is Becoming a Business Control

Next
Next

What Exploit Research Taught Me About AI Agent Security