Skip to content
AI Evaluation Engineering

Adversarial Evaluation

Targeted adversarial scenarios for RAG systems and AI agents — prompt injection, manipulated context, unsafe tool requests and sensitive-information exposure.

The Challenge

Standard functional tests confirm that a system handles expected inputs. Adversarial evaluation tests what happens when inputs are deliberately crafted to cause unsafe, unintended or harmful behaviour.

We add scoped adversarial scenarios to an evaluation engagement to test how the system responds to manipulated instructions, context, documents, tool requests and attempts to expose sensitive information.

Adversarial AI Evaluation

What Can Go Wrong

Adversarial categories we test

Direct Prompt Injection

Crafted input that overrides system instructions or bypasses guardrails.

Indirect Prompt Injection

Malicious instructions embedded in retrieved documents or context.

Manipulated Context

Poisoned or altered context that steers generation toward incorrect or harmful outputs.

System-Prompt Extraction

Attempts to reveal system prompts, configuration or internal instructions.

Sensitive-Information Exposure

Triggering the system to reveal personal, confidential or restricted data.

Unsafe Tool Requests

Crafted prompts that attempt to trigger unauthorised or dangerous tool calls.

Excessive Agency

The system takes actions beyond its intended scope or authority.

Confirmation Bypass

Attempts to skip required confirmation steps or claim false success.

Our Testing Methodology

How we approach adversarial evaluation

1

Scope Attack Surface

Identify which adversarial categories are relevant to the system and engagement.

2

Generate Scenarios

Automated attack generation provides coverage across identified categories.

3

Validate Results

Manual validation determines whether a reported result is a reproducible system failure.

4

Reusable Tests

Confirmed findings are converted into reusable tests that verify the problem does not return.

What You Get

Adversarial Report

Confirmed adversarial findings with evidence, severity and reproduction steps.

Attack Scenario Set

Categorised adversarial scenarios used during the evaluation.

Reusable Quality Test Suite

Reusable tests derived from confirmed findings, rerun after each change.

Add Adversarial Scenarios to Your Evaluation

Adversarial evaluation is not a penetration test, source-code security audit, compliance audit or certification.

Discuss Adversarial Scenarios