Skip to content
Verinika

Our Methodology

A structured, evidence-based way to evaluate the RAG systems and AI agents you deliver — from scope to reusable quality tests.

How We Work

Every engagement follows the same five steps, designed to turn uncertainty into evidence your engineers can act on.

1

Scope

We agree the intended use of the system, its architecture, the risks that matter for your client and the acceptance criteria the evaluation will be judged against.

2

Design

We design representative test cases, the datasets behind them and the rules by which each result is judged — including edge cases and the situations where the system is expected to refuse.

3

Calibrate

We decide which checks can be scored automatically and which require human judgement, then calibrate the automated graders against reviewed examples so the scores can be trusted.

4

Evaluate

We test the individual components and the complete system reproducibly, recording inputs, configurations, traces and results so every finding can be re-run.

5

Verify

We confirm each issue manually and convert confirmed failures into reusable quality tests, so your team can detect quality loss after every future system change.

Three Evaluation Dimensions

Each system in scope is evaluated across three complementary dimensions.

Retrieval and grounding

Does the system find the right evidence and stay supported by it? We evaluate retrieval quality, completeness of context, citation correctness and correct refusal when the answer is not in the sources.

Agent tasks and tool use

Does the agent complete the task through a defensible path? We evaluate outcomes, trajectories, tool selection and arguments, recovery after errors and efficiency per completed task.

Adversarial behaviour

Can the system be manipulated or made to act against its intended purpose? We add scoped scenarios for prompt injection, manipulated context, unsafe tool requests and sensitive-information exposure.

Our Principles

Evidence over opinion

Every finding is backed by reproduction steps and traces. We never report suspected issues without evidence.

Risk-proportionate testing

We allocate testing depth based on failure impact — high-risk areas get more rigorous evaluation.

Reproducibility

Our results can be independently verified. We document inputs, configurations and expected outcomes.

Honest about limitations

We clearly distinguish between what we tested and what lies outside the scope. No conclusion extends beyond our actual evaluation.

Apply this methodology to the systems you deliver

Start with one RAG system or agent workflow — white-label or alongside your delivery team.

Discuss a Partner Evaluation