AI validation in practice
What financial institutions need to test before relying on AI agents

In the latest article in our AI validation series, we examine how firms can put our AI Validation framework into practice, using two use cases to demonstrate what effective testing and oversight look like.
The Validation Challenge for AI Agents
Financial institutions are beginning to use AI agents to interpret regulation, assess internal policies and produce documentation. That creates a clear validation challenge: a system can produce a fluent, plausible answer and still be incomplete, unsupported or inconsistent.
Where firms intend to rely materially on AI agents, they need evidence that the system behaves reliably, that material conclusions can be traced to appropriate sources, that uncertainty is handled properly and that performance remains acceptable as inputs, prompts and requirements change.
Traditional model risk principles remain relevant, but AI systems differ from conventional predictive models. They may respond differently to equivalent prompts, retrieve and combine information dynamically, and carry out multi-step workflows.
The validation framework outlined in the full paper extends beyond fixed, point-in-time accuracy tests. It covers data quality and input review, behavioural testing, output evaluation, human-in-the-loop controls and ongoing monitoring.
Testing a regulatory compliance agent
Consider an AI agent used to assess regulatory requirements against an institution’s internal policies and controls. The agent may reduce the time required to review regulatory materials and identify potential compliance gaps, but regulatory interpretation is complex, context-dependent and subject to change.
That raises three immediate validation questions.
- Does the agent reach materially consistent conclusions when the same requirement is phrased in different ways? If small changes in wording produce different interpretations, that raises questions about the stability of the output.
- Can each material conclusion be traced to an approved source? A persuasive answer has limited value if a reviewer cannot identify the regulatory provision or internal policy that supports it.
- Does the system recognise when human review is required? Where information is incomplete, requirements conflict or the potential impact is material, the system should escalate rather than present an unsupported conclusion as fact.
A correct answer on one test is not sufficient. Validation should establish whether the system remains reliable when wording, context and available information change.
A different challenge for documentation and reporting agents
A documentation and reporting AI agent presents a different validation problem. Here, the principal question is whether the generated narrative remains faithful to the underlying evidence and is suitable for its intended use. Material claims, figures and references should be checked against approved sources to
identify unsupported or fabricated content.
The system’s guardrails should also be tested. Requests that fall outside its authorised scope, rely on missing information or conflict with predefined constraints should result in the appropriate response, whether that is refusal, limitation, clarification or escalation.
Human is equally important. Where outputs require approval because of their complexity, ambiguity or potential impact, validation should confirm that those cases are routed to the appropriate reviewer. Patterns in editing can provide useful evidence. Repeated corrections to the same type of output may point to a recurring weakness in the system.
Materiality matters more than the number of exceptions
Validation results must be interpreted in the context of the system’s intended use, risk profile and surrounding controls. A strong average result can still conceal a material weakness. A strong overall result can still conceal a material weakness.
One unsupported interpretation of a regulatory obligation, one fabricated material claim or one missed escalation may matter more than several minor formatting errors. Measures such as consistency, traceability and escalation rates can help, but they do not replace expert judgement about the severity and potential impact of an exception.
Validation also continues after deployment. Changes to underlying models, prompts, retrieval sources, guardrails, connected tools or intended use may alter system behaviour. Regulatory requirements and internal policies will also continue to evolve.
Monitoring should therefore look for output drift, recurring weaknesses and changes that require re-testing.
The objective is not simply to record whether individual tests have passed. It is to support a clear decision on whether the AI system is sufficiently reliable, controlled and governed for its intended use, and whether approval conditions, remediation, enhanced human oversight, usage restrictions are required.
The full paper, available below, translates these principles into specific tests, metrics and verification methods for both use cases. It also demonstrates how testing evidence can support informed decisions on approval, remediation, usage restrictions, human oversight, monitoring and revalidation.
Put AI Validation into Practice
Explore practical testing approaches, validation controls and real-world AI agent use cases for governance, risk and compliance teams.