This site is part of the Informa Connect Division of Informa PLC

This site is operated by a business or businesses owned by Informa PLC and all copyright resides with them. Informa PLC's registered office is 5 Howick Place, London SW1P 1WG. Registered in England and Wales. Number 3099067.

Risk Management
search
Model risk

How to validate AI outputs without a right answer?

Posted by on 03 September 2026
Share this article

For three decades, model validation in banking rested on an assumption so basic that nobody bothered to write it down. Run the model twice, get the same answer twice.

That assumption shaped the discipline. Traditional validation practice and much of the guidance underpinning it, including SR 11-7, EBA guidance and MaRisk, developed around systems whose logic, inputs and outputs could be independently examined and, to a meaningful extent, reproduced. Effective challenge meant an independent team testing the construction, rerunning the process and seeing where it landed. You compute the distance to a known truth, and the number carries weight because the target holds still.

A deterministic system behaves like a vending machine. Same coin, same button, same chocolate bar. If it ever hands you crisps, something has broken, and the fault is findable.

A foundation model is not a vending machine. Ask it the same question twice and you get two different paragraphs, both defensible. Nothing has broken. Variation is not a defect in the system. It is the system.

How do model risk management practices need to change in order to validate AI outputs?

  • Validation must shift from reproducibility to statistical characterisation. Foundation models produce different outputs by design, requiring assessment of distributions across multiple runs rather than single, repeatable answers.
  • Accuracy metrics alone are insufficient. Generative AI requires evaluation frameworks measuring faithfulness, groundedness, completeness, and dispersion – not just traditional precision/recall metrics.
  • Validate the entire system, not just the model. Focus on end-to-end use cases including prompts, retrieval, guardrails, human review, and business workflows, with explicit acknowledgment of residual risk.
  • Revalidation triggers must expand beyond annual cycles. Silent provider updates, prompt changes, distribution drift, scope creep, and new input types all require immediate reassessment.
  • Defensible evidence means documented distributions with decision rules. Validation findings must include evaluation set versions, model configurations, observed ranges, tail failures, and stated consequences for independent challenge.

What success looks like now

The shift is from grading the answer to characterising the behaviour.

That is not a softening. Instead of one comparison against an answer key, validators run the system repeatedly on a versioned set of cases and report a distribution.

“Faithfulness was 0.9” has no uncertainty attached. By contrast, “mean faithfulness was 0.88, with an observed standard deviation of 0.04 across twenty runs on evaluation set v3; the threshold was 0.85, failures went to human review, and every run was logged” is evidence that can be challenged. The number of runs is not universal; it should reflect the decision, risk and observed variability.

You cannot replicate what you cannot open.

The spread is part of the finding. If nineteen generated narratives read well but one omits a required counterparty, the mean conceals the run that matters. A defensible report therefore includes observed ranges, tail failures, threshold-breach frequency, worst-case outcomes and severity-weighted failure rates.

A peer-reviewed study of a banking assistant uses semantic entropy to measure how far repeated responses diverge in meaning. The method exists; the less established habit is treating dispersion, tails and consequential outliers as headline evidence.

A number becomes evidence when it travels with a definition, a repeatable procedure with the evaluation set, prompt, model version and sampling settings pinned, and a threshold with a stated consequence. Repeatability asks whether controlled reruns are comparable; reproducibility asks whether an independent party can recreate the procedure and obtain a comparable distribution.

Why accuracy alone is the wrong instrument

Generative answers broadly fail by inventing, omitting or wobbling: failures of faithfulness or groundedness, completeness, and stability or dispersion. Conventional accuracy metrics do not capture all three.

Accuracy still matters where a task has a defined target, such as intent classification; precision, recall and confusion matrices remain useful there. The mistake is treating them as sufficient for open-ended generation with no single gold answer.

The evidence is now statistical rather than exact.

Grounding is not correctness: a faithfully cited answer can still reproduce a wrong source. Nor does 0.8 mean eighty per cent safe; the unsupported fifth may carry the consequence. Completeness captures what grounding misses – the material the answer left out.

Where no gold answer exists, the ruler is often another model. That judge may share correlated failure modes and be sensitive to prompts, benchmark contamination, calibration and model updates. Model-based evaluation remains useful, but the judge needs its own validation file: defined criteria, calibration against expert review, sensitivity testing, change control and periodic challenge.

Challenging a foundation you do not own

You cannot replicate what you cannot open.

Weights, training data and fine-tuning history remain with the provider. Assessment therefore centres on input-output relationships, leaving more residual risk than with an in-house model. That residual belongs explicitly in risk appetite.

The unit of validation must therefore shift to the end-to-end use case: model, prompts, retrieval, source data, guardrails, orchestration, tools, permissions, human review, monitoring, fallbacks and the surrounding business process.

A useful resilience test is to substitute the model where technically and contractually feasible. If controls collapse when the provider or version changes, the system has material concentration or portability risk. Any model-specific dependency should be identified, tested and placed within risk appetite.

Provider supervision does not remove firms’ obligations. In July 2026, UK authorities announced the first critical-third-party designations for four global cloud and technology providers, bringing designated services under oversight by the Bank of England, PRA and FCA. The regime strengthens service resilience; designation is not authorisation and does not prove that a firm has configured, controlled or used the service safely.

The regime supervises the electricity supply. It does not establish that you have wired the building correctly.

Which events require revalidation

The old triggers were the annual cycle, material change and a performance breach. All still apply. All quietly assume you find out when the model changes.

A more honest trigger set:

  • A provider version change, announced or silent, including safety filters and system defaults
  • Any change to the prompt, the retrieval corpus or the grounding sources
  • Measured drift in the distribution, not only the mean: a stable average with a widening tail is a deterioration
  • Scope creep, where a system approved for drafting starts informing decisions
  • A new class of input the evaluation set never covered
  • A change to the judge model, because the ruler moved
  • A material change in intended use, user population, decision impact, legal obligation or regulatory classification
  • An incident, near miss, material control failure, or change to data classification, retrieval permissions, human oversight or fallback arrangements

Vendor release cycles create drift, while monitoring practice remains fragmented. Institutions therefore need explicit, risk-based triggers rather than relying on annual review to detect change.

Annual review still forces an independent look at the whole, but it cannot carry the weight alone because the events that matter arrive between inspections.

The pragmatic close

None of this argues for lighter validation. It argues for validation aimed at the right object and organised around three disciplines:

  • validate distributions rather than isolated outputs
  • validate the end-to-end system rather than only the foundation model
  • trigger revalidation through continuous change detection rather than the annual cycle alone

The rigour financial services built was never really about reproducibility. It was about showing someone who did not build the thing why it can be relied upon. That obligation has not changed. The evidence is now statistical rather than exact, and it sits at the level of the system rather than the model.

The institutions that work this out will not be the ones with the strictest policy. They will be the ones that can produce a defensible number, with its provenance and its decision rule attached, for a system whose answer is different every time.

So, to any second-line function reading this: if your provider shipped a new version tonight and told nobody, how long would it take you to notice – and what in your control environment would notice it first?

Discuss critical AI issues with risk leaders and experts at RiskMinds this November!


Reference: J. Gerhard, M. Gombert, B. Henrich and J. Smith, "Model validation of a generative-artificial-intelligence-based avatar for customer support in banking", Journal of Operational Risk 21(2), 2026, DOI: 10.21314/JOP.2026.002.

Sources for the regulatory reference: HM Treasury, "UK financial system strengthened with new safeguards for major technology providers", 10 July 2026; Bank of England, "UK financial regulators to begin overseeing Critical Third Parties announced by HM Treasury", 10 July 2026.

Share this article

Sign up for Risk Management email updates

keyboard_arrow_down