Contents

Introduction 

Third-party platforms are becoming increasingly embedded within financial institutions for managing financial crime risks such as sanctions screening, transaction monitoring and customer risk scoring. As these tools evolve from operational utilities to core decision-making tools, they directly shape regulatory outcomes in determining which transactions are flagged and escalated for review. This mirrors the shift that led to the formalisation of model risk management (MRM) frameworks, as analytical systems became core inputs to risk decisions.

However, many institutions still treat vendor solutions primarily as technology implementations, overlooking that they usually meet the definition of a model. Since these systems transform inputs (customer and transactional data) through structured logic (matching algorithms, scenarios) into outputs that drive decisions, they should fall within scope of MRM requirements around inventory, governance and periodic independent validation. Vendor model opacity does not exempt it from BAU controls.

This publication sets out a practical framework for validating third-party financial crime models. It structures validation around five core pillars: governance, representativeness and intended use, data and lineage controls, conceptual soundness and performance monitoring, illustrating their application through a sanctions screening validation use case. The objective is to provide a repeatable and structured approach that enables institutions to retain ownership, evidence control effectiveness and meet supervisory expectations even when key elements of the methodology are vendor driven.

 

Third-Party solutions are models

The “black box” problem doesn’t remove the risk

Vendor solutions often operate as "black boxes," characterized by proprietary algorithms and restricted access to underlying code and development data. While this makes validation more challenging, it does not exempt the model from scrutiny and review. If anything, opacity increases the need for discipline: institutions remain responsible for explaining what the tool is intended to do, how it is calibrated for their risk profile, what data it relies on and how it is transformed, and how performance is monitored over time. 

In supervisory reviews, reliance on vendor statements is not viewed as an adequate basis for comfort; with financial crime risk management increasingly reliant on technology for delivery, regulators expectations of firms’ oversight of these systems are heightened. Regulators now look for active ownership of the key design and configuration choices that drive operational outcomes. This includes everything from threshold calibration, segmentation logic and scenario coverage to the governance of parameter updates. Ultimately, firms must demonstrate clear ownership and understanding of the underlying data that feeds these models.

What changes once the tool is treated as a model

Once a vendor solution is formally classified as a model, it transitions from a static software asset to a dynamic risk component within the Model Inventory. This shift triggers the governance framework where ownership and risk-tiering are clearly defined, ensuring that the institution is the ultimate arbiter of the tool's performance.

Version updates, parameter shifts, and list refreshes must undergo controlled testing and receive traceable internal approvals before deployment. By maintaining a robust evidence pack that supports independent challenge, the institution can confidently stand behind the model's outputs, regardless of where the underlying code was authored.

The Strategic Value of Model Classification

  • Accountability: establishing defined ownership and risk tiering eliminates oversight gaps, ensuring that "vendor-managed" does not translate to "unmanaged."
  • Rigorous change governance: a formal classification forces vendor updates and parameter shift through a dedicated internal validation cycle, preventing unexpected changes to the risk profile.
  • Regulatory defensibility: comprehensive documentation and independent test evidence provide the proof required to withstand supervisory challenge and demonstrate active oversight.
  • Optimization of control effectiveness: proactive monitoring allows the institution to detect performance drift early, ensuring a precise balance between risk detection and operational efficiency.

 

Third-party model validation

Effective validation of third-party Financial Crime models requires shifting the focus from outputs alone to a granular assessment of the model’s internal architecture and its behaviour within the institution’s environment. Rather than treating validation as a separate exercise, each pillar should incorporate how the model is independently challenged, tested, and evidenced. This ensures that all components from vendor governance to performance are subject to robust validation.

Pillar A Pillar B Pillar C Pillar D Pillar E
Vendor selection and governance
Representativeness and intended use
Data quality requirements
Methodology and conceptual soundness
Performance and monitoring
Pillar A

Vendor selection and governance

Third-party financial crime platforms are often purchased as technology, but decisions made during vendor selection and contracting determine how much oversight is possible. Effective governance therefore starts before implementation and relies on three fundamentals: transparency around the model’s logic and limitations, contractual provisions that allow control over changes, and performance and clear internal accountability.

What good looks like in practice

  • Transparency and documentation. Institutions should have access to documentation on the solution’s methodology, configuration, segmentation, and scenario logic. Even when proprietary details are limited, vendors should provide design summaries, assumptions, approval evidence, change logs and release notes to enable independent challenge and explainability.
  • Contractual governance. Contracts should include audit rights, notification requirements for material changes, support for independent validation and clear provisions on data ownership and retention.
  • Defined change management. Vendor releases, parameter tuning, data-mapping updates, and list refreshes should follow controlled processes with impact assessments, testing and formal approvals to ensure system behaviour does not change unexpectedly.

Validation testing

  • Documentation review: assess key documentation to ensure transparency.
  • Independent challenge: confirm assumptions are understood and challengeable.
  • Minutes: review meeting minutes.
  • Governance walkthroughs: ensure changes follow internal approval and testing.
Pillar B

Representativeness and intended use

A vendor tool can be market-leading yet still misaligned with an institution’s risk profile. Representativeness ensures the solution is configured to address actual risks rather than simply manage alert volumes or rely on generic typologies. In practice, this means clearly defining what the tool is intended to do and not do within the institution’s risk appetite, ensuring monitoring and controls reflect the organization’s specific business model, products and geographic exposure rather than operating as a generic template.

What good looks like in practice

  • Clarity on intended use. State which decisions the tool supports (e.g., detection of suspicious transactions, prioritisation of alerts, investigative triage) and whether outputs are advisory or used for automated decisioning.
  • Link configuration to the risk profile. Evidence that scenarios, watchlists, thresholds and segmentation reflect the institution’s products, customer base, delivery channels and geographic exposure.
  • Make coverage explicit. Document what is in scope (products, transaction types, channels, legal entities) and how exceptions are approved and monitored.
  • Use benchmarking as a reasonableness check. Compare expected alert volumes and drivers to historical baselines and, where available, external benchmarks.

Validation Testing

  • Use-case testing: confirm outputs align with intended use.
  • Taxonomy: ensure key AML typologies are captured.
  • Benchmark analysis: compare outputs against historical data and risk profile.
  • Information transfer with vendor: integrity & completeness of transferred data.
Pillar C

Data quality requirements

For financial crime models, data is often the primary driver of outcomes rather than a background dependency. Even a well-designed algorithm can degrade if input data is incomplete or incorrectly mapped. For example, for sanction screening this could materialised through truncated names, incorrect party-role mapping, missing remittance text, or inconsistent identifiers. In addition, vendor tools typically require specific data structures and formats, meaning data must be sourced and aligned accordingly. Therefore, robust data lineage and traceability are essential to ensure data is accurately transformed, transferred, and fit for purpose.

What good looks like in practice

  • Define critical data elements (CDEs). Identify the fields that directly drive detection, such as names and aliases, nationality, address fields, transactional data, counterparty attributes and payment or remittance text.
  • Document lineage and transformations. Record how data flows from source systems into the screening or monitoring engine and what transformations occur in transit (e.g., transliteration, tokenisation, truncation or enrichment).
  • Establish interface controls. For batch or API feeds, implement completeness checks, volume reconciliations, timeliness SLAs and clear error-handling and replay procedures.
  • Monitor changes that affect outcomes. Track shifts in key data attributes such as name completeness, remittance text length, or party-role coding that may alter matching behaviour or scenario performance.

Validation Testing

  • Data quality and integrity: assess completeness, accuracy, and consistency of transaction, customer, and KYC data feeding the model.
  • End-to-end Data lineage to confirm transparency and traceability.
  • Representativeness: validate that the dataset used by the third‑party model is representative of the bank’s profile.
  • Stability and temporal consistency: perform historical stability checks on key variables to identify structural breaks or seasonality.
  • Reconciliation testing: detect losses or transformations in data flows.
Pillar D

Methodology and conceptual soundness

Even when vendors do not disclose proprietary logic, institutions are still expected to understand the conceptual basis of how model outputs are generated and what assumptions underpin them. Conceptual soundness is therefore less about inspecting code and more about being able to explain why the tool behaves as it does, under what conditions it may fail, and what controls exist to manage those limitations. This aligns with broader Model Risk Management expectations that vendor solutions remain within the institution’s governance and validation scope despite transparency constraints.

What good looks like in practice

  • Understand the conceptual logic. Institutions should understand, at a high level, how matching algorithms, scenarios, or scoring approaches produce outputs.
  • Document assumptions and limitations. Known constraints in the methodology are clearly recorded and considered in governance, tuning, and monitoring.
  • Evidence conceptual challenge. Validators should assess whether the model’s design and configuration are appropriate for the intended use and the risk profile.

Validation Testing

  • Interaction with vendor tool: review rule definitions, thresholds, segmentation logic, and parameter ranges.
  • Conceptual walkthroughs: assess the model’s conceptual robustness by analysing how outputs behave, not by relying on internal vendor logic.
  • Sensitivity analysis on rules and thresholds to evaluate robustness under extreme scenarios, low‑frequency events, and atypical transaction patterns.
  • Assumptions and limitations: assess whether the vendor’s level of transparency is sufficient to meet supervisory expectations for explainability and auditability.
Pillar E

Performance and monitoring

Validation should evaluate whether the solution performs as expected within the institution’s environment. This includes reviewing calibration choices, monitoring frameworks and explainability mechanisms to ensure the model remains stable and aligned with the institution’s risk profile over time. Effective validation extends beyond conceptual review and should incorporate risk-anchored calibration, ongoing monitoring of model behaviour and data, and outcomes analysis or back-testing.

What good looks like in practice

  • Risk-aligned calibration. Thresholds and tuning decisions reflect the institution’s risk appetite rather than purely operational considerations.
  • Performance monitoring. Key metrics such as alert volumes, detection rates and match quality are tracked over time.
  • Event-driven review and validation. In addition to periodic validation cycles, institutions should assess whether significant events such as major vendor releases, material tuning exercises, regulatory findings, or unexplained shifts in model performance warrant an interim validation or targeted review.
  • Explainability and stability. Institutions can interpret model outcomes and detect unexpected performance shifts following changes or updates.

Validation Testing

  • ROC/AUC: assess discriminatory power and classification accuracy.
  • Confusion‑matrix effectiveness testing: analyse typology‑level performance to ensure the model detects high‑risk patterns even when overall precision is low.
  • False positive and regression testing: evaluate efficiency and ensure changes do not degrade performance.
  • Detection rate testing: evaluate detection rate (suspicious cases detected vs. total suspicious cases) to assess operational efficiency.
  • Review rate: ensure alerts are actionable and operationally manageable.
  • Back-testing / replay analysis: test performance using historical cases.
  • Review of the model monitoring framework including KPIs, periodicity, thresholds and action plans.
1.

Validation scope

Grant Thornton carried out a comprehensive validation of all aspects of the sanctions screening model. This included:

  • An assessment of governance processes, including the review of documentation, approval process, meeting minutes, analysis and proof of sign off from appropriate stakeholders.
  • Model documentation was also reviewed for each component of the model to ensure supporting materials and processes are robust and understandable. The documents included model documentation, third party user manual, algorithm review, configuration settings spreadsheets, rule management documents. 
  • Data quality controls and reconciliation protocols between data sources were conducted.
  • Watch List Management, including the review of internal and external list maintenance procedures.
  • An evaluation of rule setting approach. This included a review of rule setting framework, rule approval process, volume of new rules created in recent time periods, impact analysis of potential new rule settings.
  • Inspection of algorithm configuration settings and review processes.
  • Verification that the entity carried out adequate model monitoring of volumes of alerts generated and impacts of rule settings and algorithm configurations on a regular basis. 
2.

Key challenges in the validation 

Some challenges were faced as part of this validation exercise. These were as follows:

  • Given the external nature of the solution, certain components were proprietary; however, a comprehensive understanding was achieved through detailed review of vendor documentation and engagement with relevant stakeholders, enabling the validation to be performed effectively.
  • Identifying and obtaining the most relevant datasets for testing presented a challenge, requiring careful consideration of data selection to ensure meaningful coverage. Through iterative engagement and targeted data requests, the appropriate data was successfully sourced, enabling effective validation.
  • The validation of the sanctions screening model faced significant hurdles due to its fragmented ownership across the bank. This lack of centralized oversight turned the collection of required analysis and documentation into a complex process.
  • Rates of model alerts were in normal ranges for the geographical location and regular monitoring of alert volumes were analysed. Proof of impact analysis for rule and algorithm changes were reviewed. In addition, a sample testing was also performed to further analyse impacts of rules and algorithm configurations.
  • Assessing the effectiveness of screening algorithms and rule settings required more than testing against standard data. A key challenge was creating sample records that were realistic, but also sufficiently challenging to test fuzzy matching behaviour, threshold sensitivity, and watchlist timeliness. This involved manipulating watchlist entries through character substitutions, reordered names or addresses, removal of geographic information, and focus on recent entries. Additional testing was performed under alternative algorithm settings to challenge the justification for less conservative configurations and assess whether calibration remained appropriate.
3.

Overall takeaways 

Below is a summary of the key observations identified during the review

  • Each aspect of the process is adequately documented but there is no documented procedure covering a high-level overview of the process overall. It was recommended to document the overall process.
  • Documentation exists for new rule requests and their implementation, with some detail on testing and governance. However, greater transparency is needed regarding scope, change history, review frequency, and materiality thresholds. Existing thresholds and criteria are sometimes vague.
  • Evidence of governance and reporting for the ‘lookback’ process was reviewed, which follows a false negative event. Minimum requirements and success criteria should be more explicitly defined.
  • Clearer assignment of ownership for MI production is recommended, along with complete and consistent governance documentation.
  • Adequate data quality and reconciliation checks are in place.
  • The rule‑based engine and process flow align with industry best practices. Rule settings and algorithm configurations are conceptually sound.