Skip to main content Skip to main navigation Skip to footer

Reducing False Positives in Sanctions & Watchlist Screening with AI

x

Risk screening has long had a problem with high volumes of false positives. For financial institutions managing large customer portfolios and transaction volumes, excessive false positives place enormous pressure on screening operations - and often the only way to resolve the challenge is to add headcount.

The good news is that AI can help to reduce the manual effort needed in risk screening. Hawk's AI False Positive Reduction was recently awarded the Silver Medal for Best Sanctions & Watchlist Screening Innovation by Datos Insights, with recognition given to the solution's ability to fill data gaps with the context needed for more precise matching and fewer false positives. 

In this article, we cover: 

  • Why risk screening generates huge volumes of false positives 

  • What's missing from the data institutions screen against 

  • How AI fills those gaps to clear obvious false positives 

  • How Hawk's AI False Positive Reduction works 

  • What AI means for customer and payment screening 

The Data Deficit Driving False Positives 

Screening runs on too little data. Institutions screening ultimate beneficial owners/directors and payment counterparties often work with just a name. When a counterparty resembles a sanctions or watchlist record closely enough to trigger an alert, it's rarely a true match. Most are false positives caused by common names, transliteration variations, differing international naming conventions, or modified aliases. Without secondary identifiers, it's difficult to safely clear ambiguous hits without analyst review or risking regulatory censure. 

This data deficit plays out in a few ways: 

Increases hit volumes. Fuzzy matching alone will often struggle to decipher between two entities when fields like gender or entity type aren’t available. For example, it may not be able to tell that Michelle (customer) and Michel (watchlist entry) are not the same when Michelle’s gender isn’t available in the data. This causes more irrelevant hits to be created—often hits that can be quickly dispositioned with human judgement.  

Slows hit review. Legacy systems flag potential matches but leave investigators to search manually for the information they need to make a decision. This creates delays and friction for customers, particularly in payment screening.  

Brings more scrutiny with wider coverage. Expanding detection by adding more rules only compounds the data deficit problem, forcing analysts to manually sift through and assess even more hits with missing fields, like genders, entity types, and dates of incorporation.  

Consumes data science resources for models that often fail governance. Institutions often try to build AI models to overcome data gaps, but models built in-house for screening frequently fail governance review due to weak explainability, thin documentation, or training data that's too narrow or out of date. 

Where the Data Deficit Is Most Prevalent 

Not every part of screening feels this problem equally. Two areas carry the weight of it, and neither can be fixed by adding more rules or extra manual reviews. 

Payment screening. Institutions know their own customers well: verified documents, account history, years of due diligence. They know almost nothing about the counterparty on the other side of a payment. Often, a name and a location is all a screening system has to work with.  

Corporate customer screening. Screening a corporate customer isn't just screening the entity. It means screening its beneficial owners and directors too. These individuals may have no onboarding record, no verified documents, and no direct relationship with the institution. The institution ends up screening people it hasn’t independently identified, relying only on what the corporate discloses. 

Solving the Data Challenge with AI 

Closing the data gap means giving screening systems the context they're missing. AI addresses that structural challenge directly, using predictive models to generate context that the underlying data doesn't supply on its own.  

For our Michelle example above, this means determining through linguistics that Michelle is typically a female name even though the gender field was blank when we received Michelle’s details. With this understanding, we know that Michelle shouldn’t match any male screening hits (like Michel) and can eliminate those respective hits before analyst review. 

Hawk's AI False Positive Reduction addresses data challenges by doing just that—predicting missing context, evaluating every hit for contradictions with newly derived data, and preventing clear-cut false positives from needing an analyst's time. 

How Hawk's AI False Positive Reduction Works

Diagram for Hawk's AI FPR

1. Fills in the missing context. When a payment or customer hit is generated, the solution applies predictive models to fill in what the source data leaves out:  

  • Bayesian statistical models and deep learning transformer models predict gender from a counterparty's name 
  • Multilingual language models classify whether a name refers to a person or an organization 
  • Other models identify a likely country of registration from legal form designations embedded in entity names, an entity name containing "GmbH," for example, can be associated with Germany 

2. Checks for contradictions. That new context is compared against the watchlist entry the name was flagged against. A male sanctioned entity matched against a name predicted to be female, or a German-registered entity matched against a watchlist record tied to a different country, are both signals that the hit may not be genuine. 

3. Scores the hit. Every case receives an overall match likelihood score from 1 to 100.  

  • The overall score is made up of individual scores for each attribute the model evaluates, such as gender, entity type, and country of registration 
  • This lets an analyst see not just the final score, but which specific attributes drove it 

4. Shows the evidence. Each score comes with a plain-language explanation: 

  • Every attribute used to reach that score is labelled as either supporting or refuting evidence 
  • This documents the reasoning behind the number for the institution, rather than leaving it as a black-box output 

5. Decides and routes. Institutions set their own threshold for how confident the model needs to be. Hits below the match threshold can be prevented from opening as false positives. For ambiguous hits, the AI's attribute evidence is still presented. The alert states 'Insufficient evidence to draw a system resolution,' leaving the final call to an analyst. 
 

Results From Deploying AI in Risk Screening 

The combined Hawk AI screening solution can provide significant benefits over an existing screening system: 

Greater match accuracy validated against an industry benchmark 

Hawk was tested against an incumbent screening system using the SWIFT testing framework. The results: 97.83% effectiveness and 82.03% efficiency compared to 89.59% effectiveness and 55.21% efficiency for the incumbent system which was tuned over 10 years. 

Fewer payments held on false suspicion 

One financial institution reduced payments falsely held for sanctions risk by 55% after applying Hawk's solution. The result was less disruption for customers across multiple jurisdictions, without loosening compliance obligations. 

More accurate, informed decisions 

With AI, one customer was able to derive more information for 32% of their hits, allowing them to clear a significant number of false positives and shorten average review time, making informed decisions. 

Closing Thoughts 

Data deficiency has been a major root cause of false positives. Names by themselves rarely provide enough information to screen accurately against large watchlists. AI changes that by adding the context a name is missing. It evaluates every hit flagged by existing rules and derives missing context, such as gender, entity type, or country of registration. It then compares it against the watchlist entity to separate clear matches from clear mismatches. The results are measurable: a 27-percentage-point efficiency gap over a decade-tuned incumbent system, and a 55% reduction in falsely held payments.

It also comes down to whether institutions can trust and defend what a model decides, to regulators, to internal audit, to their own risk committees. Built-in governance, data lineage, and human-in-the-loop review are what make that possible. None of these accuracy gains matter without that transparency and reliability. 

Hawk's AI False Positive Reduction solves this gap. The solution sharpens efficiency, eases friction for customers, and accelerates how probable matches get reviewed. Built-in governance and clear, human-readable explanations tackle a problem that has held back AI adoption across regulated industries: the ability to explain and defend what a model decides to regulators and internal audit teams. Learn more about Hawk's AI screening solution here

 

 


Share this page