SEMANTIC EMBEDDINGS AS A CANDIDATE-GENERATION CHANNEL FOR SANCTIONS NAME SCREENING

Authors

DOI:

https://doi.org/10.31891/csit-2026-3-2

Keywords:

sanctions screening, entity resolution, name matching, candidate generation, semantic search, vector search, dense retrieval, text embeddings, recall, AML/CFT compliance

Abstract

Sanctions name screening is an open-set identification problem with asymmetric error costs: a missed designee is more consequential than an extra alert. We evaluate OpenAI text-embedding-3-large and Google gemini-embedding-001 as candidate-generation channels over a fixed OFAC SDN individuals snapshot (7,426 entries; primary firstName + lastName only), using a controlled stress test in which the same 150 designated individuals are rendered in four Slavic-name regimes — direct Latin spelling, ICAO-style romanisation, French-convention transliteration, and a Russian form converted to Ukrainian and officially romanised — at two specificity levels, against easy and deliberately hard look-alikes, at three Matryoshka dimensions (768, 1536, 3072). Gemini led in every non-trivial regime and its margin grew with transliteration distance, reaching 34.6 pp of Recall@1 on two-part Russian-to-Ukrainian forms (entity-cluster bootstrap 95% CI 27.1–42.1 pp; exact McNemar p ≈ 1×10⁻¹⁶), with no systematic loss down to 768 dimensions, whereas OpenAI degraded with both transliteration distance and reduced dimension. Top-1 similarity did not separate transliterated positives from hard look-alikes: for OpenAI on the hardest regime the positives scored below the negatives (AUC 0.318, 95% CI 0.251–0.385). At an equal candidate budget, unioning both models' top-5 lists matched Gemini alone at k = 10 exactly (99.36%), and one designee sat at rank 60 despite a one-character surname difference. Embedding similarity is therefore a useful candidate-generation channel but not a reliable accept/reject rule; on a limited, partly synthetic corpus queried against mutable hosted APIs these are benchmark-specific findings, not a no-miss guarantee. Data, results and figures are released.

Downloads

Published

2026-09-30

How to Cite

PAVLENKO, I., & HNATUSHENKO, V. (2026). SEMANTIC EMBEDDINGS AS A CANDIDATE-GENERATION CHANNEL FOR SANCTIONS NAME SCREENING. Computer Systems and Information Technologies, (3), 17–28. https://doi.org/10.31891/csit-2026-3-2