Most AML evaluation advice is written for the buyer. This one is written for the function that will live with the decision.
The distinction matters because compliance teams are evaluated on things a procurement scorecard does not measure. Your examiner does not ask about latency. They ask whether your monitoring is calibrated to your risk assessment, whether your decisions are documented consistently, and whether you can explain a rating from eighteen months ago. Your board does not ask about scenario counts. They ask why compliance headcount needs to grow 40% next year.
So the framework below evaluates platforms against four questions a compliance function actually has to answer, then works through the nine capability areas with the compliance version of each question rather than the technical one.
Start from your program, not from a feature list
Before you look at a single platform, assemble three documents:
Your enterprise-wide risk assessment. The typologies named in it are the ones your monitoring must demonstrably cover. A vendor’s scenario library is only relevant to the extent it maps onto that document. Bring it to every demo and ask directly which scenarios cover which named risks.
Your last examination report or audit findings. These tell you where your program is already judged weak. If your findings mention documentation consistency, weight investigation workflow and audit trail heavily. If they mention tuning and calibration, weight configurability and testing. Buying a platform that does not address your open findings is an expensive way to receive the same finding twice.
Your current alert and case statistics. Monthly alert volume, false positive rate if you measure it, average time to disposition, backlog, and SAR conversion. Without these you cannot tell whether a vendor’s projection is good or terrible, and you cannot build a business case.
If you cannot produce the third document, produce it first. An evaluation without a baseline is a comparison of vendor claims to each other.
The four questions your function has to answer
Every criterion below rolls up to one of these. When you are unsure how to weight something, ask which of the four it serves.
1. What does this do to our workload? Alert volume converts directly into headcount. This is the largest financial consequence of the decision and the one most likely to land on you personally in a budget conversation.
2. Can we defend it? Every decision the platform makes or assists will eventually be reviewed by someone adversarial. Explainability and audit trail are not technical features here, they are the raw material of your examination response.
3. Do we control it, or does the vendor? A control environment you cannot change without raising a ticket is a control environment you do not own. Examiners increasingly probe this.
4. What happens when we grow or change? New product, new market, new licence condition. If each one is a project rather than a configuration change, your program will permanently trail your business.
The nine capability areas, in compliance terms
1. Functional coverage
The technical question: which modules are native?
The compliance question: how many systems will an analyst touch to close one alert, and how many places will an examiner have to look to reconstruct one customer’s history?
Every seam between two systems is a handoff you have to document and a place your evidence trail breaks. If screening sits in one tool and monitoring in another, the analyst assembles the picture manually on every alert, and you explain the integration to your examiner rather than showing them a case.
Ask specifically whether a risk rating change automatically affects monitoring, and whether a screening hit and a monitoring alert on the same customer land in the same case.
2. Configuration
The technical question: is it no-code?
The compliance question: when we need to respond to a new typology, how long until the control is live, and who has to be involved?
Test this by having a non-technical person build a rule during the demo, timed. Then ask the commercial half, in writing: do configuration changes attract professional services fees. That question is rarely asked and is the most reliable predictor of how your second year will feel.
Also ask what governance sits around the change. A platform that lets anyone alter a threshold without approval creates a different problem from the one it solves. You want maker-checker, version history, and an immutable log, because those are what turn a rule change from an operational act into a documented control decision.
3. Integration
The technical question: how many connectors?
The compliance question: how much of our IT team’s capacity does this consume, and what is the real go-live date?
Implementation time is a compliance risk, not a procurement detail. Every month between signature and production is a month your controls do not match your volume. If you are under an examination commitment or a partner bank deadline, a six-month deployment is a failed purchase however capable the platform.
Get the go-live in days, named to a comparable customer, and the engineering hours required from you as a number, both in writing.
4. Monitoring
The technical question: what is the latency?
The compliance question: can we act in time to prevent, and can we show our thresholds are calibrated to our risk assessment?
Two things matter more to a compliance function than raw speed. First, whether the platform can stop a transaction rather than only record it, because on instant rails, post-settlement detection is documentation of a loss. Second, whether thresholds can vary by customer risk band, since a single global threshold is difficult to defend against a risk-based approach your own policy describes.
Ask how detection behaves for a product or segment with no history. Threshold rules require you to know what normal looks like, and for anything newly launched you do not.
5. Screening
The technical question: which lists?
The compliance question: can we control false positives without loosening detection, and can we prove which list version was in force when a customer cleared?
Sanctions is the one control where a single miss is existential, so the temptation to tune for fewer alerts is dangerous. What you want is configurable matching, so you can match tightly on some name types and loosely on others, plus the ability to test a threshold change against historical hits before applying it. That converts tuning from a judgement call into a documented, evidenced decision.
Ask for list-version history explicitly. Examiners ask which version was in force on a given date, and many programs cannot answer.
6. Investigations
The technical question: does it have case management?
The compliance question: does it encode our procedure, and does it make our analysts consistent with each other?
Most examination findings are not about wrong decisions. They are about defensible decisions documented inconsistently. So the question is whether the workflow enforces your SOP rather than relying on analyst memory, whether statuses and SLA timers reflect your policy, and whether quality assurance reviews cases against your procedure rather than sampling five percent and hoping.
Bring your actual SOP to the demo and ask what the platform can do with it. That single request separates platforms that impose a workflow from platforms that adopt yours.
7. Reporting
The technical question: which filing formats?
The compliance question: how much analyst time goes into drafting, and can we evidence the full lifecycle of every filing?
Narrative drafting is one of the largest recurring time costs in a small function, and one of the most common sources of inconsistency between analysts. Ask whether narratives populate from case data, whether filing happens inside the case, whether the submission receipt is stored against it, and whether every draft and edit is version-retained.
Separately, ask for the management information: analyst throughput, resolution times, SLA adherence, alert-to-case conversion. You need these for your board and for your own tuning, and exporting to a spreadsheet each month is a cost.
8. Implementation and support
The technical question: what is the timeline?
The compliance question: who calibrates our rules, and who do we reach at 6pm during an examination?
Rule calibration during rollout is the step most likely to be quietly assigned to you. Establish who owns it. And weight support more heavily than a procurement process normally would, because a two-person compliance team has no internal escalation path when something breaks.
9. Segment fit
The compliance question: has this vendor been examined alongside institutions like ours?
Ask for references at your size, on your core, in your jurisdiction, and call them without the vendor present. Ask what took longer than expected, what they still do outside the platform, and what they learned in year two.
Who on your team tests what
An evaluation run entirely by one person produces one person’s blind spots.
MLRO or BSA officer: explainability, audit trail, and governance around rule changes. This is the person who will defend the program, so this person should personally attempt the two-year-old decision reconstruction.
Lead analyst: the case workspace. Have them close a routine alert in each platform and count the clicks and the system switches. They will notice in ten minutes what a procurement process misses in three meetings.
Whoever owns tuning: the rule builder and the testing tools. Have them build a real rule from your program, timed, and run the backtest.
IT or engineering: integration scope and required hours, security documentation, and data residency.
Finance or procurement: pricing structure, what triggers increases, and the cost of adding a jurisdiction. Give them the alert volume projection too, because that is the number that actually drives cost.
Building the business case
Compliance functions usually have to justify this spend to people who see it as overhead. Three framings that work better than risk avoidance alone:
Cost avoided, not cost incurred. If alert volume is projected at X and your team can absorb Y, the difference is headcount you would otherwise hire. That is a defensible number if it comes from a vendor projection against your own transaction profile rather than a marketing percentage.
Time returned to review rather than assembly. Analyst hours spent gathering data are hours not spent making decisions. Quantify current time-to-disposition and compare.
Open findings closed. If the platform addresses a documented examination finding, say so explicitly. That converts the purchase from discretionary to remedial in the eyes of a board.
Present the total cost honestly: licence, implementation, your engineering hours, implied analyst headcount, configuration change fees, and the cost of adding a jurisdiction. A business case that omits the headcount line will be reopened later.
How one platform maps to this framework
Included as a reference for the level of specificity to demand, not as a recommendation. Verify directly.
Flagright, against the compliance questions above:
Systems touched and evidence trail. Transaction monitoring, screening, risk scoring, case management, and filing run on one platform, with screening hits escalating into the same case system as monitoring alerts, and the monitoring engine reading a customer’s live risk score at transaction time to apply the threshold for that band. That addresses both the seams question and the risk-based calibration question in one architecture.
Control and governance. Validated rule creation time of 60 seconds, roughly three minutes as measured by customers, through a no-code builder, with Flagright stating that rule changes require no SQL, no engineering tickets, and no professional services fees. Around the change: maker-checker approval chains separating creation from sign-off, every rule version preserved with one-click rollback, and an immutable timestamped audit log with user attribution. Rules backtest against 90 days of historical transactions and run in shadow mode against live traffic before promotion, which turns tuning into a documented decision rather than a judgement call.
Workload. Reported 77% of alerts auto-cleared with high confidence at 94% analyst agreement, alert-to-outcome time of 4 minutes against a 38 minute baseline, and 80% faster investigation closure. Reported false positive reduction is up to 83% from threshold optimisation specifically and 93% across broader AI tooling, two different scopes, which is the kind of distinction to require of any vendor quoting a single figure.
Procedure and consistency. Your investigation SOP can be uploaded and converted into a deployed investigation agent in a stated 20 minutes, with control ranging from silent evaluation to full automation. Custom statuses, SLA timers by case type or jurisdiction or risk level, and no-code escalation routing. A QA module reviews every investigated case against your SOP, flagging missed steps, unsupported conclusions, and weak evidence, rather than sampling.
Defensibility. Every alert carries the rule and evidence behind it. AI decisions carry an annotated transaction timeline, typology citation, confidence score with contributing factors, and a specific model version, exportable in JSON or Excel for regulatory submission.
Screening control. Named, independently configurable matching algorithms including Jaro-Winkler, Levenshtein, transliteration, phonetic matching, DOB delta tolerance, tokenisation, and stopword filtering, with per-list and per-mode thresholds, simulation against historical match activity, and full list-version audit history.
Reporting. SAR and CTR filing by API direct to FinCEN and to 70+ GoAML countries from inside the case, narratives pre-filled from case data, submission receipt stored in the audit log, full version history, plus operational analytics on throughput, SLA compliance, and resolution times.
Implementation and support. Published average go-live of two weeks with rules calibration included, API-first with more than 100 native integrations including published connectivity to Jack Henry Symitar and Fiserv DNA. Reported 6 minute average support response, 24/7 coverage, dedicated CSM. ISO 27001:2022 and SOC 2 Type II certified.
Material considerations. Pricing is not published, so the business case must be built from your own quote. Cloud-native only. Screening list refresh is described as continuous without a per-list interval, so request it contractually. Uptime appears as 99.998% on product pages and 99.99% on security documentation, so ask which is contractual. If you are keeping existing case management and buying monitoring alone, confirm standalone operation and third-party alert ingestion, since the design assumes a unified stack.
Others commonly evaluated, to verify against the same questions: Unit21 and Hawk AI for combined fraud and AML with no-code configuration, Napier AI for a production sandbox approach to rule testing, Lucinity for investigation workflow depth, ComplyAdvantage as a screening and data layer, Nasdaq Verafin for US banks and credit unions, and the enterprise suites from NICE Actimize, Oracle, and SAS where scale requires them and a longer implementation is acceptable. Descriptions are drawn from public sources, are brief, and change as products evolve.
After selection: the first ninety days
The evaluation does not end at signature, and examiners will ask about this period.
Document why you chose. Retain the scorecard, the weightings, and the rationale. Vendor selection is a control decision and should be evidenced like one.
Validate before you rely. Run the new rule set in parallel or shadow mode against your existing system, compare alert populations, and document the differences and your explanation for them. This is your pre-implementation validation record.
Baseline and tune deliberately. Record alert volume, disposition rates, and time-to-decision in month one, and treat every subsequent threshold change as a documented decision with a reason, an approver, and a before-and-after measurement.
Write the governance down. Who can change a rule, who approves it, how often the rule set is reviewed, and what triggers an out-of-cycle review. A platform that logs everything still needs a policy that says who is permitted to do what.





