Socure's agents will work AML alerts, and the number they are sold on — up to 70% fewer false positives — is the same number a BSA examiner reads as suppressed alerts. In fraud, fewer false positives is a productivity win. In anti-money laundering, a drop in alert volume is indistinguishable from a drop in detection until you can prove the alerts you stopped seeing were never productive. That proof is your obligation, not your vendor's.
On 27 August, Socure announced a $156 million strategic growth investment at a $5.2 billion valuation, led by Summit Partners with Goldman Sachs Alternatives, Wells Fargo and Docusign participating — and, in the same release, the acquisition of Fravity, an agentic fraud, risk and compliance platform. If your institution is one of the top five U.S. banks, 160 public sector organisations or 600-plus fintechs already on Socure's RiskOS, this is not a feature toggle arriving in a release note. It is a change to the system that decides which of your alerts become SARs.
What Socure Actually Bought
Fravity is an Austin company whose agents do the work your L1 analysts do, and it has been folded in fast. It was founded by Kedar Samant and Rushik Upadhyay — Samant co-founded the fraud platform Simility, which PayPal bought in 2018, and Upadhyay was a chief architect on compliance at PayPal. Terms were not disclosed. The company's own domain, fravity.ai, already issues a 308 redirect to socure.com.
The product is not a scoring model. Per Biometric Update's account of the deal, Fravity's Agentic Studio lets teams "orchestrate more than 70 pre-built agents or create their own around internal policies and risk models," while an Agentic Copilot executes workflows and produces "explainable, case-ready outputs" across KYC investigations, business due diligence, transaction-monitoring alert assessment, sanctions and PEP screening, and adverse-media research.
It ships as RiskOS_Agents, and the first workflows are watchlist screening and monitoring plus know-your-business checks. Those are exactly the two surfaces your regulator examines.
The financials behind it are real. Socure reported $364 million in ARR in Q2 2026, up 63% year over year, with 133% net dollar retention and 0.01% logo churn, per the Summit Partners release; Crunchbase puts total funding above $742 million since 2012 against a $4.5 billion Series E valuation in 2021. This is not a vendor you can wait out.
The 70% Is Your Regulator's Number Too
The headline metrics come from Fravity's existing deployments, and neither the release nor any coverage says how many there were. The claim is that Fravity cut cost per case by 80%, accelerated case resolution by up to 5x, and reduced false positives by as much as 70% across its existing deployments. No customer is named. No deployment count is given. "As much as" is doing load-bearing work.
Set aside whether the number is true. Assume it is. You still cannot use it. There is no regulatory benchmark for it to beat — as Abrigo notes in its guide to threshold testing, "there is no regulatory guidance on the acceptable rate of false positives". The only accepted way to show a tuning change did not cost you coverage is above-the-line/below-the-line testing: you re-run the population at looser thresholds and sample what the new configuration would have dropped, and you keep the working papers.
The industry baseline explains why so few people defend the alert population as it stands. The Bank Policy Institute — a trade association for the largest US banks, arguing in the same report that "the current US compliance regime is broken" — surveyed 19 of its member banks and found that in 2017 they reviewed roughly 16 million alerts, filed over 640,000 SARs, and saw a median of 4% of those SARs draw a law enforcement follow-up — at a cost of 14,000 staff and about $2.4 billion. Sanctions filtering was worse: a true match rate of 0.00004%. Read that number with its provenance: it is self-reported by an interested party, and it counts follow-up inquiries back to the filing bank rather than use, which is not the same thing. FinCEN's own figures put 16% to 40% of FBI investigations in FY2024 as linked to a SAR or CTR, depending on the crime. The weaker claim survives either way and is all this argument needs: the productive share of any alert population is small, and everyone in the room knows it.
That is also precisely why the reduction is dangerous. Nobody gets examined for having too many alerts. TD Bank got examined for the opposite failure, and the shape of the finding is instructive: it failed to substantively update its transaction monitoring system despite rapid growth between 2014 and 2022, leaving roughly 80% of its transactions unscreened for suspicious activity. The bill on 10 October 2024 was $3.1 billion — a $1.43 billion criminal penalty, $452.4 million in forfeiture and a $1.3 billion civil penalty. TD's failure was inaction. But the legal theory that produced it — you owe evidence about what your system does and does not see — applies identically to a change that makes your system see less.
The Model Risk Guidance You Would Have Cited Was Rescinded in April
Here is the part your model risk team may not have caught: the federal guidance you would have used to validate RiskOS_Agents no longer exists, and its replacement deliberately excludes agentic AI.
On 17 April 2026 the OCC, Federal Reserve and FDIC issued revised model risk management guidance. OCC Bulletin 2026-13 states the scope carve-out in two sentences: "Generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance." The same bulletin rescinds OCC Bulletin 2011-12, the Model Risk Management booklet of the Comptroller's Handbook, and — most relevant here — OCC Bulletin 2021-19, the interagency statement on model risk management for bank systems supporting BSA/AML compliance. Davis Polk's read confirms the Federal Reserve side: SR 11-7 and SR 21-8, the BSA/AML letter, are both superseded.
The revised guidance is also lighter where it does apply. Sullivan & Cromwell reads it as dropping the "at least annual" validation cadence entirely and framing itself as most relevant to banking organisations with over $30 billion in total assets, though it may also reach smaller institutions with significant model risk exposure, with a request for information on AI promised at some unstated future date.
Read the practical consequence carefully, because it cuts both ways. There is no federal model risk framework that tells you how to validate an agentic alert-disposition system — and there is no federal model risk framework that tells your examiner what to accept, either. A vacuum is not a permission. It means the standard you are held to at your next exam will be constructed after the fact, from your documentation. This is the same trap as inheriting a frontier lab's own safety evaluations and calling it assurance: the absence of an external bar does not transfer the burden to somebody else.
What Part 504 Still Requires, Guidance or Not
If you are chartered or licensed in New York, none of the federal churn matters, because 3 NYCRR 504.3 is a rule, not guidance, and it already describes what you owe. Its enumerated attributes for a Transaction Monitoring Program include:
- "end-to-end, pre-and post-implementation testing of the Transaction Monitoring Program" — which means RiskOS_Agents needs a documented test before you turn it on and after.
- "documentation that articulates the institution's current detection scenarios and the underlying assumptions" — a vendor's 70+ pre-built agents are detection logic. You have to be able to write down what each one assumes.
- "protocols setting forth how alerts generated by the Transaction Monitoring Program will be investigated" — if an agent investigates and closes, your protocol has to say so explicitly.
- "on-going analysis to assess the continued relevancy of the detection scenarios" — agent behaviour drifts with the model underneath it, which no static validation catches.
The Filtering Program attributes carry the parallel obligation for watchlist screening — "on-going analysis to assess the logic and performance of the technology or tools" — and watchlist screening is the first thing Socure is shipping.
Then 504.4 closes it: a Board Resolution or Senior Officer(s) Compliance Finding filed by April 15 each year, and an obligation to "maintain for examination by the Department all records, schedules and data supporting adoption" of that finding for a period of five years. Somebody signs their name. If an agent closed 70% of an alert population, the record of why has to survive five years and be legible to a person who was not there.
Where the Agent Actually Decides
The regulated boundary is auto-disposition, not summarisation, and the industry is already crossing it. Fravity's own framing is careful — the goal is "not necessarily to remove humans, but to make that review layer faster and more scalable." Compare that with the competitive standard being set. Nasdaq Verafin, announcing its expanded agentic workforce on 10 June 2026, says agents "will be able to autonomously execute entire workflows from end-to-end, closing out false positive alerts and only surfacing alerts that require deeper, human-in-the-loop review," and claims up to a 90% reduction in sanctions alert review workload. That is where the market is going, and Socure just bought its ticket.
An agent that drafts a case narrative for a human is a productivity tool. An agent that closes an alert is a component of your SAR decisioning process. Only the second one is a change to your program — and the "explainable output" that makes the first one safe is doing much less work than the word implies.
There is now a measurement of exactly how much less. A March 2026 paper, Explainable AML Triage with LLMs: Evidence Retrieval and Counterfactual Checks, built an evidence-constrained triage framework requiring explicit citations and then perturbed the inputs to see whether the recommendation and the rationale moved coherently together. Evaluated on a synthetic benchmark rather than live casework, citation validity came in at 0.98 and evidence support at 0.88 — but counterfactual faithfulness was 0.76. Roughly one rationale in four looked well-sourced and did not actually track the decision it claimed to explain. Escalate F1 was 0.62.
That gap matters more than it looks, because of who is reading the output. Our own coverage of why vaguer AI explanations earned more deference from novice reviewers than from experts describes the exact staffing profile of an L1 alert queue. A confident, well-cited, subtly unfaithful case summary handed to a junior analyst under volume pressure does not get challenged. It gets approved — and the approval is what your five years of records will show.
The Strongest Case for Buying It Anyway
State the other side properly, because it is good. The human L1 queue you are protecting is not validated either. Nobody runs an inter-rater reliability study on alert dispositioning. Nobody measures whether analyst 7 and analyst 19 close the same alert the same way on a Friday afternoon. Even discounted for its source, the 4% follow-up rate is an indictment of the incumbent process, not of the software proposing to replace it.
Regulators have said as much. OCC Bulletin 2018-44, the joint statement on innovative BSA/AML approaches, explicitly recognises "the role of pilot programs in testing and validating the effectiveness of such approaches" and commits the agencies to early engagement. And FinCEN's proposed AML/CFT program rule, published in the Federal Register on 10 April 2026 with comments closing 9 June, moves the whole regime from a process standard to an effectiveness standard — a bank that can demonstrate better outcomes from agentic triage is better positioned under that rule than one that cannot.
So the answer is not "don't buy it." The answer is that the 70% has to become your measurement rather than Socure's marketing, and the difference between a target and a measured result is the entire distance between a pilot and an exam finding — the same gap that separates DBS's 30% credit-memo goal from anything it has actually reported.
What to Do Before RiskOS_Agents Reaches Your Queue
This Week:
- Get written confirmation of the enablement path. Ask your Socure account team, in writing, whether RiskOS_Agents can be switched on by a Socure-side configuration change, by a customer admin, or only by a contracted amendment. If the answer is either of the first two, that is a change-control gap in a regulated system — the same class of problem as a vendor acquisition quietly rewriting a data policy you thought was contractual.
- Write down the boundary. One page: which agent actions are advisory (draft, summarise, retrieve) and which are dispositive (close, clear, suppress, auto-escalate). Nothing dispositive goes live without the testing below.
- Ask for the denominator. Request the deployment count, institution types and alert populations behind the 80% / 5x / 70% figures. A vendor that will not tell you n has told you something.
This Month:
- Design the holdout before the pilot, not after. Route a randomised share of alerts through your existing process in parallel and compare disposition against the agent's, on the same population, over the same window. This is the discipline missing from most published AI wins — two retailers claimed 40% lifts this quarter and neither ran one.
- Run below-the-line sampling on the suppressed set. Pull a statistically defensible sample of the alerts the agent closed, have humans work them cold without seeing the agent's rationale, and count the productive ones. Keep the working papers.
- Log the rationale as evidence, not as a UI element. Every agent disposition needs the retrieved evidence, the model and prompt version, and the closing rationale persisted in your case management system — not in the vendor's. Five years is longer than your contract.
Before Your April 15 Certification:
- Put RiskOS_Agents through the same change-control gate as a scenario tuning. Model inventory entry, documented assumptions per agent, pre- and post-implementation test results, and a named owner.
- Brief the senior officer who signs. They are attesting to a program that now contains a component with no federal model risk framework behind it. They should know that before April, not during the exam.
- Re-test after every model version change underneath the agents. A silent upgrade to the foundation model is a change to your detection logic, and it will not arrive in a release note you recognise.
The Bottom Line
This is the second time this industry has bought its way out of an alert-volume problem. After NYDFS Part 504 landed in 2017, a generation of tuning and machine-learning triage tools sold banks the same promise — fewer alerts, same coverage — and the ones that survived examination were the ones that could produce the below-the-line sample and the working papers, not the ones with the best reduction number. The technology is far better this time. The evidentiary standard has not moved an inch.
There is a useful analogue outside finance. The FDA has cleared over 1,300 AI-enabled devices, and only three were supported by trials measuring patient outcomes. Clearance is not evidence. Neither is a $5.2 billion valuation, 133% net dollar retention, or a vendor's 70%.
Your examiner will not ask what Socure claimed. They will ask what you measured.
Continue Reading
- Visa Bought BioCatch. Get Neutrality in Writing.
- Banks Deploy Agentic AI But Most Will Fail
- 10 Banking AI Use Cases with Real ROI Benchmarks for 2026
- Claude's 10 Finance Agents: KYC From 4 Hours to 30 Seconds
- Agents Averaged 73. Only 30% Were Usable. Grade Pass/Fail.
- Descartes Bought Tai. Your Only Lever Is 60 Days.
- Stripe Bought OpenRouter. A Toggle Is Not a Contract.
