DBS just moved an agentic credit-memo drafter from 150 pilot users to roughly 1,500 corporate bankers. The headline number attached to it — a 30% cut in credit assessment time — has not happened yet.
It is a target, not a result. The bank's own announcement on 19 August 2026 describes a successful pilot and then states an aim. It does not publish what the 150 pilot users actually measured. If you are about to walk into a steering committee with "DBS got 30% on credit memos, so we should fund this," you are quoting a forecast as a benchmark — and the arithmetic underneath it is smaller than it sounds.
What DBS Actually Shipped to 1,500 Bankers
DBS deployed a set of specialist agents that turn raw source material into review-ready first drafts of corporate credit memos, and it put them in front of about 1,500 relationship managers and credit risk managers globally after a 150-person pilot. The agents cover more than 70 distinct tasks, pulling from annual reports, industry research and internal bank records, and the banker can push back for more research or fold in their own read before signing.
That is a real deployment, not a demo. It is a 10x scale-up into a control function at one of Asia's most AI-forward banks, and the framing is honest about where judgement sits: Han Kwee Juan, DBS's Group Head of Institutional Banking, described it as capturing "the knowledge and insight of our best relationship managers and credit risk managers" to "level up the quality of our credit analysis at scale." Quality at scale, not headcount.
The 30% Is a Target, Not a Result
The number every trade write-up led with is an intention. DBS aims to cut the time required for these activities "by at least 30 per cent" — future tense, no baseline, no pilot delta, no measurement window. Nobody has published what 150 bankers achieved over the pilot period, which is the only evidence that would justify a 10x expansion.
Steel-man it, because there is a decent case. "At least 30%" implies the pilot cleared 30% and the bank is being conservative in public. Banks routinely under-disclose operating metrics for competitive reasons. And DBS has earned some benefit of the doubt: it reports around SGD 1 billion in economic value from data analytics and AI in 2025, across more than 2,000 models and 430 use cases, up from SGD 180 million in 2022. This is a bank that counts.
Which is exactly why the omission matters. An organisation that publishes an annual AI value figure to the nearest hundred million did not publish a pilot delta for this one. Treat the silence as information about what the number would have looked like, or at minimum as information about how firm it is. Either way, it is not your benchmark.
30% of a 40% Task Is 12% of a Job
Do the multiplication before you build a business case on it. DBS's own release notes that preparing credit memos and related credit activities can account for up to 40% of a relationship manager's time. A 30% reduction in a task that consumes at most 40% of a role returns roughly 12% of one full-time equivalent at the ceiling — not 30%.
A task-level saving is a reduction in the hours one activity consumes. A capacity-level saving is the share of a whole role it frees. The second number is always smaller than the first, and it is the only one a CFO can spend. Twelve percent of a banker's week is about half a day. That is genuinely useful. It is not a headcount line, and it will not survive contact with a budget committee that was told "30%."
Two honest caveats cut in opposite directions. The 1,500 users include credit risk managers, whose share of time on credit work is far higher than 40% — for them the blended saving is closer to the headline. And "at least 30%" leaves room above. Against that, 40% is the ceiling for relationship managers, memo drafting is only part of credit activity, and no time saving converts to capacity at 100% efficiency; freed minutes fragment across a day and a chunk of them go to reviewing the draft you just generated.
We have watched this arithmetic break in other sectors. Walmart and Amazon both reported 40% lifts from AI shopping assistants without running a holdout. Linear's agent teams pushed 65 pull requests a week and nobody got time back — the work moved to review. The pattern is consistent enough to plan around: task-level throughput improves, role-level capacity moves much less, and the gap is where the business case dies.
Deutsche Bank Did the Same Thing Six Days Later
Two major corporate banks pushed agents into corporate banking research within a week of each other, and neither published a measured outcome. On 25 August 2026, Deutsche Bank announced it helped design and will deploy Google Cloud's new Financial Research Agent, starting in its Corporate Bank with teams serving German MidCorp clients. Marie-Jeanne Deverdun, Deutsche Bank's Chief Technology, Data and Innovation Officer, framed it as "significant potential to reduce manual research effort, improve the consistency and auditability of outputs." Potential. No percentage at all — which is arguably more honest than a target dressed as an outcome.
The product underneath is Gemini Enterprise for Financial Services, launched the same day in preview for capital markets and corporate banking, built on the Gemini Enterprise agent platform. It ships a Google-managed Financial Research agent with more than 50 foundational skills, connectors into S&P Global, Moody's, FactSet, PitchBook, MSCI, SEC EDGAR and Dun & Bradstreet, and A2A APIs for wiring into existing workflows. Deutsche Bank and CME Group are design partners.
Google's own quantified claims — a bond portfolio risk exposure analysis in "sub-5-minute execution," bond issuance presentations "from days to minutes" — are vendor statements about individual tasks in a preview product, not audited results from a production bank. Read them as the vendor's claim, because that is what they are. The same caution applies to every Google Cloud banking deployment we have covered, from HSBC's 200 use cases to Citi's wealth advisor.
What a Measured Credit-Memo Claim Looks Like
A credible claim has a baseline, an after, and a named unit of work. Taishin Bank published one: credit report production went from 20 hours to 4 hours after automating consolidation across dozens of sources, organising over 100 data points and surfacing up to 20 key risk indicators. Judgement, strategy and client management stayed with the analyst. Lee Cheng-kuo, Taishin's Chief Digital Technology Officer, put the operating principle plainly: AI is "about freeing employees from tedious work so they can do more valuable things: decision-making, creativity, and customer management."
That is a number you can interrogate. Twenty hours of what, exactly? Measured over how many reports? Compared against which analysts? You may not get every answer, but the claim has a shape that permits the question. "At least 30%" does not.
The same discipline shows up in the deployments that hold up under scrutiny. TD Bank cut mortgage pre-adjudication review from 15 hours to 3 minutes with a stated before and after. Airbnb's engineering velocity claim survived because exactly one number in it was auditable. Ask for that shape or discount the claim.
The Reviewer Becomes the Bottleneck
A "review-ready first draft" does not remove work — it relocates it from drafting to verification, and verification of plausible-looking output is slower than people think. A credit memo is a control artifact. It goes to a credit committee, it supports a lending decision, and somebody signs it. The reviewer cannot skim a fluent draft that cites 100 data points; they have to check the ones the decision turns on.
The best evidence we have on the size of that gap comes from software, where it has been measured under controlled conditions. METR ran a randomized trial with 16 experienced developers across 246 real tasks and found they took 19% longer with AI tools — while believing they had been sped up by 20%. The developers had forecast a 24% speedup going in, and still reported a speedup afterwards. The finding does not transfer directly to credit analysis, and it should not be quoted as if it does. What transfers is the shape of the error: practitioners systematically misjudge their own time savings from AI, in the optimistic direction, even after living through the opposite.
That is the specific reason a self-reported survey of your 1,500 bankers is not a measurement. Ask them and they will tell you it saved them time. Time them and you may find otherwise. This is the same trap behind the broader finding that the "saves time" AI pitch stopped clearing budget committees.
Your Regulator Will Ask for the Number
Supervisors are moving from principles to lifecycle controls, and "we aimed for 30%" is not an artifact you can hand an examiner. MAS consulted on Guidelines for AI Risk Management from 15 November 2025 to 31 January 2026 and published an AI Risk Management Toolkit with industry on 20 March 2026, covering AI oversight roles, risk materiality assessment and lifecycle controls, with an implementation workgroup tasked with developing risk frameworks for emerging AI technologies, agentic AI among them. Internationally, the Financial Stability Board consulted on a menu of 12 sound practices for responsible AI adoption on 10 June 2026.
DBS is not naive here — it runs credit analytics under a PURE framework requiring data use to be purposeful, unsurprising, respectful and explainable, and its release is careful that bankers retain responsibility for final decisions. The point is narrower: the productivity number and the control evidence come from the same measurement. If you never established a baseline for how long a memo took and how often the draft was materially wrong, you have neither the ROI case nor the model performance record, and you will be asked for both.
There is a workforce dimension you should price in too. DBS said in February 2025 it expects to reduce contract and temporary headcount by about 4,000 over three years as AI takes on more work, while creating roughly 1,000 AI-related roles — with then-CEO Piyush Gupta saying he was "struggling to create jobs" for the first time in 15 years. A 12%-of-an-FTE saving spread across 1,500 senior bankers is not that story. Do not let the two get conflated in your own business case.
What to Do Before You Fund a Scale-Up
This Week:
- Take every peer benchmark in your current AI business case and label each number target or measured. Send the list to whoever approved the case. Most of them will be targets.
- Multiply each task-level saving by the share of the role that task consumes. Put both numbers in the deck — the task saving and the FTE-equivalent — and let the committee see the difference.
- For your own pilot, write down the baseline you would need: hours per memo, memos per analyst per month, and the review rework rate. If you cannot produce those from the pilot you already ran, you did not run a measurement.
This Month:
- Instrument the review step, not just the draft step. Track time-to-approve, edit distance between draft and filed version, and the rate at which a reviewer sends a draft back. Verification cost is the number nobody budgets and everybody pays.
- Run a holdout. Hold 20% of eligible analysts off the tool for one full cycle. Self-reported time savings are not evidence, and the METR result is the reason.
- Ask your vendor which of its published figures came from a production customer with a baseline, and which came from an internal benchmark. Get the answer in writing before renewal.
Before You Approve the 10x:
- Require the pilot delta, not the pilot verdict. "Successful" is a verdict. "18% median reduction across 340 memos over 11 weeks, with a 6% rework rate" is a delta.
- Decide in advance what result would make you stop. A scale-up gate with no failure condition is a rollout with a review meeting attached — which is how most agentic banking pilots end up in production without ever proving anything.
The Bottom Line
Every technology cycle produces a peer number that gets copied faster than it gets checked. In the ERP era it was implementation timelines. In the cloud era it was the infrastructure saving that turned out to be a lift-and-shift bill increase. The agent era's version is the task-level time saving, quoted as if it were capacity, sourced from a press release that said "aims to."
DBS may well hit 30%, and if it does, that is a good outcome for a bank with 1,500 senior bankers doing credit work. It is still 12% of a job, and it is still a forecast. Your board is going to ask you for a number in a quarter. Make sure it is one you measured.
The peer benchmark is not evidence. It is a press release with a percentage in it.
Continue Reading
- Walmart and Amazon Both Said 40%. Neither Ran a Holdout.
- Agent Teams Hit 65 PRs a Week. Nobody Got Time Back.
- TD Bank's First AI Agent: 15-Hour Reviews in 3 Minutes
- Airbnb Shipped 80% More Features. One Number Is Auditable.
- Citi's AI Wealth Advisor: The Enterprise Agent Blueprint
- HSBC + Google Cloud: 200 AI Use Cases Worth $100M Each
- The 'Saves Time' AI Pitch Is Dead. Here's What Works.
- Banks Deploy Agentic AI But Most Will Fail
