Case studies
Wio Associate home task · Option 2 · Reducing manual payment investigations

Catch payment problems at the door, not in the queue.

Clearway: how Wio can find, rank and resolve payment issues earlier, and keep every customer at ease while it happens.

Prepared byJaya Sabarish Reddy Remala
RoleAssociate, CEO Office
DateOctober 2026
IncludesA working prototype
My answer on one page

Stop adding people to the end of the pipe. Fix the three places where manual work starts.

01The challenge

Payments that need a human sit in silence. Operations and Compliance work across many systems, and the customer finds out last.

02Data and research

Trace every held payment end to end by its tracking reference (UETR), and size the causes before designing anything.

03Root causes

Bad data is caught too late. Screening has no memory. Nobody owns the payment journey end to end.

04Solution and impact

Clearway: Prevent, Prioritise and Inform on one shared timeline, using rules, machine learning and an AI copilot where each fits best. People keep every decision. About half of manual hours freed at my assumed volumes.

05Execution

Discover, Design, Pilot, Scale. First 30 days: a baseline, checks at entry, one-question requests, and proactive updates.

1 · Challenge understanding

In plain words, an international payment is a parcel crossing several borders.

We pack it, label it and send it. It then passes checkpoints we do not run. If the address is wrong, the label is vague, or the sender's name looks like someone on a watch list, the parcel goes on a shelf. Someone then writes to the sender to ask what is inside. Meanwhile the sender sees "processing" and wonders where their money went.

1
InitiateApp or APIMissing details
2
ValidateFormat, limits, purposeWrong purpose
3
ScreenSanctions, AML, fraudSanctions alert
4
ReleaseSWIFT, UAEFTS, AaniRouting error
5
In transitPartner banksNo confirmation
6
CreditReceiving bankReturned
7
ReconcileStatements vs ledgerAmount gap
Steps 1 to 4 happen inside Wio, so we can see them
Steps 5 to 7 happen outside, so we only see status messages
1 · Pain points

Everyone pays for a held payment. Only the customer pays in anxiety.

Customers

Experience, trust, clarity, effort

  • Money in limbo, with no reason and no time
  • Vague requests for "documents", sometimes twice
  • Has to contact support to learn anything
  • SMEs risk late fees and supplier trust

Operations and Compliance

Effort, complexity, rework, risk

  • Most sanctions alerts turn out to be false positives
  • One case means checking five or six systems
  • Free-text questions lead to partial answers and second rounds
  • Alert fatigue raises the chance of missing a true hit

The business

Cost, growth, revenue, risk, brand

  • Cost grows with every new customer if the fix is headcount
  • Slower settlement and more support tickets
  • A brand built on word of mouth and high NPS is at stake
  • Regulatory exposure if a true hit slips through

The constraint that shapes everything: we cannot loosen controls. The answer has to be smarter controls, not fewer controls.

1 · Assumptions

What I know, and where I am filling gaps

What is public

Customers270k+ personal, 120k businesses, 10k+ families
ScaleAbout AED 61bn in assets, AED 1.24bn revenue (2025)
Cross-borderA global clearing partner for USD, EUR and GBP
RulesNew UAE AML law in force since October 2025, including a ban on tipping off
IndustryAbout 72% of Swift payment exceptions come from data and format errors

What I assumed

AssumptionValue
International payments sent per month350,000
Share needing a person6% → 21,000 cases
Mix: sanctions, data, status, recon, other45 / 25 / 15 / 10 / 5 %
Handling time per case, including questions25 min → 8,750 h
Held payments that trigger a support contact40%

I would replace every one of these in the first two weeks. The conclusion holds across a wide range, and the prototype lets us test that live.

2 · Data, insights and research

Before designing anything, I would trace every held payment end to end.

What I would look atThe question it answersWhere it likely lives
Payment events, joined by UETRWhere do payments stop, for how long, and who touched them?Payment hub, core banking, gpi tracking, status and return messages
Exception reasons and repairsWhich causes dominate, by corridor, segment and channel?Payment operations case tool, reject and return codes
Screening alerts and outcomesFalse-positive rate by rule and name pattern, and repeat alerts on the same pairScreening system audit log, analyst case notes
Requests for informationWhich questions get partial answers, and how many rounds per caseCase notes, in-app messages, email
Support contactsHow many customers ask before we tell them?CRM and chat tags, app reviews
Transfer form behaviourWhich fields get edited, fail checks or cause drop-off?Product analytics
Reconciliation breaksWhat share is just fees, and how long do breaks stay open?Reconciliation tool, partner bank statements
PeopleThe workarounds and judgement calls that data missesShadow 10 analysts for a shift. Interview 8 to 10 SME and freelancer customers.
2 · Questions and patterns

Six questions I would ask in week one

Compliance. Of last month's false positives, how many could we have closed with data we already held? What stops us remembering a cleared beneficiary?

Operations. Walk me through the last 10 cases you closed. Where did you wait, and on whom?

Engineering. Do we store the UETR and every status update for every payment in one place we can query?

Product. Which transfer-form fields get edited or fail most, and what is checked before submit?

Customer support. What is the first thing a customer asks about a held payment, and what can't you answer?

Treasury and partners. Which corridors and partner banks cause the most delays, returns and fee gaps?

Patterns I would look for

  • The few causes that make up most of the volume
  • The same customer and beneficiary flagged again after being cleared
  • False positives clustered on Arabic and South Asian name spellings
  • First-time beneficiaries and customers in their first 90 days
  • Weekend and cut-off-time effects
  • Cases needing more than one round of questions
  • Customers contacting us before we contact them
  • Analysts taking very different times on the same alert type
3 · Root cause hypotheses

Three root causes. Each one survives because it is rational for someone.

01Problems are caught downstream

Wrong IBANs, vague purposes and missing documents enter at the form. They are caught later, by screening, a partner bank or the receiving bank, when fixing them costs the most.

Why it persists · process and technologyThe form was built for a fast first send, which suited growth. Checks live in separate systems.

02Screening has no memory

Every payment is screened from scratch. Name spellings collide. Analysts lack the identifiers they need in one view, so they ask the customer.

Why it persists · technology and policyMissing a true hit is existential, while a false positive "only" costs time. So thresholds only tighten, and nobody owns the cost of noise.

03Nobody owns the whole journey

A payment's status is split across the app, payment hub, screening, tracking, reconciliation and CRM. Each team clears its own queue. The customer learns last.

Why it persists · people and dataTeam boundaries follow system boundaries. Success is measured per team ("alerts closed"), not per payment ("money arrived").

If these hold, the fix is not more people. It is checks at entry, memory in screening, and one owner of the payment from submit to arrival.

4 · What leading banks already do

Every piece exists somewhere. Nobody has joined them into one journey.

CapabilityGlobal leadersIn the UAEThe lesson for Wio
Check the name before sendingUK Confirmation of Payee runs about 1.9 million checks a day. Verification of Payee is now required for euro transfers.Aani pays by mobile number, so fewer IBANs are typed. Name checks on IBANs are still emerging.Validate account and name at submit, not after release
Track every paymentWise shows every step. 77% of its transfers arrive in under 20 seconds.Mashreq was the region's first bank on Swift gpi, with tracking in its app.Make status visible by default, so customers don't have to ask
AI on screening alertsHSBC reports 60% fewer false positives and 2 to 4 times more financial crime found.Emirates NBD automated its screening-alert investigations.Measure fewer false positives and more true hits together
Trust known patternsWise skips extra steps for trusted recipients.Mostly handled case by case.Remember cleared beneficiaries and answers already given

Where Wio can lead: incumbents bolt these on one by one. A cloud-native bank can join prevention, triage and customer messaging on one platform.

4 · Proposed solution

Clearway: three layers on one shared timeline

Prevent

Fixes root cause 01

  • Checks while the customer types: IBAN, bank code, account and name
  • Purpose suggested from history and the invoice
  • Invoice asked for up front, only when a payment is risk-flagged

Prioritise and resolve

Fixes root cause 02

  • One queue, ranked by risk, value, deadline and customer impact
  • Memory of cleared beneficiaries, with expiry
  • A copilot that gathers evidence and drafts. A person decides
  • Automatic chasing of partner banks

Inform

Fixes root cause 03

  • A parcel-style tracker for every payment
  • Every hold says what, next step and by when
  • One request, as a checklist, answered once
  • Updates sent before the customer has to ask
One timeline per payment, joined by its UETRapppayment hubscreeningtrackingreconciliationCRM

Machines gather, rank and draft. People decide. Customers always know.

4 · Prevent

The cheapest investigation is the one that never opens.

Clearway pre-flight checks catching an IBAN typo, with one-tap fixes for the IBAN, purpose and invoice

Real checks. The IBAN checksum and bank-code rules catch typos instantly, before the payment leaves.

One-tap fixes. We suggest the saved IBAN, the likely purpose and the invoice, rather than just showing an error.

Friction only where it pays. We ask for more only when risk is high, and A/B test the impact on first-send conversion.

4 · Prioritise and resolve

Rank by real risk, remember what we know, and let a copilot do the gathering.

Ranking

priority = 0.40 risk + 0.25 value + 0.20 deadline + 0.15 customer impact

Operations and Compliance set the weights together and review them monthly. A likely true match always goes to the top.

Smarter screening

  • A cleared customer and beneficiary pair is not flagged again until the list or the details change
  • Matching that understands name spellings and checks date of birth, nationality and registration
  • A model that ranks alerts first, and closes only the safest ones after it has proven itself

Investigator copilot

  • Builds the case file: KYC, history, past decisions, tracking
  • Drafts the decision rationale with its evidence
  • Drafts the customer question in plain language
  • Read-only. It cannot release, reject or message anyone

Guardrails that do not move: a person signs off every decision, everything is logged, messages on likely true matches are locked to neutral wording, models are validated by Compliance, data stays in the UAE, and true-hit detection is tested continuously.

4 · Inform

Uncertainty hurts more than waiting. Show the journey and ask once.

Four customer app screens: routine check, one quick question, delayed at a partner bank, and returned

People tolerate a wait far better when they know why and for how long. Each screen gives a status, a time, and at most one question. Wording is written with UX writing and Compliance in English and Arabic, and sensitive reviews only ever use approved neutral text.

4 · Where machine learning and AI fit

Rules where the answer is exact, ML where we rank, AI where we read and write

ProblemApproachWhat it doesWho decides
Typos and format errorsRules: IBAN checksum, bank code and country checksExact and explainable, needs no training dataCustomer fixes with one tap
Purpose and documentsClassifier on payment history, plus document AI that reads invoicesSuggests the purpose and checks invoice amount and referenceCustomer confirms
Name matchingSpelling-aware fuzzy and phonetic matching, plus secondary identifiersHandles Arabic and South Asian name variants without missing true hitsAnalyst reviews hits
Which alert firstGradient-boosted model trained on past analyst decisions, with a reason shown for every scoreRanks alerts by false-positive likelihood. Never closes one by itself in phase 1Analyst
Which payment will stallTime-to-credit model by corridor and partner bankFlags a delay before the customer notices, and drafts the chaseOperations
Case file and messagesLLM copilot with read-only tools and search over case history and policiesGathers evidence, drafts the rationale and the customer question in English and ArabicAnalyst edits and approves
Reconciliation breaksRules plus anomaly detection on fee patternsClears known fee gaps and surfaces the odd onesOperations reviews outliers

The line I would not cross: no model or LLM makes a sanctions decision. Every model is tested on past cases before it touches the queue, then monitored for drift and for any drop in true-hit detection.

4 · Technology stack

How it is built: an event backbone, a case layer, and intelligence on top

Channels
Wio appTracker, one-tap answers, push updates
Ops consoleRanked queue, case file, decisions
Leadership dashboardSTP, time-to-money, guardrails
Intelligence
Rules engineValidation, routing, auto-resolution
ML modelsRanking, stall prediction, anomalies, with a model registry and monitoring
AI copilotLLM with agent orchestration, read-only tools and retrieval, hosted in the UAE
Workflow
Case serviceOne case per UETR, deadline timers, full audit log
Messaging serviceApproved templates only, English and Arabic
Feedback loopAnalyst outcomes become training labels
Data
Event streamingEvery payment event, keyed by UETR, replayable
Operational databaseCases, decisions, memory of cleared pairs
Analytics and featuresLakehouse for baselines, features and experiments
Integrations
Payment hub and core banking
Screening system
Swift gpi, Case Management and Pre-validation
Clearing partner and CRM

Across every layer: tracing and alerting, access control, model evaluation, and data kept in the UAE. This is illustrative. In practice I would fit it to Wio's existing cloud platform rather than add new tools.

4 · The prototype

I built it, so we can discuss decisions rather than slides.

Clearway control tower with a ranked queue, a case file with copilot summary, and the customer's live phone view

What it shows

  • Prevent: a real IBAN checksum catching a typo, with one-tap fixes
  • Prioritise: a ranked queue, a copilot case file, and clearance that stays locked until the analyst confirms
  • Guardrail: on a likely true match, the copilot cannot message the customer
  • Inform: the customer's screen updates the moment the analyst acts
  • Impact: every assumption is a slider
  • How it is built: the stack, the AI uses and the guardrails

Sample data. A guided tour walks through it in ten steps.

4 · Teams and trade-offs

A cross-functional build, with every trade-off made on purpose

Who contributes

EngineeringShared timeline, validation services, case tools
Data scienceRanking, matching, experiments
Design and UX writingTracker and messages, English and Arabic
Behavioural scienceDefaults, waiting psychology, nudges
ComplianceMemory policy, model validation, approved wording
Operations and supportPlaybooks, outcome labels, quality checks

Trade-offs

TensionMy call
Speed vs safetyRank before auto-close. True-hit detection is a guardrail, never traded.
Friction now vs holds laterAsk more only when risk is high, and test it
Build vs buyBuy the rails (validation, partner messaging). Build the timeline and ranking, which are Wio's edge.
Transparency vs tipping offAlways share status. Share reasons only through approved wording.
AI power vs explainabilityThe copilot drafts with evidence. Rules and people decide.
4 · Success metrics and impact

Measure the payment, not the team

North star 1

Straight-through rate

Share of international payments completed with no manual touch

North star 2

Time-to-money when held

Median and slowest-10% time from submit to arrival, for payments a person touched

Drivers

  • Exceptions per 1,000 payments
  • Alerts per 1,000 and false-positive rate
  • Handling time
  • Question rounds per case
  • Share where we notified first
  • Support contacts per 1,000

Guardrails that must not move

  • True-hit detection
  • Quality-check disagreement
  • Regulatory deadlines met
  • Fraud and return rate
  • Form completion rate
8,750 hmanual hours a month today (assumed)
~4,600 hfreed each month
~29people's capacity moved to real risk
52%of manual work removed, no new headcount
5 · Execution roadmap

Discover → Design → Pilot → Scale, with a gate at each step

Weeks 0 to 3

Discover

  • Pull 90 days of held payments
  • Size the causes and build a baseline
  • Shadow analysts, interview customers
Gate: root causes confirmed or replaced
Weeks 3 to 6

Design

  • Entry checks and message templates
  • Ranking weights agreed with Compliance
  • Memory policy drafted
  • Copilot tested on past cases
Gate: Compliance signs off the pilot
Weeks 6 to 14

Pilot

  • One corridor and one segment
  • A/B test the entry checks
  • Copilot shadows, then assists
  • Proactive updates switched on
Gate: drivers improve, guardrails hold
Months 4 to 9

Scale

  • All corridors and incoming payments
  • Auto-close only the safest, proven lane
  • Arabic messaging, domestic rails
  • Hand over to the owning teams
Gate: run as business as usual

What I would do first, and why: the baseline, the entry checks and proactive updates. They touch the most payments, need no change to screening policy, carry the least regulatory risk, and give us the measurement everything else depends on.

5 · Collaboration plan

One dashboard, one rhythm, clear owners

Rhythm

  • Weekly, 30 minutes: held-payments review with Operations, Compliance, Engineering, Product and Support, on one dashboard
  • Every two weeks: experiment review covering what shipped, what moved and what we stop
  • Monthly: ranking and model review with Compliance

Ownership

  • One named owner for time-to-money, not just for queues
  • A decision log for every policy change
  • Analysts record outcomes when they close a case, so the models improve without extra work

What I would own

  • The baseline and cause sizing
  • Shadowing and customer interviews
  • The entry-checks spec and pilot success criteria
  • The test set that proves the copilot and ranking model before go-live
  • A weekly impact update to the CEO Office

Stronger controls without more bureaucracy: every guardrail lives inside the workflow. The confirmation step, the locked wording and the quality sample all happen in the case, never as an extra form.

5 · Risks and alternatives

What would change my mind, and what I considered instead

If the data says otherwise

Sanctions is not the biggest causeRe-order around whatever is
Compliance cannot approve the memoryLean on better identifiers and copilot speed
Entry checks hurt conversionApply them only to risk-flagged payments
The copilot is wrong or vagueKeep it in shadow mode and grow the test set
No UETR history exists todayStart capturing it now, and baseline from case tools meanwhile

Alternatives I considered

  • Outsource first-level review. Adds capacity but keeps cost rising with volume, and fixes no cause.
  • Only re-tune the screening system. Fewer false positives, but customers still can't see anything.
  • Full automation now. Too much regulatory risk before we have evidence.
  • Replace the case management system. Slow. A shared timeline plus a thin console delivers value in weeks.

The single biggest risk is that the real causes look different from my assumptions. That is why Discover comes first and has a gate.

5 · Why I can help deliver this

I have built each part of Clearway before, in other high-stakes systems.

Event backbone and exception handling

Shell, through Wipro

A Kafka pipeline carrying 115 GB a day from 200+ offshore stations with zero data loss. Bad records were quarantined with context for replay, which cut production incidents by 25%.

The same pattern as one timeline per payment, with held items that can be fixed and released

AI with a human in charge

GeneCart, NYU

An AI agent limited to a fixed set of least-privilege tools, with a person holding final authority over every change to research data.

The same guardrails as the investigator copilot

Ranking, retrieval and evaluation

NYU research platform

Raised answer quality 40%, graded on held-out test sets and re-tested on every change. Cut response time from 450 ms to under 100 ms at 3,000+ requests a second, with 99.9% uptime.

How I would build and prove the ranking model and copilot

Self-service that removes work

Shell vendor portal

Replaced a manual request queue with self-service for 500+ vendors. Support tickets fell 35% and service requests fell 45%.

The thinking behind the customer tracker and one-tap answers

And I built this prototype end to end within the task window. I enjoy finding the root cause with the people who run the process, then shipping something that measurably removes the work.

Thank you

Fewer investigations. Faster answers. No customer left wondering.

Which of my assumptions would you challenge first?

Jaya Sabarish Reddy RemalaAssociate, CEO Office · Home task, Option 2