Online Payment Platform
From Black Box to Revenue Protection Project Overview
One of the world's largest online payment platforms was losing market share to more agile competitors. It's clients processed extremely high volumes of transactions daily, but had no way to check real-time payment status. Industry standard is unforgiving: at the first sign of trouble, clients redirect traffic to competitors — and repeated incidents mean permanently lost volume, and revenue.
The Challenge
The business context
Processing several million transactions daily, across dozens of countries
Operating on thin per-transaction margins, typical for the payments industry
Experiencing a measurable, ongoing decline in market share
Customers switched to competitors at the first perceived sign of instability
A Trust Crisis
The real problem wasn't technology — it was trust.
Merchants typically use multiple payment providers simultaneously. The moment one payment platform showes any sign of trouble, they redirect traffic to competitors.
Without transparency and visibility, the platform couldn't compete, no matter how good the underlying technology and processes were.
My contribution
I turned a black-box payment platform into a transparent, trust-restoring merchant experience. As Service Designer, UX Strategist and UI Designer, I:
designed an intelligent failure-categorisation system that separated actionable issues from noise,
built a layered real-time monitoring platform, from health overview down to transaction-level detail,
introduced baseline comparison and incident-summary features that turned raw data into context,
Results achieved
Due to an NDA, I'm unable to share specific figures or client details. I left the organisation before the platform's full rollout, so the outcomes below reflect what was validated during the beta pilot, not post-launch performance.
Project overview
My role
Service Designer • UI Designer
Duration
6 months
team
1 Project owner
3 Back-end Developers
2 Front-end Developers
1 Compliance Officer
SCOPE
New archiving process,
New archiving tooling,
Governance
Research
Before designing anything, I conducted thorough diagnostics:
In-depth interviews with international merchants (e-commerce, fintech, travel)
Interviews with internal operations and engineering teams
Review of existing monitoring tools, both internal and competitors'
Analysis of transaction logs and incident reports
Stakeholder interviews with compliance, IT and management to understand constraints upstream and downstream of the visible workflow
Root cause analysis
Three Areas of Failure:
01.
The "Black Box" Problem: No Visibility
02.
Lack of Actionable Intelligence
03.
Trust Erosion Through Opacity
root cause analysis
01.
The "Black Box" Problem: No Visibility
Merchants had no way of looking into the payment flow. All they saw was:
Input: they sent a payment request
Output: success or failure
Everything in between was a black hole.

We suddenly saw more failed payments. We had no idea if it was us, them, or external factors. Within minutes we'd redirected a large share of our traffic elsewhere. Turned out later it was just a batch of expired credit cards — completely normal."
This kept happening. Merchants couldn't quickly distinguish between customer-side issues (expired cards, insufficient funds), merchant-side issues (configuration, API integration), provider-side issues (gateway failures, outages), and external issues (network, processor downtime). Making this distinction cost precious time — by then, merchants had already redirected traffic.
root cause analysis
02.
Lack of Actionable Intelligence
insight
Not all payment failures were actionable by the merchant. Merchants can't do anything about expired cards or insufficient funds — but they still spent time investigating them. This created decision fatigue: "Is this my problem, or theirs?"
Log analysis confirmed the pattern: a substantial share of failures were non-actionable (cards, currency, external blocks), a smaller share were merchant-actionable (config, integration), and a smaller share still were provider-actionable. Merchants had no way to filter these out — they had to investigate everything.
root cause analysis
03.
Trust Erosion Through Opacity
This was subtle but critical. Merchants said:

I can tolerate failures. I cannot tolerate uncertainty."
A provider who says "Yes, we had an outage, here's the root cause, here's the fix" — that merchant stayes. A provider who stays silent lost traffic immediately.
The old monitoring tools showed numbers, but no context. A failure rate on its own meant nothing to a merchant with no sense of what was normal, alarming, or catastrophic.
ServiceTransformation Strategy
Instead of isolated UI fixes, I designed a complete real-time intelligence platform built on four pillars.
01.
Real-Time Visibility
Intelligent File Handling
02.
Error Intelligence:
Actionable vs. Noise
03.
Smart Filtering
From Many Dimensions to Usability
04.
Drill-Down
From Overview to Transaction Detail
Service transformation Strategy
01.
Automating the Back-Stage: Intelligent File Handling
Core Problem:
Merchants need data on a sub-second basis, not retrospectively.
Instead of reports, I designed a living transaction stream — every transaction visible near-instantly, with volume trends and failure spikes shown in real time.
My Solution: Adaptive sampling
At low volume: show every transaction
At high volume: a statistically representative sample
Failures always shown, regardless of volume
This gave merchants both benefits at once: real-time data and performance.
Service transformation Strategy
02.
Error Intelligence: Actionable vs. Noise
This was the strategic heart of the platform.
Critical insight
Merchants didn't want to see every type of failure. They wanted to know what they could act on, and what was outside their control.
my approach
I designed an intelligent filtering system that automatically categorises failures:
Non-Actionable (Expected Noise): expired cards, insufficient funds, customer cancels, card blocked by bank
Merchant-Actionable: configuration errors, API integration issues, merchant-side timeouts, invalid request formats
Provider-Actionable: payment gateway failures, internal system errors, network infrastructure issues, processing bottlenecks
Challenge
How to show this without overwhelming the user?
My first prototypes showed three separate failure graphs. Users said: "I get three graphs at once. I don't know where to look."
My iteration: Allow for an overview of all failures, but also provide tabs dedicated to only one type of failure.
Service transformation Strategy
03.
Smart Filtering: From Many Dimensions to Usability
Merchants could filter by country/region, payment method, error type, time window, and individual transaction attributes — dozens of filterable dimensions in total.
Challenge
content here
04.
Building Prevention Into Every Touchpoint: Real-Time Validation
This was the heart of shifting the service from error-detection to error-prevention.
Every field validates as you type:
My Governance Model:
1. Content ownership
Peer review
3. Periodic audits
4. Feedback loop from Support
Content OWNERSHIP
•. Every article / topic gets a clear owner.
• The owner is responsible for accuracy and quality.
• 15% of their time allocated to content management.
PEER REVIEW
•. New or updated articles must be approved by a peer.
• Review checklist: “matches the template?”, “written in the customer’s language?”, “no duplicates?”, "up to date?"
• Right category/sub-category?
PERiodic Audits
• Each article got an expiry date
• Expiry date flagged outdated articles
• When audit was due the content owner received an automated notification.
• If audit was not performed within a certain period the article would be invisible until audit.
Feedback Loop from Support
•. A weekly meeting was set up to discuss which topics generated the most questions.
• This feeds directly back into content priorities.
• When audit was due the content owner received an automated notification.
• If audit was not performed within a certain period the article would be invisible until audit.
ServiceTransformation Strategy
Instead of isolated UI fixes, I designed a complete real-time intelligence platform built on four pillars.
01.
Real-Time Visibility
Intelligent File Handling
02.
Error Intelligence:
Actionable vs. Noise
03.
Smart Filtering
From Many Dimensions to Usability
04.
Drill-Down
From Overview to Transaction Detail
Service transformation Strategy
01.
Automating the Back-Stage: Intelligent File Handling
Core Problem:
Merchants need data on a sub-second basis, not retrospectively.
Instead of reports, I designed a living transaction stream — every transaction visible near-instantly, with volume trends and failure spikes shown in real time.
My Solution: Adaptive sampling
At low volume: show every transaction
At high volume: a statistically representative sample
Failures always shown, regardless of volume
This gave merchants both benefits at once: real-time data and performance.
Service transformation Strategy
02.
Error Intelligence: Actionable vs. Noise
This was the strategic heart of the platform.
Critical insight from research: merchants didn't want to see every type of failure. They wanted to know what they could act on, and what was outside their control.
My Approach: Intelligent Categorisation
I designed an intelligent filtering system that automatically categorises failures:
Non-Actionable (Expected Noise): expired cards, insufficient funds, customer cancels, card blocked by bank
Merchant-Actionable: configuration errors, API integration issues, merchant-side timeouts, invalid request formats
Provider-Actionable: payment gateway failures, internal system errors, network infrastructure issues, processing bottlenecks
Challenge: How to show this without overwhelming the user?
My first prototypes showed three separate failure graphs. Users said: "I get three graphs at once. I don't know where to look."
My iteration: Allow for an overview of all failures, but also provide tabs dedicated to only one type of failure.
Service transformation Strategy
03.
Smart Filtering: From Many Dimensions to Usability
Merchants could filter by country/region, payment method, error type, time window, and individual transaction attributes — dozens of filterable dimensions in total.
The Reality of 21 Metadata fields:
● 5 legally mandatory
● 7 auto-derivable from document content or the database
● 9 "nice to have" but rarely used
My Design Solution:
● Mandatory Fields: always visible
● Auto-Filled Fields: shown but greyed out, so staff can see what the service is doing on their behalf and if needed they were editable— trust!
● Optional Fields: collapsed under "Additional fields"
04.
Building Prevention Into Every Touchpoint: Real-Time Validation
This was the heart of shifting the service from error-detection to error-prevention.
Every field validates as you type:
My Governance Model:
1. Content ownership
Peer review
3. Periodic audits
4. Feedback loop from Support
Content OWNERSHIP
•. Every article / topic gets a clear owner.
• The owner is responsible for accuracy and quality.
• 15% of their time allocated to content management.
PEER REVIEW
•. New or updated articles must be approved by a peer.
• Review checklist: “matches the template?”, “written in the customer’s language?”, “no duplicates?”, "up to date?"
• Right category/sub-category?
PERiodic Audits
• Each article got an expiry date
• Expiry date flagged outdated articles
• When audit was due the content owner received an automated notification.
• If audit was not performed within a certain period the article would be invisible until audit.
Feedback Loop from Support
•. A weekly meeting was set up to discuss which topics generated the most questions.
• This feeds directly back into content priorities.
• When audit was due the content owner received an automated notification.
• If audit was not performed within a certain period the article would be invisible until audit.
Service Design Process: a Deep Dive
Exploring Alternative Service Models
I explored three fundamentally different ways of structuring the service:
Concept 1: A Guided wizard
A strict step-by-step service flow. Little flexibility.

This feels like the system doesn't trust me. Why can't I just fill in the fields I need?"
Lesson learned: Experts want control over how they move through a service. Don't force a rigid journey on people who already know the terrain.
Concept 2: All fields visible
All 47 fields exposed on a single screen — a service model with no scaffolding.
Test Result: New employees were overwhelmed — didn't know where to start, couldn't distinguish mandatory from optional fields.
Lesson learned: A service still needs to guide newcomers, even while giving experts freedom.
Concept 3: Adaptive Service Model
The service adapts to who is using it:
Beginners: guidance, helpful tooltips, mandatory-field indicators
Experts: all fields available, minimal guidance
Test Result: Promising — validated the core service model, though upload feedback and validation timing still needed refinement.
Testing & Iteration
Three Rounds of Service Validation
Round 1: Paper Prototype
(4 staff members)
Validated the core service journey
Revealed the wizard model was too restrictive
Identified the need for progress tracking within the journey
Round 2: Interactive Prototype
(8 staff members)
Tested actual interaction patterns at each touchpoint
Found the drag-and-drop entry point too subtle
Identified validation-timing issues
Round 3: beta trial
(30 days, 6 staff members)
Measured real processing time across the full service
Identified edge cases and error patterns
Collected satisfaction data
Critical Learnings
learning
Implementation & Launch
Engineering Collaboration
I delivered to the engineering team the artefacts needed to build the service, not just the screens:
Detailed interaction specs with state diagrams
A component library (exact spacing, colours, behaviour)
Edge-case documentation (40+ scenarios) grounded in the service blueprint
Technical Constraint & Solution:
Real-time duplicate detection across 3.2M records was too slow for the back-stage infrastructure. I adapted the front-stage experience from instant feedback to a "checking..." state with a 2-second debounce — preserving the feel of the service while respecting the system's real constraints.
Phased Service Rollout
Instead of a big-bang launch, we rolled out the redesigned service in stages:
Phase 1 (weeks 1-2): 3 advanced users, new documents only
Phase 2 (weeks 3-6): 12 users, a mix of old and new documents
Phase 3 (week7+): Full teams, all document types
Benefits:
Problems caught early, before they propagated across the whole service
Built internal champions who could train others — embedding the new service model socially, not just technically
Rapid iteration based on real-world use
Round 3: beta trial
(30 days, 6 staff members)
Measured real processing time across the full service
Identified edge cases and error patterns
Collected satisfaction data
Critical Learnings
learning
Trust in a Service Requires Progressive Validation.
Early beta users double-checked automatic conversions — converting files manually, then uploading them, defeating the purpose of the automation entirely.
Root Cause:
Versions of the service had made errors, and that history didn't disappear just because the system changed.
My Solution:
A "View Processing Details" feature — underneath every conversion, a clickable link to the processing log, showing exactly what the service had done on the user's behalf.
After staff had verified accuracy 2-3 times, they stopped double-checking. Trust in a service is earned incrementally, not declared.
The tool
