Online Payment Platform

From Black Box to Revenue Protection Project Overview

One of the world's largest online payment platforms was losing market share to more agile competitors. It's clients processed extremely high volumes of transactions daily, but had no way to check real-time payment status. Industry standard is unforgiving: at the first sign of trouble, clients redirect traffic to competitors — and repeated incidents mean permanently lost volume, and revenue.

The Challenge

The business context

  • Processing several million transactions daily, across dozens of countries

  • Operating on thin per-transaction margins, typical for the payments industry

  • Experiencing a measurable, ongoing decline in market share

  • Customers switched to competitors at the first perceived sign of instability

A Trust Crisis

The real problem wasn't technology — it was trust.

Merchants typically use multiple payment providers simultaneously. The moment one payment platform showes any sign of trouble, they redirect traffic to competitors.

Without transparency and visibility, the platform couldn't compete, no matter how good the underlying technology and processes were.

My contribution

I turned a black-box payment platform into a transparent, trust-restoring merchant experience. As Service Designer, UX Strategist and UI Designer, I:

  • designed an intelligent failure-categorisation system that separated actionable issues from noise,

  • built a layered real-time monitoring platform, from health overview down to transaction-level detail,

  • introduced baseline comparison and incident-summary features that turned raw data into context,

Results achieved

Due to an NDA, I'm unable to share specific figures or client details. I left the organisation before the platform's full rollout, so the outcomes below reflect what was validated during the beta pilot, not post-launch performance.

Project overview

My role

Service Designer • UI Designer

Duration

6 months

team

1 Project owner

3 Back-end Developers

2 Front-end Developers

1 Compliance Officer

SCOPE

  • New archiving process,

  • New archiving tooling,

  • Governance

Research

Before designing anything, I conducted thorough diagnostics:

  • In-depth interviews with international merchants (e-commerce, fintech, travel)

  • Interviews with internal operations and engineering teams

  • Review of existing monitoring tools, both internal and competitors'

  • Analysis of transaction logs and incident reports

  • Stakeholder interviews with compliance, IT and management to understand constraints upstream and downstream of the visible workflow

Root cause analysis

Three Areas of Failure:

01.

The "Black Box" Problem: No Visibility

02.

Lack of Actionable Intelligence

03.

Trust Erosion Through Opacity

03.

Trust Erosion Through Opacity

root cause analysis

01.

The "Black Box" Problem: No Visibility

Merchants had no way of looking into the payment flow. All they saw was:

  • Input: they sent a payment request

  • Output: success or failure

Everything in between was a black hole.

We suddenly saw more failed payments. We had no idea if it was us, them, or external factors. Within minutes we'd redirected a large share of our traffic elsewhere. Turned out later it was just a batch of expired credit cards — completely normal."

This kept happening. Merchants couldn't quickly distinguish between customer-side issues (expired cards, insufficient funds), merchant-side issues (configuration, API integration), provider-side issues (gateway failures, outages), and external issues (network, processor downtime). Making this distinction cost precious time — by then, merchants had already redirected traffic.

root cause analysis

02.

Lack of Actionable Intelligence

insight

Not all payment failures were actionable by the merchant. Merchants can't do anything about expired cards or insufficient funds — but they still spent time investigating them. This created decision fatigue: "Is this my problem, or theirs?"

Log analysis confirmed the pattern: a substantial share of failures were non-actionable (cards, currency, external blocks), a smaller share were merchant-actionable (config, integration), and a smaller share still were provider-actionable. Merchants had no way to filter these out — they had to investigate everything.

root cause analysis

03.

Trust Erosion Through Opacity

This was subtle but critical. Merchants said:

I can tolerate failures. I cannot tolerate uncertainty."

A provider who says "Yes, we had an outage, here's the root cause, here's the fix" — that merchant stayes. A provider who stays silent lost traffic immediately.

The old monitoring tools showed numbers, but no context. A failure rate on its own meant nothing to a merchant with no sense of what was normal, alarming, or catastrophic.

ServiceTransformation Strategy

Instead of isolated UI fixes, I designed a complete real-time intelligence platform built on four pillars.

01.

Real-Time Visibility

Intelligent File Handling

02.

Error Intelligence:

Actionable vs. Noise

03.

Smart Filtering

From Many Dimensions to Usability

04.

Drill-Down

From Overview to Transaction Detail

Service transformation Strategy

01.

Automating the Back-Stage: Intelligent File Handling

Core Problem:

Merchants need data on a sub-second basis, not retrospectively.

Instead of reports, I designed a living transaction stream — every transaction visible near-instantly, with volume trends and failure spikes shown in real time.

My Solution: Adaptive sampling

  • At low volume: show every transaction

  • At high volume: a statistically representative sample

  • Failures always shown, regardless of volume

This gave merchants both benefits at once: real-time data and performance.

Service transformation Strategy

02.

Error Intelligence: Actionable vs. Noise

This was the strategic heart of the platform.

Critical insight

Merchants didn't want to see every type of failure. They wanted to know what they could act on, and what was outside their control.

my approach

I designed an intelligent filtering system that automatically categorises failures:

  • Non-Actionable (Expected Noise): expired cards, insufficient funds, customer cancels, card blocked by bank

  • Merchant-Actionable: configuration errors, API integration issues, merchant-side timeouts, invalid request formats

  • Provider-Actionable: payment gateway failures, internal system errors, network infrastructure issues, processing bottlenecks

Challenge

How to show this without overwhelming the user?

My first prototypes showed three separate failure graphs. Users said: "I get three graphs at once. I don't know where to look."

My iteration: Allow for an overview of all failures, but also provide tabs dedicated to only one type of failure.

Service transformation Strategy

03.

Smart Filtering: From Many Dimensions to Usability

Merchants could filter by country/region, payment method, error type, time window, and individual transaction attributes — dozens of filterable dimensions in total.

Challenge


content here

04.

Building Prevention Into Every Touchpoint: Real-Time Validation

This was the heart of shifting the service from error-detection to error-prevention.

Every field validates as you type:

My Governance Model:

1. Content ownership

  1. Peer review

3. Periodic audits

4. Feedback loop from Support

  1. Content OWNERSHIP

•. Every article / topic gets a clear owner.

•  The owner is responsible for accuracy and quality.

•  15% of their time allocated to content management.

  1. PEER REVIEW

•. New or updated articles must be approved by a peer.

•  Review checklist: “matches the template?”, “written in the customer’s language?”, “no duplicates?”, "up to date?"

•  Right category/sub-category?

  1. PERiodic Audits

• Each article got an expiry date

•  Expiry date flagged outdated articles

•  When audit was due the content owner received an automated notification.

•  If audit was not performed within a certain period the article would be invisible until audit.

  1. Feedback Loop from Support

•. A weekly meeting was set up to discuss which topics generated the most questions.

•  This feeds directly back into content priorities.

•  When audit was due the content owner received an automated notification.

•  If audit was not performed within a certain period the article would be invisible until audit.

ServiceTransformation Strategy

Instead of isolated UI fixes, I designed a complete real-time intelligence platform built on four pillars.

01.

Real-Time Visibility

Intelligent File Handling

02.

Error Intelligence:

Actionable vs. Noise

03.

Smart Filtering

From Many Dimensions to Usability

04.

Drill-Down

From Overview to Transaction Detail

Service transformation Strategy

01.

Automating the Back-Stage: Intelligent File Handling

Core Problem:

Merchants need data on a sub-second basis, not retrospectively.

Instead of reports, I designed a living transaction stream — every transaction visible near-instantly, with volume trends and failure spikes shown in real time.

My Solution: Adaptive sampling

  • At low volume: show every transaction

  • At high volume: a statistically representative sample

  • Failures always shown, regardless of volume

This gave merchants both benefits at once: real-time data and performance.

Service transformation Strategy

02.

Error Intelligence: Actionable vs. Noise

This was the strategic heart of the platform.

Critical insight from research: merchants didn't want to see every type of failure. They wanted to know what they could act on, and what was outside their control.

My Approach: Intelligent Categorisation
I designed an intelligent filtering system that automatically categorises failures:

  • Non-Actionable (Expected Noise): expired cards, insufficient funds, customer cancels, card blocked by bank

  • Merchant-Actionable: configuration errors, API integration issues, merchant-side timeouts, invalid request formats

  • Provider-Actionable: payment gateway failures, internal system errors, network infrastructure issues, processing bottlenecks

Challenge: How to show this without overwhelming the user?

My first prototypes showed three separate failure graphs. Users said: "I get three graphs at once. I don't know where to look."

My iteration: Allow for an overview of all failures, but also provide tabs dedicated to only one type of failure.

Service transformation Strategy

03.

Smart Filtering: From Many Dimensions to Usability

Merchants could filter by country/region, payment method, error type, time window, and individual transaction attributes — dozens of filterable dimensions in total.

The Reality of 21 Metadata fields:

● 5 legally mandatory
● 7 auto-derivable from document content or the database
●  9 "nice to have" but rarely used

My Design Solution:

●  Mandatory Fields: always visible
●  Auto-Filled Fields: shown but greyed out, so staff can see what the service is doing on their behalf and if needed they were editable— trust!
●  Optional Fields: collapsed under "Additional fields"

04.

Building Prevention Into Every Touchpoint: Real-Time Validation

This was the heart of shifting the service from error-detection to error-prevention.

Every field validates as you type:

My Governance Model:

1. Content ownership

  1. Peer review

3. Periodic audits

4. Feedback loop from Support

  1. Content OWNERSHIP

•. Every article / topic gets a clear owner.

•  The owner is responsible for accuracy and quality.

•  15% of their time allocated to content management.

  1. PEER REVIEW

•. New or updated articles must be approved by a peer.

•  Review checklist: “matches the template?”, “written in the customer’s language?”, “no duplicates?”, "up to date?"

•  Right category/sub-category?

  1. PERiodic Audits

• Each article got an expiry date

•  Expiry date flagged outdated articles

•  When audit was due the content owner received an automated notification.

•  If audit was not performed within a certain period the article would be invisible until audit.

  1. Feedback Loop from Support

•. A weekly meeting was set up to discuss which topics generated the most questions.

•  This feeds directly back into content priorities.

•  When audit was due the content owner received an automated notification.

•  If audit was not performed within a certain period the article would be invisible until audit.

Service Design Process: a Deep Dive

Exploring Alternative Service Models

I explored three fundamentally different ways of structuring the service:

Concept 1: A Guided wizard

A strict step-by-step service flow. Little flexibility.

This feels like the system doesn't trust me. Why can't I just fill in the fields I need?"

Lesson learned: Experts want control over how they move through a service. Don't force a rigid journey on people who already know the terrain.

Concept 2: All fields visible

All 47 fields exposed on a single screen — a service model with no scaffolding.

Test Result: New employees were overwhelmed — didn't know where to start, couldn't distinguish mandatory from optional fields.

Lesson learned: A service still needs to guide newcomers, even while giving experts freedom.

Concept 3: Adaptive Service Model

The service adapts to who is using it:

  • Beginners: guidance, helpful tooltips, mandatory-field indicators

  • Experts: all fields available, minimal guidance

Test Result: Promising — validated the core service model, though upload feedback and validation timing still needed refinement.

Testing & Iteration

Three Rounds of Service Validation

Round 1: Paper Prototype

(4 staff members)

  • Validated the core service journey

  • Revealed the wizard model was too restrictive

  • Identified the need for progress tracking within the journey

Round 2: Interactive Prototype

(8 staff members)

  • Tested actual interaction patterns at each touchpoint

  • Found the drag-and-drop entry point too subtle

  • Identified validation-timing issues

Round 3: beta trial

(30 days, 6 staff members)

  • Measured real processing time across the full service

  • Identified edge cases and error patterns

  • Collected satisfaction data

Critical Learnings

learning

Trust in a Service Requires Progressive Validation.

Trust in a Service Requires Progressive Validation.

Early beta users double-checked automatic conversions — converting files manually, then uploading them, defeating the purpose of the automation entirely.

Early beta users double-checked automatic conversions — converting files manually, then uploading them, defeating the purpose of the automation entirely.

Root Cause:
Versions of the service had made errors, and that history didn't disappear just because the system changed.

Root Cause:
Versions of the service had made errors, and that history didn't disappear just because the system changed.

My Solution:
A "View Processing Details" feature — underneath every conversion, a clickable link to the processing log, showing exactly what the service had done on the user's behalf.

My Solution:
A "View Processing Details" feature — underneath every conversion, a clickable link to the processing log, showing exactly what the service had done on the user's behalf.

After staff had verified accuracy 2-3 times, they stopped double-checking. Trust in a service is earned incrementally, not declared.

After staff had verified accuracy 2-3 times, they stopped double-checking. Trust in a service is earned incrementally, not declared.

Implementation & Launch

Engineering Collaboration

I delivered to the engineering team the artefacts needed to build the service, not just the screens:

  • Detailed interaction specs with state diagrams

  • A component library (exact spacing, colours, behaviour)

  • Edge-case documentation (40+ scenarios) grounded in the service blueprint

Technical Constraint & Solution:
Real-time duplicate detection across 3.2M records was too slow for the back-stage infrastructure. I adapted the front-stage experience from instant feedback to a "checking..." state with a 2-second debounce — preserving the feel of the service while respecting the system's real constraints.

Phased Service Rollout

Instead of a big-bang launch, we rolled out the redesigned service in stages:

  • Phase 1 (weeks 1-2): 3 advanced users, new documents only

  • Phase 2 (weeks 3-6): 12 users, a mix of old and new documents

  • Phase 3 (week7+): Full teams, all document types

Benefits:

  • Problems caught early, before they propagated across the whole service

  • Built internal champions who could train others — embedding the new service model socially, not just technically

  • Rapid iteration based on real-world use

Round 3: beta trial

(30 days, 6 staff members)

  • Measured real processing time across the full service

  • Identified edge cases and error patterns

  • Collected satisfaction data

Critical Learnings

learning

Trust in a Service Requires Progressive Validation.

Early beta users double-checked automatic conversions — converting files manually, then uploading them, defeating the purpose of the automation entirely.

Root Cause:
Versions of the service had made errors, and that history didn't disappear just because the system changed.

My Solution:
A "View Processing Details" feature — underneath every conversion, a clickable link to the processing log, showing exactly what the service had done on the user's behalf.

After staff had verified accuracy 2-3 times, they stopped double-checking. Trust in a service is earned incrementally, not declared.

The tool