Insights · 5 min read

Payment gateway integration testing: the go-live runbook

The code is the fast part. Testing is what decides whether your first week of live traffic is calm or a fire drill. The cases to run before any payment feature ships.

By GGP Editorial

The code for a payment integration is written in a week. The testing is what decides whether the first week of live traffic is calm or a fire drill. I have watched teams launch with the happy path tested and nothing else, and the bugs did not show up until real money was moving. Here is the testing plan we run before any payment feature goes live.

Set up the sandbox properly

The gateway's sandbox is not a toy version of production. It is a production-like environment with the same endpoints and webhooks, and you should treat it that way. Connect it to a staging copy of your backend, point webhooks at a URL you can actually observe, and run the same code you will deploy. The point is to find your own bugs, not the gateway's, so make the sandbox as close to production as you can.

One practical tip: keep the gateway keys and webhook secrets in config, not in code, and never share the live secret with the sandbox. A surprising number of incidents start with a live key checked into a staging environment. Separate the environments cleanly and the testing itself gets safer.

Test the four states, then the weird ones

Every payment has a few states, and your code has to handle all of them. The obvious four are success, decline, insufficient funds, and timeout. Most teams test success and decline and stop there. The timeout is the one that matters most, because it is where double charges come from: you send a request, you get no answer, and you have to decide what to do.

Beyond the four, run the cases that only bite later. A card that passes the bank but fails 3DS verification. A refund on a charge that has not settled yet. A chargeback arriving weeks after the fact. A webhook that arrives twice or out of order. Each of these is a small test now and a support ticket later if you skip it.

Test caseWhat it checks
Success and declineThe basic two paths
Timeout and retryIdempotency, no double charge
3DS challengeExtra authentication flows
Refund and partial refundMoney moving back correctly
Duplicate or late webhookIdempotent state handling
ChargebackDispute records and alerts

Strong customer authentication is worth its own pass. In Europe PSD2 and similar rules elsewhere require 3DS challenges on many card payments, and a flow that ignores this will quietly fail on a share of real cards. Test the challenge flow, the fallback when a bank does not support it, and the error a customer sees when they abandon the challenge halfway.

Test idempotency and webhooks explicitly

Idempotency is the part you cannot just eyeball. Fire the same payment request twice with the same id and confirm the second one does not create a second charge. Then resend the same webhook twice and confirm your handler does not double-count the revenue. These two tests catch the bugs that cost real money, and they are cheap to run in the sandbox.

Reconcile in the sandbox before you go live

Run a few days of fake transactions through the sandbox and reconcile them against a test ledger. The goal is to make the month-end numbers match without a human chasing a two-dollar difference. If reconciliation is manual during testing, it will be manual in production, and manual reconciliation at scale is how finance teams end up working weekends.

A lot of this testing does not need a developer. Your finance or operations person can drive the refund, chargeback, and reconciliation cases from the gateway dashboard, and that is worth doing, because they are the ones who will live with the ledger if the integration is wrong.

The go-live runbook

Go-live is a sequence, not a switch. We do a small live charge on a real card and refund it, watch the webhook round trip end to end, then enable the payment method for a small group of users before opening it to everyone. Monitor webhook failures and the reconciliation diff daily for the first month. A bug caught the same day is a quick fix; one that sits for a week compounds.

Also plan for the gateway being down. Payment providers do have outages, and your checkout should fail in a way that tells the customer what happened and lets them retry later, rather than timing out silently. A clear error state saves a support queue on the day the provider has a bad morning. If anything in the runbook fails, stop and fix it before widening access. The cost of a rollback in payments is higher than the cost of a delayed launch.

We have moved real money through trading platforms running across Hong Kong, US, and China A-share markets, plus self-ordering and POS systems that take payments in store. That background is mostly about knowing what breaks under load and across markets, which is the part you cannot learn from a docs page. We are a China-based team working across time zones with overlap hours and a dedicated group in English and Portuguese, so a test failure in your afternoon gets looked at the same day.

Talk to us about your project

Need help applying this?

Tell us what you are building and where you are today. We typically reply within 24 hours.