Fintech & Payment Infrastructure
Designing a funds hold that stops double-spend without blocking the customer
Article
A customer sends money out of a wallet. The platform debits their balance, calls the external provider that actually moves the funds, and the call fails. So the platform does the obvious thing: it fires a compensating reversal to put the balance back. That reversal also fails — a connection reset, a queue message nobody acknowledged, a scheduler quietly disabled during an unrelated incident and never re-enabled. Now the customer is short, the platform's records say the money left, the provider's records say it never arrived, and no automated path exists to resolve the difference. Operations finds out when the customer calls.
Nothing in that sequence was anyone's mistake. Every component did what it was written to do. The design simply had one recovery mechanism, and that mechanism was itself a network request — subject to exactly the same failure modes as the request it existed to compensate for. That is the real defect, and it is architectural rather than local.
Why "deduct now, undo later" runs out of road
A compensating reversal is a second write that has to succeed in order for the first write's failure to be contained. It can be dropped. It can time out. It can race a second attempt at the same reversal. A model whose only answer to "the external call failed" is "send a second call to undo it" has no answer left for "what if the second call also fails."
Reversal paths also rot faster than the paths they protect, precisely because they almost never run. A failed-reversal queue with no escalation and nobody reading it. A retry scheduler commented out during some long-forgotten release and never restored. A handler that catches the exception, logs it at a level nobody alerts on, and returns success to its caller — so the failure is invisible rather than noisy. This is the normal condition of reversal code, not an unusual lapse. It is the branch nobody exercises, nobody writes realistic test fixtures for, and nobody watches, right up to the moment it is the only thing standing between a customer and a wrong balance.
Hardening the reversal is the intuitive response and the wrong one. It narrows the window without closing it. The window exists because the platform committed to an outcome — a final debit — before it knew what the outcome was.
Reserve first, decide the outcome later
The alternative is to make the reservation the first-class concept. When a transaction is accepted, the platform writes a durable record saying this amount is committed to this in-flight transaction, and it computes a customer's spendable balance as the ledger balance minus the sum of their active reservations. The ledger itself is not touched until the external outcome is known. On confirmed success the reservation converts into a final debit. On confirmed failure it is released, and the customer never saw a deduction to undo.
Two properties make this work, and both are easy to get wrong.
The reservation is an accounting record, not a lock. The tempting implementation is to open a database transaction, lock the account row, call the provider, and commit or roll back on the response. That turns every slow provider into a contention event: locks held for seconds, connection pools drained, throughput collapsing under exactly the load that made the slowness appear. Every database transaction in this design should be short enough to measure in milliseconds, and the external call belongs strictly outside all of them. The reservation is durable state that survives between those short transactions — which is the whole point, because the gap between "we decided to pay" and "we know whether we paid" is where the design has to hold.
Placing a reservation has to be atomic against itself. Reading the balance, summing active reservations, and inserting a new one must happen under one concurrency control — an account-row lock or a version check — or two concurrent requests both see sufficient funds and both reserve. The double-spend this mechanism exists to prevent then reappears inside the mechanism. The guard also wants a uniqueness constraint keyed on the business transaction identifier, so a retried request finds the existing reservation instead of creating a second one.
The payoff is that a customer with an in-flight transaction is not blocked. Their remaining balance is still spendable; only the reserved portion is not. Blocking the account, or refusing new transactions while one is pending, is the crude version of the same safety property and costs real usability for no additional correctness.
A timeout is not a failure, and not every dependency needs this
The state the pattern lives or dies on is the ambiguous one. A provider that does not respond has not told you it failed — it may have processed the transaction perfectly and lost only the response. Auto-releasing a reservation on timeout re-opens the double-spend window in its worst form: the customer's balance is restored, they spend it, and the provider's success confirmation arrives afterwards.
So an unresolved outcome needs to be its own state, distinct from confirmed failure, with its own resolution path: query the provider using the original reference, accept a late confirmation, de-duplicate a confirmation that arrives twice, and when none of that resolves it within the provider's own finality window, escalate to a human queue rather than guessing. Expiry should mean "a person must look at this", never "assume failure."
Structurally none of this is new. It is the authorization-and-capture pattern that card networks have run for decades, applied to a wallet ledger. What differs regionally is how often it has to be retrofitted. In Pakistan and in comparable markets, a large share of wallet, branchless-banking and microfinance platforms grew incrementally from a single mutable balance field, where every money movement is an immediate update and reversal was the only recovery primitive the original design offered. The reservation concept was never designed in, so it arrives later as a change to the most sensitive table in the system — which is a materially harder engineering problem than having started with it, and the reason the pattern reads as obvious in a design review and expensive in a codebase.
It is also not free, and should not be applied uniformly. A fast, reliably idempotent dependency that returns an unambiguous result can stay on immediate debit. A slow one, a callback-driven one, or one with no status enquiry endpoint cannot. Classify each dependency on measured response behaviour and the idempotency it actually honours, not on what its integration guide claims.
Where to start
Inventory the outbound flows that still debit before the external outcome is known, and look at what happens when their reversal path fails — not whether it exists, but whether anyone would find out. Classify each external dependency by response behaviour and idempotency. Then roll the reservation pattern out behind a per-dependency switch that defaults to the existing path, one dependency at a time, so rollback is a configuration change rather than a deployment. The mechanism is not complicated. The expensive part is retrofitting it under the pressure of an incident that has already reached a customer, rather than before.