The client, one of the world's top-three payments networks, was experiencing intermittent failures in the clusters that route fraud scoring on payment transactions. When a cluster degraded, the result was point-of-sale delays and declined transactions — the kind of failure that had already caused a nationwide, customer-facing disruption and strained relationships with the issuing banks it serves.
Their existing response was a set of scripts engineers ran by hand to reroute traffic once a cluster started failing — reactive, slow, and dependent on someone catching it in time. I was brought in to lead a $250K paid proof of concept to automate execution of those scripts, so failover could happen faster and with less manual intervention.
The original brief was to automate the manual failover script — make the existing response faster and less dependent on human intervention. What we realized early was that the data to predict failures existed before they materialized, which reframed the problem entirely. The goal shifted from faster reaction to prevention, and that shift changed the scale of value we could deliver.
We trained anomaly detection models on Splunk telemetry to identify the cluster health signatures that preceded failures, then rebuilt the failover workflow around the prediction. Instead of a script reacting after declines had begun, traffic now rerouted proactively on an early warning, before customers were affected. The client ran the system in shadow mode first, validating predictions against their own manual failovers, then moved to full automation once it had proven reliable.
After the proof, I identified the opportunity to replicate the architecture across six other services, and built the stakeholder relationships needed to position that expansion.
The engagement moved the client from reacting to failures to preventing them. Point-of-sale disruptions were caught before they reached customers, protecting the SLA commitments and issuer relationships that a public outage puts at risk. Over three years, six systems were built across the client's services — each extending the same predictive-rerouting model to a new part of the transaction-processing business.
The engagement scaled from roughly $1M to $2M to $3M as each new service came online. I structured the commercial model from scratch, including a three-year license on the IP we retained. The original contract carried a termination-for-convenience clause, so I priced it: any early exit required the client to pay more than 60% of remaining contract value. They never exercised it.