Two decades of promised modernization have delivered surprisingly little. The problem was never a shortage of will. It is that most banks are still asking whether to replace the mainframe, rather than which parts of it actually need to change.

Financial institutions worldwide continue to spend heavily maintaining core banking systems built on COBOL and mainframe architectures designed decades ago. IDC Financial Insights reported that the industry spent $36.7 billion on legacy payment systems in 2022, with spending projected to reach $57 billion by 2028. This persistence is often blamed on institutional inertia, as though banks would modernize if they were simply more willing to do so. That explanation is incomplete.

Banks have not struggled to modernize mainframes because they lack funding, skilled staff, or motivation. The deeper problem is the mainframe paradox: modernization can create as many risks as it resolves. Mainframes remain among the most reliable platforms for processing financial transactions, but they are difficult to change. The design decisions that have supported decades of stable, high-volume processing also make changes costly, complex, and risky.

Lessons From a Costly Migration

Most major core-system migration failures over the past decade, including TSB’s 2018 outage, began with a similar assumption: banks could replace a rigid legacy platform with a more flexible one without giving up reliability. In practice, that trade-off has proved difficult to achieve. The TSB failure ultimately cost more than £330 million, including customer compensation, remediation, and regulatory fines. The wider cost is harder to measure because many unsuccessful migrations receive little public attention. When they are disclosed, they often appear only in brief regulatory filings or short passages in annual reports.

The Wrong Question

The question many boards ask, “Should we replace the mainframe?” is the wrong starting point. It treats the platform as one decision and the institution’s exposure as one uniform risk. Neither is true. A large mainframe environment usually contains dozens of subsystems, each with its own purpose and risk profile. Some support essential operations and cannot be changed without careful planning. Others remain in place because removing them seems riskier than keeping them, even though few people still depend on them. The real issue is not whether to replace the mainframe. It is knowing which parts to protect, which to modernize, and which to retire. Treating the entire platform as a single decision is what creates the conditions for failures like TSB’s.

A Better Way to Triage

The more useful question is not whether to replace the mainframe. It is which legacy subsystems are too risky to leave unchanged and which can be governed through modern controls without being replaced. Answering that requires a model that regulators, risk officers, and engineers can understand and apply. In practice, this assessment usually places subsystems into three broad groups.

The first includes stable, highly critical components that change infrequently. Core ledgers often fall into this category. These systems are usually better served by added controls than by replacement. Documented APIs, stronger observability, and automated compliance checks can make the platform visible to modern governance without altering its underlying code.

The second group includes systems with a sound legacy core but a troubled integration layer. The API gateway, message bus, or batch scheduler may carry significant technical debt even when the core system remains reliable. In these cases, modernizing the connection points makes more sense than replacing the entire platform. Much of the operational risk sits at the seams.

The third group is smaller. It includes legacy components that are genuinely unsafe to leave in place, often because they contain business logic that few people still understand and regulators have begun to question. Full replacement may be justified for these systems. In most institutions, though, they represent only a small portion of the overall footprint.

A propagation-aware model can support this triage by scoring each subsystem according to its business criticality, downstream dependencies, and regulatory exposure. The score does not make the decision for leadership. It shows which decisions require attention and which do not need to be made at all.

Refusing to Triage Is the Real Cost

The money the industry spends each year maintaining mainframes is not simply the price of failing to modernize. It is often the price of failing to prioritize. Many institutions treat the entire platform as one large risk because they lack a clear way to assess each part separately. Banks that have made meaningful modernization progress in recent years did not replace everything at once. They identified which systems to wrap with new interfaces, which to modernize at specific points, and which to leave unchanged.

Legacy is not the real enemy here. Unmanaged legacy is.