I’ve been in IT long enough to remember when Windows shipped on floppy disks and enterprise architecture meant racks, not regions. In that time, I’ve seen the same cycle repeat itself over and over. A new technology shows up. A pilot gets approved. The pilot goes well enough. Then the tech hits production (and production hits back), and suddenly everyone looks surprised that reality is more complicated than the demo. You might think this is an example of a failure of technology. You’d be wrong. 

The uncomfortable truth is that most projects fail not due to a technical shortcoming, but because organizations struggle to learn and adapt quickly enough. Good technology gets implemented in the wrong place or at the wrong time. Or both. Solid architectural choices run up against poor organizational habits. A common saying among I.T. elders is “projects fail at layers 8, 9, and 10 of the OSI model: Finance, compliance, and politics.  

But what holds in today’s era of “move fast, break things” is that CIOs don’t need fewer failures, regardless of where in the tech or organizational stack they occur. They need failures that produce real learning. 

Leadership is About Working With the Team 

A lot of leaders still operate under the assumption that their job is to have all the answers and to have them the moment they’re needed. I’ve always believed the opposite. In practice, effective leadership comes from working closely with the team.

That means surrounding yourself with people you trust, clearly communicating the technical and business goals, validating that your message actually landed, asking the team what it will take to succeed, and believing the answers they give you, and then removing obstacles so they can execute. 

When those fundamentals are missing, no reporting structure or KPI framework can save you. 

KPIs are the Tip of the Iceberg 

Quarterly business reviews and KPIs provide only part of the picture. They sit at the top of the iceberg. If nothing is supporting them underneath, they become theater. 

The real causes of failure usually live in nuance and detail. A business dashboard that shows only “clean-and-green” cannot tell you that a team is afraid to raise concerns because bad news is punished, or that technical debt is piling up because everyone is rewarded for heroics instead of prevention. 

Instead, consider that organizations that can move successfully from pilot to production do so by rewarding both “no surprises” and “no heroics” behavior. Raising a concern early should be seen as a win, even if it turns out to be a false alarm. Taking the time during regular working hours to write stable code that avoids preventable after-hours outages should be recognized and rewarded. 

Fail Fast vs. Design for Failure 

Many articles will write about “fail fast” and “design for failure” architectures as if they’re the same thing. Not only are they very different concepts, but they serve equally different purposes.

Fail fast belongs in testing. That is where you explore and learn without hurting customers, consumers, or users. Fail fast is part of an iterative approach: Try something knowing it might break, see where it happens, plug the hole, repeat. Fail fast teaches teams what not to do.    

Design for failure belongs in production. That’s where resilience matters. It is expensive and ongoing, but it is how you keep real systems running when components break. Designing for failure in production means assuming things will break and building accordingly. It means your team has already asked “what happens when this fails?” before the first user hits it, not after. Organizations that skip this step don’t lack good technology. They lack the discipline to imagine failure before it arrives. 

Good organizations fail fast in testing and design for failure in production, and remain clear that one does not replace the other. 

A Simple Test for Good vs. Bad Failure 

Earlier, I said that “…CIOs need failures that produce real learning.” But how do you know if you’ve built an environment that encourages the good kind of failures?

The simplest test I know is to ask, “What did we learn?” If the answer is thoughtful and specific, it was probably a good failure. If the answer is defensive or nonexistent, it was a bad one.

A good failure teaches you something. When coupled with a blameless environment, this is where real growth happens. A bad one is where failures are reversed as quickly as possible without taking any time to investigate the “why” or “how,” often due to an environment of politics and finger-pointing that shuts down learning. 

Nobody ships perfect systems. Part of our job is to find flaws before they find us. 

The Real Goal: Learning at Production Speed 

To move reliably from pilot to production, CIOs need to focus not on technologies and tools, but instead on building and supporting teams and cultures that learn quickly, safely, and honestly from failure. The companies that excel at this attract and reward employees who demonstrate curiosity, preparedness, and humility.  

Technology keeps changing, yet organizational failure modes remain consistent. Teams that break the cycle of “bad” failures do so by intentionally building cultures where learning out loud is safe, raising a concern early signals good judgment, and “what did we learn” produces a clear answer.  

That is what scales. That is what survives contact with production.