Guest Talk News Security

Why UK Enterprises Keep Losing to Downtime, Despite Spending More on Monitoring

Amit Shingala, CEO, Motadata

More monitoring tools have improved visibility, but UK enterprises continue to suffer costly outages because detecting and responding to issues quickly remains the real challenge.

On 19 July 2024, a single faulty software update took roughly 8.5 million Windows machines offline in what has been described as the largest IT outage in history. In the UK the effects were immediate and public. Airports moved to manual check-in, GP surgeries lost access to patient records, Sky News dropped off air, and small businesses were left taking cash because card systems and ATMs had failed. Parliament debated it within days, and the government used the moment to press ahead with a Cyber Security and Resilience Bill. 

What makes that day worth remembering is not the scale of the failure but the shape of the recovery. CrowdStrike identified the problem and issued a fix in about 78 minutes. Yet disruption lingered for days, because thousands of machines had to be repaired by hand. The answer existed almost immediately. Getting it to where it was needed did not. 

UK organisations have spent heavily on monitoring over the past decade. Dashboards, metrics, traces and logs are more detailed than they have ever been, and most large IT teams can now see almost everything happening across their estate. Despite that visibility, outages keep arriving, and the ones that reach the public tend to be costly and damaging to trust. The cost of IT downtime is rarely just the minutes a service is unavailable. It is the backlog afterwards, the goodwill lost, and the staff time spent putting things back together. 

The uncomfortable truth is that visibility has largely been solved. What has not kept pace is the speed at which a problem is noticed and dealt with. More dashboards give teams more to watch. They do not, on their own, make anyone faster at spotting the one signal that matters. 

Every outage has a timeline, and the part that decides how bad it becomes is often the quietest. It is the gap between the moment something breaks and the moment a person realises it has. The CrowdStrike event was unusual because that gap was near zero. Screens went blue everywhere at once, so nobody had to be told there was a problem. Most failures are not like that. A queue backs up, a disk fills, a dependency slows, and for a while everything looks fine on the surface while the fault quietly spreads. 

“The uncomfortable truth is that visibility has largely been solved. What has not kept pace is the speed at which a problem is noticed and dealt with. More dashboards give teams more to watch. They do not, on their own, make anyone faster at spotting the one signal that matters.”

Amit Shingala, CEO, Motadata

Detecting those quieter failures is genuinely hard at modern scale. Estates are hybrid, spread across on-premises systems and several clouds, and they generate an enormous volume of alerts. When too many of those alerts are low value, the important one gets lost in the pile, and the clock that matters, the time to detection, keeps running. By the time a human notices, a small fault has often become a headline. 

The organisations that come through outages best are not usually the ones with the most tools. They are the ones that have shortened the distance between failure and response. Three habits tend to separate them. 

They treat signal quality as a priority, not an afterthought. Fewer, better alerts beat a flood of them, because a team can only act on what it can actually read. They connect detection to response, so that a confirmed problem becomes an assigned piece of work automatically rather than sitting in a channel waiting to be spotted. And they measure detection speed directly, rather than reporting uptime alone and assuming the two are the same thing. Uptime tells you what happened. Time to detection tells you how quickly you will find out next time. 

This is where the current interest in artificial intelligence earns its place, provided it is pointed at the right problem. AI is useful in IT operations when it cuts the time to detection: sorting the meaningful signal from the ordinary, surfacing the anomaly a tired team would miss at three in the morning, and connecting a symptom to its likely cause. It is far less useful when it is added for its own sake, as another layer to configure and watch. At Motadata we work on that first problem, helping IT teams shorten the distance between something going wrong and someone knowing about it. 

For UK enterprises, the next serious outage is not a question of if. It is a question of when, and of how quickly it is caught. That will be decided not by how many screens are being watched, but by how fast the right one is read.

Related posts

CIO500 Hyderabad Edition 2026 Concludes Successfully at Park Hyatt Hyderabad, Celebrating Technology Leadership and Innovation

enterpriseitworld

DESUN Hospitals Appoints Syed Kadam Murshed as Chief Information Officer to Drive Digital Healthcare Transformation

enterpriseitworld

CIO500 Chennai 2026 Kicks Off at The Leela Palace, Celebrating Over 100 Technology Leaders

enterpriseitworld