Preventing Outages in 2023: What We Can Learn from Recent Failures
February 09, 2023
Share this

"What the recent failures from Internet giants demonstrate is that the question of the next outage is not if, but when," says Dritan Suljoti, Chief Product and Technology Officer of Catchpoint, referencing the company's new white paper, Preventing Outages in 2023: What We Learned from Recent Failures. "Moreover, the downstream effect of major outages to essential Internet infrastructure, such as cloud platforms, CDNs or DNS providers, means that no company is immune, no matter how well prepared they think they are. The white paper demonstrates why it's so important for all of us to be proactive to reduce Mean Time to Repair (MTTR) when the next outage occurs."



Key lessons from the past

■ Develop an Internet Performance Monitoring strategy that allows you to monitor precisely what customers, workforce, and other users expect and build an Experience Score.

■ Monitor not only what is under your direct control, map your Internet stack to ensure you are monitoring every component of the Internet Stack relied on to deliver your content (including DNS, CDN, ISP, BGP, TCP configuration, SSL, and other cloud services, etc.).

■ Automate intelligently – design and test automation to ensure there are no bugs hiding in the code.

■ Be prepared to take fast action to remediate outages as they occur, for example, switching to a backup solution or dropping the third-party causing the issue. Develop runbooks and practice recovery.

■ Whenever change is scheduled, ensure your team is ready for any outages that may occur (intentionally or not) with a crisis call plan that includes a communication plan and templates, a plan to mitigate failures from third-parties, and a best practices monitoring and observability plan.

"Given the impact of serious outages to the bottom line, not to mention the long-tail impact to brand and reputation, amidst a landscape of increased Internet reliance alongside ever-growing Internet fragility and greater and great complexity, the need for community learnings from past failures to be shared and practical advice disseminated around stemming future major incidents and ensuring Internet Resilience is imperative," says Gerardo Dada, CMO at Catchpoint. "We believe this white paper offers an invaluable deep dive into recent outages past and key lessons learned that all of us can learn from to prevent (or mitigate the consequences of) the next major outage."

Share this

The Latest

December 18, 2024

Industry experts offer predictions on how NetOps, Network Performance Management, Network Observability and related technologies will evolve and impact business in 2025 ...

December 17, 2024

In APMdigest's 2025 Predictions Series, industry experts offer predictions on how Observability and related technologies will evolve and impact business in 2025. Part 6 covers cloud, the edge and IT outages ...

December 16, 2024

In APMdigest's 2025 Predictions Series, industry experts offer predictions on how Observability and related technologies will evolve and impact business in 2025. Part 5 covers user experience, Digital Experience Management (DEM) and the hybrid workforce ...

December 12, 2024

In APMdigest's 2025 Predictions Series, industry experts offer predictions on how Observability and related technologies will evolve and impact business in 2025. Part 4 covers logs and Observability data ...

December 11, 2024

In APMdigest's 2025 Predictions Series, industry experts offer predictions on how Observability and related technologies will evolve and impact business in 2025. Part 3 covers OpenTelemetry, DevOps and more ...

December 10, 2024

In APMdigest's 2025 Predictions Series, industry experts offer predictions on how Observability and related technologies will evolve and impact business in 2025. Part 2 covers AI's impact on Observability, including AI Observability, AI-Powered Observability and AIOps ...

December 09, 2024

The Holiday Season means it is time for APMdigest's annual list of predictions, covering IT performance topics. Industry experts — from analysts and consultants to the top vendors — offer thoughtful, insightful, and often controversial predictions on how Observability, APM, AIOps and related technologies will evolve and impact business in 2025 ...

December 05, 2024
Generative AI represents more than just a technological advancement; it's a transformative shift in how businesses operate. Companies are beginning to tap into its ability to enhance processes, innovate products and improve customer experiences. According to a new IDC InfoBrief sponsored by Endava, 60% of CEOs globally highlight deploying AI, including generative AI, as their top modernization priority to support digital business ambitions over the next two years ...
December 04, 2024

Technology leaders will invest in AI-driven customer experience (CX) strategies in the year ahead as they build more dynamic, relevant and meaningful connections with their target audiences ... As AI shifts the CX paradigm from reactive to proactive, tech leaders and their teams will embrace these five AI-driven strategies that will improve customer support and cybersecurity while providing smoother, more reliable service offerings ...

December 03, 2024

We're at a critical inflection point in the data landscape. In our recent survey of executive leaders in the data space — The State of Data Observability in 2024 — we found that while 92% of organizations now consider data reliability core to their strategy, most still struggle with fundamental visibility challenges ...