Preventing Outages During the Holiday Shopping Season
November 28, 2016

Michael Butt
BigPanda

Share this

The most destructive root cause of 75 percent of outages during big online events like Black Friday and Cyber Monday are unplanned configuration changes to a system – when IT and Ops teams find something they think might cause a problem and try to fix it immediately, unintentionally creating a much bigger issue for the web or mobile site.


The following are BigPanda's top recommendations for preventing outages during throughout the entire holiday shopping season:

- Identify the systems that are mission critical to your business. Many companies don't and try to treat their entire system as business critical – and this is a mistake. 

- Have a bulletproof plan for your critical services. Once you've identified what your critical services are, know how to keep them up with a bulletproof plan for them. For instance, if Amazon checkout goes down – you need a disaster and recovery plan for this. But if the Recommendation Engine has problems, this is not at the same level of criticality. 

- Tier your services. Having 3-5 tiers makes prioritization and response much easier, quicker and more effective when there is a problem. And make sure you have a backup and failover plan for the highest tier of your services. 

- You don't need failover for everything. IT and Ops teams who try to have failover for everything often discover that they don't have it ready for anything. 

- Don't become overly focused on the components of infrastructure. Make sure you are spending more time and focus on your services. 

- Make sure you have planned for load capacity. Not planning for the sheer volume of people visiting your web or mobile site accounts for 25 percent of outages during big online events. 

- Use a tool that allows you to consolidate your IT data. Implementing an alert correlation platform allows IT and Ops teams to separate signal from noise and focus more on the customer experience by providing a consolidated view of their IT alert data. This allows them to stop being reactive firefighters and become proactive before an issue has the chance to affect the customer.

Michael Butt is Director of Product Marketing at BigPanda.

Share this

The Latest

November 17, 2017

Just in time for the holiday shopping season, APMdigest asked experts from across the industry for their opinions on the best way to measure eCommerce performance, in terms of applications, networks and infrastructure. Part 3, the final installment, covers the customer journey ...

November 16, 2017

Just in time for the holiday shopping season, APMdigest asked experts from across the industry for their opinions on the best way to measure eCommerce performance, in terms of applications, networks and infrastructure. Part 2 covers APM and monitoring ...

November 15, 2017

As the holiday shopping season looms ahead, and online sales are positioned to challenge or even beat in-store purchases, eCommerce is on the minds of many decision makers. To help organizations decide how to gauge their eCommerce success, APMdigest compiled a list of expert opinions on the best way to measure eCommerce performance ...

November 14, 2017

More than 90 percent of respondents are concerned about data and application security in public clouds while nearly 60 percent of respondents reported that public cloud environments make it more difficult to obtain visibility into data traffic, according to a new Cloud Security survey ...

November 13, 2017

Today's technology advances have enabled end-users to operate more efficiently, and for businesses to more easily interact with customers and gather and store huge amounts of data that previously would be impossible to collect. In kind, IT departments can also collect valuable telemetry from their distributed enterprise devices to allow for many of the same benefits. But now that all this data is within reach, how can organizations make sense of it all? ...

November 09, 2017

CIOs trying to lead digital transformation at the speed needed to succeed need a mix of three scale accelerators, according to Gartner, Inc. The three scale accelerators include: digital dexterity, network effect technologies, and an industrialized digital platform ...

November 08, 2017

While the majority of IT practitioners in the UK believe their organization is equipped to support digital services, over half of them also say they face consumer-impacting incidents at least one or more times a week, sometimes costing their organizations millions in lost revenue for every hour that an application is down, according to PagerDuty's State of Digital Operations Report: United Kingdom ...

November 07, 2017

Today's IT is under considerable pressure to remain agile, responsive and scalable to meet the changing needs of business. IT infrastructure can't become a bottleneck, it must be the enabler. But as new paradigms, such as DevOps, are adopted, data center complexity increases and infrastructure constraints can block the ability to achieve these goals ...

November 06, 2017

It's 3:47am. You and the rest of the Ops team have been summoned from your peaceful slumber to mitigate an application delivery outage. Your mind races as you switch to problem solving mode. It's time to start thinking about how to make this mitigation FUN! ...

November 03, 2017

With the increased complexity of IT environments, the rising cyber threats and the growing number of IT alerts, IT organizations have come to the realization that throwing more people at IT issues doesn't solve the problem ...