Instapaper, a "read later" tool for saving web pages to read on other devices or offline, suffered an extensive outage 2 weeks ago. The site was unavailable for a day and a half, and even after restoring service, the company had to explain that its archives would be impacted for another full week. Ultimately, it was able to restore the archives sooner, but the outage garnered extensive press and social media coverage.
The cause of the outage was that an indexing file Instapaper relies on for reaching all stored links exceeded the max file size supported on the older instance of Amazon Web Services the site was first built on. You can read if you want more details .
While Instapaper hit a unique problem — a file size limitation — its experience speaks to a much larger problem: scaling a database is difficult, and never quick. That basic fact explains why outages like the one Instapaper suffered are surprisingly common.
Engineering a scaled database — and then performing the application changes needed to take advantage of that scaled out database — is tough coding work indeed. We encounter companies with full control of their source code who are petrified to make the changes needed to scale database capacity. Perhaps it's an ecommerce app, and it's too close to Black Friday. Or maybe it's just a case of attrition: the folks who really understand that code base are long gone, and the current engineers don't dare mess with the interworkings of the app.
These kinds of meltdowns are common during surge events, like the one ESPN suffered with the launch of Fantasy Football or the one Macy's suffered last Black Friday. Sometimes customers can see these events coming (e.g., they're expecting a major traffic surge on Black Friday) and sometimes they simply don't (e.g., their product gets a nod from a celebrity and all of a sudden they're swamped).
When a traffic surge takes down your site, it usually means the data tier was already fragile. Scaling the web infrastructure is pretty easy, as is scaling internet capacity. But scaling the data tier itself is where the challenges lie.
The Instapaper crisis also illustrates how the cloud alone doesn't solve the challenge of scaling the data tier. While elasticity is a hallmark of cloud services, the physics around having an application talk to multiple instances of a database remains a challenge. We've seen some customers suffer from an inflated sense of confidence that running in the cloud takes away these difficulties.
Don't wait for disaster to strike. Whether you're running on prem or in the cloud, keep a close eye on all metrics that reveal how "hot" your systems are running. Ensure your disaster recovery plan is robust — and recently tested. Better yet, don't rely on disaster recovery. Instead, run in active/active mode, where you've got multiple instances of all critical systems running in different locales, with the systems able to take on the full load if one portion fails.
Take steps now to scale your data tier and avoid these kinds of catastrophic outages. Those "Here's why we failed" engineering blog entries are no fun to write.
Michelle McLean is VP of Marketing at ScaleArc.
The Latest
Industry experts offer predictions on how NetOps, Network Performance Management, Network Observability and related technologies will evolve and impact business in 2025 ...
In APMdigest's 2025 Predictions Series, industry experts offer predictions on how Observability and related technologies will evolve and impact business in 2025. Part 6 covers cloud, the edge and IT outages ...
In APMdigest's 2025 Predictions Series, industry experts offer predictions on how Observability and related technologies will evolve and impact business in 2025. Part 5 covers user experience, Digital Experience Management (DEM) and the hybrid workforce ...
In APMdigest's 2025 Predictions Series, industry experts offer predictions on how Observability and related technologies will evolve and impact business in 2025. Part 4 covers logs and Observability data ...
In APMdigest's 2025 Predictions Series, industry experts offer predictions on how Observability and related technologies will evolve and impact business in 2025. Part 3 covers OpenTelemetry, DevOps and more ...
In APMdigest's 2025 Predictions Series, industry experts offer predictions on how Observability and related technologies will evolve and impact business in 2025. Part 2 covers AI's impact on Observability, including AI Observability, AI-Powered Observability and AIOps ...
The Holiday Season means it is time for APMdigest's annual list of predictions, covering IT performance topics. Industry experts — from analysts and consultants to the top vendors — offer thoughtful, insightful, and often controversial predictions on how Observability, APM, AIOps and related technologies will evolve and impact business in 2025 ...
Technology leaders will invest in AI-driven customer experience (CX) strategies in the year ahead as they build more dynamic, relevant and meaningful connections with their target audiences ... As AI shifts the CX paradigm from reactive to proactive, tech leaders and their teams will embrace these five AI-driven strategies that will improve customer support and cybersecurity while providing smoother, more reliable service offerings ...
We're at a critical inflection point in the data landscape. In our recent survey of executive leaders in the data space — The State of Data Observability in 2024 — we found that while 92% of organizations now consider data reliability core to their strategy, most still struggle with fundamental visibility challenges ...