Close Menu
The Financial News 247The Financial News 247
  • Home
  • News
  • Business
  • Finance
  • Companies
  • Investing
  • Markets
  • Lifestyle
  • Tech
  • More
    • Opinion
    • Climate
    • Web Stories
    • Spotlight
    • Press Release
What's On

What Your Infrastructure Chooses When You Haven’t

September 14, 2026

Streaming Services Are All Using Podcasts Differently

September 14, 2026

Chick-fil-A is one of America’s toughest franchises to get — and the startup cost may surprise you

September 13, 2026

Behind the scenes of the giant Italian restaurant, market opening at Grand Central

September 13, 2026

War between SoFi Stadium and Inglewood takes dramatic turn as judge rules $400M deal is void

September 13, 2026
Facebook X (Twitter) Instagram
The Financial News 247The Financial News 247
Demo
  • Home
  • News
  • Business
  • Finance
  • Companies
  • Investing
  • Markets
  • Lifestyle
  • Tech
  • More
    • Opinion
    • Climate
    • Web Stories
    • Spotlight
    • Press Release
The Financial News 247The Financial News 247
Home » What Your Infrastructure Chooses When You Haven’t

What Your Infrastructure Chooses When You Haven’t

By News RoomSeptember 14, 2026No Comments5 Mins Read
Facebook Twitter Pinterest LinkedIn WhatsApp Telegram Reddit Email Tumblr
Share
Facebook Twitter LinkedIn Pinterest Email

Manas Chaudhari, Tech Lead for WhatsApp’s messaging infrastructure at Meta, building distributed systems that serve billions of users daily.

​Capacity planning is taught as a number. Estimate the peak, add headroom, buy the machines. Most organizations do this part well, and much of it is automated now. It is almost never the cause of the outage.

You provision against a forecast, and the utilization graph sits comfortably under the line for months. Then the system falls over, and the incident review does not say “too much traffic.” It says one tenant. One feature. One bad deploy. One unusually expensive query.​

Capacity reduces how often you fail. It does not decide what happens once you do. Researchers who studied 21 major outages across 11 organizations found the same escape route in more than half of them, and it was shedding load: throttling requests or turning traffic off until the system could breathe.

When capacity does run short, your infrastructure decides who gets less on its own, and it is a poor decision-maker. It knows nothing about which of your customers matter. It favors whatever is loudest.​​

Right Number, Wrong Place​

​Your forecast tells you how much load to expect, and a good one sizes for the peak rather than the average. What it does not tell you is where that load lands.​

Concentration is what breaks things, and it shows up wherever workloads share a resource and one of them is hotter or more expensive than the rest. A multi-tenant database can develop a hot partition because one customer’s identifier concentrates writes on a single shard. A shared query layer meets one pathological scan holding connections everybody else waits for. Operations teams call this noisy neighbors, data modelers call it hot keys, queueing theory calls it head-of-line blocking, and your customer calls it slow. Four vocabularies for one problem, which is part of why it rarely reaches anyone as a business question.​

I have spent years building messaging infrastructure, where one write can become millions of deliveries. The daily total was never the number that mattered. A goal in a World Cup final moves a large part of the planet to their phones in the same second. Load jumps to two or three times the usual peak; fanout means each of those requests arrives multiplied, and none of it was in anyone’s forecast.​​

Loud Beats Important​

​When something fails, it retries. Retrying generates more load. So the component in the worst shape is the one asking for the most, and it gets served alongside everything else, because your infrastructure has no opinion about which requests matter. Capacity ends up being allocated in a roughly inverse proportion to health, and health has nothing to do with importance. This is worse than a coin toss, which would at least be fair. A background export nobody would miss can retry its way into the connection pool your checkout depends on, and win.​

Loud does not have to mean broken, and the healthy version is more common. Something goes viral, a wave hits one feature, and because that feature now accounts for most requests, it consumes most of the capacity. The critical path carrying a fraction of the volume and most of the revenue gets what is left. Nothing malfunctioned. Your system allocated by share of traffic, the only signal it had and share of traffic is not a measure of importance.​

Either way, the trouble outlives its cause. I have watched a system stay down long after the thing that knocked it over had been fixed, because by then the recovery traffic was the load. The original trigger was gone, and it no longer mattered. The only way out was to take work away, and taking work away means choosing what to drop.​

Decide Who Degrades First​

​Often, the decision depends on a partition boundary, a queue discipline or whichever connection pool empties first. Nobody chose that, and it usually falls on your smallest customer, the one sharing a shard with your largest.​

Two levers change that, and both cost something. Isolation bounds the blast radius, and you pay in utilization. Shedding lets you choose who degrades, and you pay in revenue from the requests you drop. Neither is exotic or hard to buy. What is missing is the input they need: a pre-agreed order of expendability.​

So the work is a decision rather than a purchase. Name who degrades first, in writing, before the incident. Which customers, which tiers, which features are expendable and in what order.​

Meta published how it does this. Product engineers decide in advance which features can be switched off, grouping them into levels aligned with named scenarios, each with a written summary of what stops working. The company rehearses it quarterly. The ordering exists before anyone needs it.​

It is an uncomfortable conversation, which is exactly why it gets deferred. Deferring it has a specific cost, and I have watched teams pay it. The conversation you skip in advance is one you end up having during the outage, with partial information, while the clock runs and everyone in the room has a different idea of what matters. The deliberation becomes part of the incident. Every minute spent establishing that the batch job is less important than checkout is a minute the outage continues, and the answer you reach under that pressure is worse than the one you would have written down calmly months earlier.​

That difference is not an engineering problem. Handling more is the easy half of scale. Deciding who gets less is the half that requires someone to say so on the record. Who is that person at your company, and when did they last write it down?

Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?

Manas Chaudhari
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email

Related News

Milo The Whale Shark Breaks World Record After Traveling 34,000 Miles

September 13, 2026

Every Powerful Leader Has One Powerful Weakness

September 11, 2026

Accountability Is The Question RTO Was Always Really About

September 11, 2026

How Supply Chain Executives Can Bring AI Into Their Operations

September 11, 2026

Why India Is Where The AI And IoT Revolution Is Being Solved

September 11, 2026

Where Quantum Technology Could Deliver Practical Value First

September 11, 2026
Add A Comment
Leave A Reply Cancel Reply

Don't Miss

Streaming Services Are All Using Podcasts Differently

News September 14, 2026

Streaming services are increasingly adding video podcasts to boost user engagement, reduce churn, and acquire…

Chick-fil-A is one of America’s toughest franchises to get — and the startup cost may surprise you

September 13, 2026

Behind the scenes of the giant Italian restaurant, market opening at Grand Central

September 13, 2026

War between SoFi Stadium and Inglewood takes dramatic turn as judge rules $400M deal is void

September 13, 2026
Stay In Touch
  • Facebook
  • Twitter
  • Pinterest
  • Instagram
  • YouTube
  • Vimeo
Our Picks

Starrett-Lehigh building reigns supreme as key Fashion Week anchor

September 13, 2026

Bari Weiss caught in ‘Cold War’ power struggle with CBS News president

September 13, 2026

Proskauer Rose law firm adds 2 more floors at 11 Times Square

September 13, 2026

Larry Ellison cancels plan to sell Oracle stock worth up to $7.5B

September 13, 2026
The Financial News 247
Facebook X (Twitter) Instagram Pinterest
  • Privacy Policy
  • Terms of use
  • Advertise
  • Contact us
© 2026 The Financial 247. All Rights Reserved.

Type above and press Enter to search. Press Esc to cancel.