Close Menu
The Financial News 247The Financial News 247
  • Home
  • News
  • Business
  • Finance
  • Companies
  • Investing
  • Markets
  • Lifestyle
  • Tech
  • More
    • Opinion
    • Climate
    • Web Stories
    • Spotlight
    • Press Release
What's On

How Shadow AI Is Skewing ROI

August 26, 2026

The Good, Bad And Ugly From The Packers’ Final Training Camp Practice

August 26, 2026

Immigrant businesses who spoke out against Mamdani supermarkets say they’re getting hit with NYC sanitation fines

August 26, 2026

The Biggest Bottleneck In Enterprise AI

August 26, 2026

Meta to limit teens to two hours daily on Facebook, Instagram

August 26, 2026
Facebook X (Twitter) Instagram
The Financial News 247The Financial News 247
Demo
  • Home
  • News
  • Business
  • Finance
  • Companies
  • Investing
  • Markets
  • Lifestyle
  • Tech
  • More
    • Opinion
    • Climate
    • Web Stories
    • Spotlight
    • Press Release
The Financial News 247The Financial News 247
Home » The Biggest Bottleneck In Enterprise AI

The Biggest Bottleneck In Enterprise AI

By News RoomAugust 26, 2026No Comments6 Mins Read
Facebook Twitter Pinterest LinkedIn WhatsApp Telegram Reddit Email Tumblr
Share
Facebook Twitter LinkedIn Pinterest Email

Suman Debnath is the director of developer relations and product at Crusoe, specializing in AI infrastructure and managed inference.

The tech industry is obsessed with new AI models. Every few months, a larger one appears. Context windows grow, benchmark scores improve, and new multimodal features expand what these systems can accomplish. Each release generates another wave of excitement, prompting organizations to ask whether the latest model will unlock their next competitive advantage.​

That focus is understandable. Foundation models have advanced at an extraordinary pace, making it easier than ever to build applications that would have seemed impossible a few years ago.​

But after years of working across cloud AI platforms, distributed machine learning frameworks, developer ecosystems and AI infrastructure, I have observed a different constraint emerging inside enterprise organizations.​ The limiting factor is no longer model capability. It is infrastructure.​

As organizations move beyond proofs-of-concept into production workloads, the conversation often changes dramatically. Success comes to depend on delivering consistent inference performance, orchestrating distributed training, managing GPU utilization, controlling costs and ensuring applications remain reliable under production traffic.​

Many organizations I’ve worked with have discovered that the hardest engineering problems begin only after the model has been selected. Infrastructure has quietly become the largest determinant of whether an AI initiative succeeds or fails.

Building AI Has Never Been Easier, But Operating AI Has Never Been Harder

Creating an AI prototype has never been more accessible.​

Developers can combine foundation models, vector databases, orchestration frameworks and cloud APIs to produce impressive demonstrations in a matter of days. Open-source tooling, managed services and increasingly capable APIs have dramatically lowered the barrier to experimentation.​

Production environments are different.​

Enterprise AI workloads expose challenges that prototypes rarely encounter. Models require distributed GPU clusters capable of serving millions of inference requests. Data pipelines must continuously ingest, transform and validate information. Memory utilization, network bandwidth, storage throughput and workload scheduling can then become just as important as model accuracy.​

Inference latency that appears insignificant during testing can become a costly operational problem when multiplied across millions of requests. Data movement between storage systems and accelerators can become a greater bottleneck than the model itself.​

These are no longer traditional infrastructure concerns managed exclusively by platform teams. They have become core product challenges.

Addressing The Infrastructure Challenge

Most of these bottlenecks are diagnosable long before they become expensive. In my experience, the teams that want to get this right must start with measurement, not procurement.​

I worked with a leading physical AI company that was struggling to scale its distributed training. The team was convinced it needed more compute. When we examined the workload, the GPUs they were already paying for were running at less than 10% utilization. The accelerators were not the constraint; they were sitting idle, waiting on a data preprocessing pipeline that could not transform and feed data quickly enough to keep them busy.​

Rather than expanding the cluster, we rebuilt the preprocessing and training pipeline on an open-source distributed computing engine—so that data loading, transformation and training could scale as coordinated stages instead of competing for the same resources. Utilization moved above 80%, and the team trained and deployed its models on a smaller footprint than it had originally requested. The capability they wanted was already in the budget. It was just trapped upstream of the GPUs.​

Not sure where to start in your own business? I’ve developed a few key best practices for addressing the infrastructure challenge:

1. Establish a baseline before you buy anything. Before approving another capacity request, I would urge leaders to ask their teams three questions: What is our actual GPU utilization across the fleet? What does a single inference request cost us at p95 latency, not at average? And where does the time go between a request arriving and a response leaving? Many organizations cannot answer these with confidence, which means they are making capital decisions on incomplete information.

2. Resist the instinct to solve utilization problems with more hardware. It’s the most common stumbling block I see: A team concludes it needs a larger cluster when the cluster it already owns is idle, waiting on a data pipeline that cannot feed the accelerators, a scheduler that cannot pack jobs efficiently or a serving configuration that leaves memory bandwidth unused. Adding GPUs to a pipeline-bound workload increases costs without improving performance. Profile the full path from raw data to returned result, then buy.​

3. Evaluate the model and the serving stack as one decision. Teams frequently select a model in isolation—and only later discover that its production economics are untenable. A somewhat smaller model, paired with disciplined batching, quantization and caching, can deliver better real-world latency and cost than a larger model deployed naively. Model selection and serving strategy are a single architectural choice, not two sequential ones.​

4. Treat capacity as a portfolio. Training tends to be bursty and relatively tolerant of scheduling delays, while inference is steady and latency-sensitive. Committing entirely to long-term reserved capacity sized for peak training demand is costly, and running everything on demand is usually worse. Match the commitment structure to the workload—and preserve the ability to move workloads as hardware generations and pricing shift. Standardizing on open orchestration and common inference interfaces is what makes that flexibility real rather than theoretical.​

5. Give unit economics an owner. Cost per 1,000 requests, or per 1 million tokens, belongs in the same review as accuracy and latency, discussed by the same team. The organizational failure mode here is more common than the technical one—infrastructure engineers sitting in a separate queue, receiving requirements after the architecture has already been settled. Embedding them in the AI product team from the start is the highest-leverage change most organizations can make.

Infrastructure Is A Strategic Capability, Not A Support Function

Customer experience depends not only on model intelligence, but on whether the underlying infrastructure can consistently deliver that intelligence with acceptable performance, reliability and cost efficiency.​

The organizations that continue treating infrastructure as a supporting function may watch their AI ambitions outgrow the platforms built to support them.​ Instead, leaders must view infrastructure as a strategic capability that evolves alongside their AI products.​

Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?

Suman Debnath
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email

Related News

How Shadow AI Is Skewing ROI

August 26, 2026

Why Software Will Not Be Replaced By AI, But Will Be Reimagined

August 26, 2026

The End Of Robotic Process Automation? Not Even Close. Here’s What’s Actually Changing

August 26, 2026

Five Factors For Effective AI-Enabled Enterprise Transformations

August 26, 2026

AI Governance Doesn’t Need A New Owner—It Needs A New Interface

August 26, 2026

SDD Fixed How AI Writes Code—Nobody’s Fixed How Humans Verify It

August 26, 2026
Add A Comment
Leave A Reply Cancel Reply

Don't Miss

The Good, Bad And Ugly From The Packers’ Final Training Camp Practice

News August 26, 2026

Bring on Minnesota.The Green Bay Packers held their 16th — and final — training camp…

Immigrant businesses who spoke out against Mamdani supermarkets say they’re getting hit with NYC sanitation fines

August 26, 2026

The Biggest Bottleneck In Enterprise AI

August 26, 2026

Meta to limit teens to two hours daily on Facebook, Instagram

August 26, 2026
Stay In Touch
  • Facebook
  • Twitter
  • Pinterest
  • Instagram
  • YouTube
  • Vimeo
Our Picks

Why Software Will Not Be Replaced By AI, But Will Be Reimagined

August 26, 2026

P.F. Chang’s at Newport Beach’s Fashion Island closed after 32 years

August 26, 2026

The End Of Robotic Process Automation? Not Even Close. Here’s What’s Actually Changing

August 26, 2026

California CEO compares controversial new tire rules to high-speed rail

August 26, 2026
The Financial News 247
Facebook X (Twitter) Instagram Pinterest
  • Privacy Policy
  • Terms of use
  • Advertise
  • Contact us
© 2026 The Financial 247. All Rights Reserved.

Type above and press Enter to search. Press Esc to cancel.