Vinod Bijlani is an AI practice leader at Hewlett Packard Enterprise.
Almost every enterprise AI conversation I have now starts with some version of the same question: Where should this actually run?
A CIO wants to know if a new fraud-detection model belongs on-premises or in the cloud. A head of data platform wants to know if last year’s “just use the API” decision still holds now that volume has 10x’d. A CISO wants to know why a vendor is pushing a public-cloud deployment when the data is regulated.
Even though these are essentially the same underlying question, I’ve found that almost nobody answers with a consistent framework. Placement decisions made by default, vendor relationships or pilot convenience deserve another look as AI becomes a production service.
Gartner forecasts worldwide AI-optimized infrastructure-as-a-service spending will reach $42 billion in 2026, with inference accounting for 55% of that spending and rising to 59% in 2027. The numbers are forecasts, but the direction is clear: Recurring inference demand changes the placement calculation.
I believe AI will be hybrid, much like enterprise cloud. Early cloud debates centered on whether everything would move to the public cloud; the more useful question became which workloads belonged where, based on cost, control and performance.
I expect AI to follow a similar path, combining hyperscalers, neoclouds, self-hosted infrastructure and edge. While these options can overlap, the workload should determine the mix.
Hyperscaler: The Generalist
Broad cloud platforms suit workloads that depend on managed models, integrated data services and enterprise support. They also suit experimentation or variable demand.
When adopting cloud platforms, check the specific service, region, capacity limits and data terms. Compliance credentials do not automatically make your application compliant. Cloud providers split duties under a shared responsibility model, and consumption pricing still needs budget controls.
Neocloud: The Specialist
Specialist GPU providers deserve consideration for compute-intensive training and sustained inference.
Evaluate the complete configuration: GPU memory, interconnects, storage throughput, capacity guarantees and support. A lower quote per GPU-hour only helps if the workload performs reliably and if your team can operate the surrounding stack, so compare equivalent service levels and commitments before assuming savings.
Self-Hosted: The Sovereign
This can suit workloads requiring direct operational control, close integration with sensitive data or predictable capacity. Sustained utilization may strengthen the economics, but factor in staffing, software, power, resilience and hardware refresh.
Ownership alone does not establish security. NIST’s zero-trust architecture rejects implicit trust based solely on location or ownership, so identity, authorization and monitoring remain essential.
Edge: The Front Line
When a manufacturing line must identify defects within a tight deadline, or a remote operation must continue without connectivity, proximity can determine placement.
NIST’s Fog Computing Conceptual Model identifies latency and sensor-data scale as challenges for centralized processing. Measure the complete journey from data capture to action—a fast model inside a slow workflow still misses the deadline. Edge systems also need updates and fleet management.
Inside The AI Design Room
With these options on the table, translate business requirements into placement decisions. Start with constraints that can rule an option out: response deadlines, connectivity, permitted data flows and operational criticality. Then compare quality, expected demand, operating capability and total cost across the remaining choices.
Follow the data through every step. The original retrieval-augmented generation (RAG) research describes models using retrieved passages to generate answers. If those passages are sent to an external model, their contents cross that boundary even when the source database stays on-premises. Include prompts, logs and outputs in the assessment.
Demand matters just as much. A marketing assistant with approved, low-sensitivity inputs and irregular usage may suit a managed cloud service; a continuously used internal assistant may justify dedicated hosted capacity or self-hosting. Neither conclusion follows from the workload’s label alone, so test realistic volumes and peak demand at the required quality and response time.
Compare cost per successfully completed business task, including retries, retrieval, data transfer and human correction alongside compute. Research, such as the study on FrugalGPT, demonstrates that selecting combinations of models can reduce inference costs while maintaining performance on evaluated tasks. Treat this as a reason to benchmark routing on your workloads, rather than assume a universal saving.
Finally, one workflow can span several environments: In a manufacturing environment, inspection runs at the edge, maintenance records are retrieved privately and an approved cloud model interprets complex problems using only data cleared to leave. Each boundary adds latency, cost and dependencies, so distribute components only where the benefit justifies it.
Set The Rules Before You Scale
CIOs can establish this discipline early, as AI use cases take shape. Three moves help.
First, make placement part of the design discussion for each new use case. Before committing to a platform, agree on the business owner, quality target, response deadline, data boundaries, expected demand and operating responsibility. Buying GPU capacity leaves different responsibilities with your team than buying a managed model endpoint.
Second, make placement review part of AI governance alongside model approval and data classification. Then revisit it when volumes, models, pricing or requirements change. Define approved fallbacks in advance—an outage must never silently route sensitive information to an unapproved service.
Third, build flexibility through a hybrid AI token factory, which I’ve written about previously. This is a shared layer for accessing, routing and governing model services across approved providers and private infrastructure. Its orchestrator should select eligible services by data policy, quality, latency, availability and cost, with common metering and audit trails.
Standard interfaces and provider adapters reduce rework when switching providers, though portability still requires testing model behavior. The business case for this is greater negotiating flexibility, controlled failover and the ability to adopt better models without rebuilding every application.
Hybrid AI becomes useful when every placement has a business reason. The question is where each part of the workflow can deliver a reliable outcome, within its data boundaries, at an acceptable cost. Understand those considerations consistently, and infrastructure choices become easier to defend and easier to revisit.
Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?

