Skip to content Skip to content
Latest
Events August 31, 2026 15 min read Beginner Verified accurate

From Metal to Model: Broadcom’s Private AI Cloud Pitch at VMware Explore 2026

From agentic AI to private cloud, business needs are evolving fast.

You come to protect. The threats are real, and the stakes are rising.

You come to build, and you have the AI and private cloud knowledge to make it happen.

You come to advance. You lead with solutions, not just ideas. This is part of your story. Your career.

You come to connect. Look around. You are part of a community of experts.

And we all come to explore.

That was the opening montage at VMware Explore 2026 in Las Vegas, and it turned out to be a decent table of contents. What followed was ninety minutes on private AI cloud, with roughly equal time given to cost, security, and getting from bare metal to a running model. Broadcom also brought up three customers who have already done the work, which is usually the more interesting part of any keynote.

Two speakers on the main stage under the VMware Explore logo
The general session, “Shaping the Future of Private AI Cloud and Agentic Innovation.”

Here are my notes.


Ram Velaga on what customers keep telling him

Ram Velaga, president of Broadcom’s Infrastructure Software Group, opened. He’s been in the software job nine months but at Broadcom fourteen years, most of that building networking silicon for very large data centers. His stated priorities: lead with technology, ship quality at scale, stay focused on operations.

He then went through four things he keeps hearing from customers. None of them will shock anyone who works in this space.

Private cloud stopped being the loser. A few years ago the press releases all read the same way: company X is going all in on public cloud. Ask those same enterprises now and the answer is more careful. There’s a place for public and a place for private, and for steady, mature workloads the numbers increasingly point back home.

Hardware is scarce and will stay scarce. This isn’t a one year supply blip. Server costs are up, driven mostly by memory, and everything Broadcom sees from the supply side says that holds for the next few years. So the question changes. Are you getting full utilization out of the hardware you already own?

Data gravity is becoming data sovereignty. Customers want to bring compute to their data rather than ship data out to someone else’s compute. And sovereignty isn’t only a national borders conversation now. It’s showing up as an internal requirement inside individual enterprises.

Nobody wants a second stack just for AI. Every organization is already working out which models to run and where. The last thing they want is a parallel set of networking, storage, and security tools that exists only to serve AI workloads.

Velaga’s answer to all four is the thing VMware has done for twenty five years, which is abstraction. Compute first, then networking, then storage, then provisioning and automation, then operations, with security running through all of it.

Slide reading VMs, Containers, AI, Same Clusters: one platform across all applications
The core claim. The VM estate you run today, the apps being rebuilt as containers, and AI models and agents on pooled GPUs, all scheduled on the same hosts rather than on new hardware.

He also made a promise that lands harder with operators than any AI feature. Seamless upgrades, virtual patching, live updates. The idea being that keeping your infrastructure current should stop costing you application uptime.


First customer: rebuilding on the way out the door

[Editor’s note: confirm this customer’s name and title from the session recording.]

The first customer up started a data center migration in 2024. Since they were moving anyway, they decided to move to VCF at the same time rather than lift and shift twice and do the work over.

They set four goals: scale, resilience, security, and a decent engineering experience. The scale answer was a standardized rack they built with their hardware partner. A fully integrated rack of network, storage, and compute, imaged and patched before it ever arrives at the facility. Nothing gets built or patched on site. That also handed them a failure domain the size of one rack. Lose a rack and the workloads don’t care.

For security they run NSX distributed firewalling with a global manager across sites, and they’ve pushed micro-segmentation through the environment. They were honest that this was the hardest piece of the whole program, with real product gaps early on that got closed through engineering escalation.

The part I’d steal is the engineering experience. Everything was redeployed rather than migrated. All infrastructure defined as code, with security compliance built into the pipelines. Application teams provision, test, and re-test without filing a ticket or waiting on a separate security review.

What they put on the slide:

  • 40 to 45 percent reduction in power consumption, mostly from moving to newer, denser hardware
  • Over 99 percent virtualized, with roughly 200 clusters deployed self-service through automation
  • No downtime to application services during the infrastructure change

Next for them is a GPU design for private AI, plus building an operations experience that matches what they gave their engineers.


Paul Turner on VCF 9.1, and an assistant that does things

Paul Turner, chief product officer for the VMware Cloud Foundation Division, took over with a reminder that none of the customer stories happen without the platform underneath.

VCF 9.1 shipped in May. The numbers he put up: 39 percent savings on the cost of a server, 40 percent reduction in security risk, zero downtime for AI operations, and adoption sitting at more than 2,000 customers running over 19 million cores.

His ask of the room was direct. VCF 9 delivers nothing sitting in a slide deck. A QR code went up linking a cloud maturity assessment, an upgrade planner, and a 9.1 adoption kit, along with a newly announced Frontier AI Security Readiness Assessment for checking your environment against AI driven threats.

Then the demo people will actually talk about: the VCF AI Assistant, coming in an upcoming 9.x release.

It opens on a dashboard with real problems on it. A vSAN security alert covering 485 hosts across 17 clusters and four vCenter instances, plus a performance alert. Drilling into the security issue, the assistant finds image drift on a cluster, runs its own pre-checks, works out which hosts can take a live patch and which need maintenance mode, and remediates. The operator approves each step, so it’s interactive rather than autonomous.

The performance walkthrough was the better one. Asked why an application was slow, the assistant ran health checks down the stack, found CPU creeping toward critical over a couple of days, and traced it back to DRS automation being switched off on August 18th during someone’s maintenance work and never switched back on. It re-enabled DRS, rebalanced the cluster, and then offered to write a compliance rule so that DRS sitting disabled for more than an hour throws an alert from now on.

That last bit is the interesting one. Finding the problem is table stakes. Turning the fix into a standing policy is where the operations hours actually go.


Kubernetes and United Airlines

Turner reminded everyone that VMware bought Heptio in 2018 and had Kubernetes embedded in vSphere by 2020, which was early by any reasonable measure. That’s now vSphere Kubernetes Service, a CNCF certified conformant distribution, along with VM Service for running VMs under Kubernetes and containers dropped straight into pods.

United Airlines came up to explain their choice. Their story is that Kubernetes adoption wasn’t an infrastructure decision at all. The developers decided to refactor into cloud native services, and infrastructure had to meet them there. United standardized on Kubernetes on premises, evaluated the options, and picked VKS because it met the requirements and happened to sit on infrastructure their teams had run since 2009.

The constraint that mattered: developers had to get the same look, feel, and tooling they get in public cloud. Same pipelines, same tools, same mental model. An application team ships to EKS or to VKS and the delivery pipeline barely notices. Their phrase for it was “just another cloud.”

For resilience they run active-active across two data centers with centralized deployment tooling and GSLB steering traffic to the healthiest site. Platform engineering wired it into the application pipelines, so app teams now control the relevant infrastructure configuration and traffic behavior alongside their own code.

Next for them: VCF Automation across the estate, and the next VCF release.


The tipping point, according to 1,800 IT leaders

Then Turner put up the survey. Radius Tech ran it with Broadcom, covering 1,800 senior IT decision makers across eight countries in North America, Europe, and Asia Pacific, in February and March 2026.

Slide titled AI Tipping Point showing 62 percent, 51 percent and 56 percent figures
  • 62 percent are very or extremely concerned about AI costs
  • 51 percent are repatriating workloads because of security
  • 56 percent run or plan to run inference and production AI on private cloud
Bar chart comparing private cloud at 56 percent against public cloud at 41 percent for production AI workloads
Private cloud 56 percent, public cloud 41 percent, with public cloud use for production AI down 15 percent year over year.

People will still burst into public cloud. But the center of gravity has moved, and the reasons are the ones Velaga opened with. Cost, and the fact that the data already lives in the data center. Bring the AI to where the applications are.


Announcing VMware Private AI Cloud

Announcement slide reading VMware Private AI Cloud

The frame for the rest of the session was three pillars: cost, security, agentic AI.

Slide titled Delivering Your Private AI Cloud with three columns for Cost, Security and Agentic AI
Cost covers efficient infrastructure, overcoming tokenomics, and governed model choice. Security covers built-in distributed firewalling, lateral threat defense, and virtual patching for workloads. Agentic AI covers runtime isolation, a curated AI marketplace, and context-aware data products.

Cost, starting with memory

The standout cost item is memory tiering. You pair DRAM with an NVMe tier, and ESX handles it transparently so the workload just sees memory.

Slide showing hardware savings comparison between DRAM-only at 211,735 dollars and NVMe-tiered at 121,384 dollars
The worked example. DRAM at $6,159 per 64GB DIMM drops from $190,650 to $98,400 once you add a $1,899 1.6TB NVMe tier. That’s 42 percent off the hardware cost of a net new server or a refresh.

If you have a refresh coming, model this before you order anything.

Turner also made the point that private cloud doesn’t have to mean on premises only.

Slide titled Rapid Adoption with VMware Cloud on AWS showing partner clouds and 98 VMware Cloud Service Providers
A fully managed option, plus Amazon Elastic VMware Service, Google Cloud VMware Engine, Microsoft Azure VMware Solution, Oracle Cloud VMware Solution, and 98 VMware Cloud Service Providers.

Then he put up one word

Slide showing the word Tokens with a red circle-and-slash over it

The argument on tokenomics is that per token pricing is impossible to forecast. On a private cloud you aren’t buying tokens, you’re buying GPUs. You size the pool, you scale it, you budget for it.

He also pushed back on model maximalism. Not every workload needs a frontier model with hundreds of billions of parameters. An internal chat application might run fine on something two orders of magnitude smaller.

Slide titled Not Every Workload Needs a Frontier AI Model showing SLMs, open-weighted models and LLMs feeding into CPU and GPU
One governed catalog. Small language models for everyday work, open weighted models for private work, large models for deep work, right sized onto CPU or GPU.

VMware AI Factory

Announcement slide reading VMware AI Factory, One Platform from Metal to Model

AI Factory is the software defined foundation under Private AI Cloud, and the goal is compressing the path from bare metal to first served model from weeks down to hours. It automates hardware provisioning, stands up the software stack, builds the Kubernetes layer, provisions object storage, and pools GPUs as a shared resource across teams rather than dedicating hardware per workload.

The private AI services sit on top of that: a unified model gallery, a multi-tenant model runtime with isolated namespaces so lines of business can share models without sharing data, an AI Gateway giving one governed consumption interface across on premises and cloud, and secure sandboxes for isolating code that agents generate.

Slide showing VMware AI Factory partner ecosystem logos including AMD, Intel, Cisco, Dell Technologies, Lenovo, Supermicro and MetalSoft

Certified VCF AI ReadyNodes come from Cisco, Dell Technologies, Lenovo, and Supermicro, with accelerator work covering both AMD Instinct with ROCm and NVIDIA. The MetalSoft partnership is the new piece. It brings heterogeneous bare metal provisioning into the VCF management console directly, taking server provisioning and repaving from weeks to minutes.

The model lifecycle demo

Sabina Anja, chief technologist for VMware Cloud Foundation, ran a scenario built around an insurance company with four teams sharing one inference platform: claims, underwriting, finance, and product.

Screenshot of the VCF Automation Private AI model gallery showing model cards for Nemotron, Gemma, Phi, Llama, DeepSeek, GLM, Kimi, Qwen and others
The model gallery inside VCF Automation, with Model Deployments, Model Routes, Guardrails and VKS Clusters in the left nav.

Making a new model available took about four minutes start to finish. Pull it from the gallery and it lands in the organization’s own OCI registry, signed and scanned, so every replica pulls internally instead of from the internet. Pick a VKS cluster with GPUs already provisioned, which in this case meant landing on the same nodes already serving claims and underwriting. Set replicas, autoscaling, quantization, and runtime flags.

Then model routes, which is the piece worth understanding properly. Applications call a permanent URL and a model name. Never a specific model version, never a server. That indirection is what lets you swap the model underneath without touching a single application. The route also carries failover models, traffic splitting for A/B, authentication, rate limits, and guardrails.

Guardrails run at the route rather than in the application, so every request gets inspected twice. On the way in, PII gets masked and unsafe prompts get blocked before they reach the model. On the way out, sensitive information gets stripped again. An application somebody wrote two years ago gets the same enforcement as one that shipped this week.

Last piece was budgets and observability. Token spend per team, usage broken out by model, and the metrics that actually matter for inference: time to first token (under a second is good), tokens per second, and GPU utilization underneath all of it.


Security, and the end of bolt-on

Slide titled Delivering Your Private AI Cloud with the Security column highlighted

Chris McCain, global field CTO, took the security segment, and his argument was blunt.

Slide titled Time to Re-Think Security contrasting bolted-on security with built-in security

An appliance based perimeter firewall is a bolt-on. It’s locking the front door of the building and calling the room safe. Once you have agents, distributed applications, and frontier AI in your threat model, security has to move to the application perimeter. Lock down services, control which agents can call which tools, inspect east-west traffic.

Turner had set this up earlier with an example worth sitting with. An autonomous, AI orchestrated attack campaign that ran four and a half days unsupervised, executing roughly 17,600 actions at about 116 actions per hour. You cannot defend against that with human speed processes and manual rule review. Software has to stop software.

VMware’s structural advantage is that every packet already passes through ESX, which is what makes a distributed firewall possible without appliances sitting in the path.

Slide showing the VMware Cloud Foundation security stack with platform hardening, vDefend, Avi Load Balancer and Live Recovery
Three layers. Hardened infrastructure, lateral security (distributed firewall, IDS/IPS, WAF/WAAP, malware prevention), and purpose built recovery with live behavioral analysis and on-demand isolated recovery environments.

McCain’s demo ran on a fictional customer:

  • A security readiness assessment that gives you a blueprint for consuming security across infrastructure, environment, and application layers, not just a score
  • Security Explorer, showing application flows in green where policy is doing its job, red where flows are still unprotected, and blue where something got blocked. In the demo, a task runner probing an application it had no business touching.
  • Policy recommendations from the security services platform, aimed squarely at the very common situation of not knowing what your application flows look like in the first place. A human reviews and approves the proposed rules.

Then the sleeper feature.

Slide titled Virtual Patching with Distributed IDS/IPS showing a CVE being mitigated across VMs and Kubernetes workloads
Threat feeds push signatures down automatically. A policy scoped to the affected workloads blocks exploitation of a specific CVE with no downtime and no change to the workload itself.

Virtual patching buys the application team weeks or months to test and schedule the real patch, while the exploit path is already closed. The IPS dashboard then shows blocked attempts against that CVE, which is how you know it’s working rather than just configured. Worth noting the slide also lists AI discovered software vulnerabilities as an input, which says something about how fast the disclosure pipeline is moving.


Two more announcements

Announcement slide reading TrueSource by Broadcom with Spring Enterprise, Trusted Artifacts and Data Services

TrueSource extends Broadcom’s Spring security work into a broader trusted supply chain offering. Spring and its dependencies, trusted Python, Node.js and Java artifacts along with secure images, and data services covering PostgreSQL, MySQL, Valkey and RabbitMQ.

Announcement slide reading AgentMinder by Broadcom with Identity, Security and Observability

AgentMinder is the governance layer for autonomous agents, covering identity, security, and observability. It sits alongside vDefend and Avi Load Balancer in the agentic AI security story.


Last customer: New Belgium Brewing

The session closed with New Belgium Brewing, the Colorado brewery behind Fat Tire. Their framing was that you can’t make world class beer without world class infrastructure, which got a laugh but is also just true.

They’re a sensor heavy manufacturing operation, and their position on that data was unambiguous. Manufacturing data is proprietary and it isn’t going in a public cloud. So VCF and private cloud enablement became the platform decision, especially as AI and machine learning objectives picked up.

Their message on security was that once you can genuinely see what’s talking to what inside your own data center, VCF 9 and the services on top of it get you to a real delivery model.

Next for them is agentic AI on VCF, so they can take corrective action faster in the manufacturing process itself.

Their closing line works as a summary of the whole ninety minutes: don’t sit on the sidelines and watch the technology go past.


What I took away

The through line across all of it was consistent, and it’s a different pitch from “we do AI too.”

Private cloud isn’t being sold as the safe, conservative option anymore. It’s being sold as the one that makes financial sense, because hardware is scarce, token pricing is unpredictable, the data is already sitting in your data center, and running a second infrastructure stack for AI is a line item nobody budgeted for.

Whether the AI assistant and AI Factory hold up in production is the question the next twelve months will answer. But the customers on stage weren’t describing pilots. They were describing 200 clusters, 45 percent less power, active-active across data centers, and micro-segmentation actually running.

That’s a different conversation than the one we were having a year ago.

Share

Leave a comment

Your email address will not be published. Required fields are marked with an asterisk.

This site uses Akismet to reduce spam. Learn how your comment data is processed.