Why AI's Future Isn't in the Cloud: OpenInfer's Play for Local, Sovereign Inference
When Behnam Bastani left Meta to start OpenInfer, he wasn’t chasing hype. He was running from a fundamental problem he’d watched compound for years: the infrastructure we’re building today won’t scale for the AI we’re about to deploy tomorrow.
“Everyone was using AI as a chatbot, as messages,” Bastani explains. “We said no, AI is going to be long-lasting, living with us. That means some pieces need to respond fast, some pieces respond slow, and they become heavily dependent on each other—what we now refer to as agents.”
This insight—that AI’s future is agentic, collaborative, and fundamentally different from the chatbot era—became the founding thesis of OpenInfer. And it’s a thesis built on years of systems-level thinking at some of tech’s most ambitious companies.
From VR Headsets to the Inference Problem
Bastani’s journey to founding OpenInfer started at Meta, where he was hired to think about the future of wearables, VR, and AR. The challenge seemed straightforward: how do you deliver a great VR experience without requiring a tethered connection to a PC?
The answer required thinking about the entire system—not just the headset, but compression, streaming, rendering, and AI working in concert. “If we look at the entire system, how compression, how streaming, how rendering, how AI works together, we can come up with a solution that brings the best experience of VR into a mobile headset,” Bastani recalls.
This led to Oculus Link and later INK, platforms that fundamentally changed how VR compute was distributed. Instead of everything happening on the headset or requiring a PC connection, they built a pipeline where some computation happened on the device, some on the PC, and AI helped mitigate latency between them.
The result? The cheapest, highest-quality VR headsets on the market. Zuckerberg announced it on stage in 2019, and it scaled to millions of users.
But Bastani wasn’t done thinking about distributed compute. At Roblox, he tackled an even harder problem: running AI inference at massive scale on CPU-only infrastructure. In 2023, when everyone said it was impossible, Bastani and his team brought voice moderation, multilingual content filtering, and moderation AI off expensive GPU clouds and onto Roblox’s own CPU-based infrastructure.
“No one thought you could do all AI inference on CPU at this scale,” Bastani says. “And that was possible because we looked at the entire stack again as a system.”
The Inference Era is Here
By 2024, a pattern became clear to Bastani: the world’s compute demand wasn’t going to be training models. It was going to be running them.
While everyone else was talking about model training and scaling foundation models, Bastani saw something different. “Usage is becoming the topic,” he explains. “And everyone was using AI as a chatbot, as messages, and we said no, AI is going to be long-lasting, living with us.”
This realization led to a bold claim: the infrastructure and software ecosystem built around today’s inference won’t scale for tomorrow’s AI. We need something fundamentally different.
“The infrastructure, the software ecosystem that is built around it, will not be able to let us scale,” Bastani says. “We need to look at the entire system in terms of how the hardware works, how low-level kernels operate, how resources are managed, how AI is processed, even how routing between different servers is done—everything built with that mission in mind of long-lasting collaborative AI inferences, aka agentic AI.”
This became the founding mission of OpenInfer: bring AI inference anywhere, on any hardware, on any infrastructure, by making AI ubiquitous.
The MVP Pivot That Shaped Everything
Building the MVP for OpenInfer wasn’t straightforward. Bastani and his co-founder initially targeted portable ecosystem devices—edge devices where inference was underserved. But the market had other ideas.
“We got pulled into the Neo Cloud direction,” Bastani explains. Neo Clouds are smaller cloud providers and compute providers who wanted to serve AI but lacked the infrastructure and capability. They could only rent GPUs. OpenInfer saw an opportunity to change that equation.
The pivot was critical. Instead of serving edge devices, OpenInfer focused on Neo Clouds, building an inference OS that could run on any hardware—old GPUs, new GPUs, CPUs, or mixtures of them all. The value prop was compelling: Neo Cloud providers could increase their profit margins by as much as 50% by offering a complete inference stack instead of just raw GPU rental.
“It reminds me of VMware,” Bastani says. “They said, ‘I don’t care about your hardware underneath it, I’m going to virtualize the workload for you.’ We’re now virtualizing inference for Neo Clouds. Don’t care if it’s a type of GPU you have, older or newer ones, a mixture of CPUs, or newer compute coming in—we virtualize it.”
Design Partners Over Guesswork
One of Bastani’s most important decisions was how to validate the problem. Rather than building in isolation, he spent significant time finding design partners—customers willing to take a risk on a startup and shape the product together.
“I’ve been on the other side where startups were pitching to me all the time and I’m like, ‘You’re solving their own problem. I’m not going to buy it,’” Bastani reflects. “So rather than us jumping into building, we spend quite a bit of time really understanding the problem.”
Today, OpenInfer works with five major Fortune 500 companies as design partners. These relationships shaped not just the product, but the entire roadmap.
The Roadmap: Where Data Lives
As OpenInfer matures, a new trend is emerging: sovereign clouds. The idea that AI should live where the data lives. Banks, robotics companies, drone manufacturers—they’re generating tons of data. Sometimes they need fast responses. Sometimes they can’t afford to send data to the cloud. Sometimes sovereignty demands it.
“The direction that AI needs to live where the data is,” Bastani explains. “But then how do you bring the compute? When you may not want to bring the whole compute, how do you split the compute? How do you make sure the compute works together?”
This is the next chapter of OpenInfer’s roadmap. And it’s a problem that requires the same systems-level thinking that shaped the company from day one.
Want to hear the full story? Listen to the complete episode with Behnam Bastani on Code Story to dive deeper into how OpenInfer is building the inference infrastructure for the agentic AI era, the hiring decisions that shaped the team, and what’s next for distributed compute. Listen now.