Getty Images

Arm chips in spotlight amid AI infrastructure cost concerns

Between a planned dual-purpose mainframe chip with IBM and the growing cost crunch as enterprises weigh private AI, Arm might have finally met its moment, experts say.

More than three decades after its founding and six months after launching a CPU design for AI infrastructure, Arm has emerged as a serious data center contender as enterprises grapple with thorny decisions around AI security, privacy and cost.

Arm began as a consumer technology, embedded in 1990s PCs and early mobile phones, branching out into edge and IoT devices over the next two decades. In 2016, Softbank acquired the company, injecting R&D funding that led to Arm forays into high-performance computing and cloud server infrastructure, most notably as the basis for AWS Graviton servers in 2018.

Along the way, more business software, programming languages and operating systems became compatible with Arm chips, further increasing adoption for applications where Arm's low power draw was attractive. Nvidia's adoption of Arm CPUs in its rack-scale AI infrastructure systems further raised the chips' profile in the industry.

"Until five years ago, it was a one-horse race: It was Intel or Intel," said Patrick Moorhead, founder, CEO and chief analyst at Moor Insights & Strategy. "Once AWS adopted Arm for Graviton, that set into motion the entire industry building open source stacks on top of Arm, and what used to be a second-class citizen, now it's native. If you're an enterprise, you're likely using Arm and you don't even know it."

Arm's AI CPU moment

Meanwhile, over the last year, agentic AI and AI infrastructure began moving out of the hyperscaler and frontier lab realm and into enterprises, as concerns about costs and governance prompted companies to consider bringing AI inference workloads in-house. Vendors such as IBM and Red Hat also propose that enterprises run a mix of lighter-weight and small language models on CPU hardware to contain AI infrastructure costs. And whether cloud-based or on-premises, AI inference and agentic AI workloads increasingly call for CPUs rather than the more powerful, expensive and scarce GPUs that AI model training had used.

These trends and Arm's trajectory fully intersected in March 2026, when Arm launched its AGI CPU for AI infrastructure, which the company's official history describes as a "defining company milestone." In the past year overall, Arm has shipped more than 1.5 billion of its Neoverse processing cores, which form the basis for the AGI CPU -- 500 million of them in the last  nine months, according to company officials on a press briefing call Aug. 20.

Furthermore, in April, the company announced an agreement with IBM to design dual-function chips for open systems and mainframes, with further details revealed this week at the Hot Chips conference in Stanford, Calif.

These developments sit "right at the heart of what's going on as AI moves toward enterprise," said Stephen Sopko, an analyst at HyperFrame Research. "And [supporting] the mainframe is all about the enterprise."

Arm IBM mainframe chip plan details turn heads

Initially, some analysts wondered whether the collaboration with Arm for mainframe integration might go the way of previous efforts to combine IBM Z and x86 processors, which materialized, but with lingering issues.

The planned new processor won't ship until IBM's next System Z refresh in 2029, but new details revealed by the companies this week surprised and intrigued some industry observers. The forthcoming chip, according to IBM and Arm's press briefing, will run both System Z's s390x and Arm Aarch64 instruction sets on the same processor cores.

Steven Dickens, founder and CEO, HyperFrame ResearchSteven Dickens

"At the moment, all the Linux packaging and tooling [supported on the mainframe] has to be ported to s390x, and IBM's got a team of 20 people who do that -- it's a full-time job," said Steven Dickens, founder and CEO at HyperFrame Research. "Once this starts to ship, there's no porting exercise to be done -- pretty much anything that runs on Arm is going to be able to run on that instruction set."

"It's a very novel way that they did it. The first time I read about it, I'm just like, 'Okay, you're putting a second chip in there.' No, they're not. 'Okay, you've got different firmware per environment.' No, they're not," Moorhead said. "I was impressed with the amount of deep integration and engineering that IBM and Arm must have done. And by the way, I do think this is a black eye for x86."

Arm plays AI power efficiency card

While mainframe integration remains a future prospect rather than a current reality, Arm's history in low-power edge and mobile environments means it can pack more chips into a smaller power footprint, a crucial consideration as enterprises consider moving AI workloads on-premises.

In a 36-kilowatt air-cooled rack, for example, Arm's AGI CPU can pack 30 1U servers, with 8,160 CPU cores, according to the company's press briefing this month. By comparison, a comparable rack of x86 CPUs would hold 17 2U servers and 4,352 CPU cores, according to the company's press materials.

In some ways, the comparisons are apples and oranges; x86 CPUs also offer more compute horsepower, even if their power draw is many times larger, but many of the most advanced rack-scale systems also include GPUs and require liquid cooling. Packing more CPU density into air-cooled systems sets up a potentially compelling selling point for Arm as enterprises look to fit more into existing on-premises data center infrastructure, Sopko said.

"If you're a procurement manager or CIO, and you've got an air-cooled data center, you're going to be looking for air-cooled [products] rather than necessarily retrofitting for liquid-cooled," Sopko said.

Moreover, Arm's low power draw could also lend itself well to physical and edge AI as they become mainstream for enterprises, Sopko said. Qualcomm has put Arm chips at the center of its new Dragonfly data center and Dragonwing edge computing hardware, touting continuity between AI infrastructure in both environments.

Sopko cited a presentation by Christiano Amon, CEO of Qualcomm, at Computex in Taipei, which explained how the same processor architecture, including Arm chips, could underpin various devices, from the edge to the data center.

"The same architecture from Qualcomm is going to go from the earbuds I'm wearing right now all the way to my computer, to my phone, to my car, to the drone flying overhead, and into the data center," Sopko said. "His point was we're going to be an agent-centric world where that agent follows you across all of these devices."

X86 isn't holding still

Still, many enterprises also already have a large estate of x86 hardware and some legacy software that still favors x86, another consideration as they look to retrofit existing infrastructure for AI workloads, Sopko said.

Major x86 vendors AMD and Intel have been improving their chips' power efficiency, including at the edge, where AMD has been developing its Ryzen and Intel its Xeon 6 and Atom processors. Last week, AMD reported that it has made significant progress toward a goal set in 2024 to deliver a 20x increase in rack efficiency for AI training and inference by 2030.

Each enterprise must evaluate the performance of Arm and x86 and their relative power efficiency for their own workloads, Moorhead said.

"There are power efficiency advantages in certain environments," he said. "It's a very debatable subject, and there are exceptions to every rule, because you can buy Arm chips with exceptional performance, and you can buy Arm chips that are more geared towards efficiency."

Complexity, memory crunch blur outlook

But Arm is here to stay, according to Sopko -- and increased density in x86 compute will likely free up data center floor space for more kinds of hardware, he said.

With the new Intel Xeon 6+ chip, you can fit that in a third the floor space … which begs the question: What are you going to do with the other two-thirds?
Stephen Sopko, Analyst, HyperFrame Research

"With the new Intel Xeon 6+ chip, you can fit that in a third the floor space of your data center [compared to previous versions]," he said. "Which begs the question: What are you going to do with the other two-thirds?"

Ultimately, the future is likely to be mixed, according to research from Omdia, a division of Informa TechTarget. In a survey of 400 respondents in North America conducted in April, 76% agreed that silicon diversity is more important now than it was two years ago.

"For enterprises, token efficiency is critical, but on-premises deployments also often have some hard constraints in terms of the overall power envelope," said Scott Sinclair, an Omdia analyst who led the survey. "It is about delivering viable options that can fit into existing environments while providing cost efficiencies."

Another major factor to consider in enterprise AI infrastructure purchasing decisions will be the global memory crunch that has vendors, including Dell and Nvidia, warning of price increases.

One analyst predicted this will further raise the appeal of CPUs, including Arm, for AI workloads.

"The shortage is also behind some of the push toward CPUs for agent work," said Mike Leone, an analyst at Moor Insights & Strategy. "When the expensive, scarce part is the memory bolted to a GPU, you want anything that doesn't need it running on a cheaper socket."

Conversely, memory cost and scarcity could drive enterprises to keep AI infrastructure in the cloud for the short term and delay a move on-premises, Sopko said.

"I think that enterprise is going to get there. It has to get there. But in so many ways right now, much of what we're seeing is that it makes sense for the enterprise to go to a [cloud vendor] and run AI workloads there," he said. "In three years, they may be running it on-premises … but for now the cost is still higher than what a lot of enterprises are willing to pay."

Beth Pariseau, senior news writer for Informa TechTarget, is an award-winning veteran of IT journalism. Have a tip? Email her or connect on LinkedIn.

Dig Deeper on Data Center