In this storyMetaNVIDIA

The Story

1 min

Meta says its in-house AI chips deliver 44 per cent lower total cost of ownership and 40 per cent better power efficiency than general-purpose Nvidia GPUs, for the specific workloads they are built to handle.

The chips belong to the Meta Training and Inference Accelerator line, known as MTIA, which the company began developing in 2023. Meta unveiled four generations in March, the 300, 400, 450 and 500, and said it would ship them on a roughly six-month cadence rather than the annual or longer pace common across the semiconductor industry.

The next chip, codenamed Iris and designated MTIA 400, is entering production this month after clearing testing. Meta said those tests ran for six weeks without surfacing major issues. The MTIA 300 is already in production handling ranking and recommendation work, while the 450 and 500 are aimed at generative image and video inference and are slated for mass deployment through 2027.

The speed comes from a modular chiplet architecture, which allows Meta to combine existing building blocks rather than redesign a chip from scratch each generation. Broadcom co-developed the programme under a partnership running to 2029, and TSMC handles fabrication, with newer parts among the first custom AI chips built on a 2-nanometre process.

The effort sits inside an enormous spending plan. Meta has guided 2026 capital expenditure to between $125 billion and $145 billion, with nearly all of the increase going towards data centres, GPUs and custom silicon. It aims to bring 7 gigawatts of compute capacity online this year and roughly double that in 2027, and has begun exploring renting spare capacity to outside customers.

Analysts characterise MTIA as a way to absorb growth and trim GPU costs at the margins rather than to displace Nvidia outright, particularly for training workloads where the CUDA software ecosystem remains dominant.

Key numbersCompany-stated
44%
Claimed Cost Of Ownership Reduction
40%
Claimed Power Efficiency Gain
~6 months
Chip Generation Cadence
$125-145 billion
2026 Capital Expenditure Guidance

Why It Matters

1 min

The qualifier at the end of that claim is doing most of the work.

Meta's figures apply to the specific workloads its chips were designed around, which at present means ranking, recommendation and similar inference tasks. Those are repetitive, well-understood and stable in shape. Building a fixed-function chip that beats a general-purpose GPU on a task you have specified in advance is not a surprising result. It is the entire reason application-specific silicon exists.

The comparison that would matter has not happened yet. Nvidia's position rests on training, where model architectures change, workloads are unpredictable, and CUDA's accumulated tooling is what engineers actually reach for. MTIA extends towards training only with the 450 and 500 generations, running through late 2027. Google and Amazon both took years to make their in-house training silicon genuinely competitive, and neither has displaced Nvidia from the workload that matters most.

So the honest reading is that Meta has solved the easier half convincingly and has not yet attempted the harder one.

What is genuinely notable is the pace. A new chip generation roughly every six months, against an industry that typically refreshes annually at best, is unusual and is the direct result of the modular chiplet approach. Meta is not designing four chips. It is designing a set of components and recombining them, which is a different and faster discipline.

The Strategic Read

1 min

The more interesting question is what runs between the chips rather than what runs on them.

Nvidia put $3.5 billion into MediaTek this month specifically so that custom accelerators built as Nvidia alternatives would still be designed around NVLink Fusion, NVLink-C2C and Nvidia's memory architecture. The logic was that if the compute die is being commoditised by customers building their own, Nvidia should own the interconnect, the memory layer and the rack integration instead. It collects either way.

Meta is precisely the customer that strategy was designed to retain. And Meta's partner is Broadcom, which is one of the few companies capable of supplying both custom silicon and the networking fabric to connect it. A Meta and Broadcom stack is the combination best placed to leave Nvidia's rack entirely rather than merely swap out its compute.

Whether it does is the thing worth watching, and it is not yet public.

The cost arithmetic also deserves proportion. Savings of 20 to 30 per cent on AI compute might amount to $2.5 billion to $5 billion a year at full deployment, against capital expenditure guided at $125 billion to $145 billion. That is real money and it is a small fraction of the spend. Nobody undertakes a four-generation silicon programme for a few per cent.

The actual motive is supply. Nvidia allocation has been rationed for two years, and owning a design means competing for TSMC capacity rather than waiting in Nvidia's queue. Meta is not primarily buying cheaper compute. It is buying the ability to decide how much it gets.

That has a consequence further down the chain. Morgan Stanley has warned of chipflation, with semiconductor and memory prices rising as hyperscaler demand absorbs supply. Every company that secures its own capacity makes the queue longer for those who cannot.

For daily, sharp analysis of the biggest moves in the Indian business and startup ecosystem, follow StartupFox.