Why On-Device AI Is Becoming the Next Battleground for Chipmakers
For most of the generative AI cycle, the competitive spotlight has been on data-centre accelerators. That is changing. As inference volumes explode, the cost, latency and privacy of sending every query to the cloud are under strain, and chipmakers are racing to put capable neural processing units (NPUs) into phones, PCs, cars and industrial devices. The contest is no longer only about who trains the largest model, but about who owns the silicon where billions of everyday AI interactions actually run.
Why the Edge Matters Now
Cloud Inference Is Hitting Physical Limits
The International Energy Agency estimates that data-centre electricity demand rose 17% in 2025, with AI-focused facilities growing about 50%, and expects total consumption to nearly double to roughly 950 TWh by 2030. Every inference task shifted to a device converts a recurring cloud cost, paid by the model provider in power and capacity, into a one-time hardware cost paid by the buyer. For hyperscalers facing grid-connection queues, offloading lightweight tasks such as summarisation, transcription and photo editing to local silicon is becoming an infrastructure strategy, not just a product feature.
The Installed Base Is the Moat
Apple disclosed an installed base of more than 2.5 billion active devices alongside its fiscal first-quarter 2026 results. That figure explains why on-device AI is strategic: whoever controls the processor, operating system and local model layer controls the default AI experience for billions of users, and the personal context that never leaves the device.
Sizing the Device Opportunity
AI PCs Cross the Majority Threshold
Gartner forecasts AI PC shipments of 143 million units in 2026, about 55% of the PC market, up from 31% in 2025 and 15.6% in 2024. It also expects 40% of software vendors to prioritise on-PC AI capabilities by end-2026, against just 2% in 2024. In business laptops, x86 held about 71% of the AI segment in 2025 versus 24% for Arm, which frames the core architecture fight: Intel and AMD defending enterprise incumbency while Arm-based designs compete on battery life and NPU efficiency.
Smartphones: Rising Penetration, Shrinking Market
Counterpoint Research expects GenAI-capable smartphones to reach 45% of global shipments in 2026, up from 36% in 2025, and 52% in 2027. Yet total shipments are forecast to fall 13.9% to 1.08 billion units, the lowest on record, because of the memory supply crisis. Counterpoint also notes that GenAI is now standard above USD 400 wholesale but has not yet become a compelling upgrade trigger. Share gains are therefore partly a function of a shrinking low-end base, a nuance often lost in headline adoption figures.
How Chipmakers Are Repositioning
Qualcomm: From Handsets to Every Edge
Qualcomm’s fiscal third-quarter 2026 results show the strategic logic. Handset revenue fell 20% to USD 5.1 billion, while automotive surged 61% to USD 1.6 billion and IoT rose 9% to USD 1.8 billion. The company now targets USD 40 billion in non-handset revenue by fiscal 2029. The edge battleground clearly extends beyond phones into cars, PCs, wearables and industrial systems, where the same low-power NPU expertise can be reused across far less cyclical end markets.
The Vertical Integration Advantage
Players that control silicon, operating system and model, notably Apple, Samsung and Google, can co-optimise memory allocation, NPU scheduling and model compression. Merchant suppliers such as Qualcomm, MediaTek, Intel and AMD must instead win through developer tooling, software runtimes and OEM breadth, making software ecosystems as decisive as raw TOPS ratings.
Memory: The Hidden Constraint
The biggest threat to on-device AI is not compute but DRAM. The Semiconductor Industry Association reported global chip sales of USD 146.8 billion in July 2026, up 135.1% year-on-year, and has endorsed a forecast of about USD 1.5 trillion for the full year, driven largely by AI infrastructure. Reuters reported that memory makers are diverting capacity toward high-bandwidth memory for AI servers, squeezing conventional DRAM, with Apple warning that rising memory prices had begun to pressure profitability. This is the central irony: the data-centre boom is starving the very devices meant to relieve it. A 3-billion-parameter model quantised to 4 bits needs roughly 1.5–2 GB of RAM before the operating system loads, so every new on-device capability carries a direct bill-of-materials cost.
Outlook: The Metrics That Will Decide Winners
Over the next three years, advantage will hinge on three measures: AI performance per watt, AI performance per gigabyte of memory, and the share of queries resolved locally rather than in the cloud. Hybrid orchestration, in which devices handle routine inference and escalate complex reasoning to the cloud, is likely to become the default architecture. Chipmakers that pair efficient NPUs with memory-light model support and mature developer ecosystems will capture the largest share of the next AI hardware cycle.