Futurism logo

Every Laptop Is an "AI PC" Now. Here's What the Number on the Box Leaves Out.

TOPS has become the headline spec for AI laptops. But the engineering decisions chipmakers actually spend money on point to a different bottleneck entirely.

By Imran ValianiPublished a day ago • 8 min read
Image edited by the author using AI.

Editorial note: I use AI tools to assist with research, drafting, fact-checking, and editing. The analysis, opinions, engineering interpretation, and final editorial decisions are my own.

When Intel launched its Lunar Lake laptop chips, one number got repeated everywhere: 48 TOPS. That's the rated speed of the chip's new NPU, its dedicated AI processor, and it was more than four times what the previous generation offered.

It was a number built for a headline.

What didn't make the headlines were two quieter engineering decisions in the same launch. Inside the NPU, Intel's own product brief lists "up to 2x bandwidth compared to previous generation." And at the system level, Intel moved up to 32GB of memory off the motherboard and onto the processor package itself.

Nobody makes that second change for a marketing slide. Putting memory on the package takes up package space, adds assembly complexity, shifts cost into the processor itself, and gives up the option to upgrade memory later. Intel did it anyway.

Intel's public explanation for the memory move centers on lower power and a smaller motherboard. It hasn't said the move was about feeding the NPU. Connecting the two decisions is my reading, not Intel's claim. But from where I sit, after 20+ years in electronics manufacturing, the pattern is hard to miss: a chipmaker spent real money on moving data, not just on making the math faster.

Which raises a fair question. If the engineering budget is going into data movement, why is TOPS, a pure compute number, the only thing anyone talks about?

Three Processors, Three Different Jobs

TOPS stands for trillions of operations per second. To understand why it's an incomplete number, it helps to know what an NPU is actually for.

A modern laptop chip has three kinds of processors inside it:

  • The CPU gives up raw efficiency for flexibility. It can run almost anything.

  • The GPU gives up some flexibility for massive parallel throughput, and it can be remarkably efficient at the work it's suited for.

  • The NPU narrows further still. It's architected mainly around one kind of math, the matrix multiplications that make neural networks run, done at low numerical precision to save power. Around that math engine sits dedicated hardware for moving and preparing data.

That narrow focus is exactly the point. The NPU is built for the AI tasks Windows now runs constantly in the background: live captions, semantic search indexing, blurring your background on a video call. Jobs that run all day, quietly, without draining the battery.

So these three aren't competitors racing on one track. They're specialists, divided by workload and power budget. Any comparison that squeezes "AI performance" into a single number owned by a single processor has already misread the architecture.

Where the "40 TOPS" Line Came From

The number now printed on flagship AI laptop spec sheets traces back to Microsoft. Microsoft set 40 TOPS as the minimum NPU rating for its Copilot+ PC badge.

Microsoft's own wording is softer than a hard cutoff. It describes 40+ TOPS as what you need for "the best" on-device AI experiences. What's actually gated behind the full hardware tier is a specific set of features, including Recall, real-time translation in Live Captions, and Windows Studio Effects. Below the line, those particular experiences aren't built to run. That's different from saying no AI works at all on lesser hardware.

As a floor, it's a real engineering requirement rather than a marketing gimmick. AMD's Ryzen AI 300 series, Intel's Core Ultra 200V, and Qualcomm's Snapdragon X series all clear it.

The trouble starts above the line.

Same Number, Different Meanings

The ratings have climbed fast. Qualcomm's flagship NPU went from 45 TOPS on the original Snapdragon X Elite to 80 on the X2 Elite, and some later X2 models now reach 85.

But the brand name on the box tells you less than you'd expect:

  • Intel's "48 TOPS" family actually ships parts rated anywhere from 40 to 48.

  • AMD's Ryzen AI naming covers chips from the same year rated at 50 and 60 TOPS, with nothing in the family name to tell you which one you're getting.

  • Apple isn't playing this game the same way. Copilot+ is a Microsoft requirement, not a universal standard, and Apple's recent M-series announcements emphasize Neural Engine architecture, GPU AI acceleration, and memory bandwidth rather than publishing a headline Neural Engine TOPS figure.

For one person buying one laptop, the fix is easy: check the exact model number, not the family name.

For a company, it's a bigger deal. In my world, when you buy components by the thousand, variance like this gets called out explicitly on a datasheet. An OEM building a product line, or an integrator promising a hundred identical machines to a client, can't treat "Core Ultra 200V" as one part number. It's a family that can ship with different NPU tiers and different thermal headroom under the same name.

Even identical numbers from different vendors don't always mean the same thing. TOPS is usually quoted at a specific numerical precision, often INT8, because lower precision packs more operations into the same silicon. Intel is explicit that its NPU is rated at INT8 and also supports higher-precision FP16. Most vendors don't state which precision their headline number reflects. Two chips both rated "50 TOPS" aren't guaranteed to be measuring the same thing.

The Clearest Gap Between Rated and Real

Apple's M4 offers the most concrete example.

Apple rates the M4's Neural Engine at 38 TOPS. Apple's own materials say "38 trillion operations per second" and never call it an INT8 figure. That label comes from press and outside researchers.

Independent researchers bypassed Apple's software framework and measured the chip directly. They found that for general matrix math, the hardware converts INT8 numbers to a higher-precision format before computing. So it doesn't run general workloads at the rate the badge implies.

The measured throughput came out to about 19 TFLOPS. One caveat on that comparison: it relies on a standard industry convention that counts INT8 operations at twice the FP16 rate, rather than directly comparing two different kinds of arithmetic. On that basis, the measured figure lands at roughly half the rated one.

A second research paper confirmed the figure. A third found the same behavior across several generations of Apple silicon.

The story has a wrinkle, and it's worth keeping. Newer measurements from the same research found that on certain data-heavy workloads, the lower-precision format did produce a real 1.85 to 1.88x speedup. The researcher attributes much of that gain to less data moving inside the chip, not to the math itself running faster. And a separate test of chatbot-style text generation on the same hardware found no speedup at all.

So lower precision does something real. It moves smaller data faster, and sometimes that matters a lot. What it doesn't reliably deliver is the flat speed multiplier a TOPS rating suggests.

The Bottleneck That Isn't on the Spec Sheet

Here's the mechanism underneath much of this.

AI workloads aren't always limited by compute. Depending on the model, the batch size, and how well the software maps the work onto the chip, the limit can be compute, how fast data moves, on-chip memory space, or the software stack itself. Engineers who design accelerators know there's no single universal bottleneck.

But for many low-batch workloads, especially the chatbot-style text generation now arriving on laptops, data movement often runs out first. A fast compute engine can sit idle waiting for data that hasn't arrived yet.

Anyone who has worked on high-speed boards knows this principle well. Shorter, wider, cleaner paths move more data. Longer paths and shared, congested buses choke throughput, no matter how fast the engine at the end could theoretically go.

Intel's product literature supports the bandwidth half of this in its own words, with that "up to 2x bandwidth" line printed next to the TOPS figure. A more specific claim, that Intel also doubled the NPU's DMA capacity (the hardware that moves data around inside the chip), comes only from press accounts of a technical briefing. And one respected independent testing outlet measured that DMA hardware directly and found it weaker than the previous generation's. I can't fully reconcile that, so I'm flagging it rather than hiding it.

Either way, the direction is clear: a bigger compute block wasn't the whole answer. The path feeding it needed to widen too.

Heat: The Constraint Nobody Advertises

There's no single industry standard for how vendors measure peak TOPS. Precision, duration, and test conditions can all vary, so figures from different companies aren't guaranteed to be comparable.

What peak figures generally don't describe is a laptop under real load, with the CPU and GPU drawing on the same power budget and memory at the same time. A thin laptop has one thermal envelope shared by all three processors.

One study shows what sustained load can do, though on a different chip than you might expect. On an iPhone 16 Pro running a sustained AI text-generation task, throughput began dropping within two runs and settled about 44 percent below its peak.

Precision matters here: that test ran on the iPhone's GPU, not its Neural Engine. The software framework used doesn't touch the NPU at all. It's real evidence that a mobile chip can throttle hard under sustained AI load, not proof about NPUs specifically. And laptops, with larger chassis and higher power budgets, shouldn't be assumed to behave the same way.

No one has published an equivalent sustained-load measurement for the laptop NPUs covered here. The mechanism is well understood. The size of the drop on any specific laptop isn't public.

None of this makes TOPS a fake number. It makes it a theoretical peak capability metric, and one that doesn't tell you how much performance a real workload will sustain once memory traffic, software placement, power, and thermals enter the picture.

What to Look for Before You Buy

This isn't a reason to write off AI PCs. Forty TOPS is a real floor, and every current Copilot+ chipmaker has models that meet or exceed it. The mistake is reading the number as a promise about how the laptop will feel.

What actually decides that:

  • The specific feature you'll use, and whether it needs just the Copilot+ minimum or much more.

  • Whether a quoted number is NPU-only or a combined platform figure padded with GPU and CPU compute.

  • The exact processor model, not the family name on the box.

One limit on this argument: Microsoft certifies specific devices before they can carry the Copilot+ badge. I couldn't confirm whether that certification tests sustained performance or only checks the peak spec. If it tests sustained performance, some of the gap described here may already be closed on certified retail laptops, even though it still exists in the underlying silicon.

The engineering behind these chips is real, and it's impressive. But as NPU compute keeps climbing, the memory system, the software, and the thermal envelope increasingly decide how much of that theoretical speed reaches a real task. That's the specification worth watching, and it's the one that isn't printed on the box.


Sources: Microsoft (Windows 11 specifications, Copilot+ requirements), Intel (Core Ultra Mobile Processors Series 2 product brief, Panther Lake launch materials), Qualcomm (Snapdragon X Elite and X2 Elite product briefs), AMD (Ryzen AI 300/400 newsroom announcements), Apple Newsroom (M4, M5, and M6 chip introductions), plus independent research on Apple's Neural Engine (github.com/maderix/ANE and three related arXiv papers) and sustained on-device inference (arXiv 2603.23640). Full citations are in the original Silicon to Software article.


ABOUT THE AUTHOR / CTA:

About the author: Imran Valiani is a Sales Director in PCB electronics manufacturing with 20+ years of industry experience. He writes about the hardware layer of technology — semiconductors, PCB manufacturing, AI infrastructure, embedded systems, and emerging electronics — at Silicon to Software.

Read more engineering analysis at SiliconToSoftware.com

artificial intelligence

About the Creator

Imran Valiani

Engineering the hardware behind modern technology. Silicon to Software covers semiconductors, PCBs, AI infrastructure, embedded systems, and emerging hardware — with practical analysis from 20+ years in electronics manufacturing.

Enjoyed the story? Support the Creator.

Subscribe for free to receive all their stories in your feed.

Subscribe For Free

Reader insights

Comments

There are no comments for this story

Be the first to respond and start the conversation.

Sign in to comment
    Written by Imran Valiani