Bitget App
Trade smarter
Buy cryptoMarketsTradeFuturesEarnAISquareMore
Four Mac Studio systems run trillion-parameter models, Apple challenges Nvidia's narrative on inference

Four Mac Studio systems run trillion-parameter models, Apple challenges Nvidia's narrative on inference

华尔街见闻华尔街见闻2026/09/24 01:01
Show original

Apple's hardware chief pointed out that after enterprises purchase devices outright, there is no need to pay by token, and the cost-effectiveness of local computing power might be superior to long-term rental of cloud GPUs. Goldman Sachs believes that if more inference tasks shift to end devices, Nvidia's share in incremental AI activity may fall short of market expectations. Currently, Nvidia's implied volatility is at its lowest since before the pandemic, and option pricing has almost no risk premium allocated for this narrative change. Therefore, Goldman Sachs recommends buying Nvidia put options and going long on Apple.

Apple is reshaping the competitive landscape of the AI hardware market with a "local inference" logic, directly targeting NVIDIA's potential incremental share in the AI inference sector.

Apple's latest M5 Ultra Mac Studio, set to deliver from September 22, supports up to 512GB of unified memory and allows for distributed inference tasks in multi-machine clusters. Apple previously demonstrated four Mac Studio machines running a one-trillion-parameter model together.

Apple hardware chief Johny Srouji told Reuters that once enterprises buy the devices, they no longer have to pay per-token cloud fees for every inference request. For teams with stable workloads, the economics of local computing power may be superior to long-term cloud GPU rentals.

The significance of this narrative for NVIDIA is not about directly losing existing orders, but rather: if more inference tasks migrate to endpoint devices or enterprise-owned machines, the share NVIDIA can capture in incremental AI inference activities may fall short of current market expectations.

Currently, Apple's stock price remains technically strong, while NVIDIA continues to fluctuate within a wide range and lacks a clear mid-term trend. Some analysts recommend using options to express a bullish view on Apple and a bearish stance on NVIDIA.

Four Mac Studio systems run trillion-parameter models, Apple challenges Nvidia's narrative on inference image 0

Product Positioning Shift: From "Using AI" to "Running AI"

Apple's positioning of the Mac Studio has undergone a fundamental change, shifting from selling machines that "can use AI" to those that "can run AI".

The core of this change lies in the structural differences of inference costs: iPhone and Mac users have already paid for the chips, so locally processed tasks no longer require extra cloud inference requests—and this same principle is now being extended to enterprises.

The M5 Ultra Mac Studio supports up to 512GB of unified memory, with this configuration arriving in late October. Large memory enables users to load models that a single standard GPU cannot accommodate, while multi-machine clusters can share the weight of trillion-parameter models.

Apple previously demonstrated 4 Mac Studios running a trillion-parameter model, making local inference technically viable as a partial substitute for some cloud GPU clusters.

However, the real test beyond Apple’s demos will be the inference speed, concurrent user capacity, and total cost of ownership compared to cloud rentals—these three are the key variables in enterprise purchase decisions.

NVIDIA Remains in the Industry Chain, but Incremental Share is Uncertain

The Mac Studio will not replace data center GPUs for all workloads. Apple itself still uses NVIDIA GPUs on Google Cloud for certain high-demand AI tasks.

But NVIDIA does not need to lose its existing orders for its narrative to shift: once more inference tasks are run on endpoint devices or enterprise-owned machines, NVIDIA's share in incremental AI activities may fall short of market assumptions.

NVIDIA's stock lacks a mid-term trend, remains confined within a broad range and is trading near the top of that range. Meanwhile, implied volatility is at its lowest since before the pandemic.

Four Mac Studio systems run trillion-parameter models, Apple challenges Nvidia's narrative on inference image 1

Low volatility usually reflects high market confidence in future moves, but a potential shift in the inference narrative is an element of uncertainty. Current option pricing has left almost no risk premium for a “migration of inference to local”.

Four Mac Studio systems run trillion-parameter models, Apple challenges Nvidia's narrative on inference image 2

For this reason, Goldman Sachs believes buying NVIDIA put options before earnings expectations are revised downward is a low-cost way to express this shift; comparatively, they prefer calls on Apple versus puts on NVIDIA.

Analysts suggest monitoring real-world inference speeds and total cost of ownership data once the 512GB configurations arrive in late October, as well as enterprise feedback on local inference solutions.

0
0

Disclaimer: The content of this article solely reflects the author's opinion and does not represent the platform in any capacity. This article is not intended to serve as a reference for making investment decisions.

Understand the market, then trade.
Bitget offers one-stop trading for cryptocurrencies, stocks, and gold.
Trade now!

You may also like

US Treasury sell-off triggers massive waves in global bond markets! Global average yields approach the 4% threshold, Japan's 10-year hits a 30-year high

On Thursday, the global government bond sell-off intensified, pushing the average yield close to 4%, a level not seen since 2007.

智通财经2026/09/24 04:06

Honda (HMC.US) firmly pursues the "ditch electric for hybrid" strategy, plans to invest $2.5 billion to build a new hybrid car factory in the US

Honda plans to build a new hybrid vehicle manufacturing plant in Ohio, USA.

智通财经2026/09/24 03:41