Is Nvidia Eating the AI Stack?

August 12, 2026

Is Nvidia Eating the AI Stack?

The trillion-parameter Nemotron 4 report opens a debate every investment committee should be having right now.


Sponsored

First a note from Mode Mobile

Some companies you only hear about after they IPO.

And some…

Eventual unicorns like Uber, Airbnb and OpenAI…

Forced the world to pay attention long before that.

Mode Mobile could be a new member to that second group.

Uber turned cars into taxis, Airbnb turned homes into hotels, and Mode Mobile is turning smartphones into EarnPhones.

With $115M+ in revenue, 3-year growth of 32,481%, and an ecosystem with more than 490M+ users, it’s what investors call a “category disruptor.”

The kind that could turn early capital into generational wealth.

They’re raising privately.

For now.

But investors can get $0.52 pre-IPO shares before their share price changes on August 14.

With a Nasdaq ticker ($MODE) secured, and early backers like Kevin Harrington from Shark Tank, the company has its eyes on potentially going public.

Their previous two raises sold out, and this one is on track to do the same.

Review the offer before August 14.



Featured Article

Is Nvidia Eating the AI Stack?

Header image

The Big Question

For two years, the investment debate around Nvidia has been simple: how long does the hardware cycle last, and when does demand soften? That framing is now obsolete. On August 11, Nvidia shipped Nemotron 3.5 Lightning, a 30-billion-parameter open model built specifically for high-volume agentic workloads, alongside NeMo Switchyard, an open-source routing library that directs each step of an agent workflow to whichever model handles it most efficiently. The same day, The Information reported that Nvidia is separately building Nemotron 4, with a flagship version expected to exceed one trillion parameters, which would make it one of the largest open-weight models ever released.

The question investment committees are now asking is not whether the hardware cycle continues. It is whether Nvidia is quietly dismantling the revenue model of every closed-API vendor it helped create.

Why Wall Street Cares

Institutional investors have spent three years valuing OpenAI at $750 billion and Anthropic at $350 billion on the assumption that proprietary frontier models represent durable, defensible revenue streams. Anthropic’s revenue reached $4.5 billion in 2025 and grew 1,100% year over year. OpenAI generated roughly $13 billion over the same period. Both companies price access to their models through API subscriptions, and that subscription model depends on enterprises having no credible self-hosted alternative.

Nvidia has now released both a fast open model and an open routing library that together, per its own benchmarks, cut task completion costs to roughly one-third of running Anthropic’s Opus 4.8 on every step of an agentic workflow. The Nemotron 3.5 Lightning model can be downloaded, modified, and deployed commercially without seeking permission from Nvidia. Switchyard routes agent traffic across open and proprietary models without requiring developers to rewrite their applications.

Portfolio managers tracking the AI infrastructure theme now face a more complicated question than simple GPU demand. The company that supplies the picks and shovels is also building the software that could reduce how much enterprises spend on the services those picks and shovels were originally built to support. That conflict of interest, or strategic brilliance depending on your view, is the debate worth having.

The Bull Case

The bull argument starts with the observation that cheaper inference does not reduce AI spending. It accelerates it. Every time the cost of a model call drops, the number of viable use cases expands. Agentic AI applications that were uneconomical at $15 per million tokens become deployable at $5. The addressable market grows faster than margin compresses.

Nvidia’s open-model strategy reinforces the hardware flywheel rather than threatening it. The Nemotron 3.5 Lightning model is available free of charge for commercial use. Switchyard is open source. Neither is a direct revenue line item. The mechanism is indirect: open models that run best on Nvidia hardware increase demand for that hardware. Self-hosting an open-weight model at scale requires accelerated compute, and Nvidia supplies the dominant share of it. Enterprises adopting Switchyard are building Nvidia-optimized AI stacks by default.

The numbers ahead of August 26 support the bull reading. Nvidia guided Q2 FY2027 revenue at $91 billion, plus or minus 2%, with gross margins of approximately 75%. Consensus FY2027 EPS sits at $8.79, up 92% from FY2026. The average analyst price target, per Barchart, is $304.32, implying roughly 47% upside from current levels. The Nemotron 4 cloud-compute budget is capped at $7 billion through fiscal 2028, according to The Information, a sum that barely registers against Nvidia’s revenue base but could generate disproportionate strategic returns if it reshapes the model ecosystem.

The Bear Case

The bear case is less about the model quality and more about the corporate relationships Nvidia is now straining. The company is, simultaneously, the primary chip supplier to OpenAI and Anthropic, a proposed guarantor of up to $250 billion in OpenAI data center financing according to the Wall Street Journal, and now a direct competitor in the open-model layer that those same companies monetize. That circular dynamic was already drawing scrutiny from institutional investors before August 11. It is harder to ignore after.

Anthropic has not signed the open-weights letter that Nvidia organized in late July, which gathered 50 signatories including OpenAI, Google, Microsoft, and Meta. Amazon also declined. The fracture lines in the open versus closed debate are not technical. They are economic. Anthropic and Amazon are sitting out because open-weight proliferation threatens their subscription revenue and their cloud inference margins respectively. As Nvidia pushes further into the open layer, the question of whether its largest customers start hedging their compute exposure to AMD or building proprietary training clusters becomes a live risk rather than a theoretical one.

There is also a product risk in the Nemotron 4 timeline. The model remains in training, and Nvidia has not confirmed either the parameter count or a release date. Employees told The Information it could be ready as early as late fall, but that is not a commitment. A slip into mid-2027 reduces the strategic urgency of the open-model announcement and hands the initiative back to Meta’s Llama series and the Chinese labs.

The Evidence

The real-world deployment data from the Switchyard release is the most instructive data point. Ramp, the corporate spend management platform, cut costs 58% and runtime 33% by routing through Switchyard rather than calling a single frontier model for every task. Cognition reported a 28% mean cost reduction in its Devin Desktop deployment. LangChain benchmarked 145 multi-turn agentic tasks and achieved 74% lower cost by routing only 7% of calls to a frontier model, accepting a 6% accuracy tradeoff. Those are not internal Nvidia metrics. They are production outcomes from named enterprise customers who were already spending real money on closed-API calls.

The model itself warrants attention independent of the routing story. Nemotron 3.5 Lightning is a mixture-of-experts architecture with 30 billion total parameters but only 3 billion active per inference step, delivering up to 4x faster output speed than comparable dense models. The 1-million-token context window makes it viable for long-running agentic sessions that would exhaust smaller context models. CrowdStrike is fine-tuning it for cybersecurity workflows. Harvey, working with Trajectory, is adapting it for legal services. Lila Sciences is customizing it for life-sciences reasoning. The enterprise adoption velocity in the first 24 hours suggests this is not a model release that will sit on Hugging Face unchecked.

Zooming out, the open-source AI debate has become a US-China contest. Chinese labs including DeepSeek, Alibaba’s Qwen family, and Moonshot AI’s Kimi have established themselves as credible open-weight competitors. Nvidia, as one of the few major US firms releasing open-source models, is positioning Nemotron as a domestic counterweight to that momentum. The geopolitical angle gives the open-model strategy a rationale that extends beyond quarterly revenue.

Sponsored

Bigger than Nvidia? Louis Navellier thinks so.

In 2016, Louis Navellier recommended Nvidia at $2.51 – split-adjusted. It went up 44,000%. He also called Apple before a 36,000% rise and Microsoft before a 60,800% climb. Now he says a new AI device coming online in Tennessee is the setup for the biggest call of his career.

He’s agreed to reveal the stock at the center of it – down to the ticker – for free.

The Mavens’ View

Sophisticated investors appear to be thinking about this in two layers, and the gap between those layers is where the interesting positioning lives.

The first layer is the pre-August 26 trade: does Nvidia beat $91 billion in revenue, hold margins near 75%, and guide Q3 above the current analyst estimate of 81% year-over-year growth? That is the conversation most active money is having. It is also largely priced in. Nvidia has beaten earnings estimates in four consecutive quarters, and the stock fell after the Q1 FY2027 report in May despite 85% revenue growth and a beat on both lines. When 43 of 47 analysts covering a stock rate it a strong buy, the risk is no longer missing estimates. It is meeting them and getting sold.

The second layer is the structural question that earnings cannot answer in a single quarter: is Nvidia building a software moat around its hardware advantage that will compound over the next three to five years, or is it overextending into a model race it cannot win against labs that do nothing else? The Mavens who have worked through this tend to conclude that Nvidia does not need to win the model race. It needs to be embedded at the routing layer, where every enterprise AI workflow passes through infrastructure it controls. Switchyard is that play. Nemotron 4 is the credibility signal that makes Switchyard’s open-model promise defensible.

Kari Briski, Nvidia’s VP of generative AI, put the strategic intent plainly: “every company and every country needs accessible frontier open models to strengthen safety and security, accelerate innovation, and provide a foundation they can rely on from one generation to the next.” That is not a product positioning statement. It is a description of where Nvidia believes the durable economic value in AI will accrue: not in any single model, but in the infrastructure layer beneath all of them.

What Investors Are Missing

The conversation has focused almost entirely on whether open models threaten closed-API revenue. That framing misses a more consequential second-order effect: Switchyard creates switching costs that do not exist in the current model ecosystem.

Today, an enterprise using Anthropic’s Claude or OpenAI’s GPT-5.6 can switch providers by changing an API key. The switching cost is low by design. Switchyard changes that calculus. An enterprise that adopts Switchyard as its routing layer, trains its cost and latency parameters against Nemotron 3.5 Lightning’s performance characteristics, and integrates that routing into its production agent workflows has built an Nvidia-optimized stack. Swapping out the routing layer means re-tuning every cost and latency tradeoff, reworking the gateway, and likely re-evaluating the hardware it runs on. That is not a trivial migration.

This is the dynamic that closed-model vendors should fear most, and it is the dynamic that markets have not yet priced. Anthropic can compete on model quality. It cannot easily compete on infrastructure lock-in with a company that sells the hardware, the inference framework, the routing library, and the open model. The question for the next 12 months is whether enterprise adoption of Switchyard is fast enough to make that lock-in measurable before the next competitive model cycle resets the conversation.

Stocks to Watch

Nvidia (NVDA). The obvious entry, but not for the obvious reason. The August 26 earnings report matters, and the guidance will move the stock. But the more important question is what adoption velocity for Switchyard looks like by the Q3 FY2027 report in November. If enterprise adoption of the routing layer is accelerating, the software moat thesis gets incrementally more credible with each quarter, independent of what any single model benchmark shows. The stock has returned roughly 2% year to date against a 7% S&P 500 gain, which means the multiple has been compressing even as the revenue line has grown at 85%. That compression resolves one way or the other on August 26.

Microsoft (MSFT). The most exposed hyperscaler to the open-model shift, in both directions. Microsoft’s Azure hosts OpenAI’s models and collects inference revenue tied to closed-API usage. It also signed the open-weights letter and has deep NIM integration across its Azure AI portfolio. If Switchyard adoption grows inside enterprise Azure deployments, Microsoft collects the compute revenue either way. The risk is that Switchyard accelerates on-premises deployments on Nvidia DGX hardware rather than Azure-hosted inference, reducing Microsoft’s share of the workflow. Watch Azure AI revenue growth in the October earnings report for early evidence.

AMD (AMD). Nvidia’s open-model push deepens the software moat in a way that raw GPU performance cannot overcome. AMD’s MI-series chips are competitive on compute per dollar in several inference benchmarks, but they do not run Switchyard or NIM natively. If enterprise AI stacks become increasingly tuned to Nvidia’s software layer, AMD’s path to meaningful inference market share narrows regardless of what its next-generation silicon delivers. This is the highest-risk position in the semiconductor peer group relative to today’s announcements.

Ramp and Cognition (private). Both companies are named design partners on the Switchyard release, and both reported material cost reductions in production deployments. As the open-weight adoption story develops, these companies will become case studies that enterprise procurement teams reference. Their vendor relationships with Nvidia are now on record, which matters when their next funding rounds are priced.

Anthropic (pre-IPO). The company is reportedly preparing for a public offering that could target a valuation between $800 billion and $1 trillion, according to reporting cited in prior coverage. The open-model competitive pressure from Nemotron changes the revenue assumptions that would support that range. Anthropic’s annual recurring revenue was $3.6 billion in early 2026. Its bull-case projection of $70 billion by 2028 depends on enterprises continuing to pay closed-API prices as the default. Switchyard and Nemotron 4 are the clearest structural threats to that assumption, and they both arrived in the same week. Investors considering the IPO should build that into their probability-weighted scenarios.


Wall St. Mavens