NVIDIA AI Factory Compute Is Becoming an Investable Asset Class

We announced partnerships with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR to establish independent financing platforms designed to mobilize over $500 billion of third-party capital to support the buildout of AI infrastructure over time.

This is a major milestone for NVIDIA and the AI industry. We have moved from an era in which companies bought chips and built data centers project by project to one in which AI factories can be financed as productive infrastructure — with repeatable platforms, long-term institutional capital and a diverse customer base that uses compute to create revenue.

AI has reached an inflection point. It is moving from research into production. AI is creating real value, and the infrastructure behind it is becoming one of the world’s most productive assets. In AI, compute is revenue.

A New Infrastructure Asset

NVIDIA compute is not just a chip. It is a complete AI factory platform including accelerated computing, networking, systems software, AI frameworks and a global developer ecosystem.

NVIDIA DSX AI factories can run the world’s broadest range of AI models, modalities and algorithms — language, vision, speech, biology, physical AI and robotics. One NVIDIA AI factory can serve many customers and many workloads. That makes it flexible and fungible.

It is also built on a globally adopted architecture used across every major cloud, and by systems makers and enterprises around the world. When needs change, the factory can be used by another customer, another cloud or another operator. This broad ecosystem gives NVIDIA compute a deep market of potential users and offtakers, helping protect residual value.

CUDA makes the factory better over time. Every generation of NVIDIA software improves the performance, efficiency and total cost of ownership of already- installed infrastructure. The hardware does not stand still: software innovation allows an AI factory to produce more intelligence at lower cost throughout its life, extending its useful economic value.

NVIDIA A100 is a powerful example. NVIDIA introduced the Ampere-based A100 in 2020, and six years later, it remains in active commercial use for AI training, fine-tuning, inference and high-performance computing. Customers continue to commit capacity for multi-year deployments, extending A100’s economic life toward a decade.

The market is also demonstrating the durability of NVIDIA compute economics. One-year H100 rental pricing rose from about $1.70 per GPU-hour in October 2025 to about $2.35 per GPU-hour in March 2026. Cross-provider on-demand median pricing rose from roughly $2.00 per GPU-hour in October 2025 to $2.70 in June 2026. Blackwell capacity commands a premium, with reported B200 cloud rates spanning approximately $5.30 to $7.05 per GPU-hour.

That is what makes NVIDIA AI factories different. Their value is not fixed at installation: CUDA continuously improves their output; the installed base remains productive well beyond its initial depreciation period; and the same standard architecture serves a deep, growing global market of AI workloads.

These are the characteristics of an investable infrastructure asset: it produces revenue, serves a broad market, improves in performance over time and can be redeployed.

Bringing Capital to AI Factories

The demand for AI infrastructure is extraordinary. But access to capital is uneven. Many great AI companies, enterprises and AI clouds have demand for compute but do not yet have access to financing at the scale or cost required to build quickly.

That is why we are partnering with the world’s leading long-term capital providers.

Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR are also among the world’s leading infrastructure investors, with deep expertise in underwriting long-lived, productive assets. Together, we are creating repeatable financing platforms to help the AI ecosystem build the factories it needs.

The platforms are designed to help qualified AI labs, enterprises and AI clouds access AI-factory infrastructure at scale. The more than $500 billion figure represents aggregate third-party capital that these platforms are designed to mobilize over time — the capital is not NVIDIA revenue, a single fund or a commitment to a single customer.

The financial institutions will independently assess each opportunity — the customer, demand, utilization, cash flow and residual value. NVIDIA provides the AI factory platform. The financial institutions provide long-term capital and financing expertise.

The Important Questions

Is this circular financing?

This initiative is designed to address that concern. We are bringing independent, long-term institutional capital into the AI infrastructure market.

The demand is real: it comes from frontier AI labs, AI-native startups, enterprises, cloud providers and countries building AI services. The capital providers independently underwrite each project — including the customer, demand, utilization, cash flow and residual value. NVIDIA provides the platform; the investors make independent financing decisions.

This is the beginning of an open capital market for AI infrastructure.

Why would NVIDIA support financing?

In some cases, NVIDIA may provide a residual-value support mechanism for up to 25% of an opportunity, assessed carefully on a project-by-project basis. That support is limited, residual-value based and designed to complement — not replace — independent underwriting.

This is substantially lower than other compute-financing arrangements. NVIDIA can provide support because NVIDIA compute is unique: it is fungible, universally adopted, software-upgradable and redeployable across a large ecosystem of customers.

Our role is to help unlock a very large pool of independent capital while maintaining disciplined risk exposure.

Can the market absorb this capacity?

The question is not whether we are building data centers. The question is whether we are building productive AI factories.

An AI factory turns energy and data into valuable intelligence. Its customers are broad: frontier AI labs, AI clouds, enterprises and nations. They are building AI because it has become useful — doing valuable work across every industry.

There is discipline in the model. Each financing partner will independently evaluate demand, utilization, cash flow and residual value. Capacity will be built around real customer economics.

Where is the return on investment?

The return is in the usefulness of AI.

Companies are using AI to write software, discover drugs, design products, serve customers, automate operations and build new services. AI factories make this possible. More compute creates better AI; better AI creates more usage; more usage creates more revenue; and more revenue drives more compute.

This is the virtuous cycle of the AI industrial revolution.

The Infrastructure of Intelligence

Every industrial revolution has been built on infrastructure: electricity, transportation, communications and computing, with every buildout enabled by external financing.

AI factories are the infrastructure of the intelligence era.

With these partnerships, NVIDIA and the world’s leading financial institutions are creating a new way to finance the infrastructure that will power this industrial revolution. We will make AI factories more accessible to the companies, industries and nations building the future.

The age of AI is here. Together, we will build the infrastructure to power it.

Why Scaling AI Compute Performance Requires a New Power Architecture

Every new generation of accelerated computing demands more from the infrastructure underneath it — more compute performance, higher rack density and more efficient, scalable power distribution. The bottleneck isn’t just wattage. It’s how power gets from the grid to the GPU. 

In traditional power delivery, electricity travels from the grid as an alternating current (AC) and gets converted multiple times, each time adding overhead and complexity as racks become denser. At the power levels that next-generation AI compute demands, even small inefficiencies compound quickly. 

800 VDC simplifies that path. By distributing power at higher voltage through a direct current (DC), fewer conversion stages stand between the grid and the accelerator — which means more of the available power reaches the compute. NVIDIA DSX reference designs are built to guide AI factories through the transition from today’s AC infrastructure through hybrid architectures and into fully native 800 VDC facilities. 

NVIDIA, Google and Microsoft have been developing the 800 VDC architecture together through the Open Compute Project (OCP), and published a joint white paper March 2026 and the LVDC Solid-State Transformer Specification v0.3 July 2026. More than 80 equipment manufacturers and infrastructure companies are already building products to this specification. 

Existing Facilities Don’t Have to Wait 

Most of today’s AI factories were designed around AC distribution. The NVIDIA MGX-compatible 800 VDC power rack, arriving in the second half of 2026, creates a hybrid architecture that brings next-generation rack-scale compute performance to facilities that are already built and operational. It’s designed to slot into existing AC infrastructure and deliver 800 VDC to compute racks within the row — no changes to the building’s electrical system required. 

“800 VDC unlocks the compute performance and power density required for AI at scale,” said Vladimir Troy, vice president of data center infrastructure at NVIDIA. “Through OCP, NVIDIA is working with more than 80 ecosystem companies to give AI factories a practical path forward — not just a future vision.” 

For site owners, this matters because the investments already made don’t have to be stranded. Land, power rights, building infrastructure — the hybrid approach preserves all of it while opening the door to higher compute density. 

A Roadmap for Every Stage of Growth 

800 VDC provides AI factories a roadmap to scale, with on-ramps at every stage of growth. 

The power rack is where most operators will start — a near-term, hybrid-compatible path into higher-density compute. For operators building out dedicated AI factory environments, the row power center — a centralized power station for a full rack row — uses an overhead 800 VDC busway to scale power distribution across multiple rack rows, supporting up to 2 megawatts per row, with availability expected in 2027. And for new facilities being designed, the DC power block — a facility-scale unit that converts grid power directly to 800 VDC in a single step — will enable direct medium-voltage conversion at massive scale: the architecture for AI infrastructure being planned for the decade ahead. 

An Open Standard Means a Real Supply Chain 

The 800 VDC architecture specifications define common interfaces so that power hardware from different vendors can work together inside the same 800 VDC facility. With 80+ companies building to the specifications, the supply chain is forming around an open standard. 

These building blocks are captured in NVIDIA DSX reference designs, giving operators a system-level blueprint to connect power architecture, rack-scale computing and facility infrastructure as they scale AI factories.

Wood Mackenzie projects $9 trillion in global AI and data infrastructure investment through 2040. The facilities that can absorb that investment will be the ones that resolved their power architecture before compute demand outran what their infrastructure could deliver. NVIDIA, Google and Microsoft are working with the broader ecosystem to make sure 800 VDC is ready when operators need it — and that existing facilities have a path to get there now. 

Read the 800 VDC OCP blog and the NVIDIA 800 VDC white paper. Learn more about the NVIDIA DSX reference architecture guide. 

NVIDIA and Local AI Community Fuel Open Source Models and Intelligent Agents

The open source ecosystem is making it easier for AI enthusiasts and developers to build, customize and run increasingly capable agents locally. 

Throughout August, NVIDIA is celebrating the partners and open source communities moving local AI forward, along with the models, applications and tools emerging across the ecosystem. That includes NVIDIA’s latest open models, software and developer tools, plus the accelerated computing, libraries and educational resources that help users get started.

It’s shaping up to be a big month for agents. Follow along for the latest developments in this special-edition NVIDIA Local AI blog series, with new updates added over the coming weeks.

Follow NVIDIA RTX Spark on X, Instagram, TikTok and Facebook — and stay informed by subscribing to the RTX AI PC newsletter. Follow NVIDIA Workstation on LinkedIn and X


Tuesday, Aug. 11, 6:00 a.m. PT 🔗

NVIDIA Introduces Nemotron 3.5 Lightning for Fast, Specialized Agentic Tasks

Today, NVIDIA expanded its Nemotron 3 model family with Nemotron 3.5 Lightning, a customizable open 30B mixture-of-experts (MoE) model for always-on agents. 

Nemotron 3.5 Lightning delivers up to 4x faster token generation and 30% faster time to completion compared to open models in its class. 

And because Nemotron 3.5 Lightning is open weights, AI enthusiasts and developers can fine-tune it with their own examples to better match specific tasks, interests and workflows. For example, they could train the model to: 

Paired with access to apps, files and other tools, these fine-tuned models can power more personalized local agentic AI experiences — from an assistant that helps manage email and calendars, to a smart-home agent that handles everyday routines, to a coding companion that works alongside developers on a local codebase.

NVIDIA collaborated with vLLM, Ollama, llama.cpp and LM Studio to provide the best local deployment experience for Nemotron 3.5 Lightning models — offering developers choice of NVFP4 and GGUF format of models. Unsloth also provides day-one support with optimized and quantized models for efficient local deployment via Unsloth Studio. 

Nemotron 3.5 Lightning runs locally on NVIDIA RTX PCs, NVIDIA DGX Spark and OEM GB10 systems, and NVIDIA Jetson, and scales up to RTX PRO workstations, NVIDIA DGX Station and GB300 deskside systems, data centers and cloud environments. With NVIDIA Blackwell systems available from Acer, ASUS, Dell Technologies, Exxact, GIGABYTE, HP, Lenovo, MSI and Supermicro, users can choose from a wide range of devices and form factors to fit their needs.

As generative AI adoption grows, enterprises are looking for ways to keep rising token costs in check without sacrificing access to frontier intelligence. NVIDIA NeMo Switchyard, an open source routing library, automatically directs each step of an agent workflow to the best-fit model based on accuracy, speed and cost. It also gives developers the flexibility to work across models and providers for different tasks. 

Internal benchmarks show that NeMo Switchyard, by routing each step across a system of models, helped maintain frontier-level task completion while reducing benchmark completion cost to roughly one-third of Opus 4.8 alone. NeMo Switchyard is available on GitHub.

Visit the Nemotron 3.5 Lightning, NeMo Switchyard and Jetson AI technical blogs to get started. And to build at the edge, start with Jetson AI Lab tutorials and discover real-world Jetson projects. Nemotron 3.5 Lightning is also available through OpenRouter, on build.nvidia.com as an NVIDIA NIM microservice, and through a broad ecosystem of NVIDIA Cloud Partners, post-training platforms, inference platforms and cloud service providers. NeMo Switchyard is available on GitHub.

NVIDIA Nemotron 3.5 Lightning and NeMo Switchyard Deliver Faster, Smarter, More Efficient Agentic AI

As AI shifts from chatbots to autonomous agents, open models are serving market demands for full control over where AI runs and how it’s deployed and evolves.

Today, NVIDIA is expanding its Nemotron 3 model family with Nemotron 3.5 Lightning, the highest-efficiency model in its class for long-running agentic AI workloads. This release follows Nemotron 3 Nano and reflects NVIDIA’s commitment to continually improving open models for greater accuracy and speed. 

Built for specialized tasks within larger multi-agent systems, Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts model, helps create smarter and more efficient agentic applications.

Also, NVIDIA is releasing NeMo Switchyard, an open source library for smart routing inside popular agent tools. Enterprises can use it to build a router based on their specific needs. When deployed, NeMo Switchyard can intelligently direct each request to the most capable and suitable model for the job without requiring developers to rewrite their applications.

Together, Nemotron 3.5 Lightning and NeMo Switchyard deliver greater control over how AI is deployed, where it runs and how efficiently it operates — across PCs, workstations, data centers and the cloud.

Nemotron 3.5 Lightning delivers frontier-level intelligence in a small, customizable open model built for high-volume agentic workflows.

Always-On Agents Need a System of Models 

Modern agentic systems — always-on agents — increasingly operate as systems of models, or model ensembles, with different models specialized for different tasks. 

NVIDIA Nemotron open models are designed for this architecture. A frontier reasoning model such as Nemotron 3 Ultra or GPT-5.6 may plan and orchestrate a workflow, while smaller specialized models like Nemotron 3.5 Lightning can perform targeted tasks such as code review, tool use, security alert monitoring and answering billing questions.

Powering High-Volume Specialized Tasks With Nemotron 3.5 Lightning

NVIDIA Nemotron 3.5 Lightning is a fully customizable open model built for high-volume tasks powering always-on agents. It was developed with contributions from the Nemotron Coalition, whose members provided evaluation methodologies, inference software and datasets to help advance the model.

The model delivers up to 4x faster output speed, leading to 30% faster agentic task completion compared with other models in its class. And because it’s open and customizable, Nemotron 3.5 Lightning can be easily post-trained with NVIDIA NeMo on an organization’s own domain data, tools and workflows to improve accuracy for specialized tasks.

PinchBench benchmarks demonstrate that Nemotron 3.5 Lightning delivers faster agentic task completion with frontier-level accuracy compared to other models in its class.

AI leaders across industries are customizing Nemotron 3.5 Lightning for their workloads, including CrowdStrike for cybersecurity, Harvey with Trajectory for legal services and CodeRabbit with Baseten for code review, helping improve accuracy for domain-specific agentic tasks. Additionally, Lila Sciences is helping to improve reasoning capabilities for agentic tasks across physical and life sciences, and Fastino Labs customized the model and is seeing leading accuracies for software development, finance and healthcare workloads. 

Enterprises have customized Nemotron 3.5 Lightning to achieve leading accuracy for their specialized task in their agentic workflows.

Nemotron 3.5 Lightning also gives organizations control over privacy and deployment. It can run on local AI systems — including NVIDIA RTX PCs, NVIDIA DGX Spark, NVIDIA DGX Station and NVIDIA Jetson — to help users maximize existing infrastructure investments, or scale across edge AI devices, NVIDIA RTX PRO workstations, data centers and cloud environments for enterprise use cases. And Nemotron 3.5 Lightning can run locally or on premises for high-volume, specialized tasks that require fast responses.

Also, as with every Nemotron launch, NVIDIA publishes as much of the training data and techniques as licensing permits, which allows for traceability, auditing and training of other models. Alongside Lightning, NVIDIA is releasing Nemotron-RL-Agentic-Terminal-Pivot, an agentic reinforcement learning dataset used to post-train it for coding agent capabilities.

More Efficient AI Apps With Model Routing 

Some models are better for coding, some for reasoning, some for lightweight tasks and some are optimized to run locally for greater privacy and efficiency. If customers rely on one default model, they might either overspend or lose quality; if they manage routing manually, it becomes integration work that can slow down a deployment.

NVIDIA NeMo Switchyard is an open source model routing library for AI agents. The technology routes prompts to the most capable and efficient model for each step of an agent workflow automatically, based on specific needs. Agent application developers can tune or modify the router with different routing algorithms to match their priorities, such as quality, latency and cost requirements. In a system of models, enterprises can create powerful AI agents with improved tokenomics. 

Internal benchmarks show that NeMo Switchyard maintains frontier-level accuracy while reducing task completion cost to nearly one-third of Opus 4.8 alone.

NVIDIA internal benchmarks show that NeMo Switchyard maintains frontier-level accuracy while reducing task completion cost to nearly one-third of Opus 4.8 alone.

NVIDIA is working with partners across the AI ecosystem to bring intelligent model routing into the tools and platforms developers already use. 

Nemotron 3.5 Lightning is available on Hugging Face, ModelScope, OpenRouter and build.nvidia.com as an NVIDIA NIM microservice as well as through a broad ecosystem of NVIDIA Cloud Partners, post-training platforms, inference platforms and cloud service providers. NeMo Switchyard is available on GitHub and coming to partner platforms soon. 

Firebird Launches CIS Region’s Largest AI Factory in Armenia

The global buildout of AI infrastructure reached a new milestone today — Firebird, an emerging AI cloud, launched the CIS region’s largest AI factory in Armenia, establishing a new AI computing hub powered by NVIDIA accelerated computing and Dell Technologies high-performance AI infrastructure

Nikol Pashinyan, prime minister of the Republic of Armenia; Zhaslan Madiyev, deputy prime minister of the Republic of Kazakhstan; and David Allen, U.S. chargé d`affaires, a.i. in Armenia, attended the AI factory opening ceremony. 

AI factories are the foundational infrastructure for the AI era, providing the computing capacity needed to train, fine-tune and deploy AI models for every domain at scale.

Building the Infrastructure to Create Intelligence at Home

While AI services are available globally, countries also need the capacity to develop and run AI for their own languages, industries and national priorities. Firebird’s AI factory brings that capacity to Armenia, giving developers, startups, enterprises, universities and public institutions the compute to build and scale AI at home.

Firebird plans to deploy more than 70,000 NVIDIA Rubin and Blackwell GPUs and 300 megawatts of AI infrastructure capacity in Armenia by the end of 2027, accelerating the country’s development as a center for AI research, advanced computing and innovation.

At this scale, energy efficiency is essential. Built on the NVIDIA DSX platform, this AI factory integrates accelerated computing, networking, power and cooling as one codesigned system. Firebird’s AI factory is designed from the ground up to turn compute into revenue. With DSX, it can run up to 40% more GPUs on the same footprint, producing more tokens per dollar and extracting more value from every megawatt of capacity.

A Magnet for Global AI, a Catalyst for Local Innovation

Firebird’s ambitions extend beyond a single site. With NVIDIA’s support, the company is pursuing an approximately 2-gigawatt AI infrastructure roadmap spanning Armenia, Kazakhstan and additional markets. 

Firebird also announced that NVIDIA intends to invest in the company, following an earlier investment by CoreWeave this year. These investments will help Firebird expand its global infrastructure and operational footprint, and support its efforts to establish the largest and most advanced compute clusters across frontier markets.

Delivered in just over six months, the Armenia AI factory demonstrates Firebird’s ability to turn ambitious infrastructure plans into operational AI capacity with exceptional speed.

Schneider Electric provides the power infrastructure supporting Firebird’s AI factory in Hrazdan, helping Firebird meet its accelerated deployment schedule by rapidly delivering and setting up critical systems, including medium- and low-voltage switchgear, three-phase uninterruptible power supply systems and rack enclosures. This keeps the power buildout moving at the pace of the compute and provides a reliable foundation to bring NVIDIA accelerated computing online at scale.

To support the facility’s thermal needs, Vertiv provided a cooling architecture combining chilled-water technology, advanced controls and Vertiv TrimCooler technology for efficient heat rejection. Vertiv’s iCOM CWM Chilled Water Manager centrally coordinates cooling resources, improving visibility, efficiency and responsiveness as demand shifts with AI workloads.

Early demand is coming from AI-native companies including Perplexity, which is working with Firebird to access high-performance AI infrastructure for its AI agent platform and answer engine. 

As AI becomes essential infrastructure worldwide, Firebird’s expansion can help make the CIS region a magnet for global companies building and running AI — and a catalyst for local developers, researchers and enterprises. 

Powered by NVIDIA’s total AI factory platform — reference architecture, accelerated computing, networking and AI software — and deployed on Dell PowerEdge servers, the new Firebird AI factory will help Armenia’s builders turn energy into intelligence and connect their innovations to the global AI economy.

 

 

GeForce NOW Shakes Up August With 26 New Games

August is here, bringing 26 new games for GeForce NOW members. 

Command the seas in World of Warships: Legends and discover what’s next in the GeForce NOW library, starting with the eight newly added games this week. 

In addition, GeForce NOW is at the QuakeCon gaming conference this week in Grapevine, Texas, with hands-on experiences awaiting attendees.

Return to QuakeCon 

Visit the NVIDIA booth at QuakeCon to experience GeForce RTX 5080-powered Ultimate cloud gaming in action, with demos showcasing stunning visuals at up to 5K 120 frames per second on an ultrawide display, as well as seamless gameplay on the Lenovo Legion Go S handheld device.

Conference attendees can see how thousands of PC games, including fan-favorite Bethesda titles, can move effortlessly across laptops, Macs, handhelds, mobile devices, TVs and more — letting members pick up where they left off on nearly any supported screen.

Gamers not at the show can try out Ultimate cloud gaming in action with a day pass and jump into Bethesda titles from any device.

All Games on Deck

World of Warships Legends on GeForce NOW
Sail the seas from nearly any screen.

World of Warships: Legends drops anchor on GeForce NOW this week, bringing free-to-play naval combat. Captains can command destroyers, cruisers and battleships across massive multiplayer battles while exploring the latest update, featuring the Pacific Hammer Campaign, a new line of U.S. destroyers and more.

Chart a course straight into the latest content across devices today without any installs needed, and check out all the games available this week:

And look forward to the games coming throughout the month:

Extra Joy From July

In addition to the dozen games announced last month, 15 more joined the GeForce NOW library. 

Mistfall Hunter didn’t make it this month. Stay tuned to GFN Thursday for the latest updates.

What are you planning to play this weekend? Let us know on X or in the comments below.

Into the Omniverse: How Open World Models Push the Frontier of Physical AI

Editor’s note: This post is part of Into the Omniverse, a series focused on how developers, 3D practitioners and enterprises can transform their workflows using the latest advancements in OpenUSD and NVIDIA Omniverse.

In July, NVIDIA joined more than 200 companies and organizations in signing “Open Weights and American AI Leadership,” an open letter arguing that AI leadership will be measured not by any single frontier model but by whether an open ecosystem reaches every sector. 

Open models, which anyone can download, inspect, modify and run on their own infrastructure, are what make that possible. Nowhere is that more crucial than in physical AI, where every deployment is a specialization problem.

Physical AI has to understand and predict consequences, not just appearances. 

To make this possible, world models learn how physical environments behave, what may happen next and which following actions make sense. They can generate physically grounded world and action data, simulate future states and provide a foundation that teams can specialize for a robot, autonomous vehicle or vision AI system.

Open world models are already being used to generate training data, test policies and specialize physical AI systems. NVIDIA Cosmos 3 brings these capabilities together in an open model family, with leading benchmark results and adoption across robotics, autonomous vehicles and vision AI.

And NVIDIA Omniverse libraries, part of NVIDIA Agent Toolkit, provides prebuilt capabilities for building simulation-ready worlds that physical AI teams can use to train, test and validate systems before real-world deployment.

World Models Are the Foundation of Physical AI

The data behind physical AI is difficult and expensive to collect at the scale required. Rare events and long-tail scenarios can be especially difficult to reproduce safely and repeatedly. 

World models enable:

 

A general model hasn’t seen a team’s particular robot, sensors or operating environment. Closing that gap requires access to model weights, a license that permits adaptation and the tools needed for post-training. 

NVIDIA Cosmos world foundation models are available under the Linux Foundation’s OpenMDW 1.1 license, enabling teams to post-train models on their own data and hardware. Specialization is where openness becomes a practical technical requirement.

Specializing a model is only part of the workflow. Teams also need environments to generate data, run simulations and test behavior. 

Omniverse libraries help developers build simulation-ready environments, while OpenUSD provides the open framework for composing, reusing and exchanging complex 3D data across digital twins, simulations and synthetic data generation workflows. Together, Omniverse and OpenUSD cut the duplicated work that can otherwise pile up every time assets, sensor configurations or environmental conditions change.

Cosmos 3: The Frontier Model

NVIDIA Cosmos 3 — a frontier open physical AI foundation omni-model built on a mixture-of-transformers architecture — combines vision reasoning, world generation and action prediction, letting developers use one model family to understand scenes, generate synthetic data, simulate future states and build specialized world action models.

Developers can use Cosmos 3 as a vision language model, as a physics-grounded world simulator that predicts future world states and generates large-scale synthetic data, or as the backbone for world action models, instead of assembling and maintaining a separate model for each capability.

The family includes Cosmos 3 Super (64B) for high-fidelity world modeling, Cosmos 3 Nano (16B) for efficient reasoning and post-training, and Cosmos 3 Edge (4B) for on-device vision reasoning and robot policy deployment. Lightweight enough to run on edge GPUs, Cosmos 3 Edge can be deployed across NVIDIA RTX GPUs, NVIDIA DGX systems and NVIDIA Jetson, including Jetson Thor platforms.

Across benchmark evaluations, Cosmos 3 ranks No. 1 on Artificial Analysis for open weights text-to-image and image-to-video generation, on PAI-Bench for world generation and in the image-to-video category of Physics-IQ. For robot policy, it ranks No. 1 on RoboLab. Cosmos 3 Super is also the highest-ranked open model on VANTAGE-Bench for vision understanding.

In addition to Cosmos, NVIDIA’s physical AI stack includes Isaac GR00T for robotics, Alpamayo for autonomous vehicles and Metropolis for vision AI. 

How Developers Are Putting Cosmos 3 to Work

Across industries, developers are building on NVIDIA Cosmos for physical AI applications: Doosan Robotics, LG Electronics, Samsung Electronics and Skild AI in robotics; Li Auto, Xiaomi and Afari in autonomous vehicles; and Centific, Fogsphere, Linker Vision, Milestone Systems and Yuan for vision AI agents powering industrial AI and smart spaces applications.

The NVIDIA Cosmos Coalition extends this work by bringing together world model builders, AI developers and physical AI leaders to contribute models, research and evaluation methods. NVIDIA recently expanded the coalition to Japan, where robotics and manufacturing leaders intend to join and develop open world models for factories, logistics, agriculture, construction, healthcare and transportation.

Together, these implementations and collaborations are establishing open world models as an adaptable foundation for physical AI across robots, autonomous vehicles and vision AI systems.

Get Plugged In

Learn more about world models, OpenUSD and physical AI development by exploring these resources:

NVIDIA Joins NSF State and Regional AI Hubs Program to Expand AI Research and Education Across the US

NVIDIA is participating in the U.S. National Science Foundation’s (NSF) State and Regional Artificial Intelligence Infrastructure Hubs program, an effort launching today to expand access to the advanced computing, data, software and expertise needed for AI-enabled research and education.

Consistent with the aims of the Genesis Mission, the program will support state and multistate groups of colleges and universities working together to strengthen America’s AI ecosystem. In partnership with private industry, philanthropic organizations and state and local governments, the program will expand the AI infrastructure, software, educational resources and technical support needed by faculty, students and researchers across the country. 

These regional hubs will help institutions share AI computing resources, accelerate scientific discovery and innovation, and prepare students to participate in the AI economy.

Expanding Access to AI Infrastructure

The State and Regional AI Infrastructure Hubs program will bring shared resources closer to the institutions and communities they serve. 

State or regional consortia can pool expertise, focus on specific local priorities, achieve economies of scale and create pathways for institutions that might otherwise remain outside the frontier of AI-enabled research and education. Flexible approaches — including on-premises infrastructure, cloud computing or a combination — will allow consortia to design resources around their regional needs and economic priorities. 

The hubs will resemble the public-private partnership between NVIDIA, NVIDIA cofounder Chris Malachowsky and the University of Florida (UF) in 2020 to turn UF into the country’s first true AI university and provide AI compute access to all Florida public universities. That initiative now serves as a national model. Since launching its university-wide initiative in 2020, UF has grown to more than 300 AI-focused faculty and embedded AI education and research across all 16 colleges. And since 2017, UF faculty and units have received more than $511 million in AI research awards.

NVIDIA has also expanded academic compute access in other ways, including as a leading contributor to the NSF-led National Artificial Intelligence Research Resource (NAIRR) pilot program on which today’s announcement is built. 

Through NAIRR, NVIDIA partnered with university research teams across the country to turn computing resources into usable scientific capacity — giving researchers the infrastructure, tools and expertise needed to move from idea to experiment to discovery. The resources also facilitated meaningful educational opportunities that gave students critical real-world skills for the AI economy. 

Preparing the AI Workforce

AI infrastructure alone is not enough. A successful national AI strategy must include efforts to build a workforce that can use advanced computing, data resources and AI tools in real scientific and industry settings.

That means pairing infrastructure with clear learning pathways. Universities, community colleges and regional partners can build degree programs, short-form certificates and stackable credentials that help learners move from foundational AI literacy into applied skills. Those pathways will help students, faculty, working adults and technical professionals use AI, including open source models and technologies, in fields like physical AI and automation, healthcare, energy, agriculture, manufacturing, quantum computing and cybersecurity. 

NVIDIA can support this work by providing training resources, educator enablement, applied learning content, technical guidance, partner platforms and access to tools that help institutions move from awareness to hands-on capability. As NVIDIA’s education and training offerings evolve, the goal remains the same: help institutions build repeatable, openly available programs that prepare learners to use AI systems, accelerated computing and data workflows responsibly and effectively.

This is how regional hubs become more than infrastructure projects. Students gain practical experience. Faculty expand their ability to teach and apply AI across disciplines. Working professionals can earn new skills without leaving the workforce. And institutions can connect training to local employer needs, research priorities and the economic opportunities that matter most to their communities.

Connecting Research, Workforce and Regional Growth

For policymakers and leaders, the hubs offer an opportunity to connect regional research and educational infrastructure with regional priorities and broader workforce and economic-development strategies.

Institutions can cultivate talent for local needs, support research connected to regional industries and build stronger relationships among universities, community colleges, employers and government. These connections will help translate AI leadership into scientific progress, new businesses and high-quality jobs.

No single organization can build this capacity alone. Sustained collaboration among government, higher education, philanthropic organizations and private industry is essential to ensure that advanced AI resources are broadly available and effectively used.

Private-sector contributions can complement public investment with technology, implementation expertise and workforce development programs. Public institutions, in turn, can help direct those capabilities toward scientific, educational and economic priorities that advance regional needs and national interest.

Learn more about the NSF State and Regional AI Infrastructure Hubs program and how NVIDIA is helping expand access to AI research and education. 

NVIDIA Alpamayo 2 Super, the Frontier Open Model for Robotaxis and Autonomous Vehicles, Now Available for Commercial Use

For robotaxis and other autonomous vehicles (AVs), the hardest problems aren’t the everyday scenarios. They’re the rare, complex situations that are difficult to anticipate and train for.

Handling these long‑tail events takes more than just object detection and motion prediction. AVs must understand the situation, reason about cause and effect, choose the right action and turn that decision into a safe, comfortable path — all in real time and in a way developers can inspect, validate and trust.

NVIDIA Alpamayo 2 Super, available now for commercial use, is part of the Alpamayo family, the most-adopted open reasoning models for autonomous driving on Hugging Face, supporting a wide range of AV-relevant capabilities within a single foundation model. 

Built on NVIDIA Cosmos 3 Super Reasoner and post‑trained with reinforcement learning, the model advances the AV ecosystem on two fronts: open commercial licensing and leading multitask capabilities for autonomous driving. 

Alpamayo 2 Super is part of NVIDIA’s growing collection of open models, datasets and tools for autonomous driving, expanding access, strengthening competition, giving developers greater control and supporting safer, more transparent AV deployment. 

Open Licensing for Production AVs

Alpamayo 2 Super is available on Hugging Face under OpenMDW‑1.1, the Linux Foundation’s permissive license for open AI model distributions. The license covers fine‑tuning, derivative models and commercial redistribution, allowing AV developers, automakers, truckmakers and suppliers to adapt Alpamayo to their own data, driving policies and deployment strategies. 

This openness lets AV researchers and companies keep control of their own data and infrastructure, as well as own the value they create through specialized models and accumulated know‑how. Such control is essential for workflows involving proprietary fleets and safety. 

Earlier Alpamayo releases were initially introduced for R&D. The OpenMDW license is now being  applied across the entire Alpamayo model family so developers can deploy any of the models commercially without requiring additional permissions. This creates a direct path from adaptation to deployment.

Open weights make that path economically viable. Teams can build on advanced reasoning without re‑training every foundation capability from scratch or paying frontier‑model costs for every task, matching the right model to the right job at the right cost. 

Alpamayo 2 Super enables frontier-scale reasoning in cloud-based development workflows, where developers can generate high-quality reasoning traces, synthetic training data and teacher outputs for model distillation. Within the Alpamayo model family, Alpamayo 2 Super delivers the highest reasoning and driving performance for multimodal autonomous driving development, while Alpamayo 1.5 and Alpamayo 1 provide more cost-efficient options for cloud-based development and model distillation.

The resulting distilled models can then be optimized for efficient, real-time inference in production vehicles. Together, the Alpamayo model family provides a cloud-to-car workflow that combines frontier-scale reasoning with scalable deployment across commercial AV fleets.

For AV programs, that means frontier‑scale reasoning in the cloud and efficient, specialized models in the vehicle — a more sustainable way to scale safe autonomy into commercial fleets.

Benchmark-Leading Reasoning at Frontier Scale

Alpamayo 2 Super ranks first on LingoQA, an autonomous driving reasoning benchmark, among nearly 40 models evaluated. In NVIDIA testing using the Lingo‑Judge metric, it outperformed Qwen2.5‑VL 72B by 17.0 points, Gemini 2.5 Pro by 15.1 points and GPT‑4o by 23.2 points, demonstrating state‑of‑the‑art reasoning for driving‑centric scenarios. Alpamayo 2 Super also ranks first across all autonomous driving benchmarks evaluated by NVIDIA, underscoring its leading performance across a broad range of AV capabilities. 

Alpamayo 2 Super offers 3x the scale of the 10‑billion‑parameter NVIDIA Alpamayo 1.5 and Alpamayo 1 models. The added capacity helps the model better generalize reasoning from sparse examples — a critical capability for the rare, multi‑agent interactions where conventional systems often struggle. 

The model reasons over full‑surround camera coverage, fusing views from the vehicle’s front, sides and rear. This 360‑degree context enables richer understanding of lane changes, merges, unprotected turns and complex intersections, where risks commonly arise.

A Multitask Foundation Model for Robotaxis and Autonomous Driving

For each driving situation, Alpamayo 2 Super can produce five tightly coupled outputs:

Together, these outputs offer insight into the model’s decision-making process. Developers can tie what the model observed to the action it selected, making decisions easier to understand, critique and validate. 

CoC traces integrate with NVIDIA Halos safety‑validation workflows and support AI safety aligned with ISO/PAS 8800 requirements, providing a stronger foundation for AV safety engineering. 

Alpamayo 2 Super can also be deployed as an autolabeler to generate CoC labels and perform visual question answering with 2D grounding on proprietary fleet data. By linking its reasoning to specific regions in camera images, the model can transform raw driving clips into richer training data, compressing annotation cycles from months to days.

Beyond planning and auto-labeling, Alpamayo 2 Super supports scene understanding, model critiquing and knowledge distillation. These multitask capabilities enable developers to use a single foundation model across more of the development stack, simplifying tooling and accelerating iteration.

An Open Ecosystem for Reasoning‑Based AVs

Alpamayo 2 Super is part of a broader family of open models, frameworks and datasets for AV development. 

Other tools in the family include: 

Alpamayo has already surpassed 500,000 downloads on Hugging Face, reinforcing its position as the most-adopted open reasoning model family for autonomous driving on the platform.

Download NVIDIA Alpamayo 2 Super on Hugging Face to explore the model, evaluate its reasoning capabilities and start building the next generation of robotaxis and autonomous vehicles.

As AI Increases Demands on Memory, Storage Steps Up

Surging AI demands are driving the need for massive datasets and context windows that burst past the confines of system memory. 

But rising needs aren’t met by simply adding more storage capacity. What’s needed is useful, grounded insights from AI factories and efficient, secure storage architectures that enable those insights. 

At this week’s Future of Memory and Storage (FMS) conference, NVIDIA is unveiling new storage advancements and showcasing how the next leap in AI depends as much on the storage infrastructure feeding accelerated computing as on the computing power itself.

The pressure on that infrastructure is intensifying as AI agents consume massive amounts of data — and GPUs can now initiate storage requests directly, generating thousands of concurrent operations.

To serve those requests, storage systems must continuously encrypt, compress, verify and reconstruct data. These critical data services can become bottlenecks when thousands of agents access storage simultaneously.

Benchmarks highlighted in this NVIDIA technical blog show that the NVIDIA Vera CPU, part of NVIDIA Vera BlueField-4 STX, delivers up to 3.21x higher throughput than an x86 CPU in a two-stage compression and encryption pipeline. This means that with Vera, storage platforms can absorb the flood of AI data more efficiently — delivering greater throughput with significantly less compute infrastructure.

With accelerated computing, storage stops being a passive place to keep data and becomes an active part of the data path. 

This upends the old economics of determining when data belongs in memory (where applications can fetch it faster) versus on a storage drive (where it can be held in cheap and plentiful space). The tradeoff was first framed 40 years ago, when the answer was measured in accessing that data in minutes. On today’s GPUs, paired with AI storage solutions from NVIDIA and partners, the same tradeoff now plays out in microseconds.

Closing the gap between AI’s needs and memory shortage depends on extreme codesign across the whole ecosystem, from memory and storage manufacturers to the software built on them. 

Open Source NVIDIA cuFile APIs Enable Interoperability for Storage Solutions

At FMS, NVIDIA announced it is open sourcing its cuFile application programming interfaces (APIs) — and the vertical storage software stack underneath them — which let GPUs, not just CPUs, read from and write to storage directly. cuFile is an open source component of NVIDIA GPUDirect Storage.

Using hundreds of thousands of GPU threads, fast high-bandwidth memory and other methodologies, cuFile enables securely accessing data from storage in just microseconds.

This represents how the industry is unifying a security-first storage stack based on Linux best practices, providing interoperability between GPUs and data. 

In addition, fast, secure access to data and storage is a foundational element to powering preventive and detective cybersecurity measures. Making cuFile openly available will help make security context, data and storage accessible at the speed AI-powered defenses need. Such open technologies support initiatives such as the new Open Secure AI Alliance.

This site is the new home for APIs that are open to contributions — with Google, Intel, NVIDIA and Meta as inaugural maintainers — and can be optimized for use across various software and hardware platforms, driving innovation and efficiency for developers and enterprises.

NVIDIA and Industry Leaders Advance New Frontier of AI Storage

In addition, NVIDIA and storage industry leaders are optimizing memory and storage solutions through an initiative called Storage-Next. The NVIDIA-driven initiative brings together storage makers, controller vendors, thermal design, cooling and orchestration operators, and standards bodies to align on how GPU-driven storage should behave — then turn these advancements into interoperable, open industry standards.

Storage-Next includes over 40 leading storage and flash vendors — including DDN, KIOXIA and Micron — each contributing to the next generation of AI storage technologies with NVIDIA.

The initiative is grounded in accelerated data access for large AI datasets. To support this, NVIDIA offers SCADA — short for scaled, accelerated data access — a framework that lets massively parallel GPUs pull only the data necessary for the application directly from storage into their own high-speed memory.

For example, DDN is integrating SCADA with Infinia, its software-defined, AI-native data intelligence platform built to eliminate storage bottlenecks at scale. 

“AI success will be defined not by how much infrastructure organizations own, but by how productively they use it,” said Sven Oehme, chief technology officer at DDN. “Our collaboration with NVIDIA is helping create a more direct, efficient connection between GPUs and data — keeping accelerated computing resources productive, speeding time to insight and enabling customers to achieve stronger business and financial returns from their AI investments.”

Storage-Next and SCADA extend NVIDIA’s longstanding work on AI storage infrastructure, including on NVIDIA Vera BlueField-4 STX — a modular, rack-scale foundation powered by the NVIDIA Vera Rubin platform, NVIDIA Vera BlueField-4 storage processors and NVIDIA Spectrum-X Ethernet networking.

NVIDIA Vera BlueField-4 STX storage processor.

Defining a new class of AI-native data platforms, NVIDIA STX uses the unified NVIDIA DOCA security stack to let enterprises enable continuous policy enforcement in the AI data path.

Plus, NVIDIA CMX Context Memory Storage provides an AI‑native context tier for long‑context, multi‑turn, agentic AI inference, built on NVIDIA STX.

SCADA Enables Fast AI Storage That Stays Secure

Speed at the storage layer comes with a catch. Letting an application talk straight to a drive is quick, but done carelessly, it can scribble over other processes’ memory — a security hole, not a feature. 

NVIDIA SCADA uses a safe, robust method to achieve scaled direct access by splitting the job in two:

It’s all part of how advancements in fast, massively parallel, efficient, secure AI storage infrastructure can feed better data to applications and AI factories — so they can produce more useful, accurate, grounded intelligence at scale. 

Join NVIDIA sessions at FMS, running Aug. 4-6 in Santa Clara, California, and learn more about NVIDIA AI storage.

See notice regarding software product information.