
AI has entered the mainstream, fueling an unprecedented demand for AI factories—specialized infrastructure designed for AI training, inference, and large-scale intelligence production. Many of these facilities will operate at gigawatt scale, requiring immense engineering and logistical efforts. Constructing a single gigawatt AI factory involves tens of thousands of workers across suppliers, architects, contractors, and engineers, assembling nearly 5 billion components and over 210,000 miles of fiber cable.
For AI factory operators, staying ahead means more than just maximizing efficiency—it’s about preventing infrastructure failures that could result in massive financial losses. According to NVIDIA, a single day of downtime in a 1-gigawatt AI factory can exceed $100 million in costs.
To support the design and optimization of AI factories, NVIDIA introduced the NVIDIA Omniverse Blueprint for AI factory planning and operations. By addressing infrastructure challenges upfront, this blueprint aims to minimize risk and accelerate deployment timelines.
Engineering AI Factories: A Simulation-First Approach
The NVIDIA Omniverse Blueprint for AI factory design and operations leverages OpenUSD libraries, allowing developers to integrate 3D data from various sources, including the facility itself, NVIDIA accelerated computing systems, and power or cooling units from providers like Schneider Electric and Vertiv.
By unifying the design and simulation of billions of components, the blueprint is intended to help engineers address complex challenges like:
Component integration and space optimization.
Cooling system performance and efficiency.
Power distribution and reliability.
Networking topology and logic.
Breaking Down Engineering Silos
One of the biggest challenges in AI factory construction is the siloed operations of different teams—power, cooling, and networking—which can lead to inefficiencies and potential failures. With the NVIDIA Omniverse Blueprint, engineers can now:
Collaborate in Full Context: Multiple disciplines can iterate in parallel, sharing live simulations that reveal how changes in one domain affect another.
Optimize Energy Usage: Real-time simulation updates enable teams to find the most efficient designs for AI workloads.
Eliminate Failure Points: By validating redundancy configurations before deployment, organizations reduce the risk of costly downtime.
Model Real-World Conditions: Predict and test how different AI workloads will impact cooling, power stability, and network congestion.
By integrating real-time simulation across disciplines, the NVIDIA Omniverse Blueprint enables engineering teams to explore different configurations, assess total cost of ownership, and optimize power utilization for improved efficiency.
Other NVIDIA Announcements
Other key announcements included the launch of NVIDIA Blackwell Ultra, the next evolution of the NVIDIA Blackwell AI factory platform, and the introduction of NVIDIA Dynamo, an open-source inference framework. Blackwell Ultra is designed to enhance training and test-time scaling inference—applying more compute during inference to improve accuracy—helping organizations accelerate AI applications such as reasoning, agentic AI, and physical AI. Meanwhile, NVIDIA Dynamo is a new AI inference-serving software aimed at maximizing token revenue for AI factories deploying reasoning AI models while ensuring optimal GPU resource utilization.
Blackwell Ultra-based products are expected to be available from partners starting in the second half of 2025. Leading companies, including Cisco, Dell Technologies, Hewlett Packard Enterprise, Lenovo, and Supermicro, will offer a range of servers built on Blackwell Ultra. Additional partners delivering these products include Aivres, ASRock Rack, ASUS, Eviden, Foxconn, GIGABYTE, Inventec, Pegatron, Quanta Cloud Technology (QCT), Wistron, and Wiwynn.
Cloud service providers Amazon Web Services, Google Cloud, Microsoft Azure, and Oracle Cloud Infrastructure, along with GPU cloud providers CoreWeave, Crusoe, Lambda, Nebius, Nscale, Yotta, and YTL, will be among the first to offer Blackwell Ultra-powered instances.
Learn more about Industrial AI's Role in the Digital Transformation of Industries