In November, chipmaker NVIDIA unveiled Eos, its data-center-scale supercomputer that promises to supercharge and optimize artificial intelligence (AI) workloads. Eos is also built to train large language models, recommender systems, and quantum simulations.
Eos arrives at a time when generative AI is revolutionizing various fields, from drug discovery to chatbots and autonomous systems. To accelerate these advancements, a robust AI infrastructure is necessary, one that can support the development of AI models at scale.
Eos’s architecture is tailor-made for AI workloads demanding ultra-low latency and high-throughput interconnectivity across a large cluster of accelerated computing nodes, making it an ideal solution for enterprises seeking to scale their AI capabilities.
READ:
Supercomputing research center equips platform with NVIDIA CUDA Quantum
NVIDIA, Microsoft to build ‘massive’ cloud AI computer
NVIDIA Eos comprises 576 NVIDIA DGX H100 systems, NVIDIA Quantum-2 InfiniBand networking, and software, promising to deliver 18.4 exaflops of FP8 AI performance. It’s a counterpart to another Eos DGX SuperPOD containing 10,752 NVIDIA H100 GPUs, utilized for MLPerf training in November.
NVIDIA Quantum-2 InfiniBand
Each DGX H100 system houses eight NVIDIA H100 Tensor Core GPUs, amounting to 4,608 H100 GPUs in Eos.
Leveraging NVIDIA Quantum-2 InfiniBand with In-Network Computing technology, Eos’ network architecture supports data transfer speeds of up to 400Gb/s, facilitating the swift movement of large datasets crucial for training complex AI models.
At its core, Eos adopts the revolutionary DGX SuperPOD architecture, powered by NVIDIA’s DGX H100 systems. This architecture is engineered to offer tightly integrated full-stack systems capable of computing at an unprecedented scale.
As organizations worldwide strive to harness the potential of AI, Eos emerges as a pivotal resource, poised to expedite the journey toward AI-driven applications that empower every sector.