InfiniBand Deep Dive: Networking for AI focused Datacentres

Learn InfiniBand fundamentals, RDMA, and NVIDIA networking stack to design and optimize AI and HPC infrastructure

$9.99 (93% OFF)
Get Course Now

About This Course

InfiniBand Deep Dive: Networking for AI Data CentresWelcome! I'm here to help you truly understand InfiniBand — the high-performance fabric powering the world's most demanding AI and HPC environments.As AI workloads explode in scale, the network is no longer an afterthought — it is the bottleneck. Slow fabrics mean idle GPUs, longer training times, and wasted investment in expensive compute. This course gives you the deep, practical knowledge to understand, deploy, and troubleshoot the technology at the heart of modern AI data centres.What you'll learn:Why traditional Ethernet and TCP/IP fall short for AI workloads — and how InfiniBand solves latency, throughput, and CPU bottleneck challengesThe full InfiniBand architecture — Physical, Link, Network, Transport, and Upper layers — with real-world analogies that make concepts stickRDMA, Zero-Copy transfers, Queue Pairs, Memory Registration, and GPUDirect RDMA — the core technologies behind high-speed GPU communicationHow the Subnet Manager works — LID assignment, topology discovery, routing table programming, and failover with Standby SMTraffic isolation using Partition Keys (PKey) — configuring Full and Limited membership across multi-tenant AI clustersQuality of Service (QoS) — assigning Service Levels (SL), mapping to Virtual Lanes (VL), and configuring bandwidth weights in OpenSMRouting algorithms in depth — MINHOP, UPDN, Fat-Tree, Adaptive Routing — and why Adaptive Routing is critical for elephant flows in AI workloadsCongestion control, Credit-Based Flow Control, credit loops, and how to prevent fabric deadlocksFabric monitoring and management at scale using NVIDIA Unified Fabric Manager (UFM), including Cyber-AI and RBACHands-on troubleshooting using ibdiagnet, ibtracert, mlxlink, ibstat, smpquery, and moreThis course is packed with visual analogies, architecture diagrams, and practical troubleshooting scenarios that make even the most complex concepts click. Whether you're a network engineer, a cloud infrastructure specialist, or an AI platform team member, this course will give you the edge to design and operate high-performance AI fabrics with confidence.No InfiniBand experience required — just bring your curiosity and your ambition.Let's get started!NOTE - For learners preparing for NCP-AIN ExamThis course covers 40–50% of the NCP-AIN exam domains, with a focused deep dive on InfiniBand. It is a valuable study companion to cover a significant portion of topics for the certification.

What you'll learn:

  • Describe the full InfiniBand architecture — from physical layer to upper-layer protocols like RDMA and GPU Direct
  • Design leaf-spine InfiniBand topologies suited for large-scale AI and HPC cluster deployments
  • Implement QoS using Service Levels (SL) and Virtual Lanes (VL) to prioritise AI and HPC workloads
  • Monitor and manage InfiniBand fabrics at scale using NVIDIA Unified Fabric Manager (UFM)
  • Diagnose and resolve common InfiniBand fabric issues using tools like ibdiagnet, ibtracert, mlxlink, and ibstat