AI Infrastructure
(AI-INFRA 100-02)
A four-course series for engineers building and operating GPU clusters
Four courses covering the AI cluster stack, from architecture vocabulary through hardware management, fabric, and provisioning. Foundations is conceptual and establishes the shared vocabulary the other three assume. The three technology courses are hands-on and independent of each other: take them in any order, or take only the one that matches the work in front of you.
Description
AI Infrastructure Foundations
A conceptual walkthrough of the stack: the GPU node, HBM and the checkpoint burst, NVLink, the four fabrics, the management plane, and the failure domains at each layer. The vocabulary and architecture the other three courses assume.
90 minutes, no lab
Redfish for Fleet Operations
Hardware management and observability through one HTTPS API across a mixed-vendor fleet: health validation before provisioning, firmware compliance and drift remediation, the scrape pipeline behind a fleet dashboard, and VirtualMedia provisioning boot.
2 hours, 7 labs
GPU Cluster Networking
RoCE v2 lossless fabric: why RDMA cannot tolerate loss, PFC and ECN on both switch and NIC, fabric telemetry with NetQ, and two live incidents diagnosed from symptom to root cause without the fault class given in advance.
2 hours, 4 labs plus one optional
AI Cluster Provisioning
Bare-metal servers to a cluster ready for ML-team handoff: host enrollment, declarative provisioning with k0rdent and Metal3/Ironic, the
GPU Operator and Network
FORMAT
Instructor-led
Each course stands alone
Foundations recommended first
Hands-on labs in the three technology courses
Lab environments hosted and preconfigured
Who Should Attend
INFRASTRUCTURE ENGINEER
You have managed servers, switches, and racks your whole career and are now standing up GPU clusters with DPUs, lossless fabric, and declarative provisioning. You need the vocabulary and then the hands-on reps, in that order.
FIELD, PRODUCT, AND OPERATIONS ROLES
You need to hold a credible technical conversation about GPU topology, RDMA fabric, and cluster automation without implementing it. Foundations is built for exactly that, and it assumes no Linux or scripting background.
Lab Requirements
Foundations: general datacenter literacy, basic IP and TCP vocabulary
Technology courses: Linux terminal, reading JSON and YAML, curl
Networking: TCP/IP, VLANs, L2 and L3 switching
Provisioning: Kubernetes and kubectl basics
No RDMA, DPU, or Redfish experience assumed