AI Infrastructure

(AI-INFRA 100-02)

A four-course series for engineers building and operating GPU clusters

Four courses covering the AI cluster stack, from architecture vocabulary through hardware management, fabric, and provisioning. Foundations is conceptual and establishes the shared vocabulary the other three assume. The three technology courses are hands-on and independent of each other: take them in any order, or take only the one that matches the work in front of you.

Description

  1. AI Infrastructure Foundations

    A conceptual walkthrough of the stack: the GPU node, HBM and the checkpoint burst, NVLink, the four fabrics, the management plane, and the failure domains at each layer. The vocabulary and architecture the other three courses assume.

    90 minutes, no lab

  2. Redfish for Fleet Operations

    Hardware management and observability through one HTTPS API across a mixed-vendor fleet: health validation before provisioning, firmware compliance and drift remediation, the scrape pipeline behind a fleet dashboard, and VirtualMedia provisioning boot.

    2 hours, 7 labs

  3. GPU Cluster Networking

    RoCE v2 lossless fabric: why RDMA cannot tolerate loss, PFC and ECN on both switch and NIC, fabric telemetry with NetQ, and two live incidents diagnosed from symptom to root cause without the fault class given in advance.

    2 hours, 4 labs plus one optional

  4. AI Cluster Provisioning

    Bare-metal servers to a cluster ready for ML-team handoff: host enrollment, declarative provisioning with k0rdent and Metal3/Ironic, the

    GPU Operator and Network

FORMAT

  • Instructor-led

  • Each course stands alone

  • Foundations recommended first

  • Hands-on labs in the three technology courses

  • Lab environments hosted and preconfigured

Who Should Attend

INFRASTRUCTURE ENGINEER

You have managed servers, switches, and racks your whole career and are now standing up GPU clusters with DPUs, lossless fabric, and declarative provisioning. You need the vocabulary and then the hands-on reps, in that order.

FIELD, PRODUCT, AND OPERATIONS ROLES

You need to hold a credible technical conversation about GPU topology, RDMA fabric, and cluster automation without implementing it. Foundations is built for exactly that, and it assumes no Linux or scripting background.

Lab Requirements

  • Foundations: general datacenter literacy, basic IP and TCP vocabulary

  • Technology courses: Linux terminal, reading JSON and YAML, curl

  • Networking: TCP/IP, VLANs, L2 and L3 switching

  • Provisioning: Kubernetes and kubectl basics

  • No RDMA, DPU, or Redfish experience assumed

Request Private training