[{"data":1,"prerenderedAt":275},["ShallowReactive",2],{"story":3},{"name":4,"created_at":5,"published_at":6,"updated_at":7,"id":8,"uuid":9,"content":10,"slug":267,"full_slug":268,"sort_by_date":25,"position":269,"tag_list":270,"is_startpage":179,"parent_id":271,"meta_data":25,"group_id":272,"first_published_at":6,"release_id":25,"lang":273,"path":25,"alternates":274,"default_full_slug":25,"translated_slugs":25},"AI-INFRA 100-01: AI Infrastructure Foundations","2026-09-18T10:58:13.501Z","2026-09-18T11:03:23.769Z","2026-09-18T11:03:23.785Z",221345765403475,"9f452448-d0ca-45fe-918a-3d3246df1250",{"seo":11,"_uid":15,"type":16,"intro":17,"title":18,"duration":13,"component":19,"course_id":20,"technology":21,"description":22,"on_schedule":179,"course_level":180,"hide_sidebar":181,"prerequisites":182,"on_demand_link":185,"lab_requirements":188,"course_objectives":221,"follow_up_courses":224,"who_should_attend":227,"additional_content":254,"on_demand_training":179,"section_below_hero":259,"hide_class_schedule_cta":181,"hide_private_training_cta":179,"section_below_hero_bg_color":13},{"_uid":12,"title":13,"plugin":14,"og_image":13,"og_title":13,"description":13,"twitter_image":13,"twitter_title":13,"og_description":13,"twitter_description":13},"867ef41d-665c-4afb-ab8c-7929189d5cb2","","seo_metatags","8fdb6453-e42e-443a-befd-e555c41d99b5","course","The shared vocabulary and architecture of an AI GPU cluster","(AI-INFRA 100-01)","training_course","AI Infrastructure Foundations","cn",{"type":23,"attrs":24,"content":26},"doc",{"backgroundColor":25},null,[27,148,155],{"type":28,"attrs":29,"content":31},"ordered_list",{"order":30},1,[32,50,64,78,92,106,120,134],{"type":33,"content":34},"list_item",[35,45],{"type":36,"attrs":37,"content":38},"paragraph",{"textAlign":25},[39],{"text":40,"type":41,"marks":42},"The workload drives the architecture","text",[43],{"type":44},"bold",{"type":36,"attrs":46,"content":47},{"textAlign":25},[48],{"text":49,"type":41},"Training and inferencing as distinct workload shapes, why a GPU rather than a CPU, and the chain of needs a single workload creates.",{"type":33,"content":51},[52,59],{"type":36,"attrs":53,"content":54},{"textAlign":25},[55],{"text":56,"type":41,"marks":57},"Memory and storage",[58],{"type":44},{"type":36,"attrs":60,"content":61},{"textAlign":25},[62],{"text":63,"type":41},"HBM against conventional DDR, what a workload reads and writes, and the checkpoint burst that shapes how AI storage is sized.",{"type":33,"content":65},[66,73],{"type":36,"attrs":67,"content":68},{"textAlign":25},[69],{"text":70,"type":41,"marks":71},"Beyond one GPU: NVLink and the rack",[72],{"type":44},{"type":36,"attrs":74,"content":75},{"textAlign":25},[76],{"text":77,"type":41},"Why one GPU is not enough, GPUs running in lockstep, NVLink as the rack-internal GPU interconnect, and the rack as the unit of deployment. Node anatomy across both rack-scale and baseboard form factors.",{"type":33,"content":79},[80,87],{"type":36,"attrs":81,"content":82},{"textAlign":25},[83],{"text":84,"type":41,"marks":85},"The four fabrics",[86],{"type":44},{"type":36,"attrs":88,"content":89},{"textAlign":25},[90],{"text":91,"type":41},"The RoCE v2 compute fabric across racks, the converged fabric carrying storage and management and tenant traffic, why the DPU exists as a distinct device from the SuperNIC, and where each fabric lives.",{"type":33,"content":93},[94,101],{"type":36,"attrs":95,"content":96},{"textAlign":25},[97],{"text":98,"type":41,"marks":99},"One cluster becomes a fleet",[100],{"type":44},{"type":36,"attrs":102,"content":103},{"textAlign":25},[104],{"text":105,"type":41},"The out-of-band network nobody touches, the two BMCs in every GPU node, and the management cluster that provisions the GPU clusters.",{"type":33,"content":107},[108,115],{"type":36,"attrs":109,"content":110},{"textAlign":25},[111],{"text":112,"type":41,"marks":113},"Provisioning and lifecycle",[114],{"type":44},{"type":36,"attrs":116,"content":117},{"textAlign":25},[118],{"text":119,"type":41},"The chain from k0rdent through Metal3/Ironic down to Redfish on each BMC, plus what the GPU Operator and Network Operator take over once a node is running Kubernetes.",{"type":33,"content":121},[122,129],{"type":36,"attrs":123,"content":124},{"textAlign":25},[125],{"text":126,"type":41,"marks":127},"Validation and failure domains",[128],{"type":44},{"type":36,"attrs":130,"content":131},{"textAlign":25},[132],{"text":133,"type":41},"UFM and Fabric Manager as continuous fabric validation, blast radius by stack layer, and where to look first when a job slows down or a node drops out.",{"type":33,"content":135},[136,143],{"type":36,"attrs":137,"content":138},{"textAlign":25},[139],{"text":140,"type":41,"marks":141},"Scale, and why AI infrastructure differs",[142],{"type":44},{"type":36,"attrs":144,"content":145},{"textAlign":25},[146],{"text":147,"type":41},"Patterns that work at ten nodes and break at a thousand, declarative operations as the default, mapping operator tasks to the component and tool that perform them, and a consolidated comparison against conventional datacenter practice.",{"type":36,"attrs":149,"content":150},{"textAlign":25},[151],{"text":152,"type":41,"marks":153},"FORMAT",[154],{"type":44},{"type":156,"content":157},"bullet_list",[158,165,172],{"type":33,"content":159},[160],{"type":36,"attrs":161,"content":162},{"textAlign":25},[163],{"text":164,"type":41},"90 minutes, instructor-led",{"type":33,"content":166},[167],{"type":36,"attrs":168,"content":169},{"textAlign":25},[170],{"text":171,"type":41},"Conceptual and architectural",{"type":33,"content":173},[174],{"type":36,"attrs":175,"content":176},{"textAlign":25},[177],{"text":178,"type":41},"No hands-on lab",false,"essentials",true,{"type":23,"content":183},[184],{"type":36},{"id":13,"url":13,"linktype":186,"fieldtype":187,"cached_url":13},"story","multilink",{"type":23,"attrs":189,"content":190},{"backgroundColor":25},[191],{"type":156,"content":192},[193,200,207,214],{"type":33,"content":194},[195],{"type":36,"attrs":196,"content":197},{"textAlign":25},[198],{"text":199,"type":41},"General datacenter literacy: servers, racks, switches",{"type":33,"content":201},[202],{"type":36,"attrs":203,"content":204},{"textAlign":25},[205],{"text":206,"type":41},"Basic networking vocabulary: IP, TCP, VLAN",{"type":33,"content":208},[209],{"type":36,"attrs":210,"content":211},{"textAlign":25},[212],{"text":213,"type":41},"Awareness that GPUs accelerate machine learning",{"type":33,"content":215},[216],{"type":36,"attrs":217,"content":218},{"textAlign":25},[219],{"text":220,"type":41},"No RDMA, DPU, or Kubernetes experience assumed",{"type":23,"content":222},[223],{"type":36},{"type":23,"content":225},[226],{"type":36},{"type":23,"attrs":228,"content":229},{"backgroundColor":25},[230,237,242,249],{"type":36,"attrs":231,"content":232},{"textAlign":25},[233],{"text":234,"type":41,"marks":235},"INFRASTRUCTURE ENGINEER",[236],{"type":44},{"type":36,"attrs":238,"content":239},{"textAlign":25},[240],{"text":241,"type":41},"You have run servers, switches, and racks for years and are now being handed GPU clusters. You need the vocabulary and the architecture before you touch the hardware: what a GPU tray is, why there are four fabrics instead of one, and why two BMCs sit in every node.",{"type":36,"attrs":243,"content":244},{"textAlign":25},[245],{"text":246,"type":41,"marks":247},"DEVOPS, DC OPS, PRODUCT, AND FIELD ROLES",[248],{"type":44},{"type":36,"attrs":250,"content":251},{"textAlign":25},[252],{"text":253,"type":41},"You need to follow and contribute to a technical conversation about GPU topology, RDMA fabric, and cluster automation without implementing it yourself. Analogies and layered diagrams carry the technical slides, and no Linux or scripting background is assumed.",{"type":23,"attrs":255,"content":256},{"backgroundColor":25},[257],{"type":36,"attrs":258},{"textAlign":25},{"type":23,"attrs":260,"content":261},{"backgroundColor":25},[262],{"type":36,"attrs":263,"content":264},{"textAlign":25},[265],{"text":266,"type":41},"A conceptual walkthrough of the AI cluster stack, from the GPU node to the fabrics to the management plane that operates a fleet of them. The goal is a working mental model and precise vocabulary, not hands-on execution: you leave able to name every component, place it in the stack, and explain what it is there for.","ai-infrastructure-foundations","training/courses/ai-infrastructure-foundations",-300,[],111882102,"5752396c-3905-4d27-96b3-f3162d73c931","default",[],1789730290007]