AI INFRASTRUCTURE · 12 CONNECTED COURSES
AI infrastructure
Hands-on operations course for new hires
Connects design and build to operational evidence, from a single server's baseline to the data center, GPUs, Kubernetes, model serving, incident response, and handover.
LEARNING ROADMAP
Understanding, design, build-out, operations, and recovery form a single flow
All 12 public courses provide complete lessons, hands-on decision activities, explanations for each option, and official sources.
- 01Lessons and labs available
The big picture of AI data centers
Covers how the rack, power, cooling, servers, network, storage, GPUs, and AI service connect into a single failure domain.
- 02Lessons and labs available
Ubuntu 24.04 server operations
Starting from the first installation screen, learn files, users, permissions, packages, systemd, networking, and storage in sequence, and judge server state from evidence gathered before and after each change.
- 03Lessons and labs available
Interconnects and AI fabrics
Connect everything from the server's internal bus to Ethernet, InfiniBand, RDMA, and the scale-out fabric, and accept link state based on evidence.
- 04Lessons and labs available
GPU hardware and host preparation
Inspect GPU servers on delivery, verify the firmware, driver, and CUDA relationships, then build a safe baseline and a rollback procedure.
- 05Lessons and labs available
AI storage
Connect the throughput, IOPS, latency, and metadata requirements of datasets, checkpoints, and model artifacts to the storage tier and to recovery testing.
- 06Lessons and labs available
Containers and registries
Understand OCI images and the container runtime, and operate reproducible builds, SBOMs, signing, registries, and air-gapped imports.
- 07Lessons and labs available
Kubernetes and K3s
Deploy container workloads as declarative state and diagnose network, storage, and failure issues on single-node and multi-node clusters.
- 08Lessons and labs available
GPU platform
Learn how Kubernetes discovers and allocates GPUs, and the verification boundaries of GPU sharing, MIG, and distributed workloads.
- 09Lessons and labs available
Operations and security
Connect RBAC, secrets, backup, restore, upgrade, patching, and policy to change management and recovery evidence.
- 10Lessons and labs available
Observability and incident response
Link the user SLO to metrics, logs, traces and events, and carry out evidence-first incident response and a postmortem.
- 11Lessons and labs available
AI model serving
Complete the production serving path from the model artifact and runtime through GPU placement, API, security, performance, capacity, and safe rollout.
- 12Lessons and labs available
AI infrastructure capstone project
Turn requirements into a design, and complete build-out, acceptance, fault injection, recovery, operations, and handover as a single evidence package.