KoreaDevKNOWLEDGE SHARING

AI INFRASTRUCTURE · 12 CONNECTED COURSES

AI infrastructure
Hands-on operations course for new hires

Connects design and build to operational evidence, from a single server's baseline to the data center, GPUs, Kubernetes, model serving, incident response, and handover.

Currently published scopePublic courses 12 items · every course includes labs, judgment activities and a three-stage assessment.
Open the Ubuntu 24.04 course

LEARNING ROADMAP

Understanding, design, build-out, operations, and recovery form a single flow

All 12 public courses provide complete lessons, hands-on decision activities, explanations for each option, and official sources.

  1. 01Lessons and labs available

    The big picture of AI data centers

    Covers how the rack, power, cooling, servers, network, storage, GPUs, and AI service connect into a single failure domain.

    4 core unitsOpen course →
  2. 02Lessons and labs available

    Ubuntu 24.04 server operations

    Starting from the first installation screen, learn files, users, permissions, packages, systemd, networking, and storage in sequence, and judge server state from evidence gathered before and after each change.

    5 core unitsOpen course →
  3. 03Lessons and labs available

    Interconnects and AI fabrics

    Connect everything from the server's internal bus to Ethernet, InfiniBand, RDMA, and the scale-out fabric, and accept link state based on evidence.

    4 core unitsOpen course →
  4. 04Lessons and labs available

    GPU hardware and host preparation

    Inspect GPU servers on delivery, verify the firmware, driver, and CUDA relationships, then build a safe baseline and a rollback procedure.

    3 core unitsOpen course →
  5. 05Lessons and labs available

    AI storage

    Connect the throughput, IOPS, latency, and metadata requirements of datasets, checkpoints, and model artifacts to the storage tier and to recovery testing.

    4 core unitsOpen course →
  6. 06Lessons and labs available

    Containers and registries

    Understand OCI images and the container runtime, and operate reproducible builds, SBOMs, signing, registries, and air-gapped imports.

    3 core unitsOpen course →
  7. 07Lessons and labs available

    Kubernetes and K3s

    Deploy container workloads as declarative state and diagnose network, storage, and failure issues on single-node and multi-node clusters.

    4 core unitsOpen course →
  8. 08Lessons and labs available

    GPU platform

    Learn how Kubernetes discovers and allocates GPUs, and the verification boundaries of GPU sharing, MIG, and distributed workloads.

    3 core unitsOpen course →
  9. 09Lessons and labs available

    Operations and security

    Connect RBAC, secrets, backup, restore, upgrade, patching, and policy to change management and recovery evidence.

    4 core unitsOpen course →
  10. 10Lessons and labs available

    Observability and incident response

    Link the user SLO to metrics, logs, traces and events, and carry out evidence-first incident response and a postmortem.

    4 core unitsOpen course →
  11. 11Lessons and labs available

    AI model serving

    Complete the production serving path from the model artifact and runtime through GPU placement, API, security, performance, capacity, and safe rollout.

    3 core unitsOpen course →
  12. 12Lessons and labs available

    AI infrastructure capstone project

    Turn requirements into a design, and complete build-out, acceptance, fault injection, recovery, operations, and handover as a single evidence package.

    3 core unitsOpen course →