Course Outline
NVIDIA AI and GPU Infrastructure Foundations
Outcome: Establish a common technical foundation and connect AI workload characteristics to GPU infrastructure choices.
AI Workloads and Accelerated Computing
- AI, machine learning, deep learning, generative AI and modern AI factory concepts.
- Training, fine-tuning and inference: different resource and performance profiles.
- Why CPUs and GPUs behave differently for parallel AI workloads.
NVIDIA GPU Architecture
- GPU generations, streaming multiprocessors, CUDA cores and Tensor Cores.
- Device memory, bandwidth, caches and data-movement considerations.
- PCIe, SXM, NVLink and NVSwitch in single- and multi-GPU systems.
- Reading device identity, topology and capabilities with NVIDIA utilities.
NVIDIA Hardware and Software Ecosystem
- Roles of CPUs, GPUs and BlueField DPUs in accelerated infrastructure.
- DGX, HGX and NVIDIA-Certified server concepts and deployment models.
- Drivers, CUDA, container runtime components, NGC and NVIDIA AI Enterprise.
- Bare metal, virtualized, Kubernetes and cloud deployment choices.
AI Infrastructure Building Blocks
- Compute, network, storage and management-plane responsibilities.
- Reference architectures and support matrices as design and validation tools.
- Foundational monitoring, orchestration, MLOps and job-scheduling concepts.
HANDS-ON LAB | Discover and Validate a GPU Node
- Inspect GPU model, driver, firmware-visible state and CUDA compatibility.
- Review GPU-to-CPU and GPU-to-GPU topology and identify potential bottlenecks.
- Create a baseline inventory for use during later virtualization and monitoring labs.
AI Server Architecture and Infrastructure Deployment
Outcome: Translate workload requirements into a supportable AI server, network and storage design.
AI Server Platforms and Topology
- PCIe and SXM platform characteristics; single- and multi-GPU design patterns.
- CPU architecture, memory, NUMA locality and PCIe lane considerations.
- NVLink/NVSwitch topology, GPU peer access and scale-up communication.
- BMC, firmware, BIOS and operating-system prerequisites for GPU servers.
Networking for AI Workloads
- Ethernet and InfiniBand roles in training and inference environments.
- RDMA, RoCE, GPUDirect RDMA and latency/bandwidth considerations.
- North-south, east-west and management network separation.
- Fabric redundancy, observability and common operational failure points.
Storage and the AI Data Path
- Dataset, checkpoint, model and inference-serving access patterns.
- Local, shared file, object and parallel storage considerations.
- Throughput, IOPS, metadata performance and GPUDirect Storage concepts.
Deployment and Operational Readiness
- Sizing compute, memory, network and storage against workload demand.
- Power, cooling, rack, cabling and environmental validation.
- Support matrices, firmware/driver alignment and acceptance testing.
- Capacity headroom, resilience and lifecycle planning.
HANDS-ON LAB | Design and Validate an AI Infrastructure Blueprint
- Map a workload profile to server, GPU, network and storage requirements.
- Review topology and identify oversubscription, NUMA or data-path risks.
- Build a concise deployment-readiness and acceptance checklist.
GPU Virtualization, NVIDIA vGPU and VFIO Passthrough
Outcome: Select and implement the appropriate GPU assignment model for virtualized AI workloads.
GPU Resource Assignment Models
- Full physical GPU assignment, time-sliced vGPU, MIG-backed vGPU and MIG comparison.
- Isolation, predictability, density, live operations and licensing trade-offs.
- Workload placement patterns for development, inference, fine-tuning and shared services.
NVIDIA vGPU Architecture and Administration
- vGPU Manager, guest driver, profile selection and frame-buffer allocation.
- Supported hypervisors, guest operating systems and compatibility validation.
- NVIDIA licence services, entitlements and operational dependencies.
- Provisioning, validation, capacity control and common failure modes.
Linux VFIO GPU Passthrough
- IOMMU concepts, PCIe device groups and platform prerequisites.
- Identifying GPU and companion devices; binding devices to vfio-pci.
- Assigning a GPU to a virtual machine and installing the guest driver.
- Validating device isolation and recovering from driver-binding conflicts.
Performance and Troubleshooting
- NUMA placement, CPU pinning, huge pages and memory-locking considerations.
- Host/guest driver compatibility and device visibility diagnostics.
- Decision criteria for passthrough versus vGPU in enterprise operations.
HANDS-ON LAB | Assign and Validate GPU Resources in a VM
- Verify IOMMU and device grouping; prepare the GPU for passthrough or vGPU assignment.
- Attach the GPU resource, install or validate guest components and confirm workload access.
- Diagnose a prepared device-visibility or compatibility fault.
Multi-Instance GPU and NVIDIA AI Enterprise
Outcome: Partition supported GPUs with MIG and deploy governed AI software components using NVIDIA AI Enterprise.
MIG Architecture and Use Cases
- Spatial GPU partitioning, hardware isolation and predictable resource allocation.
- GPU instances, compute instances, profiles and supported configuration patterns.
- MIG compared with time slicing, vGPU and full-GPU assignment.
- Suitable use cases for inference, development, smaller training and multi-tenancy.
MIG Configuration and Lifecycle
- Enabling MIG mode and creating, listing and removing instances.
- Mapping profiles to workloads and exposing devices to containers or schedulers.
- MIG-backed vGPU concepts and platform-dependent limitations.
- Reconfiguration, maintenance, persistence and troubleshooting considerations.
NVIDIA AI Enterprise
- Platform architecture, infrastructure layer and application software layer.
- Licensing, NGC access, release branches and lifecycle/support considerations.
- Drivers, NVIDIA Container Toolkit, GPU Operator and supported deployment patterns.
- Bare metal, virtualized and Kubernetes deployment paths.
Containerized AI Workloads
- Accessing approved containers and artifacts from NGC.
- Running a GPU-enabled container and validating resource visibility.
- Security, image provenance, secrets and controlled operational access.
HANDS-ON LAB | Partition a GPU and Deploy a Workload
- Create a supported MIG configuration and verify instance allocation.
- Expose a selected GPU or MIG device to a containerized workload.
- Validate isolation, utilization and basic workload health.
Monitoring, Operations, Troubleshooting and Capstone
Outcome: Operate NVIDIA AI infrastructure through repeatable monitoring, incident, change and capacity-management practices.
GPU Monitoring and Observability
- Using nvidia-smi and NVIDIA DCGM to inspect health and utilization.
- Core metrics: compute, memory, power, temperature, clocks, PCIe/NVLink and ECC.
- DCGM Exporter, Prometheus and Grafana integration patterns.
- Baselines, alert thresholds, dashboards and capacity trends.
Troubleshooting and Performance
- GPU discovery failures, driver/library mismatches and container-runtime issues.
- Xid events, ECC errors, thermal or power throttling and resource contention.
- Distinguishing infrastructure, virtualization, scheduler and workload problems.
- Evidence collection, safe remediation, escalation and post-incident review.
Operational Best Practices
- Driver, firmware and NVIDIA AI Enterprise lifecycle planning.
- Change windows, rollback readiness and compatibility validation.
- Capacity management, workload placement and service-level considerations.
- Access control, auditability, image governance and operational documentation.
- Runbooks for health checks, incident response and routine maintenance.
HANDS-ON LAB | Final Operational Capstone
- Review a scenario and select an appropriate GPU resource-assignment model.
- Validate the environment, deploy or inspect a workload and establish monitoring.
- Diagnose injected symptoms and document the evidence-based remediation path.
- Present a concise deployment and operations runbook for peer review.
Assessment and Completion
- Short knowledge checks at the end of each day.
- Completion of guided practical exercises and validation tasks.
- Participation in the final troubleshooting and operations capstone.
- NobleProg course completion certificate for participants who meet the agreed criteria.
Recommended Lab Environment
The final configuration depends on the client's target platform. A suitable lab normally includes:
- An NVIDIA data-centre GPU platform supporting the selected vGPU and/or MIG exercises.
- A supported hypervisor or Linux/KVM host with IOMMU/VFIO enabled where applicable.
- Supported Linux guest or bare-metal operating systems with administrative access.
- Required NVIDIA drivers, vGPU components and licence service access.
- NVIDIA AI Enterprise entitlement or evaluation access and an NGC account/API key.
- Docker or another supported container runtime; Kubernetes where agreed for the delivery.
- Monitoring components such as DCGM, DCGM Exporter, Prometheus and Grafana.
- Internet or registry access, or pre-staged packages and container images for restricted sites.
LAB ASSUMPTION | Hands-on vGPU, VFIO and MIG activities require compatible hardware, software and entitlements. NobleProg will validate the lab bill of materials and adapt exercises before delivery. Where a live configuration is not feasible, the affected exercise may use an instructor demonstration or prepared simulation.
Participant Deliverables
- NobleProg course materials and lab instructions.
- GPU infrastructure validation checklist.
- Sample monitoring and operational runbook structure.
- NobleProg electronic certificate of attendance or completion.
Requirements
Participants should bring a practical infrastructure background and access to the agreed laboratory environment.
- Working knowledge of Linux and command-line operations
- Familiarity with server hardware, networking and storage
- Basic understanding of virtual machines and hypervisors
- Foundational knowledge of AI, machine learning and GPU workloads
- Container and Kubernetes awareness is helpful, not mandatory
- Administrative lab access for practical exercises