Skip to content

Latest commit

 

History

History
223 lines (180 loc) · 10.1 KB

File metadata and controls

223 lines (180 loc) · 10.1 KB

Deployment

This directory contains infrastructure code for deploying AUP Learning Cloud.

Directory Structure

deploy/
├── ansible/    # Ansible playbooks for K3s cluster setup
├── k8s/        # Kubernetes components (NFS provisioner, device plugins)
├── scripts/    # Helper scripts for cluster setup
└── docs/       # Architecture diagrams

Documentation

For full deployment instructions, see the documentation site:

Quick Start

Single Node

cd ..
sudo ./auplc-installer install

Multi-Node Cluster

For SSH-preinstalled nodes, edit the Ansible inventory and multi-node values file directly. PXE remains generator-based because the controller inventory, rootfs settings, runtime overlay, and GPU policy must be generated as one consistent artifact set.

The AMD device plugin and ROCm node labeller are cluster infrastructure prerequisites owned outside AUPLC. The infrastructure owner must deploy and maintain them according to AMD's official guidance. If they are not installed, follow the pinned manual installation commands in the Kubernetes components guide. Before Helm, verify that the DaemonSets are ready and that GPU capacity is advertised.

SSH-preinstalled

Edit deploy/ansible/inventory.yml with the server and agent hostnames, IPs, k3s token, and other site settings. Keep the human template default, auplc_gpu_access_enabled: auto, unquoted on each host. auto runs Python 3 on that host to scan /sys/bus/pci/devices for vendor 0x1002 devices whose PCI class starts with 0x03. It does not use lspci or require pciutils. A match enables ROCm and the AMD GPU access package; a successful scan with no match skips both. If a scan fails, the play aborts before either is changed, and any_errors_fatal stops the play for all hosts.

Set an unquoted YAML boolean true or false only when you need to override detection. true forces ROCm and package installation, while false forces both to be skipped. Either boolean bypasses the scan. Don't quote any of these values or use alternatives such as yes and no.

For example:

k3s_cluster:
  children:
    server:
      hosts:
        controller-1:
          ansible_host: 192.0.2.10
          auplc_gpu_access_enabled: auto
    agent:
      hosts:
        gpu-worker-1:
          ansible_host: 192.0.2.11
          auplc_gpu_access_enabled: auto

Copy the human-maintained multi-node values example, then edit the copy for the site's authentication, storage, images, accelerators, and network access:

cd ..
REPO_ROOT="$(pwd)"
DEPLOY_SCRIPTS="$REPO_ROOT/skills/deploy-aup-learning-cloud/scripts"
cp runtime/values-multi-nodes.yaml.example runtime/values-multi-nodes.yaml
# Edit deploy/ansible/inventory.yml and runtime/values-multi-nodes.yaml.

python3 "$DEPLOY_SCRIPTS/validate.py" --repo "$REPO_ROOT" --topology ssh-preinstalled \
  --inventory "$REPO_ROOT/deploy/ansible/inventory.yml" \
  --values "$REPO_ROOT/runtime/values.yaml" \
  --values "$REPO_ROOT/runtime/values-multi-nodes.yaml" \
  --helm-dry-run

cd "$REPO_ROOT/deploy/ansible"
sudo ansible-playbook -i inventory.yml playbooks/pb-base.yml
sudo ansible-playbook -i inventory.yml playbooks/pb-k3s-site.yml
sudo ansible-playbook -i inventory.yml playbooks/pb-rocm.yml

kubectl rollout status -n kube-system daemonset/amdgpu-device-plugin-daemonset --timeout=5m
kubectl rollout status -n kube-system daemonset/amdgpu-labeller-daemonset --timeout=5m
kubectl get nodes -o 'custom-columns=NAME:.metadata.name,AMD_GPU:.status.allocatable.amd\.com/gpu'

cd "$REPO_ROOT"
helm upgrade --install jupyterhub ./runtime/chart \
  --namespace jupyterhub --create-namespace \
  -f runtime/values.yaml \
  -f runtime/values-multi-nodes.yaml

With --inventory alone, the validator accepts exactly one unquoted auto, true, or false value for auplc_gpu_access_enabled on every managed host. This validates the direct-edit workflow without a generated GPU resolution report. --gpu-resolution may be supplied only with --inventory; that pair is for generated artifacts, whose inventory values and resolution entries must remain strict booleans. The generator never writes auto.

The installer, Ansible GPU access role, and PXE controller install AMD's amdgpu-insecure-instinct-udev-rules package, pinned to version 30.30.4.0-2341068.24.04. Its package-owned rule sets mode 0666 only on /dev/kfd and DRM /dev/dri/renderD* nodes. It does not match /dev/dri/card*; card nodes retain the normal system policy, observed as root:video 0660.

This host permission policy is separate from Kubernetes allocation. The AMD device plugin remains the visibility boundary: only Pods that request amd.com/gpu receive allocated GPU devices, and the plugin does not change host inode ownership or mode. AUPLC Hub adds no GPU supplemental group. The tested ROCm compute path needs none: on representative GPU nodes, rocminfo succeeded as UID 12345 with only supplemental GID 100, while card nodes remained inaccessible at mode 0660. The reported agents were gfx1151 and gfx1200.

singleuser.fsGid: 100 controls shared notebook storage ownership only. It is not part of GPU access and must not be treated as a GPU group setting.

PXE-diskless

Create a fresh spec, set topology to pxe-diskless, fill the PXE network fields, and set pxe.diskless_agents_have_amd_gpus explicitly. Diskless agent hardware can't be inferred from the controller. Generation writes the canonical inventory, controller vars, runtime overlay, and GPU resolution report directly. These artifacts express the desired deployment inputs; their existence is not proof that rootfs provisioning succeeded. Review and install them before running the controller playbook, whose successful completion provisions the rootfs.

cd ..
REPO_ROOT="$(pwd)"
DEPLOY_SCRIPTS="$REPO_ROOT/skills/deploy-aup-learning-cloud/scripts"
python3 "$DEPLOY_SCRIPTS/gen_configs.py" --print-schema > spec.json
# Edit spec.json: choose pxe-diskless and fill the node, network, and PXE fields.
GENERATED_DIR="$REPO_ROOT/generated"

cd "$REPO_ROOT"
python3 "$DEPLOY_SCRIPTS/gen_configs.py" --spec spec.json --out-dir "$GENERATED_DIR"
install -m 0600 "$GENERATED_DIR/inventory.yml" "$REPO_ROOT/deploy/ansible/inventory.yml"
install -m 0644 "$GENERATED_DIR/values-basic-example.yaml" "$REPO_ROOT/runtime/values-basic-example.yaml"
python3 "$DEPLOY_SCRIPTS/validate.py" --repo "$REPO_ROOT" --topology pxe-diskless \
  --inventory "$REPO_ROOT/deploy/ansible/inventory.yml" \
  --gpu-resolution "$GENERATED_DIR/gpu-access-resolution.json" \
  --values "$REPO_ROOT/runtime/values.yaml" \
  --values "$REPO_ROOT/runtime/values-basic-example.yaml" \
  --pxe-vars "$GENERATED_DIR/pb-pxe-controller.vars.yml"

cd "$REPO_ROOT/deploy/ansible"
sudo ansible-playbook \
  -i "$GENERATED_DIR/inventory.yml" \
  playbooks/pb-pxe-controller.yml \
  -e @"$GENERATED_DIR/pb-pxe-controller.vars.yml"

kubectl rollout status -n kube-system daemonset/amdgpu-device-plugin-daemonset --timeout=5m
kubectl rollout status -n kube-system daemonset/amdgpu-labeller-daemonset --timeout=5m
kubectl get nodes -o 'custom-columns=NAME:.metadata.name,AMD_GPU:.status.allocatable.amd\.com/gpu'

cd "$REPO_ROOT"
helm upgrade --install jupyterhub ./runtime/chart \
  --namespace jupyterhub --create-namespace \
  -f runtime/values.yaml \
  -f runtime/values-basic-example.yaml

A fresh PXE rootfs receives the pinned AMD udev package during the controller playbook. A retained rootfs is accepted only when that exact package version and its unmodified package-owned rule are present, with no conflicting legacy GPU rule. Rebuild or correct a retained rootfs separately if that safety check fails.

Generator discovery failures

Error Action
Host is unreachable Restore passwordless root SSH to that inventory host, then regenerate.
lspci is missing or fails Install pciutils on the reported host and rerun generation.
Host evidence is UNKNOWN or AMD GPU BDF probes disagree Compare AMD display BDFs from lspci with vendor 0x1002 display-class devices under /sys/bus/pci/devices; fix missing or inconsistent PCI enumeration, then regenerate.
Retained PXE rootfs has the wrong AMD udev package version, a modified package rule, or a conflicting legacy GPU rule Rebuild the rootfs, or correct the package state through a separate reviewed maintenance action before rerunning the playbook.

Deployment branch boundary

This branch and these instructions do not modify or roll out any live deployment. Environment-specific deployment branches must backport the host permission and immediate artifact publication changes before their own reviewed rollout.