agentlab --help and agentlab <command> --help are authoritative; this page
is the map. Every command runs against one lab — a directory holding
agentlab.yaml (written by configure; on a terminal, a missing file starts
the form), certs/ and state/ — and talks to its kind cluster through the
cluster's own exported kubeconfig, state/kubeconfig, never your shell's
current-context.
configure and up register their lab under its clusterName in
~/.config/agentlab/labs.yaml (~/Library/Application Support/agentlab/ on
macOS), so the other commands find it from any directory, in this order:
--lab <name>(every command), refused naming the registered labs when no lab of that name is registered;- the
agentlab.yamlin the current directory: the lab itsclusterNamenames (a copy of a registered lab's file enters that lab's directory), else this directory; - a picker among the registered labs on a terminal, and off one a refusal
that names them and
--lab.
No lab is a default, not even the only one registered: a command runs against
the lab --lab, the directory or the person names, or refuses. An
agentlab.yaml without a clusterName is refused too: the file names its lab.
A lab's directory is where its agentlab.yaml, certs/ and state/ live: the
checkout for a lab you made there, ~/.local/state/<lab> for a leased lab
(agentlab-1, agentlab-2).
A lab entered from another directory is named on stderr (Lab agentlab (/path/to/lab)). configure and up skip step 3: in a directory without
agentlab.yaml they create a lab there. A clusterName another lab
directory holds is refused, naming both directories — the two would collide
on the kind cluster and its ports. down keeps the registration (the lab
still exists as a directory); a lab whose directory or agentlab.yaml is gone
is neither offered nor listed, and back once its agentlab.yaml is (the next
registration drops the entries still gone). agentlab list also shows every
kind cluster no registered lab names, marked lab directory unknown: any
command run in a lab's directory (agentlab pods) registers the lab. self-update, completion and help touch
no lab.
Nothing asks without a terminal (stdin not a TTY): no first-run question, no
lab picker, no trust or open offer. A missing agentlab.yaml is refused with
a pointer to agentlab configure --defaults, which writes the canonical lab
without a question; registered labs and no agentlab.yaml here are refused
naming them and --lab; up --trust/--open pre-answer the end of a
boot. CI, coding agents and scripts write the config first and name the lab.
The sections below are the groups agentlab --help prints, in the same order.
| Command | What it does |
|---|---|
up |
Check docker's CPUs and memory against this configuration's floors, then create the kind cluster, deploy Dex and the enabled components, and verify the OIDC chain end to end. Idempotent: unchanged re-runs are no-ops. On a terminal it ends by asking what the summary used to only describe: whether to trust the lab CA while it is untrusted, then whether to open the portal. --trust and --open (or --trust=false/--open=false) pre-answer both for scripted runs; off a terminal nothing is asked. See TLS. |
configure |
Discover this machine, then ask for the lab configuration (or keep it with --defaults) and save agentlab.yaml. Without an agentlab.yaml yet, on a terminal, it first asks one question — the detected defaults, or customize — and only Customize runs the form; up (and every command that has to create the file on the way) asks the same. With an existing file the form opens with its values. Flags below. |
trust |
Install the lab CA into the system and browser trust stores (one sudo prompt; reversible). See TLS. |
| Command | What it does |
|---|---|
open <portal|agents|prometheus> |
Open a lab URL in the browser and print it (so it can be copied where the opener finds no browser — SSH, WSL): portal is Backstage, agents the kagent UI, prometheus the lab Prometheus's web UI on the edge (https://observability.<domain>/prometheus, the path Backstage queries). The target is required; without one, the refusal names them. Refused when the target is disabled in agentlab.yaml, when the cluster is not running, or when the target does not answer yet (a short reachability probe, before any trust question, so a sudo prompt is never spent on a command that then opens nothing) — a plain fact instead of an error page in the browser. While the lab CA is untrusted, open portal and open prometheus offer agentlab trust first on a terminal and warn off one; open agents is plain HTTP on loopback and asks nothing. |
list (ls) |
Every lab registered on this machine: its name, directory, cluster state (running, exited, not created), enabled components with the chart, the portal and muster URLs (with the port suffix when the edge is not on 443), and whether its lab CA is trusted here; * marks the current directory's. Docker and the files answer, so it works while a lab is down. -o json for scripts. It never asks which lab. |
status |
The lab's live state, for the lab here (or --lab): the chart the cluster runs — its version, whether it is a dev chart (a chart directory's build or a branch's dev build), its Helm revision and which agentlab installed it from where — against the chart agentlab.yaml installs; when the two differ, the drift with both versions and which side is behind (an agentlab.yaml put back from a copy without a platform run pins an older release than runs) and the two ways to end it — agentlab platform installs the configured chart over the one in place, agentlab configure --adopt-chart writes the chart in place into the config; and the platform HelmReleases, how many are Ready and every other one with helm-controller's message. Refused with agentlab up when the cluster is not running. -o json for scripts (drift carries the line). With platform.workspaces, a workspaces line: the NFS server and its node, the NFS CSI controller with its mTLS proxy, the node plugin, the StorageClass, and whose CSIDriverConfig registers the driver with Substrate (the chart's, the lab's, or one applied by hand), worded as the platform summary words it. |
pods [-n <namespace>] |
The lab's pods across namespaces through the embedded client, shaped like kubectl get pods -A (NAMESPACE, NAME, READY, STATUS with a waiting reason such as ImagePullBackOff, RESTARTS, AGE); -n narrows to one namespace. Refused with agentlab up when the cluster is not running. |
logs <component> |
Tail a component's logs: backstage, dex, klaus-gateway, mcp-prometheus, muster or prometheus. |
login [email] |
Log in as a lab user: the headless password grant, or --browser for the real Dex login page (authorization-code flow, which asks for the user itself). Either way it prints the token claims and writes .token and kubeconfig.oidc for that user. --password overrides the one in agentlab.yaml. |
turn (--list | --template <name> <prompt>) |
One conversation with an agent as a lab user, the way the surfaces drive it: the user's Dex id_token on the kagent controller's gRPC route through the edge. --list prints the roster that user sees — ListAgents as that person (each Agent with the AgentTemplate and the Harness it pairs), or the gRPC status when the controller refuses. With --template and one prompt it creates a Session of the Agent that pairs the template with a Harness — the template's Agent that reports Ready, the one the portal picks; --harness picks the Agent pairing the template with that Harness — streams one turn, prints the answer and the terminal task state, and deletes the session. --user picks the lab user (default: the first admin). --keep leaves the session for a later --session <id> turn (--instance is the former name), and --session … --suspend suspends it to the snapshot store and resumes it before the turn, the proof that a conversation survives the snapshot. The terminal status message's metadata (a harness's usage of the turn) is printed when present. A turn that pauses for tool approval prints the request and stops; --decide approve or --decide reject (with --reason) answers every request until the task settles. --events <file> writes every streamed A2A event as one JSON object per line, the raw material of a parser fixture. --session <id> --share-with <email> has the --user share the session read-write with another lab user (--share-ttl bounds how long the share's token grants access; default until revoked), runs the turn as that user with the share token (their own Dex token plus the token, which adds to their identity), and revokes the share afterwards. |
kubectl --kubeconfig kubeconfig.oidc after login is the OIDC path — the
way to verify what a specific user can do. The kind admin context bypasses the
platform and OIDC entirely.
| Command | What it does |
|---|---|
test |
Assert RBAC for every configured user: a token from Dex, then one SelfSubjectAccessReview per expectation of each group — kubectl auth can-i, asked of the apiserver in-process with that token alone. |
The rest of the group are the lab's end-to-end proofs. Each one logs in to
Dex headlessly, drives the real components through the same paths a person
would, asserts the result and leaves nothing behind — with one documented
exception, models-test on a server that cannot delete a model (see below).
Where a command takes [email], that user runs the proof (default: the
admin). They trust certs/ca.crt directly, so they need neither agentlab trust nor any Node setting. The agent proofs run on kagent API v2, the one
line the platform ships: every agent a HelmRelease of the Generic agent chart
(1.x) rendering an AgentTemplate on the platform Harness, created through
one helper — see Agents.
| Command | What it proves |
|---|---|
platform-test [email] |
Dex → muster → mcp-kubernetes → apiserver; the APIs the lab provides itself served; the fleet's flux-multi-tenancy policy enforcing on HelmReleases in org-lab — four shapes as server-side dry runs, denied with the fleet's message where an installation denies (no tenant; a tenant installing into another namespace without a kubeconfig) and admitted where it admits (serviceAccountName: automation; kubeConfig.secretRef) (the identity proof: a viewer's kube-system Secrets Forbidden by the apiserver, a forged x-user-id changing nothing), with agents on the kagent controller route's JWT Strict policy (a call without a token refused at the edge; a valid token with a forged x-user-id attributed to the token's subject by the controller) and agent-manager writing as the caller (a viewer's create refused under the viewer's name), the atelet's image-cache policy and the two halves of Agent Substrate on one release (the atelet's image against the WorkerPool's worker image — a skew boots no golden actor) and the WorkerPool's workers Running on a node that carries its node selector (a Pending worker fails the proof with the scheduler's reason), with workspaces on the workspace-manager (registered with muster and Connected, rolled out, list_providers as the admin naming every configured provider instance), the per-server OAuth sign-in challenge, the infrastructure families (the lab's cluster as the kubernetes and prometheus members, tool-group: infrastructure, no family-less mcp-kubernetes) and, with observability on, mcp-prometheus and Backstage's metrics endpoint. Every stage is named with its outcome and duration in a summary before the verdict; a stage the lab's state cannot run fails the proof naming the stage, and only the stages agentlab.yaml leaves without a subject (agents or observability off, the 3.x line) are listed as SKIP with the reason. See The agent platform. |
vm-manager-test [email] |
The vm-manager pod as the person: GET /api/v1/host from inside the cluster refused without a token and answered with the person's Dex id_token (the token muster forwards); the x_vm-manager_* tools through muster with their annotations intact (get_host read-only, delete_vm destructive), get_host, list_images, list_networks; the MCPServer with the agent-platform tool group, Connected; then create_vm on the newest image followed with get_vm to ready (the installer boot, the installed boot, READY=1), the attestation verdicts where the image policy has golden values, the console tail, delete_vm. --skip-vm boots nothing; --vm-timeout bounds the boot (default 6m). See vm-manager. |
models-test [email] |
Managed models through muster's x_model-manager_* tools: 401 at muster without a token, then pull → ModelConfig → every write's dry run, the gitops_owned refusal and a wire_model commit opened as the person on a fake GitHub → agent turn → MCP via muster → unload → delete, one backend per run. On lmstudio the teardown is the unsupported refusal plus an unwire, since LM Studio serves no delete, and that run leaves its model downloaded. --backend picks one of platform.modelManager.backends (default: the first); --model a small, tool-calling capable model. See Models. |
serving-test [email] |
Model serving on llm-d (platform.serving): the llm-d controller, the well-known template and the models Gateway up; 401 at muster without a token, then model-manager's tools through muster as the person; the lab preset listed among the shipped ones and fitting the node (a CPU preset, the allocatable budget) → load → the LLMInferenceService Ready on the CPU runtime, no accelerator requested → the ModelConfig wired at the model's route on the models Gateway → a completion through the Gateway (401 without a token, 200 with the person's) → the agent turn on the wired ModelConfig must answer (Substrate's egress gateway trusts the lab CA through substrate.atenetEgress.upstreamTrust; with the LLM endpoint on, the ModelConfig rides the endpoint under the public name and the turn's tokens must show up in the data plane's per-model token metric) → unload, nothing left behind. --preset serves another published preset, --skip-chat skips the agent turn, --ready-timeout bounds the serve (default 20m). See Model serving on llm-d. |
pm-test [email] |
The platform-manager fixture (platform.platformManager): the lab's registry fixture as a container on the kind network → the released giantswarm-platform-manager (platform.platformManager.version, default the pinned release) as a temporary HelmRelease pinned to the lab Dex and to the fixture as its GitHub → the person signed in to it through muster → get_info naming the installed version → reconcile_capability's dry run of the fixture's installation ember, workspaces on: the Dex client muster carries https://workspace-manager.ember.agentlab.test/signin, the workspaces and workspace-manager values are kept, the commit is not refused and nothing is written → the release and the fixture removed. Refuses while platform.platformManager.enabled is false. See Platform manager. |
workspaces-test |
The workspaces proof (platform.workspaces): the NFS server, the NFS CSI controller with its mTLS proxy, the node plugin and StorageClass agentlab-workspaces in place → a 1Gi read-write-many claim of the class bound → a bare mirror seeded on it from a repository with an executable and a symbolic link → session a: a pod with its own sessions/a sub-path read-write and mirrors/ read-only, a git clone --shared of the mirror with the executable bit and the link intact, a commit, a write into the mirrors refused → session b likewise, nothing of session a visible → the whole volume mounted read-only, a write refused → the controller endpoint refused without Substrate's client certificate (the TLS alert quoted) → the actor-level mount through ate-api-server on the Substrate chart version it names: an ActorTemplate with two existing volumes, two actors on the claim's volume at sessions/a and sessions/b with the mirrors read-only, each seeing only its own session and refused a write into the mirrors, a pause and resume keeping the process and the content, both deleted with the volume and its files left, an unknown driver and a missing handle refused at create (a Substrate without existing volumes refuses the template: the skip, with ate-api-server's reason) → the namespace agentlab-workspaces-test removed (first and last) and the volume's directory gone from the export. Before that, unless --storage-only: a claude Harness Session admitted on the volume with a workspace grant the manager's key set verifies (a Session without a volume, the refusals of no grant, a foreign key and another person, the Session's turn in its own directory with the mirror read-only and its clone, its file read back on the volume); --ready-timeout bounds each wait (default 5m). See Workspaces. |
agents-test [email] |
agent-manager through muster as the signed-in user, on the Generic chart line agent-manager declares (its --agent-chart-semver, the 1.x or the 2.x line): get_info on that line and the platform Harness; create with a toolset and a git skill (refused without a toolset) → the HelmRelease as agent-manager's write (its field manager, values.toolset, the skill pinned) next to the OCIRepository on the declared line → the render (the AgentTemplate with the Harness label and the display-name and icon-url annotations, the agent's RemoteMCPServer with the toolset header) → Ready on the Harness with the installed chart on the declared line (a chart off it fails, naming both), get_agent_status agreeing with status.harnesses[] → a turn as the person answering from the skill → update → refreshSkills re-pinning to list_skills' head → delete removing the release and the render; a template no Harness admits reported failed with the reason; a viewer's create Forbidden by the apiserver; the ServiceAccount holds no RBAC of its own. |
agents-gitops-test [email] |
Editing an agent applied from git, as the portal does it: a release agent-manager composed, applied with a Flux Kustomization's provenance labels and read back managed: gitops; the change's dry run on the platform's agent-manager gitops_owned in mode apply and unsupported in mode commit (no App pin there); then against a temporary copy of the agent-manager release pinned to the lab Dex as its GitHub App and to the fake GitHub API: validate_agent gitops_owned in mode apply, force or not, a valid update with the manifests in mode commit, and update_agent mode commit opening one pull request as the person whose file is the dry run's manifest, the live release untouched, a mode apply write still refused. --github-fake-binary names another static Linux agentlab for the fake. See Agents. |
toolsets-test [email] |
Declared toolsets end to end: agent-manager requires one and writes it as values.toolset on the agent's HelmRelease, the chart renders it as the header on the agent's own RemoteMCPServer (none for preset:none; none on a release without a toolset — implicit full access), muster resolves and refuses per request, agents see their toolset through a turn on kagent, a per-server sign-in scopes a server's tools to the token, the portal's Tools step and its create path through the portal's muster backend. --model-config picks the kagent ModelConfig the throwaway agents run on (a host model without an Anthropic key); --skip-chat skips the turns that need a model to answer; --skip-portal skips the portal's create path. See Toolsets. |
skills-test [email] |
The golden boot under Substrate's egress gate: an AgentTemplate on the Go ADK Harness with one git skill pinned to a full commit of a public repository (giantswarm/agent-skills), written below the agent chart on purpose; Ready on the Harness, then one turn through the edge as the user that names the skill and answers a fact only its SKILL.md has. A failed boot prints the evidence (the Harness conditions and warnings, Substrate's ActorTemplate, actor and worker as the controller reports them, the controller, atenet and worker log lines, the versions) and boots the same template without the skill as the control. A second turn in a new Session has the agent's bash tool run git ls-remote of the source from the sandbox, without Authorization and with kagent's placeholder, to show whether the Session's egress carries the source's credential. --skill-fixture boots the skill from the lab's own private git host behind the edge (skillhost.<domain>, no GitHub token; a Secret the proof creates and never prints) and prints golden fetch: credential sent and session request: no credential from the host's record of every request, or FINDING: session request: credential sent with the request that carried it. --skill-repo and its siblings boot another fixture, --skill-secret a private one; there the probe reports the host's exit codes. --ready-timeout bounds the boot (default 10m); --model-config picks the kagent ModelConfig the agent runs on (a host model without an Anthropic key). See The skills proof. |
a2a-test [email] |
Turns over native gRPC through the edge as Swarmgeist and the portal backend drive them: the GRPCRoute and its JWT policy Accepted; no token refused at the edge (Unauthenticated; a plain POST gets 401); a forged x-user-id beside a valid token replaced; ListAgents listing the proof's Agent Ready with its display-name and icon-url annotations; CreateSession idempotent on request_id; SendStreamingMessage with the HITL extension streaming an answer; a requireApproval tool call paused at input-required → approved → completed with muster logging the call under the person, rejected → ended without the call; CancelTask on a running turn canceled server-side and a following turn answered. --ready-timeout bounds the fixture agent's boot (default 4m). See Turns through the edge. |
substrate-step-test [email] --to <release> |
Agents keep their conversation and durable volume across a Substrate release: the Substrate before the step read off the atelet, the WorkerPool and ate-api-server; an agent on the platform Harness with a conversation of three turns (a codeword, a marker its bash tool writes to /data, the codeword asked back), suspended; the substrate and substrate-crds components pinned to --to through a values overlay of the run's own and the platform upgraded in place, the workers rolled onto the new worker image; the Session resumed, answering the first turn's codeword and reading the marker back; the template's golden taken again on the new release and a new Session answering. PASS or FAIL with the cause per step (Substrate 1.8.0's refusal of a 1.7 snapshot named as the restore refusal), the agent removed and the lab back on the chart's Substrate (--stay leaves it on --to). --step-dev-image <target>=<image> adds dev images for the step alone, --ready-timeout and --retake-timeout bound the golden boot and the re-take (default 10m each), --model-config the ModelConfig. See The Substrate step proof. |
klaus-gateway-test [email] |
The Swarmgeist proof through klaus-gateway's Slack adapter: the gateway runs on the host (the 4.1 release candidate by default, --gateway-image; a local build with --gateway-binary) against the lab's public gRPC target grpcs://agentgateway.<domain>:<port> with the lab CA and the JWT policy at the edge, Slack in events mode on a fake Web API the proof serves, the person linked by a record in its OBO link store carrying the user's Dex id_token. Messages are signed Events API callbacks, Approve and Deny Block Kit clicks: @bot agents lists the fixture's AgentTemplate and hides one no Harness admits, the picker its Select button opens hides it too, and a submitted selection of it is refused with the reason; an unlinked person is asked to sign in and reaches no controller; one streamed turn under the template's display name and icon binds the thread to one Session and is attributed to the person at muster; a requireApproval binding pauses the task at input-required, Approve resumes it to completed, Deny in a second thread ends one without the call; a stop reply has the task cancelled at the controller and the thread goes on; a restart on the stores keeps the thread → Session mapping; a restart in the middle of a turn, with the edge out of the restarted gateway's reach for 45 s, still posts the answer in its thread without a reply. --gateway-port (default 18090; the admin endpoints and the in-process fake take the next two, the loopback gate in front of the edge the third), --slack-fake-binary (the static Linux agentlab the fake's container runs while the component is on; default this binary), --run-dir keeps the stores, keys and the gateway's log, --ready-timeout bounds the fixture's boot (default 10m), --model-config picks the kagent ModelConfig the fixture runs on (a host model without an Anthropic key), --harness the Harness whose admission labels the fixture carries and whose Ready entry the boot waits for (default the platform Harness kagent; a coding Harness such as claude with --model-config naming a ModelConfig it accepts). A 0.x gateway image is the documented negative. With platform.klausGateway on, the meta chart's klaus-gateway component follows: the Role scoped to the OBO link Secret and no store volume, two links written through the gateway's store package with the lab's key, the pod deleted and its replacement Ready with the same links within 60 s, one Slack turn through the pod over the in-cluster target, the fake then a container on the kind network. See The Swarmgeist proof and klaus-gateway. |
decisions-test [email] |
The proof of klaus-gateway's decisions (POST /decisions): the gateway on the host (--gateway-image, default the first release that serves them; --gateway-binary for a branch) with its reviews endpoint, a ServiceAccount token with audience klaus-gateway the lab's API server vouches for (TokenReview), the fake Slack Web API. A team decision answered by a Choose click, a person's (found by email, a direct message) in the modal with an option and own words, a team decision by a reply in its thread — each answer muster's core_mcpserver_list called as the user, attributed at muster; a tool's refusal as the status line with the decision open; a close as defaulted that refuses a later click. No agent, no model. |
beekeeper-decisions-test |
The proof of beekeeper's decisions in Slack: beekeeper serve on the host (--beekeeper-version, default the first release that puts decisions to people, its binary and CRDs; --beekeeper-binary for a branch) over the lab cluster's beekeeper resources, with its mailboxes in a throwaway Postgres container, registered in muster as the MCPServer beekeeper with the person's token forwarded; klaus-gateway and the fake Slack as decisions-test runs them, beekeeper calling the gateway with a ServiceAccount token of audience klaus-gateway. A decision for the admin filed through muster arrives as a direct message: another user's click is refused, the admin's answers it; a decision for team:agentlab is answered by another member's thread reply; one nobody answers closes with its default at its due time and refuses a late click; one its filer withdraws (note_done) loses its buttons; note.answered (via slack) and note.defaulted Events in beekeeper-agentlab. The admin's guide (local:agentlab/Guide, registered through muster) converses with them: its converse opens a direct message, their reply in its thread reaches the guide's mailbox via slack, its answer lands in the same thread; another user's send_message to the guide is refused, and with the guide's RosterEntry gone a reply gets the thread's not-delivered note. No agent, no model. |
beekeeper-central-test |
The proof of beekeeper's central resources in the local binary: beekeeper serve on the host as beekeeper-decisions-test runs it (--beekeeper-version, default the first release whose local binary reaches it; --beekeeper-binary for a branch, used by serve and both machines), with the Environment agentlab-shared and its MergeLane, and two local beekeepers, one per lab user, each with a configuration and state of its own whose central context is the lab's muster. A lease the admin's machine claims is refused to the other user's with its holder and listed there; a hold on the central lane set on one machine stops the other's merges and is listed as central; two merges into the lane take their turns through muster, the second's once the first merged and left, and the other user cannot take the first out; the admin's registered agent appears on the central roster after one watch --once; with the central instance unreachable a central claim exits 69 and the machine's own resource is still claimed. No agent, no model. |
backstage-test [email...] |
The headless Backstage sign-in and the muster hop with that user's own forwarded token, including the per-server Sign in challenge and the MCP servers page's grouping; with agents on, the Dev Portal's Agent Platform pages on kagent API v2 as a person drives them — the wizard's create path through agent-manager (skill discovery, the dry run, Deploy as the first platform-admin), the agents list per user, a streamed and a resumed conversation, HITL and Stop, edit, skills update and delete; with managed models, the Models page's model-manager read and write through the same muster hop as the person and a viewer's write refused; everything it creates removed on every path (default: every user). See Backstage. |
| Command | What it does |
|---|---|
down |
Destroy the kind cluster. certs/ is kept and the trust stores are untouched. |
untrust |
Remove exactly the lab CA from the system and browser trust stores. |
platform-down |
Remove the agent platform in the chart's ordered teardown, leaving Dex and the cluster alone. The workspace storage goes with it (its namespace, classes, CSIDriver and CRDs), and the driver's mounts and volume bytes are cleaned off the node; down cleans the node the same way before the cluster is deleted. |
| Command | What it does |
|---|---|
platform |
Install the agent platform on a running cluster: the lab Dex re-applied first (the users and the GitHub sign-in as agentlab.yaml has them; unchanged, a no-op), then one idempotent upgrade-or-install of the agent-platform chart in its lab shape through the embedded Helm (the helm upgrade --install --wait of Helm 4, in-process), then wait for every component. A re-run that would install the same chart version with the same values writes no Helm revision. up runs this when the platform is enabled. On the dev channel it first re-resolves the branch's newest build (and installs Substrate); --pin freezes the recorded build instead, --pin=false follows the branch again. Like up, it ends with the trust and portal questions on a terminal, and takes the same --trust/--open. |
certs |
Generate the lab CA and Dex server cert, re-minting only what config or policy require. --force regenerates everything and breaks a running cluster's trust. |
render |
Render every manifest from agentlab.yaml into state/ without applying anything. |
reload |
Re-render and re-apply the Dex config after editing agentlab.yaml (users, passwords, groups, the GitHub sign-in), without touching the platform. |
self-update |
Replace the binary with the latest GitHub release, once its cosign Sigstore bundle verifies (see below). --check only reports the running and the latest version, exit status 125 when a newer one exists. |
Backstage has no command of its own: it deploys with the platform
(backstage.enabled in agentlab.yaml and agentlab up), and
agentlab open portal is how you reach it.
configure probes the machine on every run and applies what it finds before
asking anything — see what it
discovers.
The flags pin a value regardless of the discovery, with or without
--defaults; nothing else in an existing file is touched.
| Flag | Effect |
|---|---|
--defaults |
Skip the form; keep the current values (or write the canonical lab) plus what the discovery finds. |
--accessible |
Prompt-per-question form mode, for screen readers and plain terminals. |
--platform[=false] |
Enable or disable the agent platform. Off gives a bare kind + Dex OIDC sandbox. |
--agents[=false] |
Enable or disable the agents runtime (kagent, part of the platform install). |
--observability[=false] |
Enable or disable the observability stack (Prometheus + mcp-prometheus). |
--backstage[=false] |
Enable or disable Backstage (implies the platform). |
--chart-version <x.y.z> |
The agent-platform chart release to install, an exact version (default: the release this agentlab was verified with, config.DefaultChartVersion). |
--chart-path <dir> |
Install the agent-platform chart from a local checkout's helm/agent-platform directory instead of the pinned release, with the checkout's connectivity chart beside it (pushed into the lab registry at the meta chart's version; refused without the sibling directory); --chart-path "" clears it. See Installing an unreleased chart. |
--chart-branch <branch> |
The dev channel: follow this agent-platform branch's newest dev build — resolved now and on every up/platform, written to chartVersion; --chart-branch "" returns to the stable channel and, unless --chart-version pins one in the same call, puts the default release in place of the branch's build the resolver left in chartVersion — a lab that leaves the dev channel does not keep its last dev build. Mutually exclusive with --chart-path; implies Substrate. See Dev channel. |
--upgrade-seed[=false] |
Seed an upgrade proof: a released chart below agentlab's floor (agent-platform 4.93.0) installs in the shape of its line instead of being refused; --upgrade-seed=false clears it. See Upgrade proofs from an older line. |
--adopt-chart |
Write the chart the running lab cluster carries into chartVersion, so agentlab.yaml says what the cluster runs — the restore step for a config put back from a copy apart from its cluster (agentlab status names the drift). A release puts the lab on the stable channel at that release (a chartPath or chartBranch the config named gives way); a dev build of the followed branch pins the dev channel at it. Refused for a chart directory's build, whose version is a placeholder, and when the lab is not running. Not with --chart-version, --chart-path or --chart-branch. |
--model-manager[=false] |
Pin managed models on or off (needs agents). Without the flag a first configure follows the host model servers the discovery finds, and a later one keeps the recorded choice: configure --defaults never turns on what the lab was configured without. Off is the chart's component off too: the rendered values state components.model-manager.enabled: false, since the meta chart runs model-manager by default from 4.24.0 (with no backend). |
--model-manager-backends ollama,lmstudio |
Pin the host model servers, in order (ollama, lemonade, lmstudio); the first is model-manager's default backend. |
--vm-manager[=false] |
Run the platform's VM provisioner (vm-manager) as a pod of the node, or turn it off; needs /dev/kvm and /dev/vhost-vsock on this machine (a machine without them turns it off on its own). |
--vm-manager-image-dir <dir> |
The image directory the vm-manager pod boots from: a vm-manager checkout's images/build after make -C images, mounted into the node at agentlab up (agentlab down && up after changing it). |
--klaus-gateway[=false] |
Run Swarmgeist (klaus-gateway) as the meta chart's in-cluster component, or turn it off: A2A on the in-cluster controller target, Slack on a placeholder Secret (its Web API the proof's fake), the OBO link store in a Secret (needs agents). See klaus-gateway. |
--github[=false] |
Register GitHub's hosted MCP server with muster as MCPServer github, signed in to as the person through an OAuth App or GitHub App client read from the Secret platform.github.secret names (default agent-platform/github-oauth-client, keys client-id and client-secret), or remove the server. agentlab never reads the client's values: the operator places the Secret. See GitHub as the person. |
--github-signin-client-id <id> |
GitHub sign-in through the lab Dex: the GitHub App's client id, which turns the sign-in on ("" turns it off). The App's callback URL is the lab Dex's own callback (<issuer>/callback); its client secret is the Secret dex/github-signin-client (key client-secret), placed by the operator's secret tooling and never read by agentlab; platform.githubSignIn.secret names another Secret in the dex namespace. See GitHub sign-in. |
--github-signin[=false] |
Turn the GitHub sign-in on (once a client id is recorded) or off; the client id and the Secret stay for the next time. |
--github-signin-orgs giantswarm,other |
GitHub sign-in: admit members of these GitHub organizations only, their teams as the token's groups (<org>:<team-slug>); empty admits any GitHub account. |
--serving[=false] |
Serve models on llm-d in the lab, or turn it off: the KServe llmisvc controller and its CRDs, the well-known runtime configs, the connectivity chart's serving slice with the models Gateway, model-manager's kserve backend and one CPU preset of the lab's (needs agents; installs cert-manager; agent-platform 4.44.0 or newer). See Model serving on llm-d. |
--workspaces[=false] |
Install the workspace storage, or turn it off: an in-cluster NFS server, the NFS CSI driver nfs.csi.k8s.io with its controller behind an Envoy mTLS proxy only Agent Substrate's API server may reach (its pod-identity certificate), and the read-write-many StorageClass agentlab-workspaces; needs agents and the host kernel's nfsd and nfs modules (available, not loaded by hand: the kernel loads them). See Workspaces. |
--workspaces-provider <fake|github|fake,github> |
The workspace-manager's provider instances: fake (the default, "" too), the lab's own GitHub run by the lab as a GitHub Enterprise-shaped instance at https://github.<domain>, nothing to register; github, a GitHub App of the lab's own on github.com, the instance once the App's ids are set and the Secret agent-platform/workspace-github carries private-key and client-secret; or both. With workspaces on, the base URL https://workspace-manager.<domain>:<gatewayPort>, its route and the Dex redirect URI <base URL>/signin are in place regardless. See Workspaces. |
--workspaces-github-app-id, --workspaces-github-client-id |
The public ids of the github instance's App (platform.workspaces.github); its private key and client secret never pass through agentlab. |
--ai-key-source <ref> |
Where the Anthropic key of the agents' default ModelConfig lives: a reference beekeeper secret copy resolves (op://<vault>/<item>/<field>, or <file>#<path> of a SOPS file), never a value. Every up and platform place it into the Secret kagent/kagent-anthropic through beekeeper (on PATH, or configure refuses), so a recreated lab carries the key without a manual step; --ai-key-source "" clears it (the key then comes from $ANTHROPIC_API_KEY, else a placeholder). See The model. |
--github-token-source <ref> |
Where the GitHub token of the portal's skill discovery and agent-manager's skill resolution lives: a reference beekeeper secret copy resolves, as for --ai-key-source, never a value. Every up and platform place it into the Secret agentlab-github-token in agent-platform and kagent through beekeeper, so a lab started without $GITHUB_TOKEN (an agent's) calls GitHub authenticated; --github-token-source "" clears it (the token then comes from $GITHUB_TOKEN, else GitHub is called unauthenticated). See Agents. |
--github-app-id <id>, --github-app-installation-id <id>, --github-app-private-key-source <ref> |
The skills GitHub App agent-manager resolves skills with (githubApp): its public ids and a reference to its private key that beekeeper secret copy resolves, never a value. Every up and platform write the Secret agentlab-github-app in agent-platform (the key placed by beekeeper) and wire agent-manager.skills.github.app. Not together with --github-token-source; "" on all three clears it. |
| Variable | Read by | Effect |
|---|---|---|
ANTHROPIC_API_KEY |
up, platform |
Becomes the Secret kagent/kagent-anthropic at deploy time while agentlab.yaml records no aiKey.source; never written to agentlab.yaml or state/. With a source the Secret is placed from it through beekeeper secret copy --to-secret instead, and without either it holds a placeholder (the ModelConfig resolves, agent turns fail until the key is placed). See Agents. |
GITHUB_TOKEN |
up, platform, backstage-test, agents-test, the rehearsal, the update check |
While agentlab.yaml records no githubToken.source, becomes the Secret agentlab-github-token (key GITHUB_TOKEN) in agent-platform and kagent at deploy time — created or updated, so a re-run rotates it — and the values name that Secret for the portal's skill discovery, agent-manager's skill resolution and the migrate Job, which then call GitHub authenticated (5000 requests an hour instead of the 60 this machine's address shares). Never written to agentlab.yaml or state/. Unset: the lab is as before, the Secret of an earlier run stays unreferenced, and the proofs print the remaining unauthenticated window before resolving skills. See Agents. |
<name> per extraModels[].apiKeyEnv |
up, platform |
The key for that model config, same handling. See Models. |
NODE_USE_SYSTEM_CA=1 |
Node >= 22.15, Claude Code | Makes Node honor the system trust store after agentlab trust. Older Node: NODE_EXTRA_CA_CERTS=$PWD/certs/ca.crt. See TLS. |
KUBECONFIG |
your shell | Never read by the lab: its embedded Helm and Kubernetes client are built from state/kubeconfig alone. KUBECONFIG=state/kubeconfig kubectl ... is the lab's view from a shell. A file it names that is a copy of this lab's admin kubeconfig (every entry kind-<clusterName> — a lab lease holds one) is refreshed by up and every other cluster-facing command once the cluster behind it was recreated, so a copy taken before down and up reaches the new cluster without a manual step; your own kubeconfig, ~/.kube/config and another lab's are never written, except that up removes a kind-<clusterName> context, cluster and user of the recreated lab whose CA is the previous cluster's (the lab's own kubeconfig, or its lease's copy, is the one to use). |
AGENTLAB_TELEMETRY_OPTOUT, DO_NOT_TRACK=1 |
every command | Disable the anonymous usage signals. See Usage data. |
AGENTLAB_TELEMETRY_TESTMODE=1 |
every command | File the signals as test data and log delivery errors, for work on the lab itself. |
AGENTLAB_NO_UPDATE_CHECK=1 |
every command | Silence the newer-release hint (below). |
agentlab self-update replaces the running binary with the latest GitHub
release for your OS and architecture — the command muster and mcp-kubernetes
have too. agentlab self-update --check only reports the running and the
latest version, with exit status 125 when a newer one exists (for scripts). A
binary without a release version (agentlab --version says dev) is
refused: reinstall it from a release or with go install. A go build from
a checkout carries Go's pseudo-version (v0.19.3-0.20260908…-8536d36) and
is treated as what it is: after the tag before it, before the tag after it.
Every release binary is verified before it is installed. CI (the
architect orb) signs each agentlab-<os>-<arch> with cosign — keyless, the
CircleCI pipeline's identity, recorded in the Rekor transparency log — and
publishes the signature next to it as agentlab-<os>-<arch>.bundle.
self-update downloads both and installs the binary only after the bundle
verifies against the Sigstore public-good trust root for a CircleCI build of
github.com/giantswarm/agentlab (the shared
selfupdate-cosign
validator, the one muster and the other Giant Swarm CLIs use). A release
without a bundle for your platform is refused before anything is downloaded; a
download that does not match its signature is refused before anything is
written. Either way the installed binary stays as it is, and the error says
why. The hint below installs nothing, so it does not need the bundle.
Every command also starts with a one-line hint on stderr while a newer release is out — devctl's per-command check, with two deliberate differences:
- A hint, never a gate. An outdated agentlab runs every command the same; nothing waits for you to update.
- It gives up fast. The GitHub round trip is capped at two seconds, and
its answer is cached for an hour under your user cache directory
(
~/.cache/agentlab/latest-release.jsonon Linux,~/Library/Caches/agentlab/on macOS). A failed attempt is remembered for ten minutes, so a machine without internet is not held up on every command, and the last known answer keeps being shown meanwhile. AGITHUB_TOKENin the environment is used when present; it only lifts GitHub's anonymous rate limit.
AGENTLAB_NO_UPDATE_CHECK=1 silences the hint (self-update itself always
works); dev builds never check.