Skip to content

docs(architecture): inject the caller's credential at egress - #139

Draft
QuentinBisson wants to merge 153 commits into
giantswarmfrom
docs/caller-credential-at-egress
Draft

QuentinBisson wants to merge 153 commits into
giantswarmfrom
docs/caller-credential-at-egress

Conversation

@QuentinBisson

@QuentinBisson QuentinBisson commented Oct 1, 2026 •

Copy link
Copy Markdown

Problem

With KAGENT_PROPAGATE_TOKEN the caller's bearer enters the actor, and keeping it from the code the actor runs depends on a second user inside the sandbox and on gVisor enforcing it (giantswarm/giantswarm#38054).

Change

A design note: the a2agateway keeps the bearer in a kagent.dev credential provider in the controller and rewrites the caller host's egress rule to a per-turn URI, so Substrate's egress gateway injects it and the actor never holds it. It covers cleanup on every way a turn ends, the provider's mTLS and store, the scope (claude harness, opt-in) and its prerequisites: giantswarm/substrate#104, per-method authorization in ate-api, and the CONNECT and Host mismatch on the plain HTTP egress listener.

@QuentinBisson

Copy link
Copy Markdown
Author

Parked: git as the person is postponed. Upstream Substrate is redesigning this area (actor JWT injection merged in agent-substrate/substrate#1960, token exchange of the actor JWT in agent-substrate/substrate#1661, the actor JWT contract in agent-substrate/substrate#1756), which likely supersedes this provider design, and this fork line has to re-pin past agent-substrate/substrate#1809 before it can follow. Kept as a draft so it can resume in place; tracking in giantswarm/giantswarm#38064.

teemow and others added 29 commits October 5, 2026 14:21
…n label

helm-controller installs OCI charts under <version>+<digest>; '+' is not a
valid label value character, so every object of the chart failed to apply
under Flux. Sanitised the way helm.sh/chart already is.
…oc branch

The fork's tag.yaml publishes the dev builds of the agent-platform's kagent
component: GitHub-hosted runners (the org has no self-hosted runners), the
registries derived from the repository owner (ghcr.io/<owner>/kagent/<image>,
oci://ghcr.io/<owner>/kagent/helm/<chart>), no PyPI/GitHub-release jobs, the
image matrix trimmed to what the chart references plus the Claude harness
(multi-arch kept), and a push to poc/agent-platform that derives the version
as <base>-dev.<branch>.<committer date>.<time>.h<sha7> — the schema
giantswarm's gitsemver emits, so Flux selects the channel with a semverFilter
on the branch part. workflow_dispatch keeps its optional version input.

The chart's defaults are stamped with the published version and image
repositories before packaging: the tag must be explicit because the chart's
fallback, .Chart.Version, carries `+<digest>` under helm-controller.

The Makefile's -X targets followed $(DOCKER_REPO); with an owner-relative
repository they missed the module, leaving the binaries without version
information. They now follow the Go module path from go.mod.
…failure

apk.cgr.dev intermittently answers some of the parallel anonymous package
fetches from hosted runners with HTTP 403, which apk reports as
"Permission denied" for a random package and which fails the wolfi-based ui
image while the alpine-based images pass. The builder keeps the finished
stages, so the step now retries the build up to three times, repeating only
the failed step; and one image's failure no longer cancels the other three.
The controller and Go ADK images are alpine:3.22 plus `apk add`; the base
layer's libcrypto3/libssl3 3.5.7-r0 carry CVE-2026-14456 (HIGH, fixed in
3.5.8-r0). `apk upgrade --no-cache` before the add brings the runtime layer to
the current Alpine packages, as kagent-dev#2756 did for the harness images.
…w, keep a ledger

The consumed branch was built and published without a test or a scan:
upstream's ci.yaml and image-scan.yaml trigger on main and release/** only,
and their jobs ask for Blacksmith runners the org does not have. Both now run
on every push to poc/agent-platform and sync/**, on every pull request to
poc/agent-platform and (the scan) weekly, on GitHub-hosted runners with a
docker/setup-buildx-action builder, with the image matrices trimmed to what
the fork publishes (controller, ui, golang-adk, claude-harness — the Python
ADK, the Codex harness and the CLI image are not offered on the platform), a
govulncheck job over the Go module graph, no paths-ignore (a docs-only pull
request must still report the required checks) and one aggregate status per
workflow — ci-ok, scan-ok — for the ruleset to require.

sync-upstream.yaml re-pins the line weekly and on demand: fast-forward the
mirror main to upstream main with the workflow token (no workflow reacts, so
upstream's CI never runs on the mirror), rebase the carried commits onto the
upstream head (scripts/fork/repin.sh, which also names commits upstream merged
in the meantime and drops them), push the candidate with the bot token so the
checks run, wait for ci-ok and scan-ok, then force-push the consumed branch
with --force-with-lease so tag.yaml publishes the dev build. A conflicting
carried commit or a red check stops the run and opens a pull request against
the consumed branch that names the cause; nothing else is pushed.

tag.yaml gains a record job: every published build's version, commit, upstream
pin and image and chart digests go into the run summary and into builds.md on
the ledger branch (scripts/fork/ledger-append.sh; re-pins get the same on
re-pins.md). FORK.md is the human half of that ledger — the pin, every carried
commit with its purpose and upstream state, what is published and how, the
automated and the manual re-pin, protection, the contribution process — and
the README opens with it.
The first sync run was refused at the mirror: "refusing to allow a GitHub App
to create or update workflow `.github/workflows/ci.yaml` without `workflows`
permission" — the workflow's own token may not update a workflow file, not
even in a mirrored upstream commit, and no `permissions:` grant changes that.
The mirror is therefore pushed with the heraldbot App's token like the
candidate and the consumed branch. An App push starts upstream's workflow files
at the mirrored refs (runners the fork does not have, registries it must not
touch), so mirror.sh records what it pushed and mirror-quiet.sh cancels those
runs right after; the ledger stays on GITHUB_TOKEN.
…arent

The second sync run reached its "checks red" step and failed there:
`pull request create failed: GraphQL: Resource not accessible by
integration (createPullRequest)`. In a checkout of a fork, gh resolves
the base repository to the parent (kagent-dev/kagent) unless --repo is
given, and the App is not installed there. Both `gh pr create` calls now
name this repository.

git-push-as.sh emits its `::add-mask::` line only under GitHub Actions:
run by hand (the manual re-pin) it printed the encoded credential to
the terminal instead of masking it.
…warm line

SUBSTRATE_REPO defaults to oci://ghcr.io/giantswarm/substrate/helm and
SUBSTRATE_VERSION to 0.0.27-dev.giantswarm.2026-09-10.19-33-37.h734ec53: the substrate and substrate-crds subcharts of
kagent and kagent-crds are the fork giantswarm/substrate's builds (upstream
v0.0.26 — the version go/go.mod pins — plus cherry-picked fixes). The Go
module pin stays upstream's release, so the version is set here rather than
derived from go.mod; both move together at a re-pin.

Signed-off-by: Timo Derstappen <timo@giantswarm.io>
The Makefile's SUBSTRATE_VERSION / SUBSTRATE_REPO are the one pin: a
substrate-pin target prints them, the e2e job reads them into its
environment and installs the line's substrate-crds and substrate charts and
its ateom-gvisor worker image (the fork's tags carry no v). kubectl-ate for the
CA/JWT bootstrap stays upstream's v0.0.26 binary — the line publishes no CLI.
Before, the job exported SUBSTRATE_VERSION=0.0.26, which also drove make
helm-version and made the fork registry look for a chart it never published.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
The consumed branch was renamed from poc/agent-platform to giantswarm (the GitHub rename retargets the default branch, the ~DEFAULT_BRANCH ruleset and open pull requests); every reference follows: the push and pull-request triggers of ci.yaml, image-scan.yaml and tag.yaml, CONSUMED in sync-upstream.yaml and scripts/fork/repin.sh, FORK.md and the README. The dev version's branch part becomes 'giantswarm' — 0.11.0-dev.giantswarm.<date>.<time>.h<sha7> — and the consumers' semverFilter moves with it (.*-dev\.giantswarm\..*).
FORK.md 'Release scheme': a release is a tag vX.Y.Z-gs.N on the consumed branch — X.Y.Z the upstream version the pin anticipates (DEV_BASE_VERSION, 0.11.0), N the line's counter from 1; first release v0.11.0-gs.1. Sorted 0.11.0-dev.giantswarm.… < 0.11.0-gs.1 < 0.11.0, so a release outranks its dev builds and never the upstream release it anticipates, and the switch to an upstream tag is a range change. Published by tag.yaml to ghcr.io/giantswarm/kagent (images + helm/ charts), digests recorded by the record job; ghcr.io rather than gsoci/the catalog because GitHub Actions has no ACR credential (the org publishes to gsoci from CircleCI) and the distinct OCI path plus the pre-release version each keep the chart out of the 0.10 wrapper's fleet range >=0.2.0 <1.0.0 at gsoci.azurecr.io/charts/giantswarm/kagent. tag.yaml's version step refuses any tag that is not v<DEV_BASE_VERSION>-gs.<N> — a bare vX.Y.Z is a name upstream will use and the mirror would then skip upstream's tag of that name. A 'Releases' table (release, tag, commit, pin, digests) is filled per tag after the agentlab proof of its dev build.
…olicy: keep)

The kagent-crds chart ships its CustomResourceDefinitions as plain
templates (there is no crds/ directory), so `helm uninstall kagent-crds`
deletes the five CRDs and, with them, every AgentTemplate, Harness,
ModelConfig, ModelProviderConfig and RemoteMCPServer in the cluster.
CRD charts conventionally mark their CRDs `helm.sh/resource-policy: keep`
so that removing or replacing the release leaves the CRDs (and the
objects behind them) in place; a CRD is then removed deliberately, with
`kubectl delete crd`, never as a side effect of a release operation.

The annotation is added where the manifests are produced rather than in
the chart files: a `+kubebuilder:metadata:annotations` marker on each of
the five v1alpha3 types, so controller-gen emits it into
go/api/config/crd/bases and `make controller-manifests` copies it into
helm/kagent-crds/templates unchanged. The chart templates stay a
verbatim copy of the generated CRDs, the manifests-check CI job keeps
passing, and no new tool (yq, sed post-processing) enters the
generation. The generated bases carry the annotation too: nothing in the
repository applies them directly (the e2e installs the chart), and to
`kubectl apply` it is an inert annotation.

Verified: `make controller-manifests` regenerates exactly the five
annotation lines in bases and templates; `helm lint helm/kagent-crds`
passes with and without `substrate.enabled`; `helm template` differs
from the previous render only in the five added lines, and the CRDs of
the kmcp-crds and substrate-crds subcharts are untouched.
…ources

A skill or plugin source on a private repository has no credential path:
the materialiser shells out to git with the process environment and no
helper, so an AgentTemplate cannot take skills from a private repository
at all.

AgentTemplate.spec.skills[].source.git and plugins[].source.git gain an
optional credentialRef, a key in a same-namespace Secret (https only,
enforced by CEL). Git speaks Basic authentication, so the key holds the
base64 of "<username>:<token>" for the URL's host; GitHub takes
"x-access-token:<token>".

The credential follows the gateway injection path every other credential
takes. CompileCredentials binds each referenced key to the source URL's
host as "Authorization: Basic <value>", an ate-secret:// URI the egress
gateway resolves on every request to that host; the actor never holds
the token, and the materialiser clones with no credential of its own.
The compiler still records one Secret-backed environment variable per
referenced key, KAGENT_ARTIFACT_CREDENTIAL_<hash of namespace, name and
key>, which reaches the actor as the inert placeholder like a model key,
so a missing Secret or key fails compilation as a reference failure
(ResolvedRefs=False) that is retried when the Secret appears. A rotation
is not a new revision: the gateway reads the Secret live.

Substrate matches credentials by hostname, so one credential serves every
source of an agent tree on that host, public repositories included, and
two different Secrets for one host are refused at compile time. A wrong
token fails the golden boot on the host's 401.
…chart at publish

The published chart's defaults named only the images' version; the two
pins a consumer needs beyond that — the index digest of the Go ADK
runtime image (a kagent.dev/v1alpha3 Harness accepts nothing but a
digest, and Helm cannot resolve one at render time) and the Substrate
worker image the chart's WorkerPool runs — were resolved by `record`
after the chart push, or known only to the Makefile, so the platform's
meta chart carried copies of both and re-pinned them by hand.

`push-helm-chart` runs after `push-images`, so the indexes exist: it now
resolves the golang-adk and claude-harness index digests the way `record`
does (`docker buildx imagetools inspect`) and stamps them into
`controller.agentImage.digest` and `runtimeImages.claudeHarness.digest`;
it stamps `substrateWorkerPool.workerImage` with the gVisor worker of the
Makefile's `SUBSTRATE_VERSION`/`SUBSTRATE_REPO` (`make -s substrate-pin`,
the same source ci.yaml reads for the e2e cluster) over the empty
default the template refuses. Every stamp is checked with yq; a missed
anchor fails the job as the existing checks do. `record` additionally
reads `helm show values` of the pushed chart and fails when its digests
disagree with the ones it writes to builds.md.

The anchors are the values keys of the two patches that follow this one
on the consumed branch (the optional default Harness, the runtime-images
ConfigMap); the workflow runs on the branch head only. FORK.md gets the
patch-table rows of all three.

Fork infrastructure — tag.yaml is fork-rewritten; never upstream.

Refs giantswarm/giantswarm#37765, giantswarm/giantswarm#37705.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
…ess for the chart's own runtime image

On the kagent API v2 an agent runs on a Harness (kagent.dev/v1alpha3): a
runtime image pinned by digest, an environment and a Substrate policy.
The chart ships the Go ADK runtime image (controller.agentImage) but
knows it by tag only, and renders no Harness for it — so every consumer
has to resolve the image's digest out of band and write its own Harness
before the first AgentTemplate can become Ready (kagent-dev#2529
is the missing build-time digest; kagent-dev#2536 proposed the
value this change uses).

controller.agentImage.digest carries the image's index digest (empty in
git; a published chart carries the digest of its build). A new `harness`
block renders one Harness — `kagent` by default, in the chart's
namespace, on the `kagent` runtime adapter — from that digest, only when
harness.create is set (default false, so nothing changes for a consumer
that renders its own Harness objects): workload.image is
<registry>/<controller.agentImage.repository>@<digest> unless
harness.image names another digest-pinned image; workerPoolRef defaults
to substrateWorkerPool.name; snapshotLocation is required (the CRD
requires the snapshot policy); env and allowedAgentTemplates.selector
are forwarded. A tag instead of a digest, a missing digest or a missing
snapshot location fail the render with the reason, instead of failing at
admission. The Harness keeps helm.sh/resource-policy: keep — it is the
runtime of every agent admitted under it.

The digest-pinned reference is built by one helper (kagent.runtimeImage)
so other templates can list the same image.

Unit tests cover the default (no Harness), the rendered object and its
image, the registry precedence, the harness.image override, the
forwarded policy and the three failures.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
…he Substrate image cache

Agent Substrate's atelet can pre-pull a pinned set of images on every
worker node (imageCache.pinnedImages), so an agent's first turn does not
wait for the pull of its runtime image. The set has to name the images
agents actually run on — this chart's runtime images at their digests —
and today only the operator of a cluster can write it, from a digest
resolved out of band, and re-write it on every upgrade.

The chart now publishes them itself: the ConfigMap `<fullname>-images`
carries one key, substrate-values.yaml, a values fragment for the
Substrate chart with atelet.imageCache.pinnedImages listing the image the
chart's Harness runs (harness.image, else the chart's own Go ADK image at
controller.agentImage.digest — so an override moves the set with it) and
the Claude runtime image at runtimeImages.claudeHarness.digest. An image
whose digest is empty is left out: a chart rendered from git lists
nothing, a published chart lists its build. The ConfigMap sits in the
release namespace, deliberately not the chart's namespaceOverride, so a
Substrate release installed next to this one can read it as values (a
Flux HelmRelease's valuesFrom) without crossing namespaces.

Unit tests cover the empty set, the two-image set and its name under a
fullnameOverride, the harness.image override and the namespace.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
SUBSTRATE_VERSION moves from the bootstrap dev build
0.0.27-dev.giantswarm.2026-09-10.19-33-37.h734ec53 to 0.0.27-gs.6, the
release of giantswarm/substrate the platform's meta chart runs today (its
worker image ghcr.io/giantswarm/substrate/ateom-gvisor:0.0.27-gs.6).
tag.yaml stamps substrateWorkerPool.workerImage from this pin, so a
release cut from the dev pin would install a dev worker as soon as the
meta chart stops forwarding the image by hand; the e2e cluster (ci.yaml)
runs the same version. Upstream pin unchanged: kagent-dev/substrate
v0.0.26 in go/go.mod, the ate-api contract.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
SUBSTRATE_VERSION moves from 0.0.27-gs.6 to 0.0.27-gs.7, the release of
giantswarm/substrate that carries giantswarm/substrate#25: ateom
re-asserts its capacity report every 10 s, so the worker record the
pod-IP repair re-creates receives its capacity within one interval
instead of staying at zero until the pod restarts (giantswarm/giantswarm#37762).
The fix lives in the ateom-gvisor image, and tag.yaml stamps
substrateWorkerPool.workerImage from this pin, so the published chart of
the next release installs ghcr.io/giantswarm/substrate/ateom-gvisor:0.0.27-gs.7;
the e2e cluster (ci.yaml) runs the same version. Inert under the
agent-platform 4.7.x meta chart, which forwards the worker image by hand;
live from 4.8.0. Upstream pin unchanged: kagent-dev/substrate v0.0.26 in
go/go.mod, the ate-api contract.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
…watch

The ConfigMap `<fullname>-images` is meant to be read as values by a
Flux HelmRelease (valuesFrom). helm-controller re-reads such a ConfigMap
on its own only when the HelmRelease spec changes or on the release's
interval, unless the ConfigMap carries the label its
--watch-configs-label-selector matches — reconcile.fluxcd.io/watch:
Enabled by default. Without it, a change of the Harness image alone (a
harness.image override, an upgrade that only moves the digests) leaves
the consumer's pinned image set stale for up to that interval.

The ConfigMap now carries the label, so a HelmRelease reading it is
upgraded as soon as an image reference changes. Any other reader ignores
the label. A unit test asserts it.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
…elector

The chart's Harness selects the AgentTemplates it admits by
harness.allowedAgentTemplates.selector, and its default selects by the
chart's own label, kagent.dev/harness: kagent. A consumer that admits
templates by a label of its own has to remove that default key: Helm
merges maps on coalesce, so an override adds keys next to the default
and the selector ends up requiring both labels. The one way to remove a
key from the merged map is a null in the consumer's values, and a null
does not survive every path values take to the chart. A Flux
HelmRelease that already exists is updated with a JSON merge patch, in
which a null means "remove the key" — the key is dropped from the
object, the chart default comes back on coalesce, and a Harness upgraded
in place selects by two labels while a freshly created one selects by
one. Every template labelled for the consumer's key alone is then
admitted by no Harness.

The Harness template now drops every matchLabels entry whose value is
the empty string before rendering the selector. An empty string is
stored by every values path, create and patch alike, and it is never a
valid selector value, so it means "not this key". A selector left with
no requirement at all fails the render: an empty label selector selects
everything, and a Harness that admits every AgentTemplate in its
namespace is not what a blanked default asks for. matchExpressions are
untouched.

Unit tests cover a blanked key next to a kept one and next to
matchExpressions, a selector kept on matchExpressions alone, and the
failure when nothing is left to select by.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
…e agentskills.io set

The Go ADK runtime loads an agent's skills through the ADK's
tool/skilltoolset/skill package, whose Parse decodes the SKILL.md frontmatter
with yaml.Decoder.KnownFields(true) into a struct that knows the
agentskills.io fields only (name, description, license, compatibility,
metadata, allowed-tools). One field outside that set failed the whole agent:

    failed to create agent: failed to load skills: invalid frontmatter:
    parse frontmatter: yaml: unmarshal errors:
      line 3: field user-invocable not found in type skill.Frontmatter

Skills are written once and run on several harnesses. Claude Code's skills
carry user-invocable, argument-hint and disable-model-invocation at the top
level; skills in the wild carry a plain version. The same SKILL.md loads on
the Python runtime, whose Frontmatter model allows extra fields
(google/adk-python, skills/models.py: extra="allow"), and in Claude Code, so
an agent that ran on the Python runtime stopped booting on the Go one.

The runtime now hands the ADK a filesystem view of the skills directory in
which every skill's SKILL.md carries only the frontmatter fields the ADK
knows (adk/pkg/skillfs). The ADK's parser, validation and resource rules stay
the single authority on what a skill is: name and description are still
required and validated, the directory name must still match, resources are
still confined to references/, assets/ and scripts/, and a SKILL.md the
filter cannot read as frontmatter (no separator, YAML that does not parse,
a document that is not a mapping) is served unchanged so the ADK reports it
as before. Fields nested under metadata are untouched.

The alternative, decoding without KnownFields in the ADK itself, belongs
upstream in google/adk-go; until it lands, the runtime is the place that
knows its skills come from other harnesses too.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
… gVisor worker's terminate collects a crashed golden actor (giantswarm/giantswarm#37773) (#32)

SUBSTRATE_VERSION 0.0.27-gs.7 -> 0.0.27-gs.9: v0.0.27-gs.9 = gs.8 + giantswarm/substrate#30, the worker (ateom-gvisor) whose TerminateWorkload skips runsc containers that are already gone and accepts a runsc delete that fails after removing its container. Before it a superseded AgentTemplate revision whose golden boot had crashed was never collected and its golden actor pinned one worker of the pool for good (gazelle 2026-09-13, two of four workers). The published chart stamps substrateWorkerPool.workerImage with the new worker; the platform's WorkerPool rolls once. FORK.md's pin row follows.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
…giantswarm.2026-09-14.19-48-17.h21027aa

Upstream 2d843e3 (kagent-dev#2802) moved kagent to Substrate v0.0.29:
the a2a gateway addresses an actor by the ate-target-actor header, which the
router of Substrate 0.0.27 does not know, so the two lines move together. The
Giant Swarm line of Substrate is re-pinned onto kagent-dev/substrate v0.0.29
in the same move (giantswarm/substrate sync/20260914-v0.0.29); this is its
dev build, for the agentlab proof of the pair. The release pin (0.0.30-gs.1)
follows the proof.

Signed-off-by: Timo Derstappen <timo@giantswarm.io>
The Substrate line's release of kagent-dev/substrate v0.0.29 plus its ten
carried patches (giantswarm/substrate FORK.md), the same tree as the dev
build 0.0.30-dev.giantswarm.2026-09-14.19-48-17.h21027aa the agentlab proof
of the three lines ran on; the published chart stamps its worker image
ghcr.io/giantswarm/substrate/ateom-gvisor:0.0.30-gs.1 into
substrateWorkerPool.workerImage.

Signed-off-by: Timo Derstappen <timo@giantswarm.io>
Add `promptCaching` (default false) and `cacheTTL` ("5m" | "1h") to
`spec.anthropic` on ModelConfig, mirroring the knobs BedrockConfig got in

When enabled, every Messages API request carries three `cache_control`
breakpoints: the last tool definition, the last system prompt block and
the last content block of the latest conversation turn. Anthropic
renders tools, then system, then messages and caches the prefix up to
each breakpoint, so an agent loop reads its whole previous history from
the cache and writes only the new turn instead of paying full input
price for the same prefix on every call. "1h" opts into the 1-hour
cache; "5m" (the CRD default) leaves the API default in place.

Go runtime: buildAnthropicParams now owns the request shaping so it can
be unit-tested, markAnthropicCacheBreakpoints sets the markers, and
usage folds cache_read/cache_creation into PromptTokenCount with
cache_read surfaced as CachedContentTokenCount so the UI shows the
effect. Python runtime: KAgentAnthropicLlm wraps the SDK client's
messages resource and marks the same three breakpoints before each
request; google-adk already maps the cache usage fields.

The Claude harness compiler treats the defaulted cacheTTL "5m" like the
Bedrock one and keeps rejecting promptCaching, since Claude Code caches
natively.

Refs kagent-dev#2787, kagent-dev#1866.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
(cherry picked from commit a0c974b)
…ests

Shallow copies would let a nested in-place mutation of a message or tool
block go unnoticed, which is exactly what these tests guard against.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
(cherry picked from commit f5679c2)
… golden boot a workload keeps failing is bounded and its cause carried up (giantswarm/giantswarm#37801) (#49)

SUBSTRATE_VERSION 0.0.30-gs.1 -> 0.0.30-gs.4: v0.0.30-gs.4 = gs.3 + giantswarm/substrate#39. The worker (ateom-gvisor) keeps each application container's last output lines and quotes them in the readiness-deadline error; ate-api counts the boots a workload itself fails (WORKLOAD_NOT_READY) and at the third crashes the golden actor and fails the ActorTemplate with GoldenActorNotReady, the bound and the last boot's error — which this controller's pair reconciler already surfaces verbatim as Ready=False ActorTemplateFailed. Before it a golden actor whose workload exited before readiness (an anonymous fetch of a private skill, a skill path missing at the pinned commit) was re-booted once a minute for as long as the template existed and the AgentTemplate read ActorTemplatePending indefinitely (gazelle 2026-09-15, about 120 boots over two hours). The published chart stamps substrateWorkerPool.workerImage with the new worker; the platform's WorkerPool rolls once. FORK.md's pin row follows.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
A runtime's network identity says nothing about the agent it runs once
agents share a pod, a ServiceAccount or an egress: on the kagent API v2
every AgentTemplate is an actor in a shared WorkerPool and every
connection leaves through one egress, so a gateway in front of the model
providers attributes all usage to that egress and knows no user. Only
the runtime knows which agent a call is made for, and only the
controller knows whose turn it is.

The controller now tells the runtime which AgentTemplate it executes
(KAGENT_AGENT_TEMPLATE, next to KAGENT_NAMESPACE; KAGENT_NAME stays the
per-Harness app name), and every model call carries that identity as
request headers: x-kagent-agent and x-kagent-agent-namespace, set by the
Go ADK's HTTP transport after a model's configured default headers so a
ModelConfig cannot name another agent. The person of the turn travels
as x-kagent-user: the controller's a2a gateway forwards it to the actor
as request metadata only when its authenticator resolved the identity
from the caller's validated token (a session carrying claims) — the
unsecure mode's X-User-Id and its default, an agent's propagated caller
id and the control plane resolve nobody and forward nothing — and the
Go ADK's A2A interceptor carries that metadata into the turn's context,
where the transport sets the header. The session owner in x-user-id and
the caller's bearer in authorization are never sent as the person: the
user's context key is its own type, so it cannot share an address with
the bearer token's key the way two pointers to zero-size structs can.
An absent header is absent, never a stand-in. The Claude harness sends
the two agent headers through ANTHROPIC_CUSTOM_HEADERS, which the
controller renders and owns; it knows no user per turn and names none.

The headers are an accounting identity the runtime asserts about itself,
never an authorization input.

Signed-off-by: Timo Derstappen <timo@giantswarm.io>
teemow and others added 24 commits October 8, 2026 08:12
The e2e job read the controller log from the Deployment once the tests
were over, which holds only the final pod. Tests that roll the controller
(withControllerEnv, the sandbox restart) lost the log of every earlier
pod. The job follows every controller pod from the moment it runs, one
file per pod, and stops the followers with the other log streams.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
Signed-off-by: Timo Derstappen <teemow@gmail.com>
giantswarm/substrate#182 lets Pause, Suspend and Resume carry a fencing token; the quiescence fence of #181 needs it. The go.mod replace and the charts' SUBSTRATE_VERSION move together.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
A quiescence claim's records fence a holder that stopped, not one frozen
past its lease while its Pause or Suspend is in flight: the successor
settles a running Actor as not quiesced and releases the claim, and the
late request then pauses or suspends an Actor the session believes is
running.

Every Pause, Suspend and fencing Resume a holder sends now carries a
Substrate fencing token: its executor id and a generation from a
PostgreSQL sequence that grows with every request. Before it releases a
claim on a running Actor, the holder sends a fenced Resume with a new
generation, which records the token without touching the runtime, so
Substrate refuses a late request with FailedPrecondition. A holder whose
own request is refused that way was superseded and leaves the claim to
its successor. Calls without a token are admitted as before.

Fixes #181

Signed-off-by: Timo Derstappen <teemow@gmail.com>
Signed-off-by: Timo Derstappen <teemow@gmail.com>
TestServerTelemetryShape read the span exporter and the metric reader
right after GetVersion and Check returned. The client has its response
before the server is done with the RPC: otelgrpc ends the server span
and records rpc.server.call.duration on the stats handler's End event,
after the response is written, so either could still be missing and the
test failed with "rpc.server.call.duration was not recorded".

Wait, bounded, for the GetVersion server span and the recorded metric
before asserting; the assertions are unchanged.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
Signed-off-by: Timo Derstappen <teemow@gmail.com>
Every Pause, Suspend and fencing Resume a quiescence holder sends carries
a fencing token. A Substrate whose Actor API predates the token refuses
any field it has no descriptor for:

  InvalidArgument: request: Invalid value: unknown field with protobuf tag 10001

so every quiesce failed and the fenced Resume before a claim's release
was retried forever: the claim was never released and the session
refused every turn after its first.

The Substrate client now detects that refusal, remembers that its
ate-api takes no token, logs one warning naming the lost fence, and
sends that request and every later one without the token. The runtime
boundary works as before the token; only a frozen holder's late Pause or
Suspend is no longer refused. ate-api exposes no version a controller
could check at startup, and Substrate is usually installed on its own,
so the chart's dependency cannot hold the pair together.

Fixes #230

Signed-off-by: Timo Derstappen <teemow@gmail.com>
Signed-off-by: Timo Derstappen <teemow@gmail.com>
Co-authored-by: giantswarm-align-files[bot] <288255335+giantswarm-align-files[bot]@users.noreply.github.com>
Co-authored-by: Quentin Bisson <quentin.bisson@gmail.com>
Nothing wrote Harness.status, so kubectl get harness showed an empty READY
column and the gRPC Harness.ready was always false although the Harness
served.

The controller derives Ready from the Harness's WorkerPool and the golden
boots of the Agents that reference it, and writes it on its own status
queue. One successful golden boot proves the Harness's workload boots, so
it reads True; a failed boot reads False with its reason while no Agent
booted, since an Agent's own configuration can fail its boot too; Unknown
while nothing has tried the workload. The status write keeps
status.capabilities.

Refs kagent-dev#2764

Signed-off-by: Timo Derstappen <teemow@gmail.com>
Signed-off-by: Timo Derstappen <teemow@gmail.com>
Helm's and Flux's readiness wait reads a Ready condition other than True
as a rollout in progress, so a Harness without Agents, Ready Unknown,
held every upgrade of the chart that ships it until the wait timed out
and the release rolled back. Before any golden boot finished the
resolved WorkerPool keeps the Harness Ready (WorkerPoolResolved); a
successful boot reads Booted, a failed one with none booted False.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
Signed-off-by: Timo Derstappen <teemow@gmail.com>
Signed-off-by: Timo Derstappen <teemow@gmail.com>
* fix(session): retry a failed quiesce once the Actor runs again

When the pause or suspend after a turn returned an error, settleIdleWork
re-read the Actor and released the claim without a snapshot as soon as it
was RUNNING again, for example after a first SuspendActor that outlived its
deadline and that Substrate then abandoned. The suspend was never retried,
so the turn kept no snapshot, and nothing on the session said why.

A RUNNING Actor now gets the pause or suspend once more, fenced with a
newer token, at the backoff settleIdleWork already uses; its outcome
settles the same way. The claim is released without a snapshot only when
the retry fails too, or the Actor is crashed or gone. The release records
why on the session as last_quiescence_failure (task, reason, message and
time), which GetSession, ListSessions and the CLI's session details show;
a later boundary with a snapshot clears it.

Signed-off-by: Timo Derstappen <teemow@gmail.com>

* docs(fork): carry the retry of a failed quiesce

Signed-off-by: Timo Derstappen <teemow@gmail.com>

---------

Signed-off-by: Timo Derstappen <teemow@gmail.com>
…ns (#227)

* ci(telemetry): fetch the contract's upstream sources once at their pins

The Telemetry Contract Check cloned three GitHub repositories on every
Weaver run: the two registries telemetry/registry/manifest.yaml depends
on (the GenAI registry re-cloning the core one through its own
manifest) and the shared policy packages, eleven clones per
`make semconv-verify`. A clone that flaked turned the check red on a
pull request's first run, and the result of the check depended on
reaching GitHub.

The three sources are now pinned in telemetry/versions.env and fetched
once by `make semconv-deps` into the git-ignored telemetry/deps: a
depth-one fetch of the pinned ref per repository, reused without
touching the network while its pin stands. The manifest names the
checkouts, semconv-check passes the policies from them, and the fetch
points a fetched registry's own git dependency at the checkout of the
same pin, refusing a pin that wants another ref, since the contract and
its dependencies must agree on the core conventions. CI caches the
directory keyed on the pins.

Weaver's `[resolve.schema_url_overrides]` does not serve here: in
v0.26.1 and v0.27.0 the loader of `registry check` and `registry
generate` clones a manifest dependency before any override is consulted;
the overrides serve the on-demand resolver of live-check and serve.

Proof: `make semconv-verify` with the checkouts present, the Weaver
container on `--network none` and a dead HTTPS proxy on the host, exits
0 after the check, the four policy fixtures, the four generations and
the drift check.

Signed-off-by: Timo Derstappen <teemow@gmail.com>

* docs(fork): ledger row for the telemetry contract's pinned sources

Signed-off-by: Timo Derstappen <teemow@gmail.com>

* ci(telemetry): fetch the pinned sources before generating, name the core pin apart

semconv-generate reads the manifest's registries from telemetry/deps, so it runs semconv-deps like semconv-check. The core registry's pin is SEMCONV_CORE_REGISTRY: the Makefile includes versions.env and names the kagent registry SEMCONV_REGISTRY.

Signed-off-by: QuentinBisson <quentin@giantswarm.io>

---------

Signed-off-by: Timo Derstappen <teemow@gmail.com>
Signed-off-by: QuentinBisson <quentin@giantswarm.io>
Co-authored-by: QuentinBisson <quentin@giantswarm.io>
Co-authored-by: Quentin Bisson <quentin.bisson@gmail.com>
…unner

TestServe polled GetProcess for the first process's COMPLETED within
require.Eventually's one second. On a loaded runner the shell's start
plus the poll take longer, so the test failed with "Condition never
satisfied" and passed on rerun. Name one bound of five seconds, like the
cleanup's stop; a passing run still returns at the first COMPLETED.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
…esourceVersion (#245)

The Harness and ModelConfig status writes were an UpdateStatus of a copy of
the informer cache, so a write the cache had not observed yet failed them
with a Conflict. Both are now a JSON merge patch of the status subresource
that carries only the fields the controller derives; the Harness patch never
carries status.capabilities.

Signed-off-by: QuentinBisson <quentin@giantswarm.io>
TestSandboxTemplateCatalog and TestSandboxKRTInformerQueueAndRestartCleanup start an
envtest API server and fail a plain go test of their package when the envtest
binaries are not set up. Skip them with a message naming KUBEBUILDER_ASSETS, and
let make test set the variable through setup-envtest, as the CI job does.

Closes #232

Signed-off-by: Timo Derstappen <teemow@gmail.com>
…r-token propagation (#254)

Substrate 1.5's egress replaces a credential header only when the request
carries it. The Claude forwarder deleted every server's compiled
Authorization and set the caller's token in its place, so a Secret-backed
Authorization was never sent without a caller token and overwrote the
caller's token with one.

A server that configures Authorization keeps its gateway placeholder; any
other takes the caller's token, or none. The adapter resolves the
compiler's credential references before forwarding. CompileCredentials
refuses a caller-owned MCP server on a host with a gateway authorization
binding when KAGENT_PROPAGATE_TOKEN is on.

Closes #253
…252)

* fix(skills): git sends the Authorization the egress gateway replaces

Substrate's egress credential effect replaces the value of a header a
request carries and leaves a request without it untouched. Git sends no
Authorization before a challenge, so a private git skill was fetched
without its credential, answered 401 and the materializer failed on a
terminal prompt.

The compiled source of a git artifact with a credentialRef is marked
authenticated, and the materializer has git send a placeholder
Authorization header for that source's URL only, configured for the one
invocation and never written to a gitconfig. Git never prompts, so a
refused credential fails with its own error. A public source sends no
header.

Signed-off-by: Timo Derstappen <teemow@gmail.com>

* docs(api): credentialRef reaches only requests that carry the header

The egress gateway replaces an Authorization header a request carries and
adds none, so the field's description and the ledger row no longer claim
the credential goes on every request to the host, or to the host's public
sources.

Signed-off-by: QuentinBisson <quentin@giantswarm.io>

---------

Signed-off-by: Timo Derstappen <teemow@gmail.com>
Signed-off-by: QuentinBisson <quentin@giantswarm.io>
Co-authored-by: QuentinBisson <quentin@giantswarm.io>
@QuentinBisson
QuentinBisson force-pushed the docs/caller-credential-at-egress branch from 1c6a54b to d32088e Compare October 8, 2026 21:21
QuentinBisson and others added 3 commits October 8, 2026 23:32
…or every client kind (#256)

* test(e2e): prove the egress credential contract through the gateway for every client kind

Each client a Secret-backed binding covers (model keys, RemoteMCPServer
headersFrom on every harness and on Claude with KAGENT_PROPAGATE_TOKEN,
git skill and plugin sources with a credentialRef) talks to a recording
upstream through Substrate's egress gateway. The upstream must see the
Secret value on every request and the placeholder on none, so a client
that sends no header fails against a gateway that replaces only present
headers.

The git upstream is git http-backend on the test host, served over HTTPS
under a cluster Service name. The CI job creates a CA per run, has the
gateway trust it (atenetEgress.upstreamTrust.caBundle) and passes it to
the tests to sign that upstream. Golden boots fetch git sources in
atespace ate-golden, so the job lets it resolve kagent Secrets.

Signed-off-by: QuentinBisson <quentin@giantswarm.io>

* fix(translator): send an Ollama Cloud model of the Python runtime to api.ollama.com

The compiler left OLLAMA_API_BASE unset for a cloud model so that the
runtime would route it, but only the Go runtime routes: the Python
runtime sent a cloud model to localhost:11434. The compiler now writes
the cloud endpoint whenever OllamaReachesCloud routes there. The Ollama
SDK sends OLLAMA_API_KEY, the gateway placeholder, as the bearer, which
a unit test now pins.

Signed-off-by: QuentinBisson <quentin@giantswarm.io>

---------

Signed-off-by: QuentinBisson <quentin@giantswarm.io>
Signed-off-by: QuentinBisson <quentin@giantswarm.io>
The turn rewrites the host's rule instead of adding one, since only the
first matching rule applies; the policy is the record of the caller, with
cleanup on every way a turn ends and a sweep at start; the plain HTTP
listener injects too, so caller hosts must be https; the store needs the
Recreate strategy; the scope is the claude harness; and the trust section
names ate-api's missing per-method authorization and the CONNECT and Host
mismatch on the plain HTTP listener.

Signed-off-by: QuentinBisson <quentin@giantswarm.io>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants