Repository navigation
docs(architecture): inject the caller's credential at egress - #139
Draft
QuentinBisson wants to merge 153 commits into
Draft
QuentinBisson wants to merge 153 commits into
QuentinBisson wants to merge 153 commits into
Conversation
Author
|
Parked: git as the person is postponed. Upstream Substrate is redesigning this area (actor JWT injection merged in agent-substrate/substrate#1960, token exchange of the actor JWT in agent-substrate/substrate#1661, the actor JWT contract in agent-substrate/substrate#1756), which likely supersedes this provider design, and this fork line has to re-pin past agent-substrate/substrate#1809 before it can follow. Kept as a draft so it can resume in place; tracking in giantswarm/giantswarm#38064. |
…n label helm-controller installs OCI charts under <version>+<digest>; '+' is not a valid label value character, so every object of the chart failed to apply under Flux. Sanitised the way helm.sh/chart already is.
…oc branch The fork's tag.yaml publishes the dev builds of the agent-platform's kagent component: GitHub-hosted runners (the org has no self-hosted runners), the registries derived from the repository owner (ghcr.io/<owner>/kagent/<image>, oci://ghcr.io/<owner>/kagent/helm/<chart>), no PyPI/GitHub-release jobs, the image matrix trimmed to what the chart references plus the Claude harness (multi-arch kept), and a push to poc/agent-platform that derives the version as <base>-dev.<branch>.<committer date>.<time>.h<sha7> — the schema giantswarm's gitsemver emits, so Flux selects the channel with a semverFilter on the branch part. workflow_dispatch keeps its optional version input. The chart's defaults are stamped with the published version and image repositories before packaging: the tag must be explicit because the chart's fallback, .Chart.Version, carries `+<digest>` under helm-controller. The Makefile's -X targets followed $(DOCKER_REPO); with an owner-relative repository they missed the module, leaving the binaries without version information. They now follow the Go module path from go.mod.
…failure apk.cgr.dev intermittently answers some of the parallel anonymous package fetches from hosted runners with HTTP 403, which apk reports as "Permission denied" for a random package and which fails the wolfi-based ui image while the alpine-based images pass. The builder keeps the finished stages, so the step now retries the build up to three times, repeating only the failed step; and one image's failure no longer cancels the other three.
The controller and Go ADK images are alpine:3.22 plus `apk add`; the base layer's libcrypto3/libssl3 3.5.7-r0 carry CVE-2026-14456 (HIGH, fixed in 3.5.8-r0). `apk upgrade --no-cache` before the add brings the runtime layer to the current Alpine packages, as kagent-dev#2756 did for the harness images.
…w, keep a ledger The consumed branch was built and published without a test or a scan: upstream's ci.yaml and image-scan.yaml trigger on main and release/** only, and their jobs ask for Blacksmith runners the org does not have. Both now run on every push to poc/agent-platform and sync/**, on every pull request to poc/agent-platform and (the scan) weekly, on GitHub-hosted runners with a docker/setup-buildx-action builder, with the image matrices trimmed to what the fork publishes (controller, ui, golang-adk, claude-harness — the Python ADK, the Codex harness and the CLI image are not offered on the platform), a govulncheck job over the Go module graph, no paths-ignore (a docs-only pull request must still report the required checks) and one aggregate status per workflow — ci-ok, scan-ok — for the ruleset to require. sync-upstream.yaml re-pins the line weekly and on demand: fast-forward the mirror main to upstream main with the workflow token (no workflow reacts, so upstream's CI never runs on the mirror), rebase the carried commits onto the upstream head (scripts/fork/repin.sh, which also names commits upstream merged in the meantime and drops them), push the candidate with the bot token so the checks run, wait for ci-ok and scan-ok, then force-push the consumed branch with --force-with-lease so tag.yaml publishes the dev build. A conflicting carried commit or a red check stops the run and opens a pull request against the consumed branch that names the cause; nothing else is pushed. tag.yaml gains a record job: every published build's version, commit, upstream pin and image and chart digests go into the run summary and into builds.md on the ledger branch (scripts/fork/ledger-append.sh; re-pins get the same on re-pins.md). FORK.md is the human half of that ledger — the pin, every carried commit with its purpose and upstream state, what is published and how, the automated and the manual re-pin, protection, the contribution process — and the README opens with it.
The first sync run was refused at the mirror: "refusing to allow a GitHub App to create or update workflow `.github/workflows/ci.yaml` without `workflows` permission" — the workflow's own token may not update a workflow file, not even in a mirrored upstream commit, and no `permissions:` grant changes that. The mirror is therefore pushed with the heraldbot App's token like the candidate and the consumed branch. An App push starts upstream's workflow files at the mirrored refs (runners the fork does not have, registries it must not touch), so mirror.sh records what it pushed and mirror-quiet.sh cancels those runs right after; the ledger stays on GITHUB_TOKEN.
…arent The second sync run reached its "checks red" step and failed there: `pull request create failed: GraphQL: Resource not accessible by integration (createPullRequest)`. In a checkout of a fork, gh resolves the base repository to the parent (kagent-dev/kagent) unless --repo is given, and the App is not installed there. Both `gh pr create` calls now name this repository. git-push-as.sh emits its `::add-mask::` line only under GitHub Actions: run by hand (the manual re-pin) it printed the encoded credential to the terminal instead of masking it.
…warm line SUBSTRATE_REPO defaults to oci://ghcr.io/giantswarm/substrate/helm and SUBSTRATE_VERSION to 0.0.27-dev.giantswarm.2026-09-10.19-33-37.h734ec53: the substrate and substrate-crds subcharts of kagent and kagent-crds are the fork giantswarm/substrate's builds (upstream v0.0.26 — the version go/go.mod pins — plus cherry-picked fixes). The Go module pin stays upstream's release, so the version is set here rather than derived from go.mod; both move together at a re-pin. Signed-off-by: Timo Derstappen <timo@giantswarm.io>
The Makefile's SUBSTRATE_VERSION / SUBSTRATE_REPO are the one pin: a substrate-pin target prints them, the e2e job reads them into its environment and installs the line's substrate-crds and substrate charts and its ateom-gvisor worker image (the fork's tags carry no v). kubectl-ate for the CA/JWT bootstrap stays upstream's v0.0.26 binary — the line publishes no CLI. Before, the job exported SUBSTRATE_VERSION=0.0.26, which also drove make helm-version and made the fork registry look for a chart it never published. Signed-off-by: Timo Derstappen <teemow@gmail.com>
The consumed branch was renamed from poc/agent-platform to giantswarm (the GitHub rename retargets the default branch, the ~DEFAULT_BRANCH ruleset and open pull requests); every reference follows: the push and pull-request triggers of ci.yaml, image-scan.yaml and tag.yaml, CONSUMED in sync-upstream.yaml and scripts/fork/repin.sh, FORK.md and the README. The dev version's branch part becomes 'giantswarm' — 0.11.0-dev.giantswarm.<date>.<time>.h<sha7> — and the consumers' semverFilter moves with it (.*-dev\.giantswarm\..*).
FORK.md 'Release scheme': a release is a tag vX.Y.Z-gs.N on the consumed branch — X.Y.Z the upstream version the pin anticipates (DEV_BASE_VERSION, 0.11.0), N the line's counter from 1; first release v0.11.0-gs.1. Sorted 0.11.0-dev.giantswarm.… < 0.11.0-gs.1 < 0.11.0, so a release outranks its dev builds and never the upstream release it anticipates, and the switch to an upstream tag is a range change. Published by tag.yaml to ghcr.io/giantswarm/kagent (images + helm/ charts), digests recorded by the record job; ghcr.io rather than gsoci/the catalog because GitHub Actions has no ACR credential (the org publishes to gsoci from CircleCI) and the distinct OCI path plus the pre-release version each keep the chart out of the 0.10 wrapper's fleet range >=0.2.0 <1.0.0 at gsoci.azurecr.io/charts/giantswarm/kagent. tag.yaml's version step refuses any tag that is not v<DEV_BASE_VERSION>-gs.<N> — a bare vX.Y.Z is a name upstream will use and the mirror would then skip upstream's tag of that name. A 'Releases' table (release, tag, commit, pin, digests) is filled per tag after the agentlab proof of its dev build.
…olicy: keep) The kagent-crds chart ships its CustomResourceDefinitions as plain templates (there is no crds/ directory), so `helm uninstall kagent-crds` deletes the five CRDs and, with them, every AgentTemplate, Harness, ModelConfig, ModelProviderConfig and RemoteMCPServer in the cluster. CRD charts conventionally mark their CRDs `helm.sh/resource-policy: keep` so that removing or replacing the release leaves the CRDs (and the objects behind them) in place; a CRD is then removed deliberately, with `kubectl delete crd`, never as a side effect of a release operation. The annotation is added where the manifests are produced rather than in the chart files: a `+kubebuilder:metadata:annotations` marker on each of the five v1alpha3 types, so controller-gen emits it into go/api/config/crd/bases and `make controller-manifests` copies it into helm/kagent-crds/templates unchanged. The chart templates stay a verbatim copy of the generated CRDs, the manifests-check CI job keeps passing, and no new tool (yq, sed post-processing) enters the generation. The generated bases carry the annotation too: nothing in the repository applies them directly (the e2e installs the chart), and to `kubectl apply` it is an inert annotation. Verified: `make controller-manifests` regenerates exactly the five annotation lines in bases and templates; `helm lint helm/kagent-crds` passes with and without `substrate.enabled`; `helm template` differs from the previous render only in the five added lines, and the CRDs of the kmcp-crds and substrate-crds subcharts are untouched.
…ources A skill or plugin source on a private repository has no credential path: the materialiser shells out to git with the process environment and no helper, so an AgentTemplate cannot take skills from a private repository at all. AgentTemplate.spec.skills[].source.git and plugins[].source.git gain an optional credentialRef, a key in a same-namespace Secret (https only, enforced by CEL). Git speaks Basic authentication, so the key holds the base64 of "<username>:<token>" for the URL's host; GitHub takes "x-access-token:<token>". The credential follows the gateway injection path every other credential takes. CompileCredentials binds each referenced key to the source URL's host as "Authorization: Basic <value>", an ate-secret:// URI the egress gateway resolves on every request to that host; the actor never holds the token, and the materialiser clones with no credential of its own. The compiler still records one Secret-backed environment variable per referenced key, KAGENT_ARTIFACT_CREDENTIAL_<hash of namespace, name and key>, which reaches the actor as the inert placeholder like a model key, so a missing Secret or key fails compilation as a reference failure (ResolvedRefs=False) that is retried when the Secret appears. A rotation is not a new revision: the gateway reads the Secret live. Substrate matches credentials by hostname, so one credential serves every source of an agent tree on that host, public repositories included, and two different Secrets for one host are refused at compile time. A wrong token fails the golden boot on the host's 401.
…chart at publish The published chart's defaults named only the images' version; the two pins a consumer needs beyond that — the index digest of the Go ADK runtime image (a kagent.dev/v1alpha3 Harness accepts nothing but a digest, and Helm cannot resolve one at render time) and the Substrate worker image the chart's WorkerPool runs — were resolved by `record` after the chart push, or known only to the Makefile, so the platform's meta chart carried copies of both and re-pinned them by hand. `push-helm-chart` runs after `push-images`, so the indexes exist: it now resolves the golang-adk and claude-harness index digests the way `record` does (`docker buildx imagetools inspect`) and stamps them into `controller.agentImage.digest` and `runtimeImages.claudeHarness.digest`; it stamps `substrateWorkerPool.workerImage` with the gVisor worker of the Makefile's `SUBSTRATE_VERSION`/`SUBSTRATE_REPO` (`make -s substrate-pin`, the same source ci.yaml reads for the e2e cluster) over the empty default the template refuses. Every stamp is checked with yq; a missed anchor fails the job as the existing checks do. `record` additionally reads `helm show values` of the pushed chart and fails when its digests disagree with the ones it writes to builds.md. The anchors are the values keys of the two patches that follow this one on the consumed branch (the optional default Harness, the runtime-images ConfigMap); the workflow runs on the branch head only. FORK.md gets the patch-table rows of all three. Fork infrastructure — tag.yaml is fork-rewritten; never upstream. Refs giantswarm/giantswarm#37765, giantswarm/giantswarm#37705. Signed-off-by: Timo Derstappen <teemow@gmail.com>
…ess for the chart's own runtime image On the kagent API v2 an agent runs on a Harness (kagent.dev/v1alpha3): a runtime image pinned by digest, an environment and a Substrate policy. The chart ships the Go ADK runtime image (controller.agentImage) but knows it by tag only, and renders no Harness for it — so every consumer has to resolve the image's digest out of band and write its own Harness before the first AgentTemplate can become Ready (kagent-dev#2529 is the missing build-time digest; kagent-dev#2536 proposed the value this change uses). controller.agentImage.digest carries the image's index digest (empty in git; a published chart carries the digest of its build). A new `harness` block renders one Harness — `kagent` by default, in the chart's namespace, on the `kagent` runtime adapter — from that digest, only when harness.create is set (default false, so nothing changes for a consumer that renders its own Harness objects): workload.image is <registry>/<controller.agentImage.repository>@<digest> unless harness.image names another digest-pinned image; workerPoolRef defaults to substrateWorkerPool.name; snapshotLocation is required (the CRD requires the snapshot policy); env and allowedAgentTemplates.selector are forwarded. A tag instead of a digest, a missing digest or a missing snapshot location fail the render with the reason, instead of failing at admission. The Harness keeps helm.sh/resource-policy: keep — it is the runtime of every agent admitted under it. The digest-pinned reference is built by one helper (kagent.runtimeImage) so other templates can list the same image. Unit tests cover the default (no Harness), the rendered object and its image, the registry precedence, the harness.image override, the forwarded policy and the three failures. Signed-off-by: Timo Derstappen <teemow@gmail.com>
…he Substrate image cache Agent Substrate's atelet can pre-pull a pinned set of images on every worker node (imageCache.pinnedImages), so an agent's first turn does not wait for the pull of its runtime image. The set has to name the images agents actually run on — this chart's runtime images at their digests — and today only the operator of a cluster can write it, from a digest resolved out of band, and re-write it on every upgrade. The chart now publishes them itself: the ConfigMap `<fullname>-images` carries one key, substrate-values.yaml, a values fragment for the Substrate chart with atelet.imageCache.pinnedImages listing the image the chart's Harness runs (harness.image, else the chart's own Go ADK image at controller.agentImage.digest — so an override moves the set with it) and the Claude runtime image at runtimeImages.claudeHarness.digest. An image whose digest is empty is left out: a chart rendered from git lists nothing, a published chart lists its build. The ConfigMap sits in the release namespace, deliberately not the chart's namespaceOverride, so a Substrate release installed next to this one can read it as values (a Flux HelmRelease's valuesFrom) without crossing namespaces. Unit tests cover the empty set, the two-image set and its name under a fullnameOverride, the harness.image override and the namespace. Signed-off-by: Timo Derstappen <teemow@gmail.com>
SUBSTRATE_VERSION moves from the bootstrap dev build 0.0.27-dev.giantswarm.2026-09-10.19-33-37.h734ec53 to 0.0.27-gs.6, the release of giantswarm/substrate the platform's meta chart runs today (its worker image ghcr.io/giantswarm/substrate/ateom-gvisor:0.0.27-gs.6). tag.yaml stamps substrateWorkerPool.workerImage from this pin, so a release cut from the dev pin would install a dev worker as soon as the meta chart stops forwarding the image by hand; the e2e cluster (ci.yaml) runs the same version. Upstream pin unchanged: kagent-dev/substrate v0.0.26 in go/go.mod, the ate-api contract. Signed-off-by: Timo Derstappen <teemow@gmail.com>
SUBSTRATE_VERSION moves from 0.0.27-gs.6 to 0.0.27-gs.7, the release of giantswarm/substrate that carries giantswarm/substrate#25: ateom re-asserts its capacity report every 10 s, so the worker record the pod-IP repair re-creates receives its capacity within one interval instead of staying at zero until the pod restarts (giantswarm/giantswarm#37762). The fix lives in the ateom-gvisor image, and tag.yaml stamps substrateWorkerPool.workerImage from this pin, so the published chart of the next release installs ghcr.io/giantswarm/substrate/ateom-gvisor:0.0.27-gs.7; the e2e cluster (ci.yaml) runs the same version. Inert under the agent-platform 4.7.x meta chart, which forwards the worker image by hand; live from 4.8.0. Upstream pin unchanged: kagent-dev/substrate v0.0.26 in go/go.mod, the ate-api contract. Signed-off-by: Timo Derstappen <teemow@gmail.com>
…watch The ConfigMap `<fullname>-images` is meant to be read as values by a Flux HelmRelease (valuesFrom). helm-controller re-reads such a ConfigMap on its own only when the HelmRelease spec changes or on the release's interval, unless the ConfigMap carries the label its --watch-configs-label-selector matches — reconcile.fluxcd.io/watch: Enabled by default. Without it, a change of the Harness image alone (a harness.image override, an upgrade that only moves the digests) leaves the consumer's pinned image set stale for up to that interval. The ConfigMap now carries the label, so a HelmRelease reading it is upgraded as soon as an image reference changes. Any other reader ignores the label. A unit test asserts it. Signed-off-by: Timo Derstappen <teemow@gmail.com>
…elector The chart's Harness selects the AgentTemplates it admits by harness.allowedAgentTemplates.selector, and its default selects by the chart's own label, kagent.dev/harness: kagent. A consumer that admits templates by a label of its own has to remove that default key: Helm merges maps on coalesce, so an override adds keys next to the default and the selector ends up requiring both labels. The one way to remove a key from the merged map is a null in the consumer's values, and a null does not survive every path values take to the chart. A Flux HelmRelease that already exists is updated with a JSON merge patch, in which a null means "remove the key" — the key is dropped from the object, the chart default comes back on coalesce, and a Harness upgraded in place selects by two labels while a freshly created one selects by one. Every template labelled for the consumer's key alone is then admitted by no Harness. The Harness template now drops every matchLabels entry whose value is the empty string before rendering the selector. An empty string is stored by every values path, create and patch alike, and it is never a valid selector value, so it means "not this key". A selector left with no requirement at all fails the render: an empty label selector selects everything, and a Harness that admits every AgentTemplate in its namespace is not what a blanked default asks for. matchExpressions are untouched. Unit tests cover a blanked key next to a kept one and next to matchExpressions, a selector kept on matchExpressions alone, and the failure when nothing is left to select by. Signed-off-by: Timo Derstappen <teemow@gmail.com>
…e agentskills.io set
The Go ADK runtime loads an agent's skills through the ADK's
tool/skilltoolset/skill package, whose Parse decodes the SKILL.md frontmatter
with yaml.Decoder.KnownFields(true) into a struct that knows the
agentskills.io fields only (name, description, license, compatibility,
metadata, allowed-tools). One field outside that set failed the whole agent:
failed to create agent: failed to load skills: invalid frontmatter:
parse frontmatter: yaml: unmarshal errors:
line 3: field user-invocable not found in type skill.Frontmatter
Skills are written once and run on several harnesses. Claude Code's skills
carry user-invocable, argument-hint and disable-model-invocation at the top
level; skills in the wild carry a plain version. The same SKILL.md loads on
the Python runtime, whose Frontmatter model allows extra fields
(google/adk-python, skills/models.py: extra="allow"), and in Claude Code, so
an agent that ran on the Python runtime stopped booting on the Go one.
The runtime now hands the ADK a filesystem view of the skills directory in
which every skill's SKILL.md carries only the frontmatter fields the ADK
knows (adk/pkg/skillfs). The ADK's parser, validation and resource rules stay
the single authority on what a skill is: name and description are still
required and validated, the directory name must still match, resources are
still confined to references/, assets/ and scripts/, and a SKILL.md the
filter cannot read as frontmatter (no separator, YAML that does not parse,
a document that is not a mapping) is served unchanged so the ADK reports it
as before. Fields nested under metadata are untouched.
The alternative, decoding without KnownFields in the ADK itself, belongs
upstream in google/adk-go; until it lands, the runtime is the place that
knows its skills come from other harnesses too.
Signed-off-by: Timo Derstappen <teemow@gmail.com>
… gVisor worker's terminate collects a crashed golden actor (giantswarm/giantswarm#37773) (#32) SUBSTRATE_VERSION 0.0.27-gs.7 -> 0.0.27-gs.9: v0.0.27-gs.9 = gs.8 + giantswarm/substrate#30, the worker (ateom-gvisor) whose TerminateWorkload skips runsc containers that are already gone and accepts a runsc delete that fails after removing its container. Before it a superseded AgentTemplate revision whose golden boot had crashed was never collected and its golden actor pinned one worker of the pool for good (gazelle 2026-09-13, two of four workers). The published chart stamps substrateWorkerPool.workerImage with the new worker; the platform's WorkerPool rolls once. FORK.md's pin row follows. Signed-off-by: Timo Derstappen <teemow@gmail.com>
…giantswarm.2026-09-14.19-48-17.h21027aa Upstream 2d843e3 (kagent-dev#2802) moved kagent to Substrate v0.0.29: the a2a gateway addresses an actor by the ate-target-actor header, which the router of Substrate 0.0.27 does not know, so the two lines move together. The Giant Swarm line of Substrate is re-pinned onto kagent-dev/substrate v0.0.29 in the same move (giantswarm/substrate sync/20260914-v0.0.29); this is its dev build, for the agentlab proof of the pair. The release pin (0.0.30-gs.1) follows the proof. Signed-off-by: Timo Derstappen <timo@giantswarm.io>
The Substrate line's release of kagent-dev/substrate v0.0.29 plus its ten carried patches (giantswarm/substrate FORK.md), the same tree as the dev build 0.0.30-dev.giantswarm.2026-09-14.19-48-17.h21027aa the agentlab proof of the three lines ran on; the published chart stamps its worker image ghcr.io/giantswarm/substrate/ateom-gvisor:0.0.30-gs.1 into substrateWorkerPool.workerImage. Signed-off-by: Timo Derstappen <timo@giantswarm.io>
Add `promptCaching` (default false) and `cacheTTL` ("5m" | "1h") to
`spec.anthropic` on ModelConfig, mirroring the knobs BedrockConfig got in
When enabled, every Messages API request carries three `cache_control`
breakpoints: the last tool definition, the last system prompt block and
the last content block of the latest conversation turn. Anthropic
renders tools, then system, then messages and caches the prefix up to
each breakpoint, so an agent loop reads its whole previous history from
the cache and writes only the new turn instead of paying full input
price for the same prefix on every call. "1h" opts into the 1-hour
cache; "5m" (the CRD default) leaves the API default in place.
Go runtime: buildAnthropicParams now owns the request shaping so it can
be unit-tested, markAnthropicCacheBreakpoints sets the markers, and
usage folds cache_read/cache_creation into PromptTokenCount with
cache_read surfaced as CachedContentTokenCount so the UI shows the
effect. Python runtime: KAgentAnthropicLlm wraps the SDK client's
messages resource and marks the same three breakpoints before each
request; google-adk already maps the cache usage fields.
The Claude harness compiler treats the defaulted cacheTTL "5m" like the
Bedrock one and keeps rejecting promptCaching, since Claude Code caches
natively.
Refs kagent-dev#2787, kagent-dev#1866.
Signed-off-by: Timo Derstappen <teemow@gmail.com>
(cherry picked from commit a0c974b)
…ests Shallow copies would let a nested in-place mutation of a message or tool block go unnoticed, which is exactly what these tests guard against. Signed-off-by: Timo Derstappen <teemow@gmail.com> (cherry picked from commit f5679c2)
… golden boot a workload keeps failing is bounded and its cause carried up (giantswarm/giantswarm#37801) (#49) SUBSTRATE_VERSION 0.0.30-gs.1 -> 0.0.30-gs.4: v0.0.30-gs.4 = gs.3 + giantswarm/substrate#39. The worker (ateom-gvisor) keeps each application container's last output lines and quotes them in the readiness-deadline error; ate-api counts the boots a workload itself fails (WORKLOAD_NOT_READY) and at the third crashes the golden actor and fails the ActorTemplate with GoldenActorNotReady, the bound and the last boot's error — which this controller's pair reconciler already surfaces verbatim as Ready=False ActorTemplateFailed. Before it a golden actor whose workload exited before readiness (an anonymous fetch of a private skill, a skill path missing at the pinned commit) was re-booted once a minute for as long as the template existed and the AgentTemplate read ActorTemplatePending indefinitely (gazelle 2026-09-15, about 120 boots over two hours). The published chart stamps substrateWorkerPool.workerImage with the new worker; the platform's WorkerPool rolls once. FORK.md's pin row follows. Signed-off-by: Timo Derstappen <teemow@gmail.com>
A runtime's network identity says nothing about the agent it runs once agents share a pod, a ServiceAccount or an egress: on the kagent API v2 every AgentTemplate is an actor in a shared WorkerPool and every connection leaves through one egress, so a gateway in front of the model providers attributes all usage to that egress and knows no user. Only the runtime knows which agent a call is made for, and only the controller knows whose turn it is. The controller now tells the runtime which AgentTemplate it executes (KAGENT_AGENT_TEMPLATE, next to KAGENT_NAMESPACE; KAGENT_NAME stays the per-Harness app name), and every model call carries that identity as request headers: x-kagent-agent and x-kagent-agent-namespace, set by the Go ADK's HTTP transport after a model's configured default headers so a ModelConfig cannot name another agent. The person of the turn travels as x-kagent-user: the controller's a2a gateway forwards it to the actor as request metadata only when its authenticator resolved the identity from the caller's validated token (a session carrying claims) — the unsecure mode's X-User-Id and its default, an agent's propagated caller id and the control plane resolve nobody and forward nothing — and the Go ADK's A2A interceptor carries that metadata into the turn's context, where the transport sets the header. The session owner in x-user-id and the caller's bearer in authorization are never sent as the person: the user's context key is its own type, so it cannot share an address with the bearer token's key the way two pointers to zero-size structs can. An absent header is absent, never a stand-in. The Claude harness sends the two agent headers through ANTHROPIC_CUSTOM_HEADERS, which the controller renders and owns; it knows no user per turn and names none. The headers are an accounting identity the runtime asserts about itself, never an authorization input. Signed-off-by: Timo Derstappen <timo@giantswarm.io>
The e2e job read the controller log from the Deployment once the tests were over, which holds only the final pod. Tests that roll the controller (withControllerEnv, the sandbox restart) lost the log of every earlier pod. The job follows every controller pod from the moment it runs, one file per pod, and stops the followers with the other log streams. Signed-off-by: Timo Derstappen <teemow@gmail.com>
Signed-off-by: Timo Derstappen <teemow@gmail.com>
giantswarm/substrate#182 lets Pause, Suspend and Resume carry a fencing token; the quiescence fence of #181 needs it. The go.mod replace and the charts' SUBSTRATE_VERSION move together. Signed-off-by: Timo Derstappen <teemow@gmail.com>
A quiescence claim's records fence a holder that stopped, not one frozen past its lease while its Pause or Suspend is in flight: the successor settles a running Actor as not quiesced and releases the claim, and the late request then pauses or suspends an Actor the session believes is running. Every Pause, Suspend and fencing Resume a holder sends now carries a Substrate fencing token: its executor id and a generation from a PostgreSQL sequence that grows with every request. Before it releases a claim on a running Actor, the holder sends a fenced Resume with a new generation, which records the token without touching the runtime, so Substrate refuses a late request with FailedPrecondition. A holder whose own request is refused that way was superseded and leaves the claim to its successor. Calls without a token are admitted as before. Fixes #181 Signed-off-by: Timo Derstappen <teemow@gmail.com>
Signed-off-by: Timo Derstappen <teemow@gmail.com>
TestServerTelemetryShape read the span exporter and the metric reader right after GetVersion and Check returned. The client has its response before the server is done with the RPC: otelgrpc ends the server span and records rpc.server.call.duration on the stats handler's End event, after the response is written, so either could still be missing and the test failed with "rpc.server.call.duration was not recorded". Wait, bounded, for the GetVersion server span and the recorded metric before asserting; the assertions are unchanged. Signed-off-by: Timo Derstappen <teemow@gmail.com>
Signed-off-by: Timo Derstappen <teemow@gmail.com>
Every Pause, Suspend and fencing Resume a quiescence holder sends carries a fencing token. A Substrate whose Actor API predates the token refuses any field it has no descriptor for: InvalidArgument: request: Invalid value: unknown field with protobuf tag 10001 so every quiesce failed and the fenced Resume before a claim's release was retried forever: the claim was never released and the session refused every turn after its first. The Substrate client now detects that refusal, remembers that its ate-api takes no token, logs one warning naming the lost fence, and sends that request and every later one without the token. The runtime boundary works as before the token; only a frozen holder's late Pause or Suspend is no longer refused. ate-api exposes no version a controller could check at startup, and Substrate is usually installed on its own, so the chart's dependency cannot hold the pair together. Fixes #230 Signed-off-by: Timo Derstappen <teemow@gmail.com>
Signed-off-by: Timo Derstappen <teemow@gmail.com>
Co-authored-by: giantswarm-align-files[bot] <288255335+giantswarm-align-files[bot]@users.noreply.github.com> Co-authored-by: Quentin Bisson <quentin.bisson@gmail.com>
Nothing wrote Harness.status, so kubectl get harness showed an empty READY column and the gRPC Harness.ready was always false although the Harness served. The controller derives Ready from the Harness's WorkerPool and the golden boots of the Agents that reference it, and writes it on its own status queue. One successful golden boot proves the Harness's workload boots, so it reads True; a failed boot reads False with its reason while no Agent booted, since an Agent's own configuration can fail its boot too; Unknown while nothing has tried the workload. The status write keeps status.capabilities. Refs kagent-dev#2764 Signed-off-by: Timo Derstappen <teemow@gmail.com>
Signed-off-by: Timo Derstappen <teemow@gmail.com>
Helm's and Flux's readiness wait reads a Ready condition other than True as a rollout in progress, so a Harness without Agents, Ready Unknown, held every upgrade of the chart that ships it until the wait timed out and the release rolled back. Before any golden boot finished the resolved WorkerPool keeps the Harness Ready (WorkerPoolResolved); a successful boot reads Booted, a failed one with none booted False. Signed-off-by: Timo Derstappen <teemow@gmail.com>
Signed-off-by: Timo Derstappen <teemow@gmail.com>
Signed-off-by: Timo Derstappen <teemow@gmail.com>
* fix(session): retry a failed quiesce once the Actor runs again When the pause or suspend after a turn returned an error, settleIdleWork re-read the Actor and released the claim without a snapshot as soon as it was RUNNING again, for example after a first SuspendActor that outlived its deadline and that Substrate then abandoned. The suspend was never retried, so the turn kept no snapshot, and nothing on the session said why. A RUNNING Actor now gets the pause or suspend once more, fenced with a newer token, at the backoff settleIdleWork already uses; its outcome settles the same way. The claim is released without a snapshot only when the retry fails too, or the Actor is crashed or gone. The release records why on the session as last_quiescence_failure (task, reason, message and time), which GetSession, ListSessions and the CLI's session details show; a later boundary with a snapshot clears it. Signed-off-by: Timo Derstappen <teemow@gmail.com> * docs(fork): carry the retry of a failed quiesce Signed-off-by: Timo Derstappen <teemow@gmail.com> --------- Signed-off-by: Timo Derstappen <teemow@gmail.com>
…ns (#227) * ci(telemetry): fetch the contract's upstream sources once at their pins The Telemetry Contract Check cloned three GitHub repositories on every Weaver run: the two registries telemetry/registry/manifest.yaml depends on (the GenAI registry re-cloning the core one through its own manifest) and the shared policy packages, eleven clones per `make semconv-verify`. A clone that flaked turned the check red on a pull request's first run, and the result of the check depended on reaching GitHub. The three sources are now pinned in telemetry/versions.env and fetched once by `make semconv-deps` into the git-ignored telemetry/deps: a depth-one fetch of the pinned ref per repository, reused without touching the network while its pin stands. The manifest names the checkouts, semconv-check passes the policies from them, and the fetch points a fetched registry's own git dependency at the checkout of the same pin, refusing a pin that wants another ref, since the contract and its dependencies must agree on the core conventions. CI caches the directory keyed on the pins. Weaver's `[resolve.schema_url_overrides]` does not serve here: in v0.26.1 and v0.27.0 the loader of `registry check` and `registry generate` clones a manifest dependency before any override is consulted; the overrides serve the on-demand resolver of live-check and serve. Proof: `make semconv-verify` with the checkouts present, the Weaver container on `--network none` and a dead HTTPS proxy on the host, exits 0 after the check, the four policy fixtures, the four generations and the drift check. Signed-off-by: Timo Derstappen <teemow@gmail.com> * docs(fork): ledger row for the telemetry contract's pinned sources Signed-off-by: Timo Derstappen <teemow@gmail.com> * ci(telemetry): fetch the pinned sources before generating, name the core pin apart semconv-generate reads the manifest's registries from telemetry/deps, so it runs semconv-deps like semconv-check. The core registry's pin is SEMCONV_CORE_REGISTRY: the Makefile includes versions.env and names the kagent registry SEMCONV_REGISTRY. Signed-off-by: QuentinBisson <quentin@giantswarm.io> --------- Signed-off-by: Timo Derstappen <teemow@gmail.com> Signed-off-by: QuentinBisson <quentin@giantswarm.io> Co-authored-by: QuentinBisson <quentin@giantswarm.io> Co-authored-by: Quentin Bisson <quentin.bisson@gmail.com>
…unner TestServe polled GetProcess for the first process's COMPLETED within require.Eventually's one second. On a loaded runner the shell's start plus the poll take longer, so the test failed with "Condition never satisfied" and passed on rerun. Name one bound of five seconds, like the cleanup's stop; a passing run still returns at the first COMPLETED. Signed-off-by: Timo Derstappen <teemow@gmail.com>
…esourceVersion (#245) The Harness and ModelConfig status writes were an UpdateStatus of a copy of the informer cache, so a write the cache had not observed yet failed them with a Conflict. Both are now a JSON merge patch of the status subresource that carries only the fields the controller derives; the Harness patch never carries status.capabilities. Signed-off-by: QuentinBisson <quentin@giantswarm.io>
TestSandboxTemplateCatalog and TestSandboxKRTInformerQueueAndRestartCleanup start an envtest API server and fail a plain go test of their package when the envtest binaries are not set up. Skip them with a message naming KUBEBUILDER_ASSETS, and let make test set the variable through setup-envtest, as the CI job does. Closes #232 Signed-off-by: Timo Derstappen <teemow@gmail.com>
…r-token propagation (#254) Substrate 1.5's egress replaces a credential header only when the request carries it. The Claude forwarder deleted every server's compiled Authorization and set the caller's token in its place, so a Secret-backed Authorization was never sent without a caller token and overwrote the caller's token with one. A server that configures Authorization keeps its gateway placeholder; any other takes the caller's token, or none. The adapter resolves the compiler's credential references before forwarding. CompileCredentials refuses a caller-owned MCP server on a host with a gateway authorization binding when KAGENT_PROPAGATE_TOKEN is on. Closes #253
…252) * fix(skills): git sends the Authorization the egress gateway replaces Substrate's egress credential effect replaces the value of a header a request carries and leaves a request without it untouched. Git sends no Authorization before a challenge, so a private git skill was fetched without its credential, answered 401 and the materializer failed on a terminal prompt. The compiled source of a git artifact with a credentialRef is marked authenticated, and the materializer has git send a placeholder Authorization header for that source's URL only, configured for the one invocation and never written to a gitconfig. Git never prompts, so a refused credential fails with its own error. A public source sends no header. Signed-off-by: Timo Derstappen <teemow@gmail.com> * docs(api): credentialRef reaches only requests that carry the header The egress gateway replaces an Authorization header a request carries and adds none, so the field's description and the ledger row no longer claim the credential goes on every request to the host, or to the host's public sources. Signed-off-by: QuentinBisson <quentin@giantswarm.io> --------- Signed-off-by: Timo Derstappen <teemow@gmail.com> Signed-off-by: QuentinBisson <quentin@giantswarm.io> Co-authored-by: QuentinBisson <quentin@giantswarm.io>
QuentinBisson
force-pushed
the
docs/caller-credential-at-egress
branch
from
October 8, 2026 21:21
1c6a54b to
d32088e
Compare
…or every client kind (#256) * test(e2e): prove the egress credential contract through the gateway for every client kind Each client a Secret-backed binding covers (model keys, RemoteMCPServer headersFrom on every harness and on Claude with KAGENT_PROPAGATE_TOKEN, git skill and plugin sources with a credentialRef) talks to a recording upstream through Substrate's egress gateway. The upstream must see the Secret value on every request and the placeholder on none, so a client that sends no header fails against a gateway that replaces only present headers. The git upstream is git http-backend on the test host, served over HTTPS under a cluster Service name. The CI job creates a CA per run, has the gateway trust it (atenetEgress.upstreamTrust.caBundle) and passes it to the tests to sign that upstream. Golden boots fetch git sources in atespace ate-golden, so the job lets it resolve kagent Secrets. Signed-off-by: QuentinBisson <quentin@giantswarm.io> * fix(translator): send an Ollama Cloud model of the Python runtime to api.ollama.com The compiler left OLLAMA_API_BASE unset for a cloud model so that the runtime would route it, but only the Go runtime routes: the Python runtime sent a cloud model to localhost:11434. The compiler now writes the cloud endpoint whenever OllamaReachesCloud routes there. The Ollama SDK sends OLLAMA_API_KEY, the gateway placeholder, as the bearer, which a unit test now pins. Signed-off-by: QuentinBisson <quentin@giantswarm.io> --------- Signed-off-by: QuentinBisson <quentin@giantswarm.io>
Signed-off-by: QuentinBisson <quentin@giantswarm.io>
The turn rewrites the host's rule instead of adding one, since only the first matching rule applies; the policy is the record of the caller, with cleanup on every way a turn ends and a sweep at start; the plain HTTP listener injects too, so caller hosts must be https; the store needs the Recreate strategy; the scope is the claude harness; and the trust section names ate-api's missing per-method authorization and the CONNECT and Host mismatch on the plain HTTP listener. Signed-off-by: QuentinBisson <quentin@giantswarm.io>
QuentinBisson
force-pushed
the
docs/caller-credential-at-egress
branch
from
October 8, 2026 21:40
d32088e to
76dc3ec
Compare
QuentinBisson
force-pushed
the
giantswarm
branch
from
October 8, 2026 22:27
57a0030 to
453cda7
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
With
KAGENT_PROPAGATE_TOKENthe caller's bearer enters the actor, and keeping it from the code the actor runs depends on a second user inside the sandbox and on gVisor enforcing it (giantswarm/giantswarm#38054).Change
A design note: the a2agateway keeps the bearer in a
kagent.devcredential provider in the controller and rewrites the caller host's egress rule to a per-turn URI, so Substrate's egress gateway injects it and the actor never holds it. It covers cleanup on every way a turn ends, the provider's mTLS and store, the scope (claude harness, opt-in) and its prerequisites: giantswarm/substrate#104, per-method authorization in ate-api, and the CONNECT andHostmismatch on the plain HTTP egress listener.