Skip to content

feat(api): Harness caller credentials the egress gateway sets as the turn's caller - #147

Draft
QuentinBisson wants to merge 177 commits into
giantswarmfrom
feat/caller-credential-provider
Draft

QuentinBisson wants to merge 177 commits into
giantswarmfrom
feat/caller-credential-provider

Conversation

@QuentinBisson

@QuentinBisson QuentinBisson commented Oct 2, 2026 •

Copy link
Copy Markdown

Problem

A coding agent must git clone and call the GitHub API as the person who sent the turn, without a GitHub token, or the caller's bearer, in the sandbox. Caller routes (#126) do it for git only, through a loopback forwarder, an env var and an agentgateway route.

Change

Not proven in agentlab yet. Supersedes #126; towards giantswarm/giantswarm#38056, sign-in link follow-up giantswarm/giantswarm#38107.

teemow and others added 30 commits September 25, 2026 09:28
…y=disabled without listing its tools

A RemoteMCPServer that authenticates every caller has no credential to
offer the controller: agents forward the caller token at run time
(KAGENT_PROPAGATE_TOKEN), but the controller's own discovery request
carries no bearer, the server answers 401 and the CR stays
Accepted=False (DiscoveryFailed: Unauthorized) although every agent that
references it works. spec.headersFrom is not a way out: its values are
resolved into each agent's config and applied last by the runtime, so a
controller credential there would replace the caller token in every
agent.

The kagent.dev/discovery=disabled label now opts a RemoteMCPServer out
of tool discovery: the reconciler skips the listing, reports
Accepted=True with reason DiscoveryDisabled, clears
status.discoveredTools and keeps the server in the catalog with no tools
and no LastConnected. The controller also reacts to label changes, so
the opt-out and its removal apply at once rather than on the next
periodic refresh.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
(cherry picked from commit 7c0e532)
…f the catalog

main honors the label nowhere: kagent-dev#2565 removed
predicates.DiscoveryDisabledPredicate together with the old MCPServer
tool controller, and the catalog reconciler restored in kagent-dev#2644 watches
every kmcp MCPServer unfiltered, so the opt-out the docs describe for
the agentgateway setup was a no-op on main.

The MCPServer catalog reconciler now skips discovery for a labeled
server and drops its catalog projection. That is what the 0.10.x
predicate achieved, minus two stale-row cases: a label added after
discovery left the row in place, and the deletion of a labeled server
never removed it. The controller filters no events, so a label change
re-enters Reconcile at once.

A RemoteMCPServer stays in the catalog with the DiscoveryDisabled
status on purpose: it is kagent's own CR that agents reference. An
MCPServer is kmcp's resource that kagent only projects, has no kagent
status to explain an empty row, and AgentTemplates cannot reference it
directly.

Rebased onto main after kagent-dev#2774 (simplify database client operations):
the reconciler's deleteCatalog is gone, the catalog store's
DeleteToolServer with the MCPServer group kind is what the not-found
path calls now, so the opt-out path calls the same.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
(cherry picked from commit bf0b26e)
…n label

helm-controller installs OCI charts under <version>+<digest>; '+' is not a
valid label value character, so every object of the chart failed to apply
under Flux. Sanitised the way helm.sh/chart already is.
…oc branch

The fork's tag.yaml publishes the dev builds of the agent-platform's kagent
component: GitHub-hosted runners (the org has no self-hosted runners), the
registries derived from the repository owner (ghcr.io/<owner>/kagent/<image>,
oci://ghcr.io/<owner>/kagent/helm/<chart>), no PyPI/GitHub-release jobs, the
image matrix trimmed to what the chart references plus the Claude harness
(multi-arch kept), and a push to poc/agent-platform that derives the version
as <base>-dev.<branch>.<committer date>.<time>.h<sha7> — the schema
giantswarm's gitsemver emits, so Flux selects the channel with a semverFilter
on the branch part. workflow_dispatch keeps its optional version input.

The chart's defaults are stamped with the published version and image
repositories before packaging: the tag must be explicit because the chart's
fallback, .Chart.Version, carries `+<digest>` under helm-controller.

The Makefile's -X targets followed $(DOCKER_REPO); with an owner-relative
repository they missed the module, leaving the binaries without version
information. They now follow the Go module path from go.mod.
…failure

apk.cgr.dev intermittently answers some of the parallel anonymous package
fetches from hosted runners with HTTP 403, which apk reports as
"Permission denied" for a random package and which fails the wolfi-based ui
image while the alpine-based images pass. The builder keeps the finished
stages, so the step now retries the build up to three times, repeating only
the failed step; and one image's failure no longer cancels the other three.
The controller and Go ADK images are alpine:3.22 plus `apk add`; the base
layer's libcrypto3/libssl3 3.5.7-r0 carry CVE-2026-14456 (HIGH, fixed in
3.5.8-r0). `apk upgrade --no-cache` before the add brings the runtime layer to
the current Alpine packages, as kagent-dev#2756 did for the harness images.
…w, keep a ledger

The consumed branch was built and published without a test or a scan:
upstream's ci.yaml and image-scan.yaml trigger on main and release/** only,
and their jobs ask for Blacksmith runners the org does not have. Both now run
on every push to poc/agent-platform and sync/**, on every pull request to
poc/agent-platform and (the scan) weekly, on GitHub-hosted runners with a
docker/setup-buildx-action builder, with the image matrices trimmed to what
the fork publishes (controller, ui, golang-adk, claude-harness — the Python
ADK, the Codex harness and the CLI image are not offered on the platform), a
govulncheck job over the Go module graph, no paths-ignore (a docs-only pull
request must still report the required checks) and one aggregate status per
workflow — ci-ok, scan-ok — for the ruleset to require.

sync-upstream.yaml re-pins the line weekly and on demand: fast-forward the
mirror main to upstream main with the workflow token (no workflow reacts, so
upstream's CI never runs on the mirror), rebase the carried commits onto the
upstream head (scripts/fork/repin.sh, which also names commits upstream merged
in the meantime and drops them), push the candidate with the bot token so the
checks run, wait for ci-ok and scan-ok, then force-push the consumed branch
with --force-with-lease so tag.yaml publishes the dev build. A conflicting
carried commit or a red check stops the run and opens a pull request against
the consumed branch that names the cause; nothing else is pushed.

tag.yaml gains a record job: every published build's version, commit, upstream
pin and image and chart digests go into the run summary and into builds.md on
the ledger branch (scripts/fork/ledger-append.sh; re-pins get the same on
re-pins.md). FORK.md is the human half of that ledger — the pin, every carried
commit with its purpose and upstream state, what is published and how, the
automated and the manual re-pin, protection, the contribution process — and
the README opens with it.
The first sync run was refused at the mirror: "refusing to allow a GitHub App
to create or update workflow `.github/workflows/ci.yaml` without `workflows`
permission" — the workflow's own token may not update a workflow file, not
even in a mirrored upstream commit, and no `permissions:` grant changes that.
The mirror is therefore pushed with the heraldbot App's token like the
candidate and the consumed branch. An App push starts upstream's workflow files
at the mirrored refs (runners the fork does not have, registries it must not
touch), so mirror.sh records what it pushed and mirror-quiet.sh cancels those
runs right after; the ledger stays on GITHUB_TOKEN.
…arent

The second sync run reached its "checks red" step and failed there:
`pull request create failed: GraphQL: Resource not accessible by
integration (createPullRequest)`. In a checkout of a fork, gh resolves
the base repository to the parent (kagent-dev/kagent) unless --repo is
given, and the App is not installed there. Both `gh pr create` calls now
name this repository.

git-push-as.sh emits its `::add-mask::` line only under GitHub Actions:
run by hand (the manual re-pin) it printed the encoded credential to
the terminal instead of masking it.
…warm line

SUBSTRATE_REPO defaults to oci://ghcr.io/giantswarm/substrate/helm and
SUBSTRATE_VERSION to 0.0.27-dev.giantswarm.2026-09-10.19-33-37.h734ec53: the substrate and substrate-crds subcharts of
kagent and kagent-crds are the fork giantswarm/substrate's builds (upstream
v0.0.26 — the version go/go.mod pins — plus cherry-picked fixes). The Go
module pin stays upstream's release, so the version is set here rather than
derived from go.mod; both move together at a re-pin.

Signed-off-by: Timo Derstappen <timo@giantswarm.io>
The Makefile's SUBSTRATE_VERSION / SUBSTRATE_REPO are the one pin: a
substrate-pin target prints them, the e2e job reads them into its
environment and installs the line's substrate-crds and substrate charts and
its ateom-gvisor worker image (the fork's tags carry no v). kubectl-ate for the
CA/JWT bootstrap stays upstream's v0.0.26 binary — the line publishes no CLI.
Before, the job exported SUBSTRATE_VERSION=0.0.26, which also drove make
helm-version and made the fork registry look for a chart it never published.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
The consumed branch was renamed from poc/agent-platform to giantswarm (the GitHub rename retargets the default branch, the ~DEFAULT_BRANCH ruleset and open pull requests); every reference follows: the push and pull-request triggers of ci.yaml, image-scan.yaml and tag.yaml, CONSUMED in sync-upstream.yaml and scripts/fork/repin.sh, FORK.md and the README. The dev version's branch part becomes 'giantswarm' — 0.11.0-dev.giantswarm.<date>.<time>.h<sha7> — and the consumers' semverFilter moves with it (.*-dev\.giantswarm\..*).
FORK.md 'Release scheme': a release is a tag vX.Y.Z-gs.N on the consumed branch — X.Y.Z the upstream version the pin anticipates (DEV_BASE_VERSION, 0.11.0), N the line's counter from 1; first release v0.11.0-gs.1. Sorted 0.11.0-dev.giantswarm.… < 0.11.0-gs.1 < 0.11.0, so a release outranks its dev builds and never the upstream release it anticipates, and the switch to an upstream tag is a range change. Published by tag.yaml to ghcr.io/giantswarm/kagent (images + helm/ charts), digests recorded by the record job; ghcr.io rather than gsoci/the catalog because GitHub Actions has no ACR credential (the org publishes to gsoci from CircleCI) and the distinct OCI path plus the pre-release version each keep the chart out of the 0.10 wrapper's fleet range >=0.2.0 <1.0.0 at gsoci.azurecr.io/charts/giantswarm/kagent. tag.yaml's version step refuses any tag that is not v<DEV_BASE_VERSION>-gs.<N> — a bare vX.Y.Z is a name upstream will use and the mirror would then skip upstream's tag of that name. A 'Releases' table (release, tag, commit, pin, digests) is filled per tag after the agentlab proof of its dev build.
…cator

The chart renders AUTH_MODE and AUTH_USER_ID_CLAIM on the controller
(controller.auth.mode, controller.auth.userIdClaim) and
docs/OIDC_PROXY_AUTH_ARCHITECTURE.md describes the two modes, but since
the legacy runtime went (kagent-dev#2595) nothing reads either variable: the
controller calls app.Run with empty Options, so it always runs the
UnsecureAuthenticator -- any caller names themselves with X-User-Id --
and ProxyAuthenticator has no caller at all. A deployment that sets
controller.auth.mode: trusted-proxy behind oauth2-proxy or an API
gateway gets a controller that admits everyone, and nothing tells it so.

app.AuthenticatorFromEnv selects the built-in authenticator AUTH_MODE
names and the controller's main passes it as Options.Authenticator:

- unset or "unsecure": UnsecureAuthenticator, as today;
- "trusted-proxy": ProxyAuthenticator reading the claim AUTH_USER_ID_CLAIM
  names ("sub" by default) from the bearer the proxy validated; a request
  without a bearer is refused whatever X-User-Id says, and the agent-call
  branch (X-Agent-Name set) is unchanged;
- anything else fails startup naming the valid modes, instead of silently
  running unsecure.

The helper is exported so that a library consumer, which cannot import
the internal auth package, can run core's modes too; Options and Run are
untouched, so a consumer that supplies its own authenticator sees no
change. gRPC callers are covered as before: the authentication
interceptor already maps the authorization, x-user-id and x-agent-name
metadata onto the headers the authenticator reads. The docs table drops
the --auth-mode flags, which went with the legacy runtime.
…olicy: keep)

The kagent-crds chart ships its CustomResourceDefinitions as plain
templates (there is no crds/ directory), so `helm uninstall kagent-crds`
deletes the five CRDs and, with them, every AgentTemplate, Harness,
ModelConfig, ModelProviderConfig and RemoteMCPServer in the cluster.
CRD charts conventionally mark their CRDs `helm.sh/resource-policy: keep`
so that removing or replacing the release leaves the CRDs (and the
objects behind them) in place; a CRD is then removed deliberately, with
`kubectl delete crd`, never as a side effect of a release operation.

The annotation is added where the manifests are produced rather than in
the chart files: a `+kubebuilder:metadata:annotations` marker on each of
the five v1alpha3 types, so controller-gen emits it into
go/api/config/crd/bases and `make controller-manifests` copies it into
helm/kagent-crds/templates unchanged. The chart templates stay a
verbatim copy of the generated CRDs, the manifests-check CI job keeps
passing, and no new tool (yq, sed post-processing) enters the
generation. The generated bases carry the annotation too: nothing in the
repository applies them directly (the e2e installs the chart), and to
`kubectl apply` it is an inert annotation.

Verified: `make controller-manifests` regenerates exactly the five
annotation lines in bases and templates; `helm lint helm/kagent-crds`
passes with and without `substrate.enabled`; `helm template` differs
from the previous render only in the five added lines, and the CRDs of
the kmcp-crds and substrate-crds subcharts are untouched.
…Releases table

Images and charts published to ghcr.io/giantswarm/kagent by run 34542852237;
the digests are the index digests a Harness pins. The release images are
rebuilt from the tree of the dev build h0ac5240, so their digests differ from
that build's; consumers of the release pin these.
…ources

A skill or plugin source on a private repository has no credential path:
the materialiser shells out to git with the process environment and no
helper, so an AgentTemplate cannot take skills from a private repository
at all.

AgentTemplate.spec.skills[].source.git and plugins[].source.git gain an
optional credentialRef, a key in a same-namespace Secret (https only,
enforced by CEL). Git speaks Basic authentication, so the key holds the
base64 of "<username>:<token>" for the URL's host; GitHub takes
"x-access-token:<token>".

The credential follows the gateway injection path every other credential
takes. CompileCredentials binds each referenced key to the source URL's
host as "Authorization: Basic <value>", an ate-secret:// URI the egress
gateway resolves on every request to that host; the actor never holds
the token, and the materialiser clones with no credential of its own.
The compiler still records one Secret-backed environment variable per
referenced key, KAGENT_ARTIFACT_CREDENTIAL_<hash of namespace, name and
key>, which reaches the actor as the inert placeholder like a model key,
so a missing Secret or key fails compilation as a reference failure
(ResolvedRefs=False) that is retried when the Secret appears. A rotation
is not a new revision: the gateway reads the Secret live.

Substrate matches credentials by hostname, so one credential serves every
source of an agent tree on that host, public repositories included, and
two different Secrets for one host are refused at compile time. A wrong
token fails the golden boot on the host's 401.
…Releases table

Signed-off-by: Timo Derstappen <teemow@gmail.com>
…ead of leaving it submitted

A turn the gateway could not deliver to the AgentInstance runtime left its
task TASK_STATE_SUBMITTED for good: SendStreamingMessage stored the task in
prepareSend, then a failing dial, a failing ingester start, or a runtime
stream that errored before its first event only returned the error. The
instance kept the task as its active one, every later message was refused
as a conflict, and every client saw a turn that never ends.

A dispatch that fails after the task was stored now records the task as
TASK_STATE_FAILED with the cause as its status message, publishes that
status update to the caller before the stream ends with the error, and,
being a terminal transition, releases the instance's active task. The
runtime holds nothing for such a task, so no snapshot is taken: the failure
is recorded even when the runtime is the unreachable part.

A stream that fails before its first event may also have lost only the
response, so a dispatch run first asks the runtime for the task: one that
answers with it is left to finish it, and the existing recovery finds the
task there; one that does not know the task, or cannot be asked, never
started the turn. A run started by a subscription never fails the task,
since it may be racing the dispatch about to deliver it.

reconcileActiveTask treats a submitted task the runtime does not know as
failed once it is older than the dispatch grace period (one minute), rather
than refusing every new message on the instance; a younger one may still be
in flight. The unary SendMessage keeps returning the submitted task with a
runtime error, since a lost response does not mean the runtime did not take
the task.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
… release, v0.11.0-gs.4 (ca399a3)

The carried-commits table gains the row for fix(a2agateway) (#15; upstream branch upstream/a2agateway-dispatch-failure, giantswarm/giantswarm#37742 row 33); the Releases table records v0.11.0-gs.4 with the digests of run 34688265253.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
…chart at publish

The published chart's defaults named only the images' version; the two
pins a consumer needs beyond that — the index digest of the Go ADK
runtime image (a kagent.dev/v1alpha3 Harness accepts nothing but a
digest, and Helm cannot resolve one at render time) and the Substrate
worker image the chart's WorkerPool runs — were resolved by `record`
after the chart push, or known only to the Makefile, so the platform's
meta chart carried copies of both and re-pinned them by hand.

`push-helm-chart` runs after `push-images`, so the indexes exist: it now
resolves the golang-adk and claude-harness index digests the way `record`
does (`docker buildx imagetools inspect`) and stamps them into
`controller.agentImage.digest` and `runtimeImages.claudeHarness.digest`;
it stamps `substrateWorkerPool.workerImage` with the gVisor worker of the
Makefile's `SUBSTRATE_VERSION`/`SUBSTRATE_REPO` (`make -s substrate-pin`,
the same source ci.yaml reads for the e2e cluster) over the empty
default the template refuses. Every stamp is checked with yq; a missed
anchor fails the job as the existing checks do. `record` additionally
reads `helm show values` of the pushed chart and fails when its digests
disagree with the ones it writes to builds.md.

The anchors are the values keys of the two patches that follow this one
on the consumed branch (the optional default Harness, the runtime-images
ConfigMap); the workflow runs on the branch head only. FORK.md gets the
patch-table rows of all three.

Fork infrastructure — tag.yaml is fork-rewritten; never upstream.

Refs giantswarm/giantswarm#37765, giantswarm/giantswarm#37705.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
…ess for the chart's own runtime image

On the kagent API v2 an agent runs on a Harness (kagent.dev/v1alpha3): a
runtime image pinned by digest, an environment and a Substrate policy.
The chart ships the Go ADK runtime image (controller.agentImage) but
knows it by tag only, and renders no Harness for it — so every consumer
has to resolve the image's digest out of band and write its own Harness
before the first AgentTemplate can become Ready (kagent-dev#2529
is the missing build-time digest; kagent-dev#2536 proposed the
value this change uses).

controller.agentImage.digest carries the image's index digest (empty in
git; a published chart carries the digest of its build). A new `harness`
block renders one Harness — `kagent` by default, in the chart's
namespace, on the `kagent` runtime adapter — from that digest, only when
harness.create is set (default false, so nothing changes for a consumer
that renders its own Harness objects): workload.image is
<registry>/<controller.agentImage.repository>@<digest> unless
harness.image names another digest-pinned image; workerPoolRef defaults
to substrateWorkerPool.name; snapshotLocation is required (the CRD
requires the snapshot policy); env and allowedAgentTemplates.selector
are forwarded. A tag instead of a digest, a missing digest or a missing
snapshot location fail the render with the reason, instead of failing at
admission. The Harness keeps helm.sh/resource-policy: keep — it is the
runtime of every agent admitted under it.

The digest-pinned reference is built by one helper (kagent.runtimeImage)
so other templates can list the same image.

Unit tests cover the default (no Harness), the rendered object and its
image, the registry precedence, the harness.image override, the
forwarded policy and the three failures.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
…he Substrate image cache

Agent Substrate's atelet can pre-pull a pinned set of images on every
worker node (imageCache.pinnedImages), so an agent's first turn does not
wait for the pull of its runtime image. The set has to name the images
agents actually run on — this chart's runtime images at their digests —
and today only the operator of a cluster can write it, from a digest
resolved out of band, and re-write it on every upgrade.

The chart now publishes them itself: the ConfigMap `<fullname>-images`
carries one key, substrate-values.yaml, a values fragment for the
Substrate chart with atelet.imageCache.pinnedImages listing the image the
chart's Harness runs (harness.image, else the chart's own Go ADK image at
controller.agentImage.digest — so an override moves the set with it) and
the Claude runtime image at runtimeImages.claudeHarness.digest. An image
whose digest is empty is left out: a chart rendered from git lists
nothing, a published chart lists its build. The ConfigMap sits in the
release namespace, deliberately not the chart's namespaceOverride, so a
Substrate release installed next to this one can read it as values (a
Flux HelmRelease's valuesFrom) without crossing namespaces.

Unit tests cover the empty set, the two-image set and its name under a
fullnameOverride, the harness.image override and the namespace.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
SUBSTRATE_VERSION moves from the bootstrap dev build
0.0.27-dev.giantswarm.2026-09-10.19-33-37.h734ec53 to 0.0.27-gs.6, the
release of giantswarm/substrate the platform's meta chart runs today (its
worker image ghcr.io/giantswarm/substrate/ateom-gvisor:0.0.27-gs.6).
tag.yaml stamps substrateWorkerPool.workerImage from this pin, so a
release cut from the dev pin would install a dev worker as soon as the
meta chart stops forwarding the image by hand; the e2e cluster (ci.yaml)
runs the same version. Upstream pin unchanged: kagent-dev/substrate
v0.0.26 in go/go.mod, the ate-api contract.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
…Releases table (#19)

The Releases table records v0.11.0-gs.5 — the head after #17 (the publish-time stamps, controller.agentImage.digest with the optional default Harness, the runtime-images ConfigMap) and #18 (Substrate re-pinned to 0.0.27-gs.6) — with the digests of run 34696441320; the first published chart of the line that carries the golang-adk digest and the Substrate worker image of its build. Cut for giantswarm/giantswarm#37765.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
QuentinBisson and others added 15 commits October 1, 2026 17:45
… may reach

Signed-off-by: QuentinBisson <quentin@giantswarm.io>
Drop the normalisation the CRD pattern already refuses, and mark the
FORK.md row interim: upstream moves egress to Agent.spec.egress.
The runtime keys a conversation's session on x-user-id, and the gateway
forwarded the caller there. On a turn an AgentInstance share authorizes,
the caller is a visitor: the runtime found no session of the visitor's in
the instance's context and started an empty one, so the agent answered a
collaborator without the conversation so far.

Send the instance's creator as x-user-id when a share of the routed
instance authorizes the turn. x-kagent-user still names the visitor, so
model usage stays on the person of the turn.

Signed-off-by: Timo Derstappen <teemow@gmail.com>
Signed-off-by: Timo Derstappen <teemow@gmail.com>
)

* fix(a2agateway): run a turn a share of the instance authorizes in its owner's session

The runtime keys its session on x-user-id and the context id. The
runtime dialer forwarded the visitor as x-user-id, so a turn a share of
the instance authorized found no session under the visitor and started
an empty one in the same context: the collaborator never saw the
conversation's history. The dialer now names the share's owner, as the
HTTP A2A handler already does for a session share. The visitor stays
the person of the turn (x-kagent-user) and keeps their own credentials.

Refs: giantswarm/giantswarm#38089
Signed-off-by: Jose Armesto <github@armesto.net>

* docs(fork): ledger row for the share owner's session on the runtime dialer

Signed-off-by: Jose Armesto <github@armesto.net>

* docs(fork): name giantswarm/giantswarm#37742 as the upstream exit of the share owner's session

Refs: giantswarm/giantswarm#38089
Signed-off-by: Jose Armesto <github@armesto.net>

---------

Signed-off-by: Jose Armesto <github@armesto.net>
The Go ADK lists the tools of every MCP server at startup to classify its
MCP App tools. A runtime boot has no caller, so a server that takes the
caller's Authorization (KAGENT_PROPAGATE_TOKEN, or Authorization among
the allowed headers) can only refuse that list: it answers 401, the ADK
logs an error and falls back to lazy discovery at the first turn.

Skip the startup listing for such a server when the boot context yields
no Authorization for it, log the deferral at debug, and keep the lazy
discovery, which runs at the first turn with the caller's token. A
server with a static Authorization header is still classified at
startup.

Fixes #99

Signed-off-by: Timo Derstappen <teemow@gmail.com>
Signed-off-by: Timo Derstappen <teemow@gmail.com>
Signed-off-by: Timo Derstappen <teemow@gmail.com>
One Harness per entry with the policy shape of harness.* plus a runtime:
a claude entry defaults to the chart's claude-harness image at its
published digest, a kagent entry to the Go ADK image. Names must be distinct
from the chart's own Harness; the selector helper is shared.
CreateAgentInstanceShareRequest takes an optional ttl. The share records
expires_at (migration 000003, a nullable column) and token resolution
treats an expired share like an unknown token. The owner still lists an
expired share so it can be revoked.

Signed-off-by: QuentinBisson <quentin@giantswarm.io>
…ists

Signed-off-by: QuentinBisson <quentin@giantswarm.io>
Signed-off-by: QuentinBisson <quentin@giantswarm.io>
The controller sets expires_at from its own clock, and the token lookup now compares it with the controller's time instead of Postgres now(), so skew between the two cannot move a share's expiry. The ttl error says it must not be negative, since zero means no expiry.

Signed-off-by: QuentinBisson <quentin@giantswarm.io>
KAGENT_SHARE_MAX_TTL (controller.shareMaxTTL) bounds the ttl a share may
request; a share created without one receives the cap. Zero keeps shares
unbounded.
…ream does

The share payload no longer carries expires_at, so a column changed by a rollback or by hand cannot make a share undecodable, and token resolution compares the column with the database's now(). The ttl comment says an unset ttl takes the configured maximum. FORK.md records the upstream names of the cap and that a cap set later leaves earlier shares unbounded.

Signed-off-by: QuentinBisson <quentin@giantswarm.io>
@QuentinBisson

Copy link
Copy Markdown
Author

Parked: git as the person is postponed. Upstream Substrate is redesigning this area (actor JWT injection merged in agent-substrate/substrate#1960, token exchange of the actor JWT in agent-substrate/substrate#1661, the actor JWT contract in agent-substrate/substrate#1756), which likely supersedes this provider design, and this fork line has to re-pin past agent-substrate/substrate#1809 before it can follow. Kept as a draft so it can resume in place; tracking in giantswarm/giantswarm#38064.

QuentinBisson and others added 7 commits October 2, 2026 15:02
…am is cut

The run that dispatched a turn stopped following it when the stream to the
runtime failed or ended before the task quiesced, even when the runtime
still held the task, so the task reached its store only once a client
resubscribed or the next message reconciled it. A dispatching run whose
stream fails or ends early on a runtime that answers for the task now
resubscribes and keeps ingesting, at most three times in a row without a
new event, then fails with the stream's own error. A lost runtime and a
cancel still end the turn first, and a run that observes a task still ends
with its stream.

Signed-off-by: QuentinBisson <quentin@giantswarm.io>
…way sets the turn caller's token on a host

A caller credential names a host, the token broker's audience and the scheme
(Bearer, or Basic with a username for git). It compiles into the revision's
gateway bindings as ate-secret://kagent.dev/caller/<audience>/..., next to
the model, MCP and git-source ones, so the same conflict and passthrough
refusals hold, and its host joins the egress allowlist.

Signed-off-by: QuentinBisson <quentin@giantswarm.io>
…the turn's caller

The A2A gateway records the session of each turn it dispatches, from the
moment the turn reaches the runtime until its run stops following the task.
A caller credential provider on its own mTLS listener, accepting only the
egress gateway's SPIFFE ID, answers FetchSecret for the actor's running
turn: it exchanges the caller's bearer at the token broker (RFC 8693, HTTP
Basic client) once per turn and audience, and answers max_age 0. Outside a
turn, without a bearer, or for a caller the broker holds no grant for, the
gateway refuses the request.

Signed-off-by: QuentinBisson <quentin@giantswarm.io>
…credential provider

The controller gets a Substrate servicedns serving certificate, the
podidentity trust bundle the egress gateway's client certificate chains to,
the token broker's client secret, and a caller-credentials port on its
Service.

Signed-off-by: QuentinBisson <quentin@giantswarm.io>
The provider's refusals reach the agent as the egress gateway's 403 body, so
they name the remedy (no turn, no sign-in, an expired sign-in, no grant) and
leave the broker's error and URL to the log. Concurrent requests of a turn
share one exchange, which no longer holds the registry or a new turn on the
actor; a refusal is answered again for 30 s; nothing exchanged for a turn
that ended is answered. invalid_request is no caller refusal, and a token
issued without expires_in lasts a minute.

Signed-off-by: QuentinBisson <quentin@giantswarm.io>
A turn's caller is known only to the replica that dispatched it, and the
egress gateway reaches any replica behind the controller Service.

Signed-off-by: QuentinBisson <quentin@giantswarm.io>
Signed-off-by: QuentinBisson <quentin@giantswarm.io>
@teemow

teemow commented Oct 7, 2026

Copy link
Copy Markdown
Member

Heads-up from an agent working #178 (step 2, a draft PR follows): its Vertex access-token binding edits the same two files as commit 56870b9 here, in adjacent hunks — go/core/internal/egress/credentials.go (the authority check right after your caller-URI branch in CanonicalCredentials) and go/core/internal/translator/credentials.go (bind's URI and modelCredentialTarget, right after your caller loop). No semantic overlap; whichever lands second rebases.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants