Repository navigation
feat(api): Harness caller credentials the egress gateway sets as the turn's caller - #147
QuentinBisson wants to merge 177 commits into
Conversation
…y=disabled without listing its tools A RemoteMCPServer that authenticates every caller has no credential to offer the controller: agents forward the caller token at run time (KAGENT_PROPAGATE_TOKEN), but the controller's own discovery request carries no bearer, the server answers 401 and the CR stays Accepted=False (DiscoveryFailed: Unauthorized) although every agent that references it works. spec.headersFrom is not a way out: its values are resolved into each agent's config and applied last by the runtime, so a controller credential there would replace the caller token in every agent. The kagent.dev/discovery=disabled label now opts a RemoteMCPServer out of tool discovery: the reconciler skips the listing, reports Accepted=True with reason DiscoveryDisabled, clears status.discoveredTools and keeps the server in the catalog with no tools and no LastConnected. The controller also reacts to label changes, so the opt-out and its removal apply at once rather than on the next periodic refresh. Signed-off-by: Timo Derstappen <teemow@gmail.com> (cherry picked from commit 7c0e532)
…f the catalog main honors the label nowhere: kagent-dev#2565 removed predicates.DiscoveryDisabledPredicate together with the old MCPServer tool controller, and the catalog reconciler restored in kagent-dev#2644 watches every kmcp MCPServer unfiltered, so the opt-out the docs describe for the agentgateway setup was a no-op on main. The MCPServer catalog reconciler now skips discovery for a labeled server and drops its catalog projection. That is what the 0.10.x predicate achieved, minus two stale-row cases: a label added after discovery left the row in place, and the deletion of a labeled server never removed it. The controller filters no events, so a label change re-enters Reconcile at once. A RemoteMCPServer stays in the catalog with the DiscoveryDisabled status on purpose: it is kagent's own CR that agents reference. An MCPServer is kmcp's resource that kagent only projects, has no kagent status to explain an empty row, and AgentTemplates cannot reference it directly. Rebased onto main after kagent-dev#2774 (simplify database client operations): the reconciler's deleteCatalog is gone, the catalog store's DeleteToolServer with the MCPServer group kind is what the not-found path calls now, so the opt-out path calls the same. Signed-off-by: Timo Derstappen <teemow@gmail.com> (cherry picked from commit bf0b26e)
…n label helm-controller installs OCI charts under <version>+<digest>; '+' is not a valid label value character, so every object of the chart failed to apply under Flux. Sanitised the way helm.sh/chart already is.
…oc branch The fork's tag.yaml publishes the dev builds of the agent-platform's kagent component: GitHub-hosted runners (the org has no self-hosted runners), the registries derived from the repository owner (ghcr.io/<owner>/kagent/<image>, oci://ghcr.io/<owner>/kagent/helm/<chart>), no PyPI/GitHub-release jobs, the image matrix trimmed to what the chart references plus the Claude harness (multi-arch kept), and a push to poc/agent-platform that derives the version as <base>-dev.<branch>.<committer date>.<time>.h<sha7> — the schema giantswarm's gitsemver emits, so Flux selects the channel with a semverFilter on the branch part. workflow_dispatch keeps its optional version input. The chart's defaults are stamped with the published version and image repositories before packaging: the tag must be explicit because the chart's fallback, .Chart.Version, carries `+<digest>` under helm-controller. The Makefile's -X targets followed $(DOCKER_REPO); with an owner-relative repository they missed the module, leaving the binaries without version information. They now follow the Go module path from go.mod.
…failure apk.cgr.dev intermittently answers some of the parallel anonymous package fetches from hosted runners with HTTP 403, which apk reports as "Permission denied" for a random package and which fails the wolfi-based ui image while the alpine-based images pass. The builder keeps the finished stages, so the step now retries the build up to three times, repeating only the failed step; and one image's failure no longer cancels the other three.
The controller and Go ADK images are alpine:3.22 plus `apk add`; the base layer's libcrypto3/libssl3 3.5.7-r0 carry CVE-2026-14456 (HIGH, fixed in 3.5.8-r0). `apk upgrade --no-cache` before the add brings the runtime layer to the current Alpine packages, as kagent-dev#2756 did for the harness images.
…w, keep a ledger The consumed branch was built and published without a test or a scan: upstream's ci.yaml and image-scan.yaml trigger on main and release/** only, and their jobs ask for Blacksmith runners the org does not have. Both now run on every push to poc/agent-platform and sync/**, on every pull request to poc/agent-platform and (the scan) weekly, on GitHub-hosted runners with a docker/setup-buildx-action builder, with the image matrices trimmed to what the fork publishes (controller, ui, golang-adk, claude-harness — the Python ADK, the Codex harness and the CLI image are not offered on the platform), a govulncheck job over the Go module graph, no paths-ignore (a docs-only pull request must still report the required checks) and one aggregate status per workflow — ci-ok, scan-ok — for the ruleset to require. sync-upstream.yaml re-pins the line weekly and on demand: fast-forward the mirror main to upstream main with the workflow token (no workflow reacts, so upstream's CI never runs on the mirror), rebase the carried commits onto the upstream head (scripts/fork/repin.sh, which also names commits upstream merged in the meantime and drops them), push the candidate with the bot token so the checks run, wait for ci-ok and scan-ok, then force-push the consumed branch with --force-with-lease so tag.yaml publishes the dev build. A conflicting carried commit or a red check stops the run and opens a pull request against the consumed branch that names the cause; nothing else is pushed. tag.yaml gains a record job: every published build's version, commit, upstream pin and image and chart digests go into the run summary and into builds.md on the ledger branch (scripts/fork/ledger-append.sh; re-pins get the same on re-pins.md). FORK.md is the human half of that ledger — the pin, every carried commit with its purpose and upstream state, what is published and how, the automated and the manual re-pin, protection, the contribution process — and the README opens with it.
The first sync run was refused at the mirror: "refusing to allow a GitHub App to create or update workflow `.github/workflows/ci.yaml` without `workflows` permission" — the workflow's own token may not update a workflow file, not even in a mirrored upstream commit, and no `permissions:` grant changes that. The mirror is therefore pushed with the heraldbot App's token like the candidate and the consumed branch. An App push starts upstream's workflow files at the mirrored refs (runners the fork does not have, registries it must not touch), so mirror.sh records what it pushed and mirror-quiet.sh cancels those runs right after; the ledger stays on GITHUB_TOKEN.
…arent The second sync run reached its "checks red" step and failed there: `pull request create failed: GraphQL: Resource not accessible by integration (createPullRequest)`. In a checkout of a fork, gh resolves the base repository to the parent (kagent-dev/kagent) unless --repo is given, and the App is not installed there. Both `gh pr create` calls now name this repository. git-push-as.sh emits its `::add-mask::` line only under GitHub Actions: run by hand (the manual re-pin) it printed the encoded credential to the terminal instead of masking it.
…warm line SUBSTRATE_REPO defaults to oci://ghcr.io/giantswarm/substrate/helm and SUBSTRATE_VERSION to 0.0.27-dev.giantswarm.2026-09-10.19-33-37.h734ec53: the substrate and substrate-crds subcharts of kagent and kagent-crds are the fork giantswarm/substrate's builds (upstream v0.0.26 — the version go/go.mod pins — plus cherry-picked fixes). The Go module pin stays upstream's release, so the version is set here rather than derived from go.mod; both move together at a re-pin. Signed-off-by: Timo Derstappen <timo@giantswarm.io>
The Makefile's SUBSTRATE_VERSION / SUBSTRATE_REPO are the one pin: a substrate-pin target prints them, the e2e job reads them into its environment and installs the line's substrate-crds and substrate charts and its ateom-gvisor worker image (the fork's tags carry no v). kubectl-ate for the CA/JWT bootstrap stays upstream's v0.0.26 binary — the line publishes no CLI. Before, the job exported SUBSTRATE_VERSION=0.0.26, which also drove make helm-version and made the fork registry look for a chart it never published. Signed-off-by: Timo Derstappen <teemow@gmail.com>
The consumed branch was renamed from poc/agent-platform to giantswarm (the GitHub rename retargets the default branch, the ~DEFAULT_BRANCH ruleset and open pull requests); every reference follows: the push and pull-request triggers of ci.yaml, image-scan.yaml and tag.yaml, CONSUMED in sync-upstream.yaml and scripts/fork/repin.sh, FORK.md and the README. The dev version's branch part becomes 'giantswarm' — 0.11.0-dev.giantswarm.<date>.<time>.h<sha7> — and the consumers' semverFilter moves with it (.*-dev\.giantswarm\..*).
FORK.md 'Release scheme': a release is a tag vX.Y.Z-gs.N on the consumed branch — X.Y.Z the upstream version the pin anticipates (DEV_BASE_VERSION, 0.11.0), N the line's counter from 1; first release v0.11.0-gs.1. Sorted 0.11.0-dev.giantswarm.… < 0.11.0-gs.1 < 0.11.0, so a release outranks its dev builds and never the upstream release it anticipates, and the switch to an upstream tag is a range change. Published by tag.yaml to ghcr.io/giantswarm/kagent (images + helm/ charts), digests recorded by the record job; ghcr.io rather than gsoci/the catalog because GitHub Actions has no ACR credential (the org publishes to gsoci from CircleCI) and the distinct OCI path plus the pre-release version each keep the chart out of the 0.10 wrapper's fleet range >=0.2.0 <1.0.0 at gsoci.azurecr.io/charts/giantswarm/kagent. tag.yaml's version step refuses any tag that is not v<DEV_BASE_VERSION>-gs.<N> — a bare vX.Y.Z is a name upstream will use and the mirror would then skip upstream's tag of that name. A 'Releases' table (release, tag, commit, pin, digests) is filled per tag after the agentlab proof of its dev build.
…cator The chart renders AUTH_MODE and AUTH_USER_ID_CLAIM on the controller (controller.auth.mode, controller.auth.userIdClaim) and docs/OIDC_PROXY_AUTH_ARCHITECTURE.md describes the two modes, but since the legacy runtime went (kagent-dev#2595) nothing reads either variable: the controller calls app.Run with empty Options, so it always runs the UnsecureAuthenticator -- any caller names themselves with X-User-Id -- and ProxyAuthenticator has no caller at all. A deployment that sets controller.auth.mode: trusted-proxy behind oauth2-proxy or an API gateway gets a controller that admits everyone, and nothing tells it so. app.AuthenticatorFromEnv selects the built-in authenticator AUTH_MODE names and the controller's main passes it as Options.Authenticator: - unset or "unsecure": UnsecureAuthenticator, as today; - "trusted-proxy": ProxyAuthenticator reading the claim AUTH_USER_ID_CLAIM names ("sub" by default) from the bearer the proxy validated; a request without a bearer is refused whatever X-User-Id says, and the agent-call branch (X-Agent-Name set) is unchanged; - anything else fails startup naming the valid modes, instead of silently running unsecure. The helper is exported so that a library consumer, which cannot import the internal auth package, can run core's modes too; Options and Run are untouched, so a consumer that supplies its own authenticator sees no change. gRPC callers are covered as before: the authentication interceptor already maps the authorization, x-user-id and x-agent-name metadata onto the headers the authenticator reads. The docs table drops the --auth-mode flags, which went with the legacy runtime.
…olicy: keep) The kagent-crds chart ships its CustomResourceDefinitions as plain templates (there is no crds/ directory), so `helm uninstall kagent-crds` deletes the five CRDs and, with them, every AgentTemplate, Harness, ModelConfig, ModelProviderConfig and RemoteMCPServer in the cluster. CRD charts conventionally mark their CRDs `helm.sh/resource-policy: keep` so that removing or replacing the release leaves the CRDs (and the objects behind them) in place; a CRD is then removed deliberately, with `kubectl delete crd`, never as a side effect of a release operation. The annotation is added where the manifests are produced rather than in the chart files: a `+kubebuilder:metadata:annotations` marker on each of the five v1alpha3 types, so controller-gen emits it into go/api/config/crd/bases and `make controller-manifests` copies it into helm/kagent-crds/templates unchanged. The chart templates stay a verbatim copy of the generated CRDs, the manifests-check CI job keeps passing, and no new tool (yq, sed post-processing) enters the generation. The generated bases carry the annotation too: nothing in the repository applies them directly (the e2e installs the chart), and to `kubectl apply` it is an inert annotation. Verified: `make controller-manifests` regenerates exactly the five annotation lines in bases and templates; `helm lint helm/kagent-crds` passes with and without `substrate.enabled`; `helm template` differs from the previous render only in the five added lines, and the CRDs of the kmcp-crds and substrate-crds subcharts are untouched.
…Releases table Images and charts published to ghcr.io/giantswarm/kagent by run 34542852237; the digests are the index digests a Harness pins. The release images are rebuilt from the tree of the dev build h0ac5240, so their digests differ from that build's; consumers of the release pin these.
…ources A skill or plugin source on a private repository has no credential path: the materialiser shells out to git with the process environment and no helper, so an AgentTemplate cannot take skills from a private repository at all. AgentTemplate.spec.skills[].source.git and plugins[].source.git gain an optional credentialRef, a key in a same-namespace Secret (https only, enforced by CEL). Git speaks Basic authentication, so the key holds the base64 of "<username>:<token>" for the URL's host; GitHub takes "x-access-token:<token>". The credential follows the gateway injection path every other credential takes. CompileCredentials binds each referenced key to the source URL's host as "Authorization: Basic <value>", an ate-secret:// URI the egress gateway resolves on every request to that host; the actor never holds the token, and the materialiser clones with no credential of its own. The compiler still records one Secret-backed environment variable per referenced key, KAGENT_ARTIFACT_CREDENTIAL_<hash of namespace, name and key>, which reaches the actor as the inert placeholder like a model key, so a missing Secret or key fails compilation as a reference failure (ResolvedRefs=False) that is retried when the Secret appears. A rotation is not a new revision: the gateway reads the Secret live. Substrate matches credentials by hostname, so one credential serves every source of an agent tree on that host, public repositories included, and two different Secrets for one host are refused at compile time. A wrong token fails the golden boot on the host's 401.
…Releases table Signed-off-by: Timo Derstappen <teemow@gmail.com>
…ead of leaving it submitted A turn the gateway could not deliver to the AgentInstance runtime left its task TASK_STATE_SUBMITTED for good: SendStreamingMessage stored the task in prepareSend, then a failing dial, a failing ingester start, or a runtime stream that errored before its first event only returned the error. The instance kept the task as its active one, every later message was refused as a conflict, and every client saw a turn that never ends. A dispatch that fails after the task was stored now records the task as TASK_STATE_FAILED with the cause as its status message, publishes that status update to the caller before the stream ends with the error, and, being a terminal transition, releases the instance's active task. The runtime holds nothing for such a task, so no snapshot is taken: the failure is recorded even when the runtime is the unreachable part. A stream that fails before its first event may also have lost only the response, so a dispatch run first asks the runtime for the task: one that answers with it is left to finish it, and the existing recovery finds the task there; one that does not know the task, or cannot be asked, never started the turn. A run started by a subscription never fails the task, since it may be racing the dispatch about to deliver it. reconcileActiveTask treats a submitted task the runtime does not know as failed once it is older than the dispatch grace period (one minute), rather than refusing every new message on the instance; a younger one may still be in flight. The unary SendMessage keeps returning the submitted task with a runtime error, since a lost response does not mean the runtime did not take the task. Signed-off-by: Timo Derstappen <teemow@gmail.com>
… release, v0.11.0-gs.4 (ca399a3) The carried-commits table gains the row for fix(a2agateway) (#15; upstream branch upstream/a2agateway-dispatch-failure, giantswarm/giantswarm#37742 row 33); the Releases table records v0.11.0-gs.4 with the digests of run 34688265253. Signed-off-by: Timo Derstappen <teemow@gmail.com>
…chart at publish The published chart's defaults named only the images' version; the two pins a consumer needs beyond that — the index digest of the Go ADK runtime image (a kagent.dev/v1alpha3 Harness accepts nothing but a digest, and Helm cannot resolve one at render time) and the Substrate worker image the chart's WorkerPool runs — were resolved by `record` after the chart push, or known only to the Makefile, so the platform's meta chart carried copies of both and re-pinned them by hand. `push-helm-chart` runs after `push-images`, so the indexes exist: it now resolves the golang-adk and claude-harness index digests the way `record` does (`docker buildx imagetools inspect`) and stamps them into `controller.agentImage.digest` and `runtimeImages.claudeHarness.digest`; it stamps `substrateWorkerPool.workerImage` with the gVisor worker of the Makefile's `SUBSTRATE_VERSION`/`SUBSTRATE_REPO` (`make -s substrate-pin`, the same source ci.yaml reads for the e2e cluster) over the empty default the template refuses. Every stamp is checked with yq; a missed anchor fails the job as the existing checks do. `record` additionally reads `helm show values` of the pushed chart and fails when its digests disagree with the ones it writes to builds.md. The anchors are the values keys of the two patches that follow this one on the consumed branch (the optional default Harness, the runtime-images ConfigMap); the workflow runs on the branch head only. FORK.md gets the patch-table rows of all three. Fork infrastructure — tag.yaml is fork-rewritten; never upstream. Refs giantswarm/giantswarm#37765, giantswarm/giantswarm#37705. Signed-off-by: Timo Derstappen <teemow@gmail.com>
…ess for the chart's own runtime image On the kagent API v2 an agent runs on a Harness (kagent.dev/v1alpha3): a runtime image pinned by digest, an environment and a Substrate policy. The chart ships the Go ADK runtime image (controller.agentImage) but knows it by tag only, and renders no Harness for it — so every consumer has to resolve the image's digest out of band and write its own Harness before the first AgentTemplate can become Ready (kagent-dev#2529 is the missing build-time digest; kagent-dev#2536 proposed the value this change uses). controller.agentImage.digest carries the image's index digest (empty in git; a published chart carries the digest of its build). A new `harness` block renders one Harness — `kagent` by default, in the chart's namespace, on the `kagent` runtime adapter — from that digest, only when harness.create is set (default false, so nothing changes for a consumer that renders its own Harness objects): workload.image is <registry>/<controller.agentImage.repository>@<digest> unless harness.image names another digest-pinned image; workerPoolRef defaults to substrateWorkerPool.name; snapshotLocation is required (the CRD requires the snapshot policy); env and allowedAgentTemplates.selector are forwarded. A tag instead of a digest, a missing digest or a missing snapshot location fail the render with the reason, instead of failing at admission. The Harness keeps helm.sh/resource-policy: keep — it is the runtime of every agent admitted under it. The digest-pinned reference is built by one helper (kagent.runtimeImage) so other templates can list the same image. Unit tests cover the default (no Harness), the rendered object and its image, the registry precedence, the harness.image override, the forwarded policy and the three failures. Signed-off-by: Timo Derstappen <teemow@gmail.com>
…he Substrate image cache Agent Substrate's atelet can pre-pull a pinned set of images on every worker node (imageCache.pinnedImages), so an agent's first turn does not wait for the pull of its runtime image. The set has to name the images agents actually run on — this chart's runtime images at their digests — and today only the operator of a cluster can write it, from a digest resolved out of band, and re-write it on every upgrade. The chart now publishes them itself: the ConfigMap `<fullname>-images` carries one key, substrate-values.yaml, a values fragment for the Substrate chart with atelet.imageCache.pinnedImages listing the image the chart's Harness runs (harness.image, else the chart's own Go ADK image at controller.agentImage.digest — so an override moves the set with it) and the Claude runtime image at runtimeImages.claudeHarness.digest. An image whose digest is empty is left out: a chart rendered from git lists nothing, a published chart lists its build. The ConfigMap sits in the release namespace, deliberately not the chart's namespaceOverride, so a Substrate release installed next to this one can read it as values (a Flux HelmRelease's valuesFrom) without crossing namespaces. Unit tests cover the empty set, the two-image set and its name under a fullnameOverride, the harness.image override and the namespace. Signed-off-by: Timo Derstappen <teemow@gmail.com>
SUBSTRATE_VERSION moves from the bootstrap dev build 0.0.27-dev.giantswarm.2026-09-10.19-33-37.h734ec53 to 0.0.27-gs.6, the release of giantswarm/substrate the platform's meta chart runs today (its worker image ghcr.io/giantswarm/substrate/ateom-gvisor:0.0.27-gs.6). tag.yaml stamps substrateWorkerPool.workerImage from this pin, so a release cut from the dev pin would install a dev worker as soon as the meta chart stops forwarding the image by hand; the e2e cluster (ci.yaml) runs the same version. Upstream pin unchanged: kagent-dev/substrate v0.0.26 in go/go.mod, the ate-api contract. Signed-off-by: Timo Derstappen <teemow@gmail.com>
…Releases table (#19) The Releases table records v0.11.0-gs.5 — the head after #17 (the publish-time stamps, controller.agentImage.digest with the optional default Harness, the runtime-images ConfigMap) and #18 (Substrate re-pinned to 0.0.27-gs.6) — with the digests of run 34696441320; the first published chart of the line that carries the golang-adk digest and the Substrate worker image of its build. Cut for giantswarm/giantswarm#37765. Signed-off-by: Timo Derstappen <teemow@gmail.com>
… may reach Signed-off-by: QuentinBisson <quentin@giantswarm.io>
Drop the normalisation the CRD pattern already refuses, and mark the FORK.md row interim: upstream moves egress to Agent.spec.egress.
The runtime keys a conversation's session on x-user-id, and the gateway forwarded the caller there. On a turn an AgentInstance share authorizes, the caller is a visitor: the runtime found no session of the visitor's in the instance's context and started an empty one, so the agent answered a collaborator without the conversation so far. Send the instance's creator as x-user-id when a share of the routed instance authorizes the turn. x-kagent-user still names the visitor, so model usage stays on the person of the turn. Signed-off-by: Timo Derstappen <teemow@gmail.com>
Signed-off-by: Timo Derstappen <teemow@gmail.com>
) * fix(a2agateway): run a turn a share of the instance authorizes in its owner's session The runtime keys its session on x-user-id and the context id. The runtime dialer forwarded the visitor as x-user-id, so a turn a share of the instance authorized found no session under the visitor and started an empty one in the same context: the collaborator never saw the conversation's history. The dialer now names the share's owner, as the HTTP A2A handler already does for a session share. The visitor stays the person of the turn (x-kagent-user) and keeps their own credentials. Refs: giantswarm/giantswarm#38089 Signed-off-by: Jose Armesto <github@armesto.net> * docs(fork): ledger row for the share owner's session on the runtime dialer Signed-off-by: Jose Armesto <github@armesto.net> * docs(fork): name giantswarm/giantswarm#37742 as the upstream exit of the share owner's session Refs: giantswarm/giantswarm#38089 Signed-off-by: Jose Armesto <github@armesto.net> --------- Signed-off-by: Jose Armesto <github@armesto.net>
The Go ADK lists the tools of every MCP server at startup to classify its MCP App tools. A runtime boot has no caller, so a server that takes the caller's Authorization (KAGENT_PROPAGATE_TOKEN, or Authorization among the allowed headers) can only refuse that list: it answers 401, the ADK logs an error and falls back to lazy discovery at the first turn. Skip the startup listing for such a server when the boot context yields no Authorization for it, log the deferral at debug, and keep the lazy discovery, which runs at the first turn with the caller's token. A server with a static Authorization header is still classified at startup. Fixes #99 Signed-off-by: Timo Derstappen <teemow@gmail.com>
Signed-off-by: Timo Derstappen <teemow@gmail.com>
Signed-off-by: Timo Derstappen <teemow@gmail.com>
One Harness per entry with the policy shape of harness.* plus a runtime: a claude entry defaults to the chart's claude-harness image at its published digest, a kagent entry to the Go ADK image. Names must be distinct from the chart's own Harness; the selector helper is shared.
CreateAgentInstanceShareRequest takes an optional ttl. The share records expires_at (migration 000003, a nullable column) and token resolution treats an expired share like an unknown token. The owner still lists an expired share so it can be revoked. Signed-off-by: QuentinBisson <quentin@giantswarm.io>
…ists Signed-off-by: QuentinBisson <quentin@giantswarm.io>
Signed-off-by: QuentinBisson <quentin@giantswarm.io>
The controller sets expires_at from its own clock, and the token lookup now compares it with the controller's time instead of Postgres now(), so skew between the two cannot move a share's expiry. The ttl error says it must not be negative, since zero means no expiry. Signed-off-by: QuentinBisson <quentin@giantswarm.io>
KAGENT_SHARE_MAX_TTL (controller.shareMaxTTL) bounds the ttl a share may request; a share created without one receives the cap. Zero keeps shares unbounded.
…ream does The share payload no longer carries expires_at, so a column changed by a rollback or by hand cannot make a share undecodable, and token resolution compares the column with the database's now(). The ttl comment says an unset ttl takes the configured maximum. FORK.md records the upstream names of the cap and that a cap set later leaves earlier shares unbounded. Signed-off-by: QuentinBisson <quentin@giantswarm.io>
|
Parked: git as the person is postponed. Upstream Substrate is redesigning this area (actor JWT injection merged in agent-substrate/substrate#1960, token exchange of the actor JWT in agent-substrate/substrate#1661, the actor JWT contract in agent-substrate/substrate#1756), which likely supersedes this provider design, and this fork line has to re-pin past agent-substrate/substrate#1809 before it can follow. Kept as a draft so it can resume in place; tracking in giantswarm/giantswarm#38064. |
…am is cut The run that dispatched a turn stopped following it when the stream to the runtime failed or ended before the task quiesced, even when the runtime still held the task, so the task reached its store only once a client resubscribed or the next message reconciled it. A dispatching run whose stream fails or ends early on a runtime that answers for the task now resubscribes and keeps ingesting, at most three times in a row without a new event, then fails with the stream's own error. A lost runtime and a cancel still end the turn first, and a run that observes a task still ends with its stream. Signed-off-by: QuentinBisson <quentin@giantswarm.io>
…way sets the turn caller's token on a host A caller credential names a host, the token broker's audience and the scheme (Bearer, or Basic with a username for git). It compiles into the revision's gateway bindings as ate-secret://kagent.dev/caller/<audience>/..., next to the model, MCP and git-source ones, so the same conflict and passthrough refusals hold, and its host joins the egress allowlist. Signed-off-by: QuentinBisson <quentin@giantswarm.io>
…the turn's caller The A2A gateway records the session of each turn it dispatches, from the moment the turn reaches the runtime until its run stops following the task. A caller credential provider on its own mTLS listener, accepting only the egress gateway's SPIFFE ID, answers FetchSecret for the actor's running turn: it exchanges the caller's bearer at the token broker (RFC 8693, HTTP Basic client) once per turn and audience, and answers max_age 0. Outside a turn, without a bearer, or for a caller the broker holds no grant for, the gateway refuses the request. Signed-off-by: QuentinBisson <quentin@giantswarm.io>
…credential provider The controller gets a Substrate servicedns serving certificate, the podidentity trust bundle the egress gateway's client certificate chains to, the token broker's client secret, and a caller-credentials port on its Service. Signed-off-by: QuentinBisson <quentin@giantswarm.io>
The provider's refusals reach the agent as the egress gateway's 403 body, so they name the remedy (no turn, no sign-in, an expired sign-in, no grant) and leave the broker's error and URL to the log. Concurrent requests of a turn share one exchange, which no longer holds the registry or a new turn on the actor; a refusal is answered again for 30 s; nothing exchanged for a turn that ended is answered. invalid_request is no caller refusal, and a token issued without expires_in lasts a minute. Signed-off-by: QuentinBisson <quentin@giantswarm.io>
A turn's caller is known only to the replica that dispatched it, and the egress gateway reaches any replica behind the controller Service. Signed-off-by: QuentinBisson <quentin@giantswarm.io>
Signed-off-by: QuentinBisson <quentin@giantswarm.io>
b985f86 to
8a8fdf3
Compare
a5186dd to
f7bf3da
Compare
|
Heads-up from an agent working #178 (step 2, a draft PR follows): its Vertex access-token binding edits the same two files as commit 56870b9 here, in adjacent hunks — |
Problem
A coding agent must
git cloneand call the GitHub API as the person who sent the turn, without a GitHub token, or the caller's bearer, in the sandbox. Caller routes (#126) do it for git only, through a loopback forwarder, an env var and an agentgateway route.Change
Harness.spec.substrate.callerCredentials[](hostname, broker audience,BearerorBasicwith a username) compiles to anate-secret://kagent.dev/caller/...gateway binding, so every client in the sandbox gets the person's token from Substrate's egress gateway.kagent.devcredential provider (mTLS, the gateway's SPIFFE ID only) for an actor whose dispatched turn is running. It exchanges the caller's bearer at the token broker (RFC 8693) once per turn and audience, shared by concurrent requests, and answersmax_age: 0(feat(substrate): a credential provider bounds how long the egress gateway reuses its answer agentgateway-upstream#37, feat(credprovider): a provider bounds how long the injector reuses its answer substrate#110). Refusals name the remedy (no turn, no sign-in, expired sign-in, no grant) and nothing of the platform, because they reach the agent as the gateway's 403 body (fix(substrate): answer a credential provider's failure by its class agentgateway-upstream#39, feat(atenet): tell the actor why a credential was refused substrate#112).controller.substrate.callerCredentialswith more than one controller replica. The egress gateway's provider entry is feat(chart): further credential providers for the egress gateway substrate#104.Not proven in agentlab yet. Supersedes #126; towards giantswarm/giantswarm#38056, sign-in link follow-up giantswarm/giantswarm#38107.