Skip to content

feat(chart): further credential providers for the egress gateway - #104

Draft
QuentinBisson wants to merge 143 commits into
giantswarmfrom
feat/egress-credential-providers
Draft

QuentinBisson wants to merge 143 commits into
giantswarmfrom
feat/egress-credential-providers

Conversation

@QuentinBisson

@QuentinBisson QuentinBisson commented Oct 1, 2026 •

Copy link
Copy Markdown

Problem

The egress gateway resolves ate-secret:// URIs through a list of credential providers keyed by URI authority, but the chart renders only the bundled kubernetes.io one. kagent needs a second provider that answers for the caller of the current turn, so the caller's bearer is injected at egress and never enters the actor (giantswarm/giantswarm#38054).

Change

credentialProvider.additionalProviders adds providers on both egress listeners. uriAuthority must be a lowercase DNS name and host must be <service>.<namespace>.svc:<port>, the name the provider's servicedns.podcert.ate.dev serving certificate carries; the gateway presents its pod identity, which the provider must require. The bundled provider cannot be replaced and an authority cannot appear twice. The default render is unchanged.

EItanya and others added 30 commits September 20, 2026 13:30
Build and publish versioned binaries, container images, and Helm charts from release tags.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Resolve atelet discovery and identity from the pod namespace, centralize install defaults, and allow explicitly selected local clusters to run without Pod Certificates. Keep authenticated transport as the default.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Add a configurable end-to-end workflow deadline and propagate it through lease acquisition. Apply released worker assignments to the cache immediately so subsequent scheduling sees the completed pause.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Parse PKCS1 RSA and SEC1 EC keys alongside PKCS8 keys, including regression coverage for RSA bundles.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Require the agentgateway E2E lane, reuse the installed control plane for microVM demos, wait for asset storage initialization, and accommodate runtime startup and counter persistence behavior in E2E checks.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Package the control plane, workers, PostgreSQL, RustFS, and CRDs as Helm charts. Keep manifests and generated RBAC aligned, add Helm E2E checks, and include current scheduling, sandbox permissions, and agentgateway configuration.

Co-authored-by: Jet Chiang <jetjiang.ez@gmail.com>
Co-authored-by: Keith Mattix II <keithmattix2@gmail.com>
Signed-off-by: Jet Chiang <pokyuen.jetchiang-ext@solo.io>
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Allow an external PostgreSQL instance and a configurable schema, validate connection settings, and pass the schema to the API server.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Wire the API server snapshot backend and S3 settings to the chart storage configuration.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Configure trace, metric, and log endpoints independently, expose trace sampling, and route agentgateway access logs through the collector logs pipeline.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Isolate sandbox asset download tests from pause image pulls, explicitly advance the CA file timestamp, and disable VCS stamping for license checks in temporary verification worktrees.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Delete application containers before the pause container so their shared sandbox remains available throughout teardown. Cover the deletion order with a regression test.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Rebuild the fork on upstream while preserving features, require agentgateway runtime validation, and use a guarded push. Delete task-owned clusters and disposable assets before finishing while preserving shared resources and recovery data.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Select agentgateway expectations in the Helm test job, align the chart sandbox assets with the canonical manifest, and enable the CONNECT tunnel logging used by egress validation. This retains upstream gVisor checkpoint and restore fixes and closes configuration gaps between Helm and manifest installations.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Expose ateApi.extraArgs so installations can configure API flags such as the template resync interval without editing the deployment template.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
A layer pull can share a singleflight call with retirement and return without unpacking the removed layer. Distinguish pull results from retirement results and retry after retirement completes. Cover the interleaving with a deterministic concurrency test.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
The egress ext_proc server listens on loopback. Probe the metrics readiness endpoint so Kubernetes can observe readiness through the pod IP, matching the upstream manifests.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
global.imageRegistry redirects every image at once for air-gapped
mirrors. Component images now resolve through the same
registry/repository split every kagent-family chart uses:
image.registry was one string carrying its path
(ghcr.io/kagent-dev/substrate) and is now the registry host only,
joined onto image.repository, so one global value (which overrides
image.registry) redirects the whole family. This is a breaking change
for values files that put a full prefix in image.registry: rendered
silently they would produce a doubled prefix failing only at pod start,
so the render fails instead, naming the split. A default render is
byte-identical to main.

Single-string images.* references (postgres, rustfs, aws-cli,
agentgateway) have their registry segment replaced by the containerd
rule (first path segment with a dot or colon), preserving repository
paths either way. global.imagePullSecrets merges (union) into every pod
spec, which previously had no pull-secret surface at all.
global.imagePullPolicy replaces the hardcoded IfNotPresent values as a
fallback, via substrate.imagePullPolicy.

Verified: a default render is byte-identical to main; the mirror knob
redirects all 9 images with paths preserved; the old-shape registry
fails loudly at template time; pull secrets land on all 9 pod specs;
the pullPolicy fallback fires.

Signed-off-by: Jonathan Jamroga <jjamroga@gmail.com>
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
* Add Kubernetes credential sidecar for agentgateway egress

Resolve ate-secret URIs from Kubernetes Secrets with injector-only mTLS and
default-deny atespace-to-namespace grants. Deploy the provider alongside
ateapi so it can reuse the existing Pod identity, certificate, and Service.

Wire the provider into agentgateway's HTTPS interception route, with
optional Helm and Kustomize configuration and namespace-scoped RBAC setup.
Pin the nightly containing credential-provider configuration support; its
RPC and URI scheme still need alignment before end-to-end injection works.

Co-authored-by: Yufan Su <yufans@google.com>
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>

* Exercise credential injection through the rendered AGW route

Add an opt-in Docker integration test covering actor authentication, TLS interception, provider mTLS, namespace denial, and cleartext denial using the Helm egress configuration.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>

* Pin AGW with the current credential-provider protocol

Use the published multi-architecture build supporting FetchSecret and ate-secret URIs. Document the tested integration and keep Docker pull diagnostics out of the integration test container ID.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>

* Exercise Kubernetes credential injection through real actors in Helm E2E

Validate exact header injection and namespace, RBAC, and cleartext denials against a local HTTPS origin. Remove fixture workers before their namespace to avoid delayed namespace finalization.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>

* Deploy the Kubernetes credential provider independently of ateapi

Follow the upstream package and Deployment layout with a dedicated ServiceAccount and Service. Consolidate gateway integration coverage into the cluster E2E suite, including cache isolation across atespaces.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>

* Test credential injection through the installed HTTP route

Install the provider's get-only Secret RBAC with the chart and standalone manifests. Enable injection on HTTP alongside HTTPS so the E2E can use the installed gateway configuration without modifying ConfigMaps or provisioning origin certificates.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>

* Always install credential injection and HTTPS interception

Configure the provider and MITM gateway on every Helm install, and run credential and MITM E2E alongside the standard suites using the initial configuration.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>

---------

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Co-authored-by: Yufan Su <yufans@google.com>
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Build the nested provider package with the release images and publish it under the kubernetes-secrets basename expected by the Helm chart.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
…ng to ghcr.io/giantswarm/substrate, upstream sync, ledger (#1)

* ci(fork): run upstream's suites on the giantswarm branch and its pull requests

The consumed branch of the Giant Swarm line is giantswarm, not main (main is
the upstream mirror): pr-workflow and helm-e2e trigger on pushes to it,
govulncheck on pushes to it, on pull requests and weekly on the default branch.

Signed-off-by: Timo Derstappen <timo@giantswarm.io>

* ci(fork): publish the line to ghcr.io/giantswarm/substrate and sync it with upstream

publish.yaml: every push to giantswarm publishes the six images the agent
platform runs (ateapi, atecontroller, atelet, atenet, podcertcontroller,
ateom-gvisor; linux/amd64+arm64 with ko), a digest-true mirror of the
agentgateway image the chart deploys, and both charts with their image
defaults stamped to this registry, as dev versions
<next upstream patch>-dev.giantswarm.<date>.<time>.h<sha7>; a v* tag publishes
a release. Every own image is Trivy-scanned before the charts are pushed;
.trivyignore holds time-boxed acceptances only.

sync-upstream.yaml (weekly + dispatch, taylorbot token): force-mirrors upstream
main (upstream rebases its main onto agent-substrate), probes whether the
carried patches still rebase onto it, and on a pin input re-pins the branch —
rebase, build, test, force-push — or opens a hand-over pull request naming the
conflicting patch.

Signed-off-by: Timo Derstappen <timo@giantswarm.io>

* docs(fork): divergence ledger, CODEOWNERS and README pointer for the Giant Swarm line

FORK.md records the pin (v0.0.26, kagent's go.mod replace), the carried
patches with their upstream state, the re-pin procedure, what is published
where and under which versions, the assets that are not images (gVisor runsc,
micro-VM), the consumers, and how to contribute (upstream first, DCO).
Team Bumblebee owns the fork (giantswarm/giantswarm#37757).

Signed-off-by: Timo Derstappen <timo@giantswarm.io>

* ci(fork): make the stamping guard fail under errexit

A negated grep never trips set -e; the check that no upstream registry
reference survives the stamping now fails the job explicitly.

Signed-off-by: Timo Derstappen <timo@giantswarm.io>

* docs(fork): CODEOWNERS at the root in the org's generated shape

giantswarm/github's align-files writes a root CODEOWNERS for every repository
of a team file and overwrites one it recognises as generated; keeping the
same content here means the align run finds nothing to change.

Signed-off-by: Timo Derstappen <timo@giantswarm.io>

* ci(fork): license headers on the fork's own files

hack/verify/boilerplate.sh (part of run-tests) requires the Apache header on
every source file; the fork's files carry it as The Agent Substrate Authors.

Signed-off-by: Timo Derstappen <timo@giantswarm.io>

---------

Signed-off-by: Timo Derstappen <timo@giantswarm.io>
* docs(fork): the sync workflow's push needs a token with the workflow scope

The first sync-upstream run was refused: a personal access token without the
workflow scope cannot push commits that touch .github/workflows, which every
mirror of upstream main and every re-pin does. Recorded in the re-pin
procedure, together with the manual fallback and the hand-made bootstrap of
the main mirror.

Signed-off-by: Timo Derstappen <timo@giantswarm.io>

* ci(fork): the sync workflow pushes as the HeraldBot GitHub App

A classic personal access token without the workflow scope cannot push commits
that touch .github/workflows, which every mirror of upstream main and every
re-pin does (the first run was refused). The org's HeraldBot App is installed
on every repository with contents and workflows write access; the workflow
mints its token per run and both rulesets admit the App as a bypass actor.

Signed-off-by: Timo Derstappen <timo@giantswarm.io>

---------

Signed-off-by: Timo Derstappen <timo@giantswarm.io>
…over PRs in this repository (#3)

A mirror push carries upstream's workflow files, whose main triggers run
upstream's suites here for nothing; the sync now cancels them right after the
push (actions: write). gh pr create inside a fork checkout targets the parent
repository unless --repo names this one. FORK.md: bypass actors are the
HeraldBot App (both branches) and repository admins (the line); cherry-picks
of upstream commits are rebase-merged so a re-pin drops them by patch identity.

Signed-off-by: Timo Derstappen <timo@giantswarm.io>
…he golden boot fetches before readyz (#4)

* Let an actor's egress through while it resumes

An actor's tunneled egress was armed only after every container served
readyz, and atenet's egress gateway refused a CONNECT from any actor that
was not RUNNING — a state the control plane commits only after the
worker's Run returned, that is, after readyz. A workload that fetches
what it needs to become ready — a model, a skill repository, packages —
therefore could not: atunnel dropped the intercepted connection on the
floor, the fetch died with "Broken pipe", readyz never came, and the
actor was re-run every minute with one worker pinned to it. Nothing was
logged at either hop; the workload's own error was the only trace. An
ActorTemplate whose workload fetches at start-up never got its golden
snapshot.

The certificate the tunnel authenticates with is minted for exactly this
placement before the workload starts (prepareActorEgress; the credential
broker mints only for a worker's current assignment), so the identity is
there from the first packet. Readiness gates ingress, not egress.

  * ateom (gVisor and micro-VM): arm egress right after the actor
    network is set up and the failure cleanup is registered, before the
    first container starts; ingress still waits for readyz.
  * atenet egress: a RESUMING actor is placed on a worker and may
    tunnel. SUSPENDED, PAUSED, CRASHED and DELETING are still refused: a
    certificate within its lifetime must not tunnel for an actor that
    has left its worker. The refusal is logged with the actor, its
    state and the destination.
  * atunnel: a dropped egress connection (no active egress, expired
    certificate) is logged with the peer address.

* docs(fork): ledger row for the egress-while-resuming patch (#4)
…ne of agentgateway (#5)

* fork(publish): atenet-router and atenet-egress run the Giant Swarm line of agentgateway

images.agentgateway is stamped to a release of giantswarm/agentgateway-upstream
(AGENTGATEWAY_IMAGE, v1.5.1-gs.1 = upstream v1.5.0 rebuilt and scanned on the
line) instead of a crane mirror of kagent-dev's build of unknown source; the
image is recorded by digest with the own images and scanned report-only (its
scan gates its own publish). FORK.md: the pin section names the agentgateway
the chart runs and the convergence rule with the line.

Signed-off-by: Timo Derstappen <teemow@gmail.com>

* fork(publish): anchor the agentgateway stamp on the key, strip the tag at the last colon; FORK.md notes the two lines' tag conventions

Signed-off-by: Timo Derstappen <teemow@gmail.com>

---------

Signed-off-by: Timo Derstappen <teemow@gmail.com>
… authorization that admits a resuming actor (#7)

v1.5.1-gs.1 was upstream v1.5.0 as built on the line: no CONNECT-time actor check at all.
v1.5.1-gs.2 carries agentgateway#3237 (every egress CONNECT authorized against ate-api: UID,
then state) plus the line's patch admitting a RESUMING actor, so the golden boot of a workload
that fetches before readyz passes the third gate without giving up the authorization.
FORK.md: the pin row and the egress-while-resuming row name the dataplane the line now runs.
…e — the line's dataplane (agentgateway#3237) reads it there (#9)

* chart: authorize the egress actor as a frontend policy at CONNECT time

agentgateway#3237 moved the actor check on an egress tunnel from a route
policy on the inner HTTP listener to a frontend policy on the CONNECT
itself (substrateEgress under frontendPolicies) and dropped the
route-level field. A dataplane carrying it refuses the previous config
("unknown field `substrateEgress`") and never becomes ready.

The egress config now declares the policy where that dataplane reads it,
and images.agentgateway pins a build that carries #3237. The two move
together: the previous build accepts only the route-level shape, and a
build without #3237 knows neither the frontend field nor the check.

Same change as kagent-dev#28 for a later dataplane, where the
policy carries the name agentgateway#3318 gave it,
substrateEgressActorResolution.

* ci(fork): the publish refuses a drift between the chart's images.agentgateway and AGENTGATEWAY_IMAGE; ledger row for the frontend-policy egress config
Allow ate-api-server to read an external PostgreSQL connection string from a Secret while keeping schema configuration in the existing ConfigMap. Add Helm unit coverage for bundled, literal, Secret-backed, and invalid setups.

(cherry picked from commit 1872249)
(cherry picked from commit 41097da)
QuentinBisson and others added 15 commits October 1, 2026 10:14
…he 2026-09-30 audit (#95)

Every carried row names its #37742 row and the upstream item that now covers,
supersedes or overlaps it; wrong commit SHAs, closed pull requests still
marked open and branches that no longer exist are corrected.

Signed-off-by: QuentinBisson <quentin@giantswarm.io>
rootfs and upper were 0700. On gVisor the container's / takes that mode,
so a non-root process could not search its own root. 0755 is what
containerd and every image root use; the 0700 actor directory above
keeps other actors out. The micro-VM guest already sees 0755.

(cherry picked from commit 916f3c6)
… at 0700

A restore composes the rootfs of a bundle whose directories already exist,
and MkdirAll leaves their mode alone, so only the explicit chmod opens the
root. The test pre-creates them at 0700 and expects 0755.
…agent-substrate#1910)

Follow-up to agent-substrate#1906. Removing a directory entry needs write permission on
the directory that holds it, and with all capabilities dropped uid 0
gets no exemption. A non-root container can create a subdirectory in its
`0777` durable dir; that subdirectory belongs to the container's uid,
typically `0755`, so root cannot empty it. `resetActorDirs` fails after
every checkpoint of such an actor and the suspend never completes. Chmod
first, as the bundle dir does, is no way out: chmod needs ownership or
`CAP_FOWNER`.

Why atelet and not ateom, which already holds the capability: an ateom
that dies before cleanup leaves the files behind and moves the hang to
the next resume on that node. atelet owns the actor directories and
needs it on the crash path either way.

Manifest only, so no unit test. Exercised on kind with the micro-VM
class: uid 65532 writes a durable dir, then suspend and resume. Without
the capability the same run hangs in `SUSPENDING`.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR (none
needed)

(cherry picked from commit afb62e1)
Created 0700, a non-root container could not write its own volume. Now
0777, as Kubernetes creates an emptyDir: one actor's containers can run
as different users, and the 0700 actor directory above keeps other
actors out. A restore keeps the mode the snapshot recorded; volumes
created earlier stay 0700, which is fine while every container runs as
root.

(cherry picked from commit 895ec3c)
atelet writes the OCI spec before ateom mounts the rootfs, so it cannot
read the image's /etc/passwd from a mount. Image.ReadFile reads a file
from the cached layers top-down, honoring recorded whiteouts and opaque
dirs. Capped at 1 MiB.

(cherry picked from commit 1e0466d)
Lets a test run it against a temp dir. No behavior change.

(cherry picked from commit fa8578d)
Resolve USER as Docker does, against the image's own /etc/passwd and
/etc/group: uid:gid as is, a bare uid with its login group (gid 0 without
an entry), names looked up or Run/Restore fails, group memberships as
supplementary groups. No USER means root. The pause container follows the
same rule.

On gVisor every snapshot taken before this needs re-taking, whatever the
image: runsc restore checks Process.User, and the pause image sets
USER 65535.

(cherry picked from commit 5925173)
Start the process in the image's WORKDIR instead of "/", and create the
directory in the actor's upper when the image lacks it. runsc restore
checks Process.Cwd as well as Process.User; docs/upgrade.md says what an
upgrade across this has to do.

(cherry picked from commit d41a2fc)
credentialProvider.additionalProviders adds providers to the egress
gateway's credentialProviders list beside the bundled kubernetes.io one,
on both the HTTPS and the HTTP listener. The gateway dials each with its
pod identity and verifies a servicedns.podcert.ate.dev serving
certificate. An entry needs uriAuthority and host, cannot replace
kubernetes.io and cannot name an authority twice. The default render is
unchanged.
@teemow

teemow commented Oct 1, 2026

Copy link
Copy Markdown
Member

From Timo's agent: #106 adds the chart's otel env to the egress ext-proc container in atenet-egress.yaml (the container's env, not your ConfigMap hunks) and one FORK.md carried-patch row. Is it OK to merge ahead of this PR?

@teemow

teemow commented Oct 1, 2026

Copy link
Copy Markdown
Member

From Timo's agent: two more for the same issue, #107 (serverboot honors OTEL_TRACES_EXPORTER=none) and #108 (ate-controller turns ateom's exporters off without an endpoint). Each adds one FORK.md carried-patch row at the end of the table and leaves your files otherwise alone. OK to merge them ahead of this PR as well?

uriAuthority must be a lowercase DNS name and host must be
<service>.<namespace>.svc:<port>, the name a servicedns serving
certificate carries, so a value can neither inject YAML into the gateway
config nor name a provider the gateway can never match or verify. The
values and README say which client identity the provider must require.
@teemow

teemow commented Oct 2, 2026

Copy link
Copy Markdown
Member

Question from Team Bumblebee's agents (written by an agent on Timo's behalf): may the agent merge #106, #107 and #108 (the lab OTLP endpoint for agent-platform#463; green, they touch FORK.md and #106 touches atenet-egress.yaml, files of this draft), or do you want them rebased after this PR lands?

@QuentinBisson

Copy link
Copy Markdown
Author

Parked: git as the person is postponed. Upstream Substrate is redesigning this area (actor JWT injection merged in agent-substrate#1960, token exchange of the actor JWT in agent-substrate#1661, the actor JWT contract in agent-substrate#1756), which likely supersedes this provider design, and this fork line has to re-pin past agent-substrate#1809 before it can follow. Kept as a draft so it can resume in place; tracking in giantswarm/giantswarm#38064.

@teemow

teemow commented Oct 2, 2026

Copy link
Copy Markdown
Member

From Timo's agent: while this draft is parked, is it OK for you if #106, #107 and #108 land first on the same files (FORK.md rows, the egress ext-proc env, the otel comment in values.yaml and the README row), and you rebase when you resume?

@QuentinBisson

Copy link
Copy Markdown
Author

sure

@teemow

teemow commented Oct 6, 2026

Copy link
Copy Markdown
Member

Written by an agent: #145 (a carry of kagent-dev#47) also edits the credentialProviders block in atenet-egress.yaml. It turns the bundled k8s.io entry into a range over k8s.io and google-access-token.k8s.io and touches nothing of this PR. May #145 land first, with whichever lands second taking the small rebase, or would you rather fold it into this PR?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants