Repository navigation
feat(chart): further credential providers for the egress gateway - #104
QuentinBisson wants to merge 143 commits into
Conversation
Build and publish versioned binaries, container images, and Helm charts from release tags. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Resolve atelet discovery and identity from the pod namespace, centralize install defaults, and allow explicitly selected local clusters to run without Pod Certificates. Keep authenticated transport as the default. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Add a configurable end-to-end workflow deadline and propagate it through lease acquisition. Apply released worker assignments to the cache immediately so subsequent scheduling sees the completed pause. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Parse PKCS1 RSA and SEC1 EC keys alongside PKCS8 keys, including regression coverage for RSA bundles. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Require the agentgateway E2E lane, reuse the installed control plane for microVM demos, wait for asset storage initialization, and accommodate runtime startup and counter persistence behavior in E2E checks. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Package the control plane, workers, PostgreSQL, RustFS, and CRDs as Helm charts. Keep manifests and generated RBAC aligned, add Helm E2E checks, and include current scheduling, sandbox permissions, and agentgateway configuration. Co-authored-by: Jet Chiang <jetjiang.ez@gmail.com> Co-authored-by: Keith Mattix II <keithmattix2@gmail.com> Signed-off-by: Jet Chiang <pokyuen.jetchiang-ext@solo.io> Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Allow an external PostgreSQL instance and a configurable schema, validate connection settings, and pass the schema to the API server. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Wire the API server snapshot backend and S3 settings to the chart storage configuration. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Configure trace, metric, and log endpoints independently, expose trace sampling, and route agentgateway access logs through the collector logs pipeline. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Isolate sandbox asset download tests from pause image pulls, explicitly advance the CA file timestamp, and disable VCS stamping for license checks in temporary verification worktrees. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Delete application containers before the pause container so their shared sandbox remains available throughout teardown. Cover the deletion order with a regression test. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Rebuild the fork on upstream while preserving features, require agentgateway runtime validation, and use a guarded push. Delete task-owned clusters and disposable assets before finishing while preserving shared resources and recovery data. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Select agentgateway expectations in the Helm test job, align the chart sandbox assets with the canonical manifest, and enable the CONNECT tunnel logging used by egress validation. This retains upstream gVisor checkpoint and restore fixes and closes configuration gaps between Helm and manifest installations. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Expose ateApi.extraArgs so installations can configure API flags such as the template resync interval without editing the deployment template. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
A layer pull can share a singleflight call with retirement and return without unpacking the removed layer. Distinguish pull results from retirement results and retry after retirement completes. Cover the interleaving with a deterministic concurrency test. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
The egress ext_proc server listens on loopback. Probe the metrics readiness endpoint so Kubernetes can observe readiness through the pod IP, matching the upstream manifests. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
global.imageRegistry redirects every image at once for air-gapped mirrors. Component images now resolve through the same registry/repository split every kagent-family chart uses: image.registry was one string carrying its path (ghcr.io/kagent-dev/substrate) and is now the registry host only, joined onto image.repository, so one global value (which overrides image.registry) redirects the whole family. This is a breaking change for values files that put a full prefix in image.registry: rendered silently they would produce a doubled prefix failing only at pod start, so the render fails instead, naming the split. A default render is byte-identical to main. Single-string images.* references (postgres, rustfs, aws-cli, agentgateway) have their registry segment replaced by the containerd rule (first path segment with a dot or colon), preserving repository paths either way. global.imagePullSecrets merges (union) into every pod spec, which previously had no pull-secret surface at all. global.imagePullPolicy replaces the hardcoded IfNotPresent values as a fallback, via substrate.imagePullPolicy. Verified: a default render is byte-identical to main; the mirror knob redirects all 9 images with paths preserved; the old-shape registry fails loudly at template time; pull secrets land on all 9 pod specs; the pullPolicy fallback fires. Signed-off-by: Jonathan Jamroga <jjamroga@gmail.com> Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
* Add Kubernetes credential sidecar for agentgateway egress Resolve ate-secret URIs from Kubernetes Secrets with injector-only mTLS and default-deny atespace-to-namespace grants. Deploy the provider alongside ateapi so it can reuse the existing Pod identity, certificate, and Service. Wire the provider into agentgateway's HTTPS interception route, with optional Helm and Kustomize configuration and namespace-scoped RBAC setup. Pin the nightly containing credential-provider configuration support; its RPC and URI scheme still need alignment before end-to-end injection works. Co-authored-by: Yufan Su <yufans@google.com> Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io> * Exercise credential injection through the rendered AGW route Add an opt-in Docker integration test covering actor authentication, TLS interception, provider mTLS, namespace denial, and cleartext denial using the Helm egress configuration. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io> * Pin AGW with the current credential-provider protocol Use the published multi-architecture build supporting FetchSecret and ate-secret URIs. Document the tested integration and keep Docker pull diagnostics out of the integration test container ID. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io> * Exercise Kubernetes credential injection through real actors in Helm E2E Validate exact header injection and namespace, RBAC, and cleartext denials against a local HTTPS origin. Remove fixture workers before their namespace to avoid delayed namespace finalization. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io> * Deploy the Kubernetes credential provider independently of ateapi Follow the upstream package and Deployment layout with a dedicated ServiceAccount and Service. Consolidate gateway integration coverage into the cluster E2E suite, including cache isolation across atespaces. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io> * Test credential injection through the installed HTTP route Install the provider's get-only Secret RBAC with the chart and standalone manifests. Enable injection on HTTP alongside HTTPS so the E2E can use the installed gateway configuration without modifying ConfigMaps or provisioning origin certificates. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io> * Always install credential injection and HTTPS interception Configure the provider and MITM gateway on every Helm install, and run credential and MITM E2E alongside the standard suites using the initial configuration. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io> --------- Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io> Co-authored-by: Yufan Su <yufans@google.com> Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Build the nested provider package with the release images and publish it under the kubernetes-secrets basename expected by the Helm chart. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
…ng to ghcr.io/giantswarm/substrate, upstream sync, ledger (#1) * ci(fork): run upstream's suites on the giantswarm branch and its pull requests The consumed branch of the Giant Swarm line is giantswarm, not main (main is the upstream mirror): pr-workflow and helm-e2e trigger on pushes to it, govulncheck on pushes to it, on pull requests and weekly on the default branch. Signed-off-by: Timo Derstappen <timo@giantswarm.io> * ci(fork): publish the line to ghcr.io/giantswarm/substrate and sync it with upstream publish.yaml: every push to giantswarm publishes the six images the agent platform runs (ateapi, atecontroller, atelet, atenet, podcertcontroller, ateom-gvisor; linux/amd64+arm64 with ko), a digest-true mirror of the agentgateway image the chart deploys, and both charts with their image defaults stamped to this registry, as dev versions <next upstream patch>-dev.giantswarm.<date>.<time>.h<sha7>; a v* tag publishes a release. Every own image is Trivy-scanned before the charts are pushed; .trivyignore holds time-boxed acceptances only. sync-upstream.yaml (weekly + dispatch, taylorbot token): force-mirrors upstream main (upstream rebases its main onto agent-substrate), probes whether the carried patches still rebase onto it, and on a pin input re-pins the branch — rebase, build, test, force-push — or opens a hand-over pull request naming the conflicting patch. Signed-off-by: Timo Derstappen <timo@giantswarm.io> * docs(fork): divergence ledger, CODEOWNERS and README pointer for the Giant Swarm line FORK.md records the pin (v0.0.26, kagent's go.mod replace), the carried patches with their upstream state, the re-pin procedure, what is published where and under which versions, the assets that are not images (gVisor runsc, micro-VM), the consumers, and how to contribute (upstream first, DCO). Team Bumblebee owns the fork (giantswarm/giantswarm#37757). Signed-off-by: Timo Derstappen <timo@giantswarm.io> * ci(fork): make the stamping guard fail under errexit A negated grep never trips set -e; the check that no upstream registry reference survives the stamping now fails the job explicitly. Signed-off-by: Timo Derstappen <timo@giantswarm.io> * docs(fork): CODEOWNERS at the root in the org's generated shape giantswarm/github's align-files writes a root CODEOWNERS for every repository of a team file and overwrites one it recognises as generated; keeping the same content here means the align run finds nothing to change. Signed-off-by: Timo Derstappen <timo@giantswarm.io> * ci(fork): license headers on the fork's own files hack/verify/boilerplate.sh (part of run-tests) requires the Apache header on every source file; the fork's files carry it as The Agent Substrate Authors. Signed-off-by: Timo Derstappen <timo@giantswarm.io> --------- Signed-off-by: Timo Derstappen <timo@giantswarm.io>
* docs(fork): the sync workflow's push needs a token with the workflow scope The first sync-upstream run was refused: a personal access token without the workflow scope cannot push commits that touch .github/workflows, which every mirror of upstream main and every re-pin does. Recorded in the re-pin procedure, together with the manual fallback and the hand-made bootstrap of the main mirror. Signed-off-by: Timo Derstappen <timo@giantswarm.io> * ci(fork): the sync workflow pushes as the HeraldBot GitHub App A classic personal access token without the workflow scope cannot push commits that touch .github/workflows, which every mirror of upstream main and every re-pin does (the first run was refused). The org's HeraldBot App is installed on every repository with contents and workflows write access; the workflow mints its token per run and both rulesets admit the App as a bypass actor. Signed-off-by: Timo Derstappen <timo@giantswarm.io> --------- Signed-off-by: Timo Derstappen <timo@giantswarm.io>
…over PRs in this repository (#3) A mirror push carries upstream's workflow files, whose main triggers run upstream's suites here for nothing; the sync now cancels them right after the push (actions: write). gh pr create inside a fork checkout targets the parent repository unless --repo names this one. FORK.md: bypass actors are the HeraldBot App (both branches) and repository admins (the line); cherry-picks of upstream commits are rebase-merged so a re-pin drops them by patch identity. Signed-off-by: Timo Derstappen <timo@giantswarm.io>
…he golden boot fetches before readyz (#4) * Let an actor's egress through while it resumes An actor's tunneled egress was armed only after every container served readyz, and atenet's egress gateway refused a CONNECT from any actor that was not RUNNING — a state the control plane commits only after the worker's Run returned, that is, after readyz. A workload that fetches what it needs to become ready — a model, a skill repository, packages — therefore could not: atunnel dropped the intercepted connection on the floor, the fetch died with "Broken pipe", readyz never came, and the actor was re-run every minute with one worker pinned to it. Nothing was logged at either hop; the workload's own error was the only trace. An ActorTemplate whose workload fetches at start-up never got its golden snapshot. The certificate the tunnel authenticates with is minted for exactly this placement before the workload starts (prepareActorEgress; the credential broker mints only for a worker's current assignment), so the identity is there from the first packet. Readiness gates ingress, not egress. * ateom (gVisor and micro-VM): arm egress right after the actor network is set up and the failure cleanup is registered, before the first container starts; ingress still waits for readyz. * atenet egress: a RESUMING actor is placed on a worker and may tunnel. SUSPENDED, PAUSED, CRASHED and DELETING are still refused: a certificate within its lifetime must not tunnel for an actor that has left its worker. The refusal is logged with the actor, its state and the destination. * atunnel: a dropped egress connection (no active egress, expired certificate) is logged with the peer address. * docs(fork): ledger row for the egress-while-resuming patch (#4)
… and the third gate in the egress dataplane (#6)
…ne of agentgateway (#5) * fork(publish): atenet-router and atenet-egress run the Giant Swarm line of agentgateway images.agentgateway is stamped to a release of giantswarm/agentgateway-upstream (AGENTGATEWAY_IMAGE, v1.5.1-gs.1 = upstream v1.5.0 rebuilt and scanned on the line) instead of a crane mirror of kagent-dev's build of unknown source; the image is recorded by digest with the own images and scanned report-only (its scan gates its own publish). FORK.md: the pin section names the agentgateway the chart runs and the convergence rule with the line. Signed-off-by: Timo Derstappen <teemow@gmail.com> * fork(publish): anchor the agentgateway stamp on the key, strip the tag at the last colon; FORK.md notes the two lines' tag conventions Signed-off-by: Timo Derstappen <teemow@gmail.com> --------- Signed-off-by: Timo Derstappen <teemow@gmail.com>
… authorization that admits a resuming actor (#7) v1.5.1-gs.1 was upstream v1.5.0 as built on the line: no CONNECT-time actor check at all. v1.5.1-gs.2 carries agentgateway#3237 (every egress CONNECT authorized against ate-api: UID, then state) plus the line's patch admitting a RESUMING actor, so the golden boot of a workload that fetches before readyz passes the third gate without giving up the authorization. FORK.md: the pin row and the egress-while-resuming row name the dataplane the line now runs.
…e — the line's dataplane (agentgateway#3237) reads it there (#9) * chart: authorize the egress actor as a frontend policy at CONNECT time agentgateway#3237 moved the actor check on an egress tunnel from a route policy on the inner HTTP listener to a frontend policy on the CONNECT itself (substrateEgress under frontendPolicies) and dropped the route-level field. A dataplane carrying it refuses the previous config ("unknown field `substrateEgress`") and never becomes ready. The egress config now declares the policy where that dataplane reads it, and images.agentgateway pins a build that carries #3237. The two move together: the previous build accepts only the route-level shape, and a build without #3237 knows neither the frontend field nor the check. Same change as kagent-dev#28 for a later dataplane, where the policy carries the name agentgateway#3318 gave it, substrateEgressActorResolution. * ci(fork): the publish refuses a drift between the chart's images.agentgateway and AGENTGATEWAY_IMAGE; ledger row for the frontend-policy egress config
Allow ate-api-server to read an external PostgreSQL connection string from a Secret while keeping schema configuration in the existing ConfigMap. Add Helm unit coverage for bundled, literal, Secret-backed, and invalid setups. (cherry picked from commit 1872249)
(cherry picked from commit 41097da)
…he 2026-09-30 audit (#95) Every carried row names its #37742 row and the upstream item that now covers, supersedes or overlaps it; wrong commit SHAs, closed pull requests still marked open and branches that no longer exist are corrected. Signed-off-by: QuentinBisson <quentin@giantswarm.io>
rootfs and upper were 0700. On gVisor the container's / takes that mode, so a non-root process could not search its own root. 0755 is what containerd and every image root use; the 0700 actor directory above keeps other actors out. The micro-VM guest already sees 0755. (cherry picked from commit 916f3c6)
… at 0700 A restore composes the rootfs of a bundle whose directories already exist, and MkdirAll leaves their mode alone, so only the explicit chmod opens the root. The test pre-creates them at 0700 and expects 0755.
…agent-substrate#1910) Follow-up to agent-substrate#1906. Removing a directory entry needs write permission on the directory that holds it, and with all capabilities dropped uid 0 gets no exemption. A non-root container can create a subdirectory in its `0777` durable dir; that subdirectory belongs to the container's uid, typically `0755`, so root cannot empty it. `resetActorDirs` fails after every checkpoint of such an actor and the suspend never completes. Chmod first, as the bundle dir does, is no way out: chmod needs ownership or `CAP_FOWNER`. Why atelet and not ateom, which already holds the capability: an ateom that dies before cleanup leaves the files behind and moves the hang to the next resume on that node. atelet owns the actor directories and needs it on the crash path either way. Manifest only, so no unit test. Exercised on kind with the micro-VM class: uid 65532 writes a durable dir, then suspend and resume. Without the capability the same run hangs in `SUSPENDING`. - [x] Tests pass - [x] Appropriate changes to documentation are included in the PR (none needed) (cherry picked from commit afb62e1)
Created 0700, a non-root container could not write its own volume. Now 0777, as Kubernetes creates an emptyDir: one actor's containers can run as different users, and the 0700 actor directory above keeps other actors out. A restore keeps the mode the snapshot recorded; volumes created earlier stay 0700, which is fine while every container runs as root. (cherry picked from commit 895ec3c)
atelet writes the OCI spec before ateom mounts the rootfs, so it cannot read the image's /etc/passwd from a mount. Image.ReadFile reads a file from the cached layers top-down, honoring recorded whiteouts and opaque dirs. Capped at 1 MiB. (cherry picked from commit 1e0466d)
Lets a test run it against a temp dir. No behavior change. (cherry picked from commit fa8578d)
Resolve USER as Docker does, against the image's own /etc/passwd and /etc/group: uid:gid as is, a bare uid with its login group (gid 0 without an entry), names looked up or Run/Restore fails, group memberships as supplementary groups. No USER means root. The pause container follows the same rule. On gVisor every snapshot taken before this needs re-taking, whatever the image: runsc restore checks Process.User, and the pause image sets USER 65535. (cherry picked from commit 5925173)
Start the process in the image's WORKDIR instead of "/", and create the directory in the actor's upper when the image lacks it. runsc restore checks Process.Cwd as well as Process.User; docs/upgrade.md says what an upgrade across this has to do. (cherry picked from commit d41a2fc)
credentialProvider.additionalProviders adds providers to the egress gateway's credentialProviders list beside the bundled kubernetes.io one, on both the HTTPS and the HTTP listener. The gateway dials each with its pod identity and verifies a servicedns.podcert.ate.dev serving certificate. An entry needs uriAuthority and host, cannot replace kubernetes.io and cannot name an authority twice. The default render is unchanged.
|
From Timo's agent: #106 adds the chart's otel env to the egress |
|
From Timo's agent: two more for the same issue, #107 (serverboot honors |
uriAuthority must be a lowercase DNS name and host must be <service>.<namespace>.svc:<port>, the name a servicedns serving certificate carries, so a value can neither inject YAML into the gateway config nor name a provider the gateway can never match or verify. The values and README say which client identity the provider must require.
|
Parked: git as the person is postponed. Upstream Substrate is redesigning this area (actor JWT injection merged in agent-substrate#1960, token exchange of the actor JWT in agent-substrate#1661, the actor JWT contract in agent-substrate#1756), which likely supersedes this provider design, and this fork line has to re-pin past agent-substrate#1809 before it can follow. Kept as a draft so it can resume in place; tracking in giantswarm/giantswarm#38064. |
|
sure |
90dc3ba to
e27ae95
Compare
|
Written by an agent: #145 (a carry of kagent-dev#47) also edits the |
Problem
The egress gateway resolves
ate-secret://URIs through a list of credential providers keyed by URI authority, but the chart renders only the bundledkubernetes.ioone. kagent needs a second provider that answers for the caller of the current turn, so the caller's bearer is injected at egress and never enters the actor (giantswarm/giantswarm#38054).Change
credentialProvider.additionalProvidersadds providers on both egress listeners.uriAuthoritymust be a lowercase DNS name andhostmust be<service>.<namespace>.svc:<port>, the name the provider'sservicedns.podcert.ate.devserving certificate carries; the gateway presents its pod identity, which the provider must require. The bundled provider cannot be replaced and an authority cannot appear twice. The default render is unchanged.