Skip to content

feat(helm): add PodMonitors for the workloads that serve metrics - #46

Open
QuentinBisson wants to merge 21 commits into
kagent-dev:mainfrom
QuentinBisson:feat/chart-podmonitors
Open

QuentinBisson wants to merge 21 commits into
kagent-dev:mainfrom
QuentinBisson:feat/chart-podmonitors

Conversation

@QuentinBisson

Copy link
Copy Markdown

Problem

Six workloads in the chart serve a Prometheus endpoint on a pod port, and nothing collects any of them.

ate-api-server, atelet, the atenet router, the atenet egress and the k8s credential provider start the metrics server themselves in serverboot. ate-controller serves controller-runtime's own metrics server on :8080, its default, because its ctrl.Options sets no Metrics field and only BindAddress: "0" disables that server. The router and the egress each add a second port for their agentgateway container's stats.

Three of those pods carry prometheus.io/* annotations. A Prometheus reads those only when it is configured for pod-annotation discovery. A Prometheus Operator installation discovers PodMonitor objects instead and ignores the annotations, so on such an installation every one of these endpoints is served and none is collected. That is every instrument in docs/metrics/registry/metrics.yaml, and the controller_runtime_* and workqueue_* families that docs/metrics/substrate.yaml lists under bridged_metric_families.

Change

One PodMonitor per workload that serves metrics, behind metrics.podMonitor.enabled, off by default because the object does not apply on a cluster without the Prometheus Operator CRDs. interval and labels are configurable, so an installation can set whatever its podMonitorSelector matches on.

The port name and the pod label are pinned per workload, because no two metrics sources agree on the port name and the credential provider's pod label is the substrate.fullname helper's output rather than the bare component name. A wrong name or label yields a monitor that matches nothing and reports no error, so the unit tests pin both per workload.

podcertificate-controller gets no monitor. It emits no metrics, which docs/metrics/substrate.yaml records under blind_spots.

helm-verify gains a helm-unittest step. The repo had no way to run chart unit tests, so the new suite would not have run in CI.

The committed manifests under manifests/ate-install/ are unchanged, since the template renders nothing at the default values. make verify-helm-template, helm lint and the new suite all pass locally.

EItanya and others added 20 commits September 22, 2026 14:21
Build and publish versioned binaries, container images, and Helm charts from release tags.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Resolve atelet discovery and identity from the pod namespace, centralize install defaults, and allow explicitly selected local clusters to run without Pod Certificates. Keep authenticated transport as the default.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Add a configurable end-to-end workflow deadline and propagate it through lease acquisition. Apply released worker assignments to the cache immediately so subsequent scheduling sees the completed pause.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Parse PKCS1 RSA and SEC1 EC keys alongside PKCS8 keys, including regression coverage for RSA bundles.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Require the agentgateway E2E lane, reuse the installed control plane for microVM demos, wait for asset storage initialization, and accommodate runtime startup and counter persistence behavior in E2E checks.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Package the control plane, workers, PostgreSQL, RustFS, and CRDs as Helm charts. Keep manifests and generated RBAC aligned, add Helm E2E checks, and include current scheduling, sandbox permissions, and agentgateway configuration.

Co-authored-by: Jet Chiang <jetjiang.ez@gmail.com>
Co-authored-by: Keith Mattix II <keithmattix2@gmail.com>
Signed-off-by: Jet Chiang <pokyuen.jetchiang-ext@solo.io>
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Allow an external PostgreSQL instance and a configurable schema, validate connection settings, and pass the schema to the API server.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Wire the API server snapshot backend and S3 settings to the chart storage configuration.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Configure trace, metric, and log endpoints independently, expose trace sampling, and route agentgateway access logs through the collector logs pipeline.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Isolate sandbox asset download tests from pause image pulls, explicitly advance the CA file timestamp, and disable VCS stamping for license checks in temporary verification worktrees.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Delete application containers before the pause container so their shared sandbox remains available throughout teardown. Cover the deletion order with a regression test.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Rebuild the fork on upstream while preserving features, require agentgateway runtime validation, and use a guarded push. Delete task-owned clusters and disposable assets before finishing while preserving shared resources and recovery data.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Select agentgateway expectations in the Helm test job, align the chart sandbox assets with the canonical manifest, and enable the CONNECT tunnel logging used by egress validation. This retains upstream gVisor checkpoint and restore fixes and closes configuration gaps between Helm and manifest installations.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Expose ateApi.extraArgs so installations can configure API flags such as the template resync interval without editing the deployment template.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
A layer pull can share a singleflight call with retirement and return without unpacking the removed layer. Distinguish pull results from retirement results and retry after retirement completes. Cover the interleaving with a deterministic concurrency test.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
The egress ext_proc server listens on loopback. Probe the metrics readiness endpoint so Kubernetes can observe readiness through the pod IP, matching the upstream manifests.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
global.imageRegistry redirects every image at once for air-gapped
mirrors. Component images now resolve through the same
registry/repository split every kagent-family chart uses:
image.registry was one string carrying its path
(ghcr.io/kagent-dev/substrate) and is now the registry host only,
joined onto image.repository, so one global value (which overrides
image.registry) redirects the whole family. This is a breaking change
for values files that put a full prefix in image.registry: rendered
silently they would produce a doubled prefix failing only at pod start,
so the render fails instead, naming the split. A default render is
byte-identical to main.

Single-string images.* references (postgres, rustfs, aws-cli,
agentgateway) have their registry segment replaced by the containerd
rule (first path segment with a dot or colon), preserving repository
paths either way. global.imagePullSecrets merges (union) into every pod
spec, which previously had no pull-secret surface at all.
global.imagePullPolicy replaces the hardcoded IfNotPresent values as a
fallback, via substrate.imagePullPolicy.

Verified: a default render is byte-identical to main; the mirror knob
redirects all 9 images with paths preserved; the old-shape registry
fails loudly at template time; pull secrets land on all 9 pod specs;
the pullPolicy fallback fires.

Signed-off-by: Jonathan Jamroga <jjamroga@gmail.com>
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Retain upstream secret URI syntax and client CA rotation while adding Helm deployment, configurable injector identity, strict default-deny authorization, and credential injection coverage.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Build the nested provider package with the release images and publish it under the kubernetes-secrets basename expected by the Helm chart.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Six workloads serve a Prometheus endpoint on a pod port and nothing
collects any of them. ate-api-server, atelet, the atenet router, the
atenet egress and the k8s credential provider start the metrics server
in serverboot; ate-controller serves controller-runtime's own on :8080,
its default, since its ctrl.Options sets no Metrics field. The router
and the egress each add a second port for their agentgateway container.

Three of those pods carry prometheus.io/* annotations, which a Prometheus
reads only when it is configured for pod-annotation discovery. A
Prometheus Operator installation discovers PodMonitors instead and
ignores the annotations, so on such an installation every one of these
endpoints is served and none is collected.

Add one PodMonitor per workload, off by default because the object needs
the monitoring.coreos.com/v1 CRDs. The port name and the pod label are
pinned per workload: no two metrics sources agree on the port name, and
the credential provider's label is the substrate.fullname helper's output
rather than the bare component name. A wrong name or label yields a
monitor that matches nothing and reports no error.

podcertificate-controller gets no monitor. It emits no metrics, which
docs/metrics/substrate.yaml records under blind_spots.

Run the new chart unit tests in helm-verify, which had no way to run them.

Signed-off-by: Quentin Bisson <quentin@giantswarm.io>
@QuentinBisson

Copy link
Copy Markdown
Author

@EItanya I had to force push

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants