feat(helm): add PodMonitors for the workloads that serve metrics - #46
Open
QuentinBisson wants to merge 21 commits into
Open
QuentinBisson wants to merge 21 commits into
QuentinBisson wants to merge 21 commits into
Conversation
Build and publish versioned binaries, container images, and Helm charts from release tags. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Resolve atelet discovery and identity from the pod namespace, centralize install defaults, and allow explicitly selected local clusters to run without Pod Certificates. Keep authenticated transport as the default. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Add a configurable end-to-end workflow deadline and propagate it through lease acquisition. Apply released worker assignments to the cache immediately so subsequent scheduling sees the completed pause. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Parse PKCS1 RSA and SEC1 EC keys alongside PKCS8 keys, including regression coverage for RSA bundles. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Require the agentgateway E2E lane, reuse the installed control plane for microVM demos, wait for asset storage initialization, and accommodate runtime startup and counter persistence behavior in E2E checks. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Package the control plane, workers, PostgreSQL, RustFS, and CRDs as Helm charts. Keep manifests and generated RBAC aligned, add Helm E2E checks, and include current scheduling, sandbox permissions, and agentgateway configuration. Co-authored-by: Jet Chiang <jetjiang.ez@gmail.com> Co-authored-by: Keith Mattix II <keithmattix2@gmail.com> Signed-off-by: Jet Chiang <pokyuen.jetchiang-ext@solo.io> Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Allow an external PostgreSQL instance and a configurable schema, validate connection settings, and pass the schema to the API server. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Wire the API server snapshot backend and S3 settings to the chart storage configuration. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Configure trace, metric, and log endpoints independently, expose trace sampling, and route agentgateway access logs through the collector logs pipeline. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Isolate sandbox asset download tests from pause image pulls, explicitly advance the CA file timestamp, and disable VCS stamping for license checks in temporary verification worktrees. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Delete application containers before the pause container so their shared sandbox remains available throughout teardown. Cover the deletion order with a regression test. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Rebuild the fork on upstream while preserving features, require agentgateway runtime validation, and use a guarded push. Delete task-owned clusters and disposable assets before finishing while preserving shared resources and recovery data. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Select agentgateway expectations in the Helm test job, align the chart sandbox assets with the canonical manifest, and enable the CONNECT tunnel logging used by egress validation. This retains upstream gVisor checkpoint and restore fixes and closes configuration gaps between Helm and manifest installations. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Expose ateApi.extraArgs so installations can configure API flags such as the template resync interval without editing the deployment template. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
A layer pull can share a singleflight call with retirement and return without unpacking the removed layer. Distinguish pull results from retirement results and retry after retirement completes. Cover the interleaving with a deterministic concurrency test. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
The egress ext_proc server listens on loopback. Probe the metrics readiness endpoint so Kubernetes can observe readiness through the pod IP, matching the upstream manifests. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
global.imageRegistry redirects every image at once for air-gapped mirrors. Component images now resolve through the same registry/repository split every kagent-family chart uses: image.registry was one string carrying its path (ghcr.io/kagent-dev/substrate) and is now the registry host only, joined onto image.repository, so one global value (which overrides image.registry) redirects the whole family. This is a breaking change for values files that put a full prefix in image.registry: rendered silently they would produce a doubled prefix failing only at pod start, so the render fails instead, naming the split. A default render is byte-identical to main. Single-string images.* references (postgres, rustfs, aws-cli, agentgateway) have their registry segment replaced by the containerd rule (first path segment with a dot or colon), preserving repository paths either way. global.imagePullSecrets merges (union) into every pod spec, which previously had no pull-secret surface at all. global.imagePullPolicy replaces the hardcoded IfNotPresent values as a fallback, via substrate.imagePullPolicy. Verified: a default render is byte-identical to main; the mirror knob redirects all 9 images with paths preserved; the old-shape registry fails loudly at template time; pull secrets land on all 9 pod specs; the pullPolicy fallback fires. Signed-off-by: Jonathan Jamroga <jjamroga@gmail.com> Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Retain upstream secret URI syntax and client CA rotation while adding Helm deployment, configurable injector identity, strict default-deny authorization, and credential injection coverage. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Build the nested provider package with the release images and publish it under the kubernetes-secrets basename expected by the Helm chart. Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Six workloads serve a Prometheus endpoint on a pod port and nothing collects any of them. ate-api-server, atelet, the atenet router, the atenet egress and the k8s credential provider start the metrics server in serverboot; ate-controller serves controller-runtime's own on :8080, its default, since its ctrl.Options sets no Metrics field. The router and the egress each add a second port for their agentgateway container. Three of those pods carry prometheus.io/* annotations, which a Prometheus reads only when it is configured for pod-annotation discovery. A Prometheus Operator installation discovers PodMonitors instead and ignores the annotations, so on such an installation every one of these endpoints is served and none is collected. Add one PodMonitor per workload, off by default because the object needs the monitoring.coreos.com/v1 CRDs. The port name and the pod label are pinned per workload: no two metrics sources agree on the port name, and the credential provider's label is the substrate.fullname helper's output rather than the bare component name. A wrong name or label yields a monitor that matches nothing and reports no error. podcertificate-controller gets no monitor. It emits no metrics, which docs/metrics/substrate.yaml records under blind_spots. Run the new chart unit tests in helm-verify, which had no way to run them. Signed-off-by: Quentin Bisson <quentin@giantswarm.io>
QuentinBisson
force-pushed
the
feat/chart-podmonitors
branch
from
September 22, 2026 18:22
b38511f to
81c23d9
Compare
Author
|
@EItanya I had to force push |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Six workloads in the chart serve a Prometheus endpoint on a pod port, and nothing collects any of them.
ate-api-server,atelet, the atenet router, the atenet egress and the k8s credential provider start the metrics server themselves inserverboot.ate-controllerserves controller-runtime's own metrics server on:8080, its default, because itsctrl.Optionssets noMetricsfield and onlyBindAddress: "0"disables that server. The router and the egress each add a second port for their agentgateway container's stats.Three of those pods carry
prometheus.io/*annotations. A Prometheus reads those only when it is configured for pod-annotation discovery. A Prometheus Operator installation discoversPodMonitorobjects instead and ignores the annotations, so on such an installation every one of these endpoints is served and none is collected. That is every instrument indocs/metrics/registry/metrics.yaml, and thecontroller_runtime_*andworkqueue_*families thatdocs/metrics/substrate.yamllists underbridged_metric_families.Change
One
PodMonitorper workload that serves metrics, behindmetrics.podMonitor.enabled, off by default because the object does not apply on a cluster without the Prometheus Operator CRDs.intervalandlabelsare configurable, so an installation can set whatever itspodMonitorSelectormatches on.The port name and the pod label are pinned per workload, because no two metrics sources agree on the port name and the credential provider's pod label is the
substrate.fullnamehelper's output rather than the bare component name. A wrong name or label yields a monitor that matches nothing and reports no error, so the unit tests pin both per workload.podcertificate-controllergets no monitor. It emits no metrics, whichdocs/metrics/substrate.yamlrecords underblind_spots.helm-verifygains ahelm-unitteststep. The repo had no way to run chart unit tests, so the new suite would not have run in CI.The committed manifests under
manifests/ate-install/are unchanged, since the template renders nothing at the default values.make verify-helm-template,helm lintand the new suite all pass locally.