Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 11 additions & 2 deletions apps/monitoring/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,8 +14,17 @@ existing **Amazon Managed Service for Prometheus (AMP)** workspace.
IAM role managed through Crossplane, scoped to writing metrics to the selected
AMP workspace.

Collected metrics cover node readiness, workload replicas, pod status and
restarts, container CPU and memory usage, and collector health.
Collected metrics cover node readiness and pressure conditions, node pool
labels, allocatable capacity, workload replicas and rollout progress, HPA
state, pod phase, restarts, waiting and exit reasons, container CPU and memory
usage against requests and limits, CPU throttling, OOM kills, kubelet pod
counts and evictions, and collector health. The node root cgroup is kept from
cAdvisor so whole-node usage is available without node-exporter.

The Alloy keep-list and the kube-state-metrics allowlist are the contract with
the Grafana dashboards in the `infrastructure` repo
(`observability/dashboards/grafana/src/dashboards/kubernetes/`). Add a metric to
both when a panel needs it; drop it from both when nothing reads it.
Comment on lines +24 to +27

## Configuration

Expand Down
2 changes: 1 addition & 1 deletion apps/monitoring/chart/Chart.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@ apiVersion: v2
name: monitoring
description: Kubernetes metrics sent to the existing Amazon Managed Service for Prometheus workspace
type: application
version: 0.1.2
version: 0.1.3
dependencies:
- name: alloy
repository: https://grafana.github.io/helm-charts
Expand Down
60 changes: 54 additions & 6 deletions apps/monitoring/chart/files/config.alloy
Original file line number Diff line number Diff line change
Expand Up @@ -74,16 +74,62 @@ prometheus.scrape "alloy" {
}

prometheus.relabel "v1" {
// Only the metrics consumed by the overview and rollout checks reach AMP.
// Only metrics the Kubernetes dashboards and alerts read reach AMP
// (infrastructure: observability/dashboards/grafana/src/dashboards/kubernetes/).
// kube_* names must also be on the kube-state-metrics metricAllowlist in values.yaml.
rule {
source_labels = ["__name__"]
regex = "up|scrape_samples_scraped|scrape_samples_post_metric_relabeling|scrape_duration_seconds|kube_node_info|kube_node_status_condition|kube_node_status_allocatable|kube_pod_info|kube_pod_status_phase|kube_pod_container_status_restarts_total|kube_pod_container_resource_requests|kube_pod_container_resource_limits|kube_deployment_spec_replicas|kube_deployment_status_replicas_available|kube_statefulset_replicas|kube_statefulset_status_replicas_ready|kube_daemonset_status_desired_number_scheduled|kube_daemonset_status_number_ready|kubelet_node_name|container_cpu_usage_seconds_total|container_memory_working_set_bytes|process_resident_memory_bytes|process_cpu_seconds_total|prometheus_remote_storage_samples_pending|prometheus_remote_storage_samples_failed_total|prometheus_remote_storage_samples_retried_total|prometheus_remote_storage_samples_total|prometheus_remote_storage_queue_highest_sent_timestamp_seconds"
regex = string.join([
// Scrape health
"up",
"scrape_samples_post_metric_relabeling",
// kube-state-metrics: nodes
"kube_node_info",
"kube_node_labels",
"kube_node_status_condition",
"kube_node_status_allocatable",
// kube-state-metrics: pods
"kube_pod_info",
"kube_pod_status_phase",
"kube_pod_container_status_restarts_total",
"kube_pod_container_status_waiting_reason",
"kube_pod_container_status_last_terminated_reason",
"kube_pod_container_resource_requests",
"kube_pod_container_resource_limits",
// kube-state-metrics: workloads + autoscaling
"kube_deployment_spec_replicas",
"kube_deployment_status_replicas_available",
"kube_deployment_status_replicas_updated",
"kube_statefulset_replicas",
"kube_statefulset_status_replicas_ready",
"kube_statefulset_status_replicas_updated",
"kube_horizontalpodautoscaler_spec_max_replicas",
"kube_horizontalpodautoscaler_status_current_replicas",
"kube_horizontalpodautoscaler_status_desired_replicas",
// cAdvisor: usage, throttling, OOM kills (per container and node root cgroup)
"container_cpu_usage_seconds_total",
"container_memory_working_set_bytes",
"container_cpu_cfs_periods_total",
"container_cpu_cfs_throttled_periods_total",
"container_oom_events_total",
// kubelet
"kubelet_running_pods",
"kubelet_evictions",
// Alloy self: collector memory and remote-write health
"process_resident_memory_bytes",
"prometheus_remote_storage_samples_pending",
"prometheus_remote_storage_samples_failed_total",
"prometheus_remote_storage_samples_retried_total",
"prometheus_remote_storage_queue_highest_sent_timestamp_seconds",
], "|")
action = "keep"
}
// Discard cgroup aggregates; keep real Kubernetes containers only.
// Discard intermediate cgroup aggregates (pod sandboxes, kubepods slices) so
// per-container sums count each container once. The node root cgroup
// (id="/") is kept: it is whole-node usage without needing node-exporter.
rule {
source_labels = ["__name__", "container"]
regex = "container_(cpu_usage_seconds_total|memory_working_set_bytes);(|POD)"
source_labels = ["__name__", "container", "id"]
regex = "container_.*;(|POD);/.+"
action = "drop"
}
rule {
Expand All @@ -93,9 +139,11 @@ prometheus.relabel "v1" {
replacement = "$1"
}
// Preserve series identity (including cgroup id and cpu); no arbitrary pod labels.
// reason: waiting / last-terminated reason. horizontalpodautoscaler: HPA name.
// label_karpenter_sh_nodepool: from kube_node_labels. eviction_signal: kubelet_evictions.
rule {
action = "labelkeep"
regex = "__name__|job|instance|source|node|namespace|pod|uid|container|id|cpu|condition|status|phase|resource|unit|deployment|statefulset|daemonset|component_id|component_path|remote_name|url"
regex = "__name__|job|instance|source|node|namespace|pod|uid|container|id|cpu|condition|status|phase|reason|resource|unit|deployment|statefulset|horizontalpodautoscaler|label_karpenter_sh_nodepool|eviction_signal|component_id|component_path|remote_name|url"
}
forward_to = [prometheus.remote_write.amp.receiver]
}
Expand Down
24 changes: 21 additions & 3 deletions apps/monitoring/chart/values.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -92,22 +92,40 @@ kube-state-metrics:
autosharding:
enabled: false
prometheusScrape: false
collectors: [nodes, pods, deployments, statefulsets, daemonsets]
# Every metric below must also be on the Alloy keep-list in files/config.alloy,
# and every label it introduces on the Alloy labelkeep, or it never reaches AMP.
# No DaemonSets run on Auto Mode (system daemons are AWS-managed), so that
# collector is off.
collectors: [nodes, pods, deployments, statefulsets, horizontalpodautoscalers]
metricAllowlist:
# Nodes: inventory, Ready/pressure conditions, allocatable (capacity denominators), pool label.
- kube_node_info
- kube_node_labels
- kube_node_status_condition
- kube_node_status_allocatable
# Pods: placement, phase, restarts, waiting/exit reasons, requests and limits.
- kube_pod_info
- kube_pod_status_phase
- kube_pod_container_status_restarts_total
- kube_pod_container_status_waiting_reason
- kube_pod_container_status_last_terminated_reason
- kube_pod_container_resource_requests
- kube_pod_container_resource_limits
# Workloads: desired vs available/ready, and updated for rollout progress.
- kube_deployment_spec_replicas
- kube_deployment_status_replicas_available
- kube_deployment_status_replicas_updated
- kube_statefulset_replicas
- kube_statefulset_status_replicas_ready
- kube_daemonset_status_desired_number_scheduled
- kube_daemonset_status_number_ready
- kube_statefulset_status_replicas_updated
# Autoscaling: current vs desired vs ceiling.
- kube_horizontalpodautoscaler_spec_max_replicas
- kube_horizontalpodautoscaler_status_current_replicas
- kube_horizontalpodautoscaler_status_desired_replicas
# kube_node_labels only carries labels named here (as label_<sanitised_key>).
# karpenter.sh/nodepool → label_karpenter_sh_nodepool: system / general-purpose / frontend.
metricLabelsAllowlist:
- nodes=[karpenter.sh/nodepool]
nodeSelector:
kubernetes.io/os: linux
tolerations: []
Expand Down