fix: bound Prometheus histogram memory between scrapes - #363
Open
korkin25 wants to merge 1 commit into
Open
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Prometheus Core 1.2.1 stores each distribution observation until the next scrape.
With telemetry enabled and no scraper, routine database queries keep growing the
raw-sample ETS table: a synthetic run increased its memory from 181,384 bytes at
1,000 events to 16,015,816 bytes at 100,000 events.
Aggregate histograms when observations arrive, retaining one bucket/count/sum row
per metric and label combination. Keep Core's scalar handlers and exporter, the
existing inclusive bucket boundaries and fractional sums, and the metrics
endpoint's authorization and response format. Memory depends on distinct series
and configured buckets rather than observations between scrapes.
Validation of the identical collector/reporter/test implementation: nine regression
tests passed, including one million events without scraping, concurrent writers
and scrapes, exact boundaries and sums, malformed measurements, and collector
restart/detachment. The no-scrape test retained one aggregate and zero raw samples,
using 1,449 ETS words at both 100,000 and 1,000,000 events. Compilation, strict Credo
and compile-connected xref checks passed for that implementation. No dependencies
are added or upgraded.