Skip to content

feat: add TDigest, every quantile of a stream from one sketch - #7642

Merged
DenizAltunkapan merged 2 commits into
TheAlgorithms:masterfrom
alxkm:feat/t-digest
Oct 9, 2026
Merged

DenizAltunkapan merged 2 commits into
TheAlgorithms:masterfrom
alxkm:feat/t-digest

Conversation

@alxkm

@alxkm alxkm commented Oct 9, 2026

Copy link
Copy Markdown
Member

Adds the t-digest: a small, mergeable sketch that answers any quantile of a stream, and its CDF, from one structure. It is the general counterpart of P2QuantileEstimator, which tracks a single quantile chosen in advance in constant memory.

The sketch is a list of centroids, each holding a mean and a weight, kept sorted by mean. What makes it more than a histogram is the rule that decides how large a centroid may grow: a centroid absorbs points only while it spans at most one unit of the scale function

k(q) = d / (2 pi) * asin(2q - 1)

where d is the compression. asin is steep near q = 0 and q = 1, so centroids in the tails are forced to stay tiny, often a single point, while those near the median are allowed to grow large. The rank error is of the order of 1 / d around the median and shrinks towards the ends, which is the trade latency percentiles want: p99.9 comes out nearly exact from a sketch of a few kilobytes, independent of the stream length.

Operation Cost
add O(log d) amortised
quantile, cdf O(d)
merge O(d log d)
memory O(d)

A few decisions that are worth a reviewer's attention:

  • Insertions are buffered and folded into the centroid list in batches, which is what keeps add cheap. Queries fold the pending buffer in before answering, so they mutate internal state; the class is documented as not thread-safe rather than pretending otherwise.
  • The minimum and maximum are tracked separately and returned exactly for q = 0 and q = 1. Between them, quantile and cdf interpolate linearly across centroids, each treated as sitting at the centre of the weight it carries, with min and max closing the two ends.
  • Digests merge the way partial sums do, so sketches built on separate shards combine into one. Merging a digest into itself gives the same result as merging an equal copy.
  • Bad input is refused with an IllegalArgumentException: non-finite samples, invalid weights, quantiles outside [0, 1], a NaN passed to cdf, and a compression below 10, where the scale function stops leaving room for a useful number of centroids. Querying an empty digest is an IllegalStateException.

TDigestTest runs 52 cases across 25 methods. Accuracy is asserted against the true rank rather than eyeballed: on 100 000 uniform samples every tested quantile lands within 0.01 of the requested rank; on 200 000 Gaussian samples the rank error stays under a bound that narrows towards the tails, from 0.01 at the median to 0.0005 at q = 0.001 and q = 0.999. Three shards merged answer within 0.01 of the true rank, with total weight, minimum and maximum exact. The centroid count stays at or below the compression after 200 000 samples for compressions of 20, 100 and 500, with the centroids sorted and their weights summing to the total. Weighted, sorted, constant and heavily duplicated input, and wildly disparate weights, are covered as well.

Reference: T. Dunning, O. Ertl, Computing extremely accurate quantiles using t-digests.

Checklist

  • I have read CONTRIBUTING.md.
  • This pull request is all my own work -- I have not plagiarized it.
  • All filenames are in PascalCase.
  • All functions and variable names follow Java naming conventions.
  • All new algorithms have a URL in their comments that points to Wikipedia or other similar explanations.
  • All new algorithms include a corresponding test class that validates their functionality.
  • All new code is formatted with clang-format -i --style=file path/to/your/file.java

P2QuantileEstimator tracks one quantile chosen in advance, in constant memory. A t-digest keeps a small list of centroids, each a mean and a weight, sorted by mean, and answers any quantile and the CDF afterwards from the same sketch. A centroid may absorb points only while it spans at most one unit of the scale function k(q) = d / (2 pi) * asin(2q - 1), and asin is steep at both ends, so centroids in the tails stay at one or a few points while those near the median are allowed to grow. The rank error is of the order of 1 / d around the median and shrinks towards q = 0 and q = 1, which is the trade latency percentiles want: p99.9 comes out nearly exact from a sketch of a few kilobytes.

Insertions are buffered and folded into the centroid list in batches, so add is O(log d) amortised; quantile and cdf are O(d) and interpolate linearly between centroids, each treated as sitting at the centre of the weight it carries. The minimum and maximum are tracked separately and returned exactly for q = 0 and q = 1. Two digests merge the way partial sums do, so sketches built on separate shards combine into one. Non-finite samples, invalid weights, quantiles outside [0, 1] and a compression below 10 are refused; queries fold the pending buffer in first, so the class is documented as not thread-safe.

Tests: on 100 000 uniform samples every tested quantile lands within 0.01 of the requested rank; on 200 000 Gaussian samples the rank error stays under a bound that narrows towards the tails, from 0.01 at the median to 0.0005 at q = 0.001 and q = 0.999; three shards merged answer within 0.01 of the true rank, with total weight, minimum and maximum exact; the centroid count stays at or below the compression after 200 000 samples for compressions of 20, 100 and 500; and weighted, sorted, constant and heavily duplicated input are covered.
Signed-off-by: alxkm <19151554+alxkm@users.noreply.github.com>
@codecov-commenter

codecov-commenter commented Oct 9, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 98.01980% with 4 lines in your changes missing coverage. Please review.
✅ Project coverage is 82.17%. Comparing base (cb6f438) to head (fe31b7b).

Files with missing lines Patch % Lines
...main/java/com/thealgorithms/streaming/TDigest.java 98.01% 2 Missing and 2 partials ⚠️
Additional details and impacted files
@@             Coverage Diff              @@
##             master    #7642      +/-   ##
============================================
+ Coverage     82.04%   82.17%   +0.13%     
- Complexity     8286     8363      +77     
============================================
  Files           840      841       +1     
  Lines         25872    26074     +202     
  Branches       5053     5086      +33     
============================================
+ Hits          21227    21427     +200     
- Misses         3856     3858       +2     
  Partials        789      789              

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@DenizAltunkapan
DenizAltunkapan enabled auto-merge (squash) October 9, 2026 09:13
@DenizAltunkapan
DenizAltunkapan merged commit 5eb9d66 into TheAlgorithms:master Oct 9, 2026
13 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants