Skip to content

Series length (T) as an axis in design-state admissibility — and a saturation question #114

Description

@rabujamra

Chris and I were discussing short-series behaviour and he pointed me toward the repo. Two findings from pushing on very small T, and a question I'd genuinely value Tom's view on.

processbehavior 0.2.0 from PyPI, pandas 2.3.3, Python 3.11.15. Docs read: quickstart, coffee shop, design-state lineage, API reference. Everything below is reproducible — seed 7 for the synthetic series, default_rng(20260903) with 20,000 replicates for the sampling study, 400,000 for the null distribution in §2, and the CMS PY2024 Shared Savings public use file for the real series. Happy to contribute the scripts.

A related observation about signal reporting at small T touches the "results an analyst would mistake for correct output" language in SECURITY.md, so I am sending that one through the private advisory channel rather than including it here.


1. The design state is constant in T; the estimator is not

The lineage classifies replication structure — cell sizes N_kt, complete versus incomplete grid. Series length is a separate property, and ADS 2 is ADS 2 whether T = 4 or T = 400.

A stable simulated process, mean 5.50, sd 1.20 — true natural process limits span a width of 7.20. Feeding the first T points:

T limit width share of true ADS recommended
3 2.78 39% 2 X
4 3.16 44% 2 X
6 3.14 44% 2 X
8 4.37 61% 2 X
12 4.64 64% 2 X
20 5.34 74% 2 X
30 5.73 80% 2 X
60 6.64 92% 2 X
150 6.84 95% 2 X

The design state, the recommendation and the analysis menu are constant down the whole column.

The sampling distribution behind that. MR̄/d₂ under a stable normal process with σ fixed at 1.0, 20,000 replicates per length:

n median p10 p90 p90/p10 RSD
3 0.889 0.337 1.797 5.34× 59%
4 0.925 0.429 1.689 3.94× 50%
8 0.969 0.592 1.454 2.45× 34%
12 0.976 0.671 1.351 2.01× 27%
25 0.989 0.769 1.245 1.62× 19%

At four points the estimate spans a near-fourfold range between its 10th and 90th percentile on a process whose variation never moved.

The design report at T = 4:

  Min cell size: 1 | K: 0 | T: 4
  Structure: Complete structure
  Available analyses (ADS 2):
    Primary: Histogram, Xbar, S, X *, mR
    Methods: Capability, Loss Function, Maximum Information

T: 4 is computed and printed. Nothing downstream consumes it. "Complete structure" is correct on the replication axis, and Capability and Loss Function are offered from four points.


2. Why I was at this end of the range, and what the real series did

My series are Medicare shared-savings settlements. Each unit is an accountable care organisation, each observation is an annual figure, and I get three or four points per unit. They also trend by construction, because medical costs rise and the benchmark years are built to reflect that.

476 ACOs, per-capita expenditure across three benchmark years plus the performance year:

mean level, BY1 -> PY                        $10,162 -> $10,184 -> $10,886 -> $12,622
year-over-year steps that increase           85.4%
median drift share |mean step| / mean|step|  1.000

MR̄ here is largely estimating the slope of medical cost growth rather than variation around a level. That is a property of the estimator at this length rather than anything the library does wrong — but it means the number the library correctly computes is not the number the analyst thinks they asked for.

And a drift test cannot help, for a reason worth stating. With three differences, an independent series is monotone with probability 2/4! = 1/12 ≈ 8.3%. That is above 5%, so the 95th percentile of the null distribution for drift share is 1.000 — the maximum attainable value. No four-point series can exceed it. Simulation at 400,000 replicates gives 8.3%, matching 1/12. The 95th percentile falls to 0.77 at T = 5 and 0.59 at T = 6, so this is specific to the shortest series rather than general.

So the two failure modes point in opposite directions — narrow limits and false alarms on a short stable series, wide limits and silence on a short trending one — and in both cases the design state reads as healthy.


3. A caveat on my own test

Those four points are not four observations of one process. Benchmark years are constructed inputs to a contractual benchmark, risk-adjusted separately and partially restated. A purist would say I fed the package something it should never have been given — and that objection makes the series worse than short, not better.

It was accepted without comment, and a user without the domain background would not know they had done anything wrong.


4. The question

This is the part I'd most value Tom's view on, and it is not rhetorical.

When the diagnostics that would report non-stationarity are disabled by the same shortness that makes the estimate imprecise, is there a defensible short-series variation estimator — or is the correct answer that four points does not carry a variation measure at all?

I've assumed the latter and gated my own work accordingly, but that is a judgement, not a result. Wheeler's own guidance is that useful limits can come from 5–6 points and firm up by 15–20, so a reasonable person could argue the estimate should be returned with its 3.94× interval attached and the reader left to discount it. My objection is that a number with a 3.94× interval gets quoted without the interval — which is an argument about institutions, not statistics.

Happy to move this to a Discussion if that is the better venue for a methodology question.


5. A suggestion, offered tentatively

Less "add a concept" than "connect two things already present": T is computed and printed in the design report, and SignalConfig.min_observations already exists with a default of 20.

  • an adequacy flag alongside the design state, derived from T rather than N_kt
  • below some threshold, a warning or recommended_chart = None with a stated reason, in the style of the existing ads_reason
  • optionally a monotone-trend flag, since a trending series changes what MR̄ means independently of length

The reason for putting it inside the design state rather than beside it: the package's distinguishing claim is that structure determines admissibility before anything is computed. Series length is a structural property, and it currently sits outside the gate everything else passes through.


6. What I liked, said plainly

The formulate() / execute() separation is the best decision in the API. Making "what does this structure permit" a first-class inspectable object, before computation, is the right decomposition — and it is precisely why the T question stood out, because the architecture is already shaped to hold that kind of constraint.

The ODS → ADS distinction is subtle and correct. Detecting incompleteness on raw data and then reporting what survives cleansing is careful work that most tools would collapse into one number.

Validating against Bishop's published reference results on every commit is the standard I would most want other analytical libraries to copy. Most demonstrate internal consistency; checking against an external authority is harder and different.

Activity

  1. cnicholas commented on Sep 10, 2026

    @cnicholas
    Owner

    Additional feedback from @rabujamra

    Few notes I wanted to share (which should make it cleaner also).

    First, min_observations isn't just unenforced, it's unreachable. SignalDetector._get_min_observations() returns from a hardcoded per-rule table rather than reading the config, so SignalConfig(min_observations=30) doesn't change anything for the eight built-in rules. That's in signals/detector.py.

    Second - and this is the useful one - most of the information needed for the fix I suggested is already computed. The rule loop already builds rules_skipped[rule_name] = f'needs {min_obs} observations, have {len(filtered_data)}' and passes it to _build_result, with your own comment saying it's there so the result can report partial evaluation. It just doesn't reach the "✓ No signals detected" string. So "report the effective rule set" looks like a reporting change rather than a design change.
    Also, the advisory is blocked on me, not you - GitHub wants affected versions and a severity before it can publish. I'll put 0.2.0 and Low. Still no CVE or credit needed from my side.

    Separately I owe you actual feedback on the package rather than just bug reports - the formulate() / execute() split is the best decision in the API, and I'll write properly on that when I'm not buried.

  2. added a commit that references this issue on Sep 10, 2026
  3. cnicholas commented on Sep 10, 2026

    @cnicholas
    Owner

    @rabujamra,
    The bug is fixed on main and ships in 0.2.1. SignalConfig.min_observations was unreachable: the detector read a hardcoded per-rule table and never consulted the config. And, as you spotted, the detector already recorded which rules it skipped for length, but the summary ignored that. Both are addressed in #118:

    • The per-rule table is now SignalDetector.RULE_MIN_OBSERVATIONS, the structural minimum. A rule below it is skipped, as before.
    • min_observations is now the advisory threshold. Below it the detector emits a ProcessBehaviorWarning naming both numbers, and the result is marked partial even when every runnable rule ran.
    • The summary for a partial evaluation now reads ⚠ Partial evaluation in X: 2 of 8 rules applicable at n=4 (below min_observations=20); skipped: rule_3 (needs 5), ... No signals from the rules evaluated. The checkmark line is reserved for a complete evaluation.
    • SignalResult gains rules_evaluated, rules_applicable, n_observations, below_min_observations and evaluation_note. Your four-point ramp is now a test.

    Which rules run, and what they flag, is unchanged, and the 280 reference assertions are untouched.

    On the methodology question. I am meeting with Tom this weekend to work through it. My current position, which I want his view on before it goes in the docs: four points do carry a variation measure, but with a stated precision, and the library's job is to attach that qualifier so it cannot be quietly dropped. Series length is a precision property rather than a replication-structure property, so I intend to keep it beside the design state rather than inside it. ADS 2 stays ADS 2 whether T is 4 or 400. What changes is that the design report will say so, in a line like

    Structure: Complete structure
    Series length: T=4, short (limits provisional; useful limits need 5-6 points, firm at 15-20) with a matching study.series_length property, a warning from formulate() below the shortest band, and a docs page that states the position and reproduces your sampling study. That is the second half of this issue and will land as a separate PR once the band edges are settled.

    I am deliberately not adding a monotone-trend flag. Your section 2 shows it has no power at T=4, and I would rather not ship another silent false negative.

    When the scripts are ready, a PR against validation/ or a link here both work. I would like to credit you on the docs page if you are comfortable with that.

  4. self-assigned this
    on Sep 10, 2026
  5. rabujamra commented on Sep 10, 2026

    @rabujamra
    ContributorAuthor

    Thanks for turning that around so fast, and for #118 — the partial-evaluation
    summary naming which rules ran is the part I'd have wanted most, and reserving the
    checkmark for a complete evaluation is a better fix than the warning I asked for.

    Scripts attached. short_series_sampling.py is self-contained — numpy, pandas,
    processbehavior, no data files — and regenerates every simulated number in the
    issue above. Happy to send it as a PR against validation/ instead if you'd
    rather have it in the tree; say which and I'll open one.

    processbehavior 0.2.1 | numpy 2.4.4 | pandas 2.3.3
    
    1. LIMIT WIDTH vs SERIES LENGTH
       one stable process, mean 5.50 sd 1.20, true limit span 7.20
    
          T  limit width  share of true    ADS  recommended
          3         2.78            39%      2            X
          4         3.16            44%      2            X
          6         3.14            44%      2            X
          8         4.37            61%      2            X
         12         4.64            64%      2            X
         20         5.34            74%      2            X
         30         5.73            80%      2            X
         60         6.64            92%      2            X
        150         6.84            95%      2            X
    
    2. SAMPLING DISTRIBUTION OF MRbar/d2, sigma fixed at 1.0
       20000 replicates per length
    
          n   median      p10      p90   p90/p10    RSD
          3    0.889    0.337    1.797     5.34x    59%
          4    0.925    0.429    1.689     3.94x    50%
          5    0.942    0.482    1.585     3.29x    44%
          8    0.969    0.592    1.454     2.45x    34%
         12    0.976    0.671    1.351     2.01x    27%
         25    0.989    0.769    1.245     1.62x    19%
    
    3. THE DRIFT TEST'S NULL
       400000 replicates of T iid standard normal draws
    
          T  P(monotone)    theory 2/T! 95th pct share
          4      0.0831         0.0833          1.000
          5      0.0164         0.0167          0.771
          6      0.0027         0.0028          0.581
    

    Run on 0.2.1, and I re-ran it against 0.2.0 — every number in all three
    sections is identical.
    The fix changed the reporting, not the estimator or the
    design-state output.

    Three reproducibility notes.

    The n=5 row. My table above skipped it, but the original loop computed it. It
    has to stay in the script or the RNG stream shifts and the n≥8 rows move in the
    third decimal. Seed default_rng(20260903), lengths in the order
    3, 4, 5, 8, 12, 25.

    precision=12 on every formulate() call. At the default of 3 the limits are
    rounded before the width is taken, which moves section 1 at these scales.

    A correction to my own arithmetic above. I gave the right number at T=4 (1/12)
    but stated the formula imprecisely. For T iid observations from a continuous
    distribution all T! orderings are equally likely, so the probability of a monotone
    ordering is 2/T!, not 2/(T-1)!. Simulation on iid standard normal draws gives
    0.0831, 0.0164 and 0.0027 at T = 4, 5 and 6 against 0.0833, 0.0167 and 0.0028.

    On the methodology question, I think the distinction you're drawing is cleaner than
    the gate I suggested: ADS describes replication structure; series length describes
    the precision of what that structure permits us to estimate. Agreed as well on not
    adding the monotone-trend flag — at T=4 the null result above gives it no useful
    rejection region.

    If it would help you and Tom with the band edges, I'm happy to extend the sampling
    table to every n from 3 through 30. And yes, absolutely comfortable with credit on
    the docs page — thank you.

    short_series_sampling.py

  6. cnicholas commented on Sep 10, 2026

    @cnicholas
    Owner

    Ran it against the 0.2.1 wheel in a clean environment: every number in all three sections matches.

    Yes to a PR against validation/, keep the filename.

    Yes to the n 3 through 30 extension, and if it can land before Saturday it goes straight into the discussion with Tom on the band edges.

    Credit confirmed. Thank you for corrections as well; the docs page will carry the corrected form.

  7. rabujamra commented on Sep 10, 2026

    @rabujamra
    ContributorAuthor

    n = 3 through 30, attached as short_series_bands.py. Self-contained, numpy plus
    processbehavior, no data files.

    MRbar/d2 under an iid standard normal process, sigma = 1.0
    100000 replicates per length, independent stream per n (default_rng([20260903, n]))
    
       n      p5     p10     p25     p50     p75     p90     p95  p90/p10     RSD   se(p10)   se(p90)
    -------------------------------------------------------------------------------------------------
       3   0.234   0.337   0.560   0.895   1.326   1.810   2.138    5.37x      59%    0.0017    0.0047
       4   0.331   0.430   0.635   0.928   1.293   1.676   1.928    3.90x      49%    0.0016    0.0035
       5   0.392   0.488   0.683   0.948   1.261   1.593   1.805    3.27x      44%    0.0015    0.0032
       6   0.436   0.528   0.710   0.951   1.239   1.526   1.716    2.89x      40%    0.0015    0.0027
       7   0.475   0.565   0.738   0.962   1.221   1.484   1.659    2.63x      36%    0.0014    0.0025
       8   0.509   0.595   0.760   0.971   1.210   1.453   1.607    2.44x      34%    0.0013    0.0023
       9   0.533   0.618   0.775   0.974   1.196   1.418   1.562    2.30x      31%    0.0013    0.0021
      10   0.556   0.637   0.786   0.974   1.185   1.396   1.528    2.19x      30%    0.0012    0.0020
      11   0.578   0.656   0.798   0.977   1.177   1.375   1.502    2.10x      28%    0.0012    0.0018
      12   0.595   0.670   0.809   0.981   1.170   1.360   1.477    2.03x      27%    0.0012    0.0017
      13   0.606   0.680   0.816   0.983   1.165   1.344   1.458    1.98x      26%    0.0011    0.0017
      14   0.623   0.694   0.825   0.984   1.160   1.330   1.440    1.92x      25%    0.0011    0.0016
      15   0.634   0.703   0.829   0.983   1.153   1.315   1.420    1.87x      24%    0.0011    0.0015
      16   0.646   0.714   0.836   0.987   1.150   1.310   1.410    1.83x      23%    0.0010    0.0015
      17   0.655   0.722   0.840   0.986   1.145   1.300   1.398    1.80x      23%    0.0010    0.0014
      18   0.665   0.730   0.846   0.986   1.139   1.288   1.383    1.77x      22%    0.0010    0.0014
      19   0.675   0.739   0.850   0.987   1.136   1.280   1.370    1.73x      21%    0.0009    0.0013
      20   0.683   0.745   0.855   0.988   1.133   1.274   1.360    1.71x      21%    0.0010    0.0013
      21   0.688   0.749   0.857   0.989   1.132   1.266   1.350    1.69x      20%    0.0009    0.0012
      22   0.696   0.756   0.862   0.989   1.127   1.258   1.342    1.67x      20%    0.0009    0.0012
      23   0.702   0.761   0.863   0.989   1.124   1.253   1.333    1.65x      19%    0.0009    0.0012
      24   0.709   0.766   0.869   0.990   1.123   1.249   1.326    1.63x      19%    0.0009    0.0012
      25   0.715   0.771   0.872   0.991   1.120   1.243   1.321    1.61x      18%    0.0008    0.0011
      26   0.718   0.775   0.873   0.991   1.118   1.238   1.315    1.60x      18%    0.0009    0.0011
      27   0.724   0.780   0.875   0.992   1.116   1.234   1.307    1.58x      18%    0.0008    0.0011
      28   0.730   0.784   0.879   0.993   1.115   1.229   1.301    1.57x      17%    0.0008    0.0011
      29   0.734   0.787   0.881   0.992   1.111   1.224   1.294    1.55x      17%    0.0008    0.0010
      30   0.738   0.790   0.883   0.993   1.110   1.220   1.288    1.54x      17%    0.0008    0.0010
    

    Three notes on how to read it.

    Each row stands on its own. The first script draws every length from one
    stream, which is why its n=5 row has to stay in the loop. That is fine for
    reproducing a fixed table and wrong for a table someone will cut thresholds out
    of, so this one gives each length its own stream, default_rng([20260903, n]).
    Rows therefore agree with the table in the issue to within Monte Carlo error
    rather than exactly — the script prints that comparison, max difference 0.013 at
    n=3 and under 0.009 from n=5 up.

    Monte Carlo standard errors are in the last two columns, at 100,000
    replicates. The operational point: a threshold placed where two adjacent n differ
    by less than their MC error is a threshold placed on nothing. At n=3 the se on
    p90 is 0.005 against a p90 of 1.810, so the resolution is fine everywhere in this
    range — but it is worth having the number rather than assuming it.

    The shape, stated without proposing anything. The p90/p10 spread falls
    steeply from 5.37× at n=3 to 2.44× at n=8, then slowly: 2.03× at 12, 1.71× at 20,
    1.54× at 30. Most of the improvement available is spent by roughly n=8–10, and
    after that it is a long shallow tail.

    Which functional should govern an adequacy band — p90/p10, RSD, a coverage
    probability, something else — and where the cuts fall are yours and Tom's to
    decide. I have deliberately not proposed any. Happy to compute where a criterion
    lands once you have picked one, including quantiles I have not printed here.

    I will open the PR against validation/ keeping the filenames.

    short_series_bands.py

  8. added 2 commits that reference this issue on Sep 10, 2026
  9. cnicholas commented on Sep 16, 2026

    @cnicholas
    Owner

    @rabujamra, Tom and I got together to review the question you raised. We generated a synthetic dataset based on your prior post.
    file: aco_per_capita_expenditure.csv

    There are 24 ACOs in the file, 96 obs in total.

    How Tom read your data. He built VAS for analytic studies which include statistical process control applications where data usually arrive quickly and action follows quickly. SPC detection rules were originally designed for SPC applications. VAS and Process Behavior(PB) which is the software that implements the VAS methodology rests on the scientific method, the search for nonrandom patterns in the data and the nontrivial replication of results to build credibility, not on proscribed rules about how many points a chart needs, or what constitutes identification of an assignable cause. When only a few time periods are used it raises the level of uncertainty, but this does not make the analysis uninformative.

Two things follow that I think are important for you to consider.

    First, from Tom's perspective your data has a great deal of nontrivial replication. For example, formulated as one study, the same rising pattern over time is reproduced across nearly every organization. That agreement across independent subgroups is the finding, and it does not depend on the process behavior limits.
     
    Second, the question Tom would put to you at the beginning of your study is whether this is an “enumerative study”, where action is taken on the frame that produced the data, or whether this is an “analytic study”, where you will be making predictions into the future and taking actions based on the predictions. If the latter, which is likely, the uncertainty associated with these predictions cannot be measured, or controlled, using probability theory because the predictions are extrapolations into the future. For probability theory to apply the assumption “the future behaves like the past” must be true in fact which in most analytic studies is not true. This is why the nontrivial replication is the confirmatory methodology to build credibility and not probability calculations that depend upon the precision of the standard deviation (sigma) behind the behavior limits. His book is largely about this distinction.

    Please look at the six charts attached below that are part of the output produced by Process Behavior which clearly demonstrate these principles.

    Image Image Image Image Image Image

    Here is code to generate two charts. My guess from "476 of 476" is that you ran one four point study per organization. Run one study instead, with the organization as a factor and the year as time, and look at these in order:

    Note: the chart will be busy with all the observations plotted so you might want to subset to 100. You can easily zoom in with the plotly controls.

    import processbehavior as pb

    study = pb.ProcessBehavior.read_csv("aco_per_capita_expenditure.csv").formulate(
    response="PER CAPITA EXPENDITURE", factors=["ACO"], time="YEAR"
    )
    print(study.design()) # ADS 2 ... Series length: T=4. Sigma for X/mR rests on 3 moving ranges.

    1. Every organization on one individuals chart, one scale, lane by lane

    (the first analysis in the app's sequence)

    combined = study.execute(chart="X", by=[], companion=True)
    combined.plot(chart="X")

    ...versus one small chart per organization

    per_aco = study.execute(chart="X", by=["ACO"], companion=True)
    per_aco.plot(chart="X", facet=True, ncols=4)

    2. The R2 residual on an individuals chart: the year to year change itself

    (the last analysis in the app's sequence)

    r2 = study.execute(chart="X", value="R2", by=[], companion=True)
    r2.plot(chart="X")

    Please let us know if you have any additional questions.

  10. rabujamra commented on Sep 17, 2026

    @rabujamra
    ContributorAuthor

    Thanks Chris — and thanks to you both for building the dataset, that made this concrete quickly.

    I ran the combined and per-ACO formulations. Before I go further on Tom's substantive argument, there's one implementation point I wanted to check with you.

    With by=[], the combined series has 96 observations and the mR chart reports 95 moving ranges. The within-ACO ranges the data can supply number 72 — sum over ACOs of (n_i − 1), which is N − K = 96 − 24. The remaining 23 correspond to the 23 transitions between successive ACO blocks, computed from the last year of one ACO to the first year of the next. For example, the reported mR at ACO-002/2021 is |11837.21 − 11500.14| = 337.07, which spans ACO-001 2024 → ACO-002 2021.

    On this dataset that materially changes the calculation: mean mR is 1443.208 over all 95 ranges versus 933.792 over the 72 within-ACO ranges only. All three mR points currently flagged beyond the limits are first-year boundary ranges (ACO-019, ACO-021 and ACO-024, each at 2021).

    Two things I checked so this isn't about input ordering. First, a copy of the file sorted by (YEAR, ACO) instead of (ACO, YEAR) produces identical output — ProcessBehavior establishes its own combined ordering, so this isn't a case of badly sorted input. Second, by=["ACO"] reports exactly 72 ranges and is unaffected. The attached script takes its boundary pairs from the library's own X-chart sequence rather than from the CSV, and reproduces the same 23 transitions and the same values on sorted, reordered, and deliberately unbalanced (2–4 years per ACO) copies.

    This struck me as possibly related to #121 one layer down: the plotting change stops a line being drawn across a lane boundary, and the mR calculation may still difference across that same boundary.

    Separately, on Tom's point — the distinction he draws between the short-series question and replication across organizations is the part I want to sit with, and I'll come back to it properly once I understand what the combined formulation is doing. It reframed the question for me more than it settled it.

    Does that look like intended behavior? Happy to open a separate issue with the reproducible check if that's more useful than a comment here.

    mr_lane_boundaries.py

  11. cnicholas commented on Sep 18, 2026

    @cnicholas
    Owner

    @rabujamra, thank you for this. It is a careful check and you have it exactly right: the combined chart differences across the transition from one ACO to the next, and on your data those 23 transitions lift the average moving range by about half. That is intended, and I want to give you the straight version of why, including the part I discovered rather than designed.

    How the combined chart is built. With by=[] the points are not ordered purely by time, the way a textbook individuals chart would be. They are ordered by ACO first and then by year, so each organization occupies its own lane at one common scale and the same pattern can be read lane by lane. I chose that ordering for the picture, because it is how Tom reads data. I will be honest that I did not think through what it meant for the moving range at the time. Your question made me look.

    What it does to the arithmetic. The moving range runs across the whole sequence in that order, so the step from one organization to the next contributes a range like any other. With many periods per subgroup those transitions are a handful among hundreds and barely move sigma. With four periods per subgroup they are a quarter of the ranges, and they dominate. So your combined limits are wide, and they are wide precisely because of the between-organization differences.

    Tom's reference does the same thing. I checked against his Minitab output on the test database (PM SDS 2, 8 cells × 100 periods, 800 observations, 799 ranges, worksheet in cell order). Differencing consecutive rows gives 792 within-cell ranges and 7 ranges at the cell transitions. Including the 7 transitions, the mR center is 0.868 and the individuals limits are 235.47 and 240.09, which match his slide to the last digit. Excluding them, the center would be 0.841 and the limits would move to 235.54 and 240.02, which do not match. The library reproduces his reference only because it keeps the transitions.

    Why the transitions are not artefacts. The combined chart's job is to ask whether these subgroups belong to one process. A jump at the transition between two organizations is evidence on exactly that question. It is not an artefact of sort order; it is the level difference between two organizations, observed. On your data the chart is saying, loudly, that 24 organizations are not one process. That is the same finding Tom pointed you to from the pattern side: the rising shape replicates in nearly every lane, while the levels do not agree. The three flagged ranges are the three largest level shifts between neighbouring organizations. They are signals, and they arrived before you had asked the question.

    You are right that this sits one layer below #121. That change was visual only: the line no longer connects across a lane boundary, so the eye is not led to read the jump as a within-organization movement. The range itself stays in the calculation, deliberately.

    Three views of the same study. The library gives you the question you want to ask, rather than one chart with one set of limits.

    1. The process as a whole. Every ACO, four observations each, one set of limits across the whole frame. This is the chart you ran, and it asks "is this one process?"
    combined = study.execute(chart="X", by=[], companion=True)
    combined.plot(chart="X")
    1. The same lanes, each with its own center line and limits. This asks "how does each organization behave on its own, side by side at one scale?" The moving range within each phase is what sets that phase's limits, so the transitions no longer enter.
    phased = study.execute(chart="X", by=[], phased=True, companion=True)
    phased.plot(chart="X")
    Image
    1. Drilling in, one chart per organization. This is the 72-range version you already confirmed.
    per_aco = study.execute(chart="X", by=["ACO"], companion=True)
    per_aco.plot(chart="X", facet=True, ncols=4)

    No separate issue needed; this comment is the record, and I will add a note to the docs so the next reader finds it without having to reverse-engineer the sequence as you did. Your script was a clean way to establish it, and I appreciate the care.

    On Tom's larger point, take the time you need. That reframing is the part of this that matters most, and I would rather hear your considered reaction than a quick one.

  12. rabujamra commented on Sep 19, 2026

    @rabujamra
    ContributorAuthor

    Chris — thank you for this, and for going and checking against Tom's reference rather than just answering me. The "part I discovered rather than designed" line is a generous way to put it and I appreciate it.

    I took the ordering question seriously enough to run it on the 24-ACO file rather than reason about it abstractly. Script attached: mr_permutation_invariance.py, one run, about a second. It holds every observation fixed and varies only lane order.

    Your arithmetic checks exactly.

    24 organisations x 4 obs = 96 obs, 95 moving ranges
      within-ACO     72 ranges   mean    933.79   <- invariant to lane order
      boundaries     23 ranges   mean   3037.90   <- varies with lane adjacency
      boundary/within mean ratio                     3.25x
    
      mRbar as loaded                              1443.21
      mRbar within-only                             933.79
      increase from including boundaries              54.6%
    

    "About half" was right.

    The permutation moved my reading toward yours, not away from it.

    Across 20,000 random lane orderings the X chart never once returned zero signals — the count stayed between 3 and 8, and zero occurred in 0.00% of orderings. So the substantive conclusion, that these 24 organisations are not behaving as one process, is robust to lane ordering on this dataset. That is evidence for your reading of the combined chart and I would rather report it than bury it.

    What is not invariant is the boundary-signal structure.

    ordering              mRbar    X-limit width   X signals   mR signals
      as loaded           1443.2         7677.9          4           3
      sorted by level     1272.5         6769.9          7           0
      interleaved         1497.6         7967.4          3           6
      20k random mean     1413.4         7519.4       3.92        4.48
                 range  1269-1545     6753-8221        3-7         1-9
    

    Using D4 = 3.267 for the mR upper limit, the flagged boundary transitions are:

      as loaded       ACO-018|ACO-019, ACO-020|ACO-021, ACO-023|ACO-024
      sorted by level (none)
      interleaved     ACO-018|ACO-024, ACO-023|ACO-007, ACO-009|ACO-006,
                      ACO-002|ACO-022, ACO-010|ACO-005, ACO-020|ACO-008
    

    Worth checking those against what companion=True actually flags in your own output — if your mR rule differs from the classical D4 the labels will move, though the spread across orderings should not.

    The observations and the 72 within-ACO ranges never change. One thing worth saying plainly: the shipped ACO-001..024 order sits at the 77th percentile of random orderings by mRbar. It is not adversarial. It is arbitrary, which is the part that interests me.

    And a boundary range may not be measuring quite what it looks like.

    A boundary moving range changes organisation and resets time in the same number: it steps from organisation i's 2024 observation to organisation j's 2021 observation. Every one of the 24 lanes rises here.

      mean within-lane first-to-last change         2586.9
      mean boundary |last_i - first_j|              3037.9
      mean adjacent pure-level gap |mean_i-mean_j|  1833.9
      boundary / pure-level ratio                      1.66x
      difference from pure-level comparison            39.6% of mean boundary magnitude
    

    That comparison is descriptive. I am not claiming an additive decomposition — with absolute values and heterogeneous trajectories the two pieces do not cleanly separate. The point is only that the boundary construction differs systematically from a pure between-organisation level comparison.

    Where that leaves me.

    The narrow statement I think the data supports: the overall conclusion is robust to lane ordering; the number and identity of the boundary signals are conditional on the traversal. The boundary signals look like evidence along the chosen path rather than invariant pairwise properties of the organisations.

    Which leaves me with a question rather than a complaint: how should a boundary signal be read? If it is a claim about the traversal, then the phased=True view you pointed me to is the right default for anyone asking about individual organisations, and the docs note you are planning would cover it. If it is meant as a claim about the organisations themselves, the time reset above is doing some of the work.

    Happy to run anything else that would be useful, and happy to be told I have mis-framed it.

    On Tom's larger point — I am taking you at your word and giving it the time it deserves. I will come back on the enumerative/analytic question separately, because I do not think it is the same conversation as this one and I would rather not blur them.

    mr_permutation_invariance.py

  13. cnicholas commented on Sep 19, 2026

    @cnicholas
    Owner

    @rabujamra, thank you. I ran mr_permutation_invariance.py against the file and every number matches, including the 77th percentile. The library's mR limit is the classical D4, and companion=True flags exactly the three transitions you list for the shipped order, at the first-year points of ACO-019, ACO-021 and ACO-024, plus the same four X points.

    Your narrow statement is the right one, and I would put it the same way. The combined chart's conclusion survives any lane order, which is the question that chart exists to answer. The identity of a boundary signal does not, because it is a claim about the traversal: the two lanes, in that order, do not join into one stable stream. And you are right that on a rising process the boundary step carries the trend reversal along with the level gap, so it is systematically larger than a pure level comparison. That is a property of the lane-major sequence on trending data rather than of the two organisations.

    So the answer to your question is that a boundary signal should be read as evidence along the path, and claims about the organisations should come from the charts that have no path. Two of them are already there:

    1. Which organisations differ. The recentred R5 chart by organisation, an Xbar chart of each organisation's level with limits set from R2, flags 8 of the 24. The points are order-free, since each is an organisation's mean. The limits are not entirely, because in ADS 2 they are set from R2, and R2 is built from the same lane-major sequence, so one of every lane's four R2 values is a boundary half-difference. Applying your method, across 200 random lane orders the limit width ranged from 2,921 to 3,428; the eight flagged under the shipped order were flagged under every order, with one or two marginal organisations joining in a minority of them. (Edited: my first version of this paragraph said the set was invariant to ordering. That was too strong.)
    levels = study.execute(chart="Xbar", by=["ACO"], value="R5", recentered=True)
    levels.plot()
    1. How each organisation behaves. The phased view, which you already have, with each lane's limits from its own ranges.

    The docs note I owe you will say exactly this: the combined chart answers "one process?"; boundary signals are path evidence; organisation-level claims come from the R5 chart and the phased view.

    Your script belongs in the repo beside the other two. Please open a PR against validation/, same as #119, keeping the filename, so it lands with you as the author.

    On the enumerative versus analytic question, agreed, it is a separate conversation, and it deserves its own thread when you are ready.

  14. cnicholas commented on Sep 19, 2026

    @cnicholas
    Owner

    @rabujamra, the docs note has landed, and #130 is merged with your script and the ACO file beside it under validation/.

    The chart-types guide now has three sections that come directly from this thread: Three Views of the Same Study, which gives the combined, phased and stratified views with the question each one answers; Reading a Boundary Signal, which says that a signal at a transition is evidence about the traversal and explains why the transitions stay in the calculation; and Which Claims Are Order-Free, a table of what survives reordering the lanes, including the R5 limits caveat you already folded into your script's scope note. It is also the first time the phased view is documented in the user guide, which it should have been from the start.

    The validation README lists all three of your scripts with their PR numbers, and the permutation check runs from the repo root against the committed file.

    Thank you for pushing on this until the claims were stated at the right strength. Closing this one; the enumerative-versus-analytic thread is yours to open whenever you're ready.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions