Repository navigation
Series length (T) as an axis in design-state admissibility — and a saturation question #114
Description
Activity
Additional feedback from @rabujamra
Few notes I wanted to share (which should make it cleaner also).
First, min_observations isn't just unenforced, it's unreachable. SignalDetector._get_min_observations() returns from a hardcoded per-rule table rather than reading the config, so SignalConfig(min_observations=30) doesn't change anything for the eight built-in rules. That's in signals/detector.py.
Second - and this is the useful one - most of the information needed for the fix I suggested is already computed. The rule loop already builds rules_skipped[rule_name] = f'needs {min_obs} observations, have {len(filtered_data)}' and passes it to _build_result, with your own comment saying it's there so the result can report partial evaluation. It just doesn't reach the "✓ No signals detected" string. So "report the effective rule set" looks like a reporting change rather than a design change.
Also, the advisory is blocked on me, not you - GitHub wants affected versions and a severity before it can publish. I'll put 0.2.0 and Low. Still no CVE or credit needed from my side.Separately I owe you actual feedback on the package rather than just bug reports - the formulate() / execute() split is the best decision in the API, and I'll write properly on that when I'm not buried.
- added a commit that references this issue
on Sep 10, 2026 @rabujamra,
The bug is fixed on main and ships in 0.2.1. SignalConfig.min_observations was unreachable: the detector read a hardcoded per-rule table and never consulted the config. And, as you spotted, the detector already recorded which rules it skipped for length, but the summary ignored that. Both are addressed in #118:- The per-rule table is now SignalDetector.RULE_MIN_OBSERVATIONS, the structural minimum. A rule below it is skipped, as before.
- min_observations is now the advisory threshold. Below it the detector emits a ProcessBehaviorWarning naming both numbers, and the result is marked partial even when every runnable rule ran.
- The summary for a partial evaluation now reads ⚠ Partial evaluation in X: 2 of 8 rules applicable at n=4 (below min_observations=20); skipped: rule_3 (needs 5), ... No signals from the rules evaluated. The checkmark line is reserved for a complete evaluation.
- SignalResult gains rules_evaluated, rules_applicable, n_observations, below_min_observations and evaluation_note. Your four-point ramp is now a test.
Which rules run, and what they flag, is unchanged, and the 280 reference assertions are untouched.
On the methodology question. I am meeting with Tom this weekend to work through it. My current position, which I want his view on before it goes in the docs: four points do carry a variation measure, but with a stated precision, and the library's job is to attach that qualifier so it cannot be quietly dropped. Series length is a precision property rather than a replication-structure property, so I intend to keep it beside the design state rather than inside it. ADS 2 stays ADS 2 whether T is 4 or 400. What changes is that the design report will say so, in a line like
Structure: Complete structure
Series length: T=4, short (limits provisional; useful limits need 5-6 points, firm at 15-20) with a matching study.series_length property, a warning from formulate() below the shortest band, and a docs page that states the position and reproduces your sampling study. That is the second half of this issue and will land as a separate PR once the band edges are settled.I am deliberately not adding a monotone-trend flag. Your section 2 shows it has no power at T=4, and I would rather not ship another silent false negative.
When the scripts are ready, a PR against validation/ or a link here both work. I would like to credit you on the docs page if you are comfortable with that.
Thanks for turning that around so fast, and for #118 — the partial-evaluation
summary naming which rules ran is the part I'd have wanted most, and reserving the
checkmark for a complete evaluation is a better fix than the warning I asked for.Scripts attached.
short_series_sampling.pyis self-contained — numpy, pandas,
processbehavior, no data files — and regenerates every simulated number in the
issue above. Happy to send it as a PR againstvalidation/instead if you'd
rather have it in the tree; say which and I'll open one.processbehavior 0.2.1 | numpy 2.4.4 | pandas 2.3.3 1. LIMIT WIDTH vs SERIES LENGTH one stable process, mean 5.50 sd 1.20, true limit span 7.20 T limit width share of true ADS recommended 3 2.78 39% 2 X 4 3.16 44% 2 X 6 3.14 44% 2 X 8 4.37 61% 2 X 12 4.64 64% 2 X 20 5.34 74% 2 X 30 5.73 80% 2 X 60 6.64 92% 2 X 150 6.84 95% 2 X 2. SAMPLING DISTRIBUTION OF MRbar/d2, sigma fixed at 1.0 20000 replicates per length n median p10 p90 p90/p10 RSD 3 0.889 0.337 1.797 5.34x 59% 4 0.925 0.429 1.689 3.94x 50% 5 0.942 0.482 1.585 3.29x 44% 8 0.969 0.592 1.454 2.45x 34% 12 0.976 0.671 1.351 2.01x 27% 25 0.989 0.769 1.245 1.62x 19% 3. THE DRIFT TEST'S NULL 400000 replicates of T iid standard normal draws T P(monotone) theory 2/T! 95th pct share 4 0.0831 0.0833 1.000 5 0.0164 0.0167 0.771 6 0.0027 0.0028 0.581Run on 0.2.1, and I re-ran it against 0.2.0 — every number in all three
sections is identical. The fix changed the reporting, not the estimator or the
design-state output.Three reproducibility notes.
The n=5 row. My table above skipped it, but the original loop computed it. It
has to stay in the script or the RNG stream shifts and the n≥8 rows move in the
third decimal. Seeddefault_rng(20260903), lengths in the order
3, 4, 5, 8, 12, 25.precision=12on everyformulate()call. At the default of 3 the limits are
rounded before the width is taken, which moves section 1 at these scales.A correction to my own arithmetic above. I gave the right number at T=4 (1/12)
but stated the formula imprecisely. For T iid observations from a continuous
distribution all T! orderings are equally likely, so the probability of a monotone
ordering is2/T!, not2/(T-1)!. Simulation on iid standard normal draws gives
0.0831, 0.0164 and 0.0027 at T = 4, 5 and 6 against 0.0833, 0.0167 and 0.0028.On the methodology question, I think the distinction you're drawing is cleaner than
the gate I suggested: ADS describes replication structure; series length describes
the precision of what that structure permits us to estimate. Agreed as well on not
adding the monotone-trend flag — at T=4 the null result above gives it no useful
rejection region.If it would help you and Tom with the band edges, I'm happy to extend the sampling
table to every n from 3 through 30. And yes, absolutely comfortable with credit on
the docs page — thank you.Ran it against the 0.2.1 wheel in a clean environment: every number in all three sections matches.
Yes to a PR against validation/, keep the filename.
Yes to the n 3 through 30 extension, and if it can land before Saturday it goes straight into the discussion with Tom on the band edges.
Credit confirmed. Thank you for corrections as well; the docs page will carry the corrected form.
n = 3 through 30, attached as
short_series_bands.py. Self-contained, numpy plus
processbehavior, no data files.MRbar/d2 under an iid standard normal process, sigma = 1.0 100000 replicates per length, independent stream per n (default_rng([20260903, n])) n p5 p10 p25 p50 p75 p90 p95 p90/p10 RSD se(p10) se(p90) ------------------------------------------------------------------------------------------------- 3 0.234 0.337 0.560 0.895 1.326 1.810 2.138 5.37x 59% 0.0017 0.0047 4 0.331 0.430 0.635 0.928 1.293 1.676 1.928 3.90x 49% 0.0016 0.0035 5 0.392 0.488 0.683 0.948 1.261 1.593 1.805 3.27x 44% 0.0015 0.0032 6 0.436 0.528 0.710 0.951 1.239 1.526 1.716 2.89x 40% 0.0015 0.0027 7 0.475 0.565 0.738 0.962 1.221 1.484 1.659 2.63x 36% 0.0014 0.0025 8 0.509 0.595 0.760 0.971 1.210 1.453 1.607 2.44x 34% 0.0013 0.0023 9 0.533 0.618 0.775 0.974 1.196 1.418 1.562 2.30x 31% 0.0013 0.0021 10 0.556 0.637 0.786 0.974 1.185 1.396 1.528 2.19x 30% 0.0012 0.0020 11 0.578 0.656 0.798 0.977 1.177 1.375 1.502 2.10x 28% 0.0012 0.0018 12 0.595 0.670 0.809 0.981 1.170 1.360 1.477 2.03x 27% 0.0012 0.0017 13 0.606 0.680 0.816 0.983 1.165 1.344 1.458 1.98x 26% 0.0011 0.0017 14 0.623 0.694 0.825 0.984 1.160 1.330 1.440 1.92x 25% 0.0011 0.0016 15 0.634 0.703 0.829 0.983 1.153 1.315 1.420 1.87x 24% 0.0011 0.0015 16 0.646 0.714 0.836 0.987 1.150 1.310 1.410 1.83x 23% 0.0010 0.0015 17 0.655 0.722 0.840 0.986 1.145 1.300 1.398 1.80x 23% 0.0010 0.0014 18 0.665 0.730 0.846 0.986 1.139 1.288 1.383 1.77x 22% 0.0010 0.0014 19 0.675 0.739 0.850 0.987 1.136 1.280 1.370 1.73x 21% 0.0009 0.0013 20 0.683 0.745 0.855 0.988 1.133 1.274 1.360 1.71x 21% 0.0010 0.0013 21 0.688 0.749 0.857 0.989 1.132 1.266 1.350 1.69x 20% 0.0009 0.0012 22 0.696 0.756 0.862 0.989 1.127 1.258 1.342 1.67x 20% 0.0009 0.0012 23 0.702 0.761 0.863 0.989 1.124 1.253 1.333 1.65x 19% 0.0009 0.0012 24 0.709 0.766 0.869 0.990 1.123 1.249 1.326 1.63x 19% 0.0009 0.0012 25 0.715 0.771 0.872 0.991 1.120 1.243 1.321 1.61x 18% 0.0008 0.0011 26 0.718 0.775 0.873 0.991 1.118 1.238 1.315 1.60x 18% 0.0009 0.0011 27 0.724 0.780 0.875 0.992 1.116 1.234 1.307 1.58x 18% 0.0008 0.0011 28 0.730 0.784 0.879 0.993 1.115 1.229 1.301 1.57x 17% 0.0008 0.0011 29 0.734 0.787 0.881 0.992 1.111 1.224 1.294 1.55x 17% 0.0008 0.0010 30 0.738 0.790 0.883 0.993 1.110 1.220 1.288 1.54x 17% 0.0008 0.0010Three notes on how to read it.
Each row stands on its own. The first script draws every length from one
stream, which is why its n=5 row has to stay in the loop. That is fine for
reproducing a fixed table and wrong for a table someone will cut thresholds out
of, so this one gives each length its own stream,default_rng([20260903, n]).
Rows therefore agree with the table in the issue to within Monte Carlo error
rather than exactly — the script prints that comparison, max difference 0.013 at
n=3 and under 0.009 from n=5 up.Monte Carlo standard errors are in the last two columns, at 100,000
replicates. The operational point: a threshold placed where two adjacent n differ
by less than their MC error is a threshold placed on nothing. At n=3 the se on
p90 is 0.005 against a p90 of 1.810, so the resolution is fine everywhere in this
range — but it is worth having the number rather than assuming it.The shape, stated without proposing anything. The p90/p10 spread falls
steeply from 5.37× at n=3 to 2.44× at n=8, then slowly: 2.03× at 12, 1.71× at 20,
1.54× at 30. Most of the improvement available is spent by roughly n=8–10, and
after that it is a long shallow tail.Which functional should govern an adequacy band — p90/p10, RSD, a coverage
probability, something else — and where the cuts fall are yours and Tom's to
decide. I have deliberately not proposed any. Happy to compute where a criterion
lands once you have picked one, including quantiles I have not printed here.I will open the PR against
validation/keeping the filenames.- added 2 commits that reference this issue
on Sep 10, 2026 @rabujamra, Tom and I got together to review the question you raised. We generated a synthetic dataset based on your prior post.
file: aco_per_capita_expenditure.csvThere are 24 ACOs in the file, 96 obs in total.
How Tom read your data. He built VAS for analytic studies which include statistical process control applications where data usually arrive quickly and action follows quickly. SPC detection rules were originally designed for SPC applications. VAS and Process Behavior(PB) which is the software that implements the VAS methodology rests on the scientific method, the search for nonrandom patterns in the data and the nontrivial replication of results to build credibility, not on proscribed rules about how many points a chart needs, or what constitutes identification of an assignable cause. When only a few time periods are used it raises the level of uncertainty, but this does not make the analysis uninformative. Two things follow that I think are important for you to consider.
First, from Tom's perspective your data has a great deal of nontrivial replication. For example, formulated as one study, the same rising pattern over time is reproduced across nearly every organization. That agreement across independent subgroups is the finding, and it does not depend on the process behavior limits.
Second, the question Tom would put to you at the beginning of your study is whether this is an “enumerative study”, where action is taken on the frame that produced the data, or whether this is an “analytic study”, where you will be making predictions into the future and taking actions based on the predictions. If the latter, which is likely, the uncertainty associated with these predictions cannot be measured, or controlled, using probability theory because the predictions are extrapolations into the future. For probability theory to apply the assumption “the future behaves like the past” must be true in fact which in most analytic studies is not true. This is why the nontrivial replication is the confirmatory methodology to build credibility and not probability calculations that depend upon the precision of the standard deviation (sigma) behind the behavior limits. His book is largely about this distinction.Please look at the six charts attached below that are part of the output produced by Process Behavior which clearly demonstrate these principles.
Here is code to generate two charts. My guess from "476 of 476" is that you ran one four point study per organization. Run one study instead, with the organization as a factor and the year as time, and look at these in order:
Note: the chart will be busy with all the observations plotted so you might want to subset to 100. You can easily zoom in with the plotly controls.
import processbehavior as pb
study = pb.ProcessBehavior.read_csv("aco_per_capita_expenditure.csv").formulate(
response="PER CAPITA EXPENDITURE", factors=["ACO"], time="YEAR"
)
print(study.design()) # ADS 2 ... Series length: T=4. Sigma for X/mR rests on 3 moving ranges.1. Every organization on one individuals chart, one scale, lane by lane
(the first analysis in the app's sequence)
combined = study.execute(chart="X", by=[], companion=True)
combined.plot(chart="X")...versus one small chart per organization
per_aco = study.execute(chart="X", by=["ACO"], companion=True)
per_aco.plot(chart="X", facet=True, ncols=4)2. The R2 residual on an individuals chart: the year to year change itself
(the last analysis in the app's sequence)
r2 = study.execute(chart="X", value="R2", by=[], companion=True)
r2.plot(chart="X")Please let us know if you have any additional questions.
Thanks Chris — and thanks to you both for building the dataset, that made this concrete quickly.
I ran the combined and per-ACO formulations. Before I go further on Tom's substantive argument, there's one implementation point I wanted to check with you.
With by=[], the combined series has 96 observations and the mR chart reports 95 moving ranges. The within-ACO ranges the data can supply number 72 — sum over ACOs of (n_i − 1), which is N − K = 96 − 24. The remaining 23 correspond to the 23 transitions between successive ACO blocks, computed from the last year of one ACO to the first year of the next. For example, the reported mR at ACO-002/2021 is |11837.21 − 11500.14| = 337.07, which spans ACO-001 2024 → ACO-002 2021.
On this dataset that materially changes the calculation: mean mR is 1443.208 over all 95 ranges versus 933.792 over the 72 within-ACO ranges only. All three mR points currently flagged beyond the limits are first-year boundary ranges (ACO-019, ACO-021 and ACO-024, each at 2021).
Two things I checked so this isn't about input ordering. First, a copy of the file sorted by (YEAR, ACO) instead of (ACO, YEAR) produces identical output — ProcessBehavior establishes its own combined ordering, so this isn't a case of badly sorted input. Second, by=["ACO"] reports exactly 72 ranges and is unaffected. The attached script takes its boundary pairs from the library's own X-chart sequence rather than from the CSV, and reproduces the same 23 transitions and the same values on sorted, reordered, and deliberately unbalanced (2–4 years per ACO) copies.
This struck me as possibly related to #121 one layer down: the plotting change stops a line being drawn across a lane boundary, and the mR calculation may still difference across that same boundary.
Separately, on Tom's point — the distinction he draws between the short-series question and replication across organizations is the part I want to sit with, and I'll come back to it properly once I understand what the combined formulation is doing. It reframed the question for me more than it settled it.
Does that look like intended behavior? Happy to open a separate issue with the reproducible check if that's more useful than a comment here.
@rabujamra, thank you for this. It is a careful check and you have it exactly right: the combined chart differences across the transition from one ACO to the next, and on your data those 23 transitions lift the average moving range by about half. That is intended, and I want to give you the straight version of why, including the part I discovered rather than designed.
How the combined chart is built. With
by=[]the points are not ordered purely by time, the way a textbook individuals chart would be. They are ordered by ACO first and then by year, so each organization occupies its own lane at one common scale and the same pattern can be read lane by lane. I chose that ordering for the picture, because it is how Tom reads data. I will be honest that I did not think through what it meant for the moving range at the time. Your question made me look.What it does to the arithmetic. The moving range runs across the whole sequence in that order, so the step from one organization to the next contributes a range like any other. With many periods per subgroup those transitions are a handful among hundreds and barely move sigma. With four periods per subgroup they are a quarter of the ranges, and they dominate. So your combined limits are wide, and they are wide precisely because of the between-organization differences.
Tom's reference does the same thing. I checked against his Minitab output on the test database (PM SDS 2, 8 cells × 100 periods, 800 observations, 799 ranges, worksheet in cell order). Differencing consecutive rows gives 792 within-cell ranges and 7 ranges at the cell transitions. Including the 7 transitions, the mR center is 0.868 and the individuals limits are 235.47 and 240.09, which match his slide to the last digit. Excluding them, the center would be 0.841 and the limits would move to 235.54 and 240.02, which do not match. The library reproduces his reference only because it keeps the transitions.
Why the transitions are not artefacts. The combined chart's job is to ask whether these subgroups belong to one process. A jump at the transition between two organizations is evidence on exactly that question. It is not an artefact of sort order; it is the level difference between two organizations, observed. On your data the chart is saying, loudly, that 24 organizations are not one process. That is the same finding Tom pointed you to from the pattern side: the rising shape replicates in nearly every lane, while the levels do not agree. The three flagged ranges are the three largest level shifts between neighbouring organizations. They are signals, and they arrived before you had asked the question.
You are right that this sits one layer below #121. That change was visual only: the line no longer connects across a lane boundary, so the eye is not led to read the jump as a within-organization movement. The range itself stays in the calculation, deliberately.
Three views of the same study. The library gives you the question you want to ask, rather than one chart with one set of limits.
- The process as a whole. Every ACO, four observations each, one set of limits across the whole frame. This is the chart you ran, and it asks "is this one process?"
combined = study.execute(chart="X", by=[], companion=True) combined.plot(chart="X")
- The same lanes, each with its own center line and limits. This asks "how does each organization behave on its own, side by side at one scale?" The moving range within each phase is what sets that phase's limits, so the transitions no longer enter.
phased = study.execute(chart="X", by=[], phased=True, companion=True) phased.plot(chart="X")
- Drilling in, one chart per organization. This is the 72-range version you already confirmed.
per_aco = study.execute(chart="X", by=["ACO"], companion=True) per_aco.plot(chart="X", facet=True, ncols=4)
No separate issue needed; this comment is the record, and I will add a note to the docs so the next reader finds it without having to reverse-engineer the sequence as you did. Your script was a clean way to establish it, and I appreciate the care.
On Tom's larger point, take the time you need. That reframing is the part of this that matters most, and I would rather hear your considered reaction than a quick one.
Chris — thank you for this, and for going and checking against Tom's reference rather than just answering me. The "part I discovered rather than designed" line is a generous way to put it and I appreciate it.
I took the ordering question seriously enough to run it on the 24-ACO file rather than reason about it abstractly. Script attached:
mr_permutation_invariance.py, one run, about a second. It holds every observation fixed and varies only lane order.Your arithmetic checks exactly.
24 organisations x 4 obs = 96 obs, 95 moving ranges within-ACO 72 ranges mean 933.79 <- invariant to lane order boundaries 23 ranges mean 3037.90 <- varies with lane adjacency boundary/within mean ratio 3.25x mRbar as loaded 1443.21 mRbar within-only 933.79 increase from including boundaries 54.6%"About half" was right.
The permutation moved my reading toward yours, not away from it.
Across 20,000 random lane orderings the X chart never once returned zero signals — the count stayed between 3 and 8, and zero occurred in 0.00% of orderings. So the substantive conclusion, that these 24 organisations are not behaving as one process, is robust to lane ordering on this dataset. That is evidence for your reading of the combined chart and I would rather report it than bury it.
What is not invariant is the boundary-signal structure.
ordering mRbar X-limit width X signals mR signals as loaded 1443.2 7677.9 4 3 sorted by level 1272.5 6769.9 7 0 interleaved 1497.6 7967.4 3 6 20k random mean 1413.4 7519.4 3.92 4.48 range 1269-1545 6753-8221 3-7 1-9Using D4 = 3.267 for the mR upper limit, the flagged boundary transitions are:
as loaded ACO-018|ACO-019, ACO-020|ACO-021, ACO-023|ACO-024 sorted by level (none) interleaved ACO-018|ACO-024, ACO-023|ACO-007, ACO-009|ACO-006, ACO-002|ACO-022, ACO-010|ACO-005, ACO-020|ACO-008Worth checking those against what
companion=Trueactually flags in your own output — if your mR rule differs from the classical D4 the labels will move, though the spread across orderings should not.The observations and the 72 within-ACO ranges never change. One thing worth saying plainly: the shipped ACO-001..024 order sits at the 77th percentile of random orderings by mRbar. It is not adversarial. It is arbitrary, which is the part that interests me.
And a boundary range may not be measuring quite what it looks like.
A boundary moving range changes organisation and resets time in the same number: it steps from organisation i's 2024 observation to organisation j's 2021 observation. Every one of the 24 lanes rises here.
mean within-lane first-to-last change 2586.9 mean boundary |last_i - first_j| 3037.9 mean adjacent pure-level gap |mean_i-mean_j| 1833.9 boundary / pure-level ratio 1.66x difference from pure-level comparison 39.6% of mean boundary magnitudeThat comparison is descriptive. I am not claiming an additive decomposition — with absolute values and heterogeneous trajectories the two pieces do not cleanly separate. The point is only that the boundary construction differs systematically from a pure between-organisation level comparison.
Where that leaves me.
The narrow statement I think the data supports: the overall conclusion is robust to lane ordering; the number and identity of the boundary signals are conditional on the traversal. The boundary signals look like evidence along the chosen path rather than invariant pairwise properties of the organisations.
Which leaves me with a question rather than a complaint: how should a boundary signal be read? If it is a claim about the traversal, then the
phased=Trueview you pointed me to is the right default for anyone asking about individual organisations, and the docs note you are planning would cover it. If it is meant as a claim about the organisations themselves, the time reset above is doing some of the work.Happy to run anything else that would be useful, and happy to be told I have mis-framed it.
On Tom's larger point — I am taking you at your word and giving it the time it deserves. I will come back on the enumerative/analytic question separately, because I do not think it is the same conversation as this one and I would rather not blur them.
@rabujamra, thank you. I ran
mr_permutation_invariance.pyagainst the file and every number matches, including the 77th percentile. The library's mR limit is the classical D4, andcompanion=Trueflags exactly the three transitions you list for the shipped order, at the first-year points of ACO-019, ACO-021 and ACO-024, plus the same four X points.Your narrow statement is the right one, and I would put it the same way. The combined chart's conclusion survives any lane order, which is the question that chart exists to answer. The identity of a boundary signal does not, because it is a claim about the traversal: the two lanes, in that order, do not join into one stable stream. And you are right that on a rising process the boundary step carries the trend reversal along with the level gap, so it is systematically larger than a pure level comparison. That is a property of the lane-major sequence on trending data rather than of the two organisations.
So the answer to your question is that a boundary signal should be read as evidence along the path, and claims about the organisations should come from the charts that have no path. Two of them are already there:
- Which organisations differ. The recentred R5 chart by organisation, an Xbar chart of each organisation's level with limits set from R2, flags 8 of the 24. The points are order-free, since each is an organisation's mean. The limits are not entirely, because in ADS 2 they are set from R2, and R2 is built from the same lane-major sequence, so one of every lane's four R2 values is a boundary half-difference. Applying your method, across 200 random lane orders the limit width ranged from 2,921 to 3,428; the eight flagged under the shipped order were flagged under every order, with one or two marginal organisations joining in a minority of them. (Edited: my first version of this paragraph said the set was invariant to ordering. That was too strong.)
levels = study.execute(chart="Xbar", by=["ACO"], value="R5", recentered=True) levels.plot()
- How each organisation behaves. The phased view, which you already have, with each lane's limits from its own ranges.
The docs note I owe you will say exactly this: the combined chart answers "one process?"; boundary signals are path evidence; organisation-level claims come from the R5 chart and the phased view.
Your script belongs in the repo beside the other two. Please open a PR against
validation/, same as #119, keeping the filename, so it lands with you as the author.On the enumerative versus analytic question, agreed, it is a separate conversation, and it deserves its own thread when you are ready.
@rabujamra, the docs note has landed, and #130 is merged with your script and the ACO file beside it under
validation/.The chart-types guide now has three sections that come directly from this thread: Three Views of the Same Study, which gives the combined, phased and stratified views with the question each one answers; Reading a Boundary Signal, which says that a signal at a transition is evidence about the traversal and explains why the transitions stay in the calculation; and Which Claims Are Order-Free, a table of what survives reordering the lanes, including the R5 limits caveat you already folded into your script's scope note. It is also the first time the phased view is documented in the user guide, which it should have been from the start.
The validation README lists all three of your scripts with their PR numbers, and the permutation check runs from the repo root against the committed file.
Thank you for pushing on this until the claims were stated at the right strength. Closing this one; the enumerative-versus-analytic thread is yours to open whenever you're ready.
Chris and I were discussing short-series behaviour and he pointed me toward the repo. Two findings from pushing on very small T, and a question I'd genuinely value Tom's view on.
processbehavior 0.2.0from PyPI, pandas 2.3.3, Python 3.11.15. Docs read: quickstart, coffee shop, design-state lineage, API reference. Everything below is reproducible — seed 7 for the synthetic series,default_rng(20260903)with 20,000 replicates for the sampling study, 400,000 for the null distribution in §2, and the CMS PY2024 Shared Savings public use file for the real series. Happy to contribute the scripts.A related observation about signal reporting at small T touches the "results an analyst would mistake for correct output" language in SECURITY.md, so I am sending that one through the private advisory channel rather than including it here.
1. The design state is constant in T; the estimator is not
The lineage classifies replication structure — cell sizes N_kt, complete versus incomplete grid. Series length is a separate property, and ADS 2 is ADS 2 whether T = 4 or T = 400.
A stable simulated process, mean 5.50, sd 1.20 — true natural process limits span a width of 7.20. Feeding the first T points:
The design state, the recommendation and the analysis menu are constant down the whole column.
The sampling distribution behind that. MR̄/d₂ under a stable normal process with σ fixed at 1.0, 20,000 replicates per length:
At four points the estimate spans a near-fourfold range between its 10th and 90th percentile on a process whose variation never moved.
The design report at T = 4:
T: 4is computed and printed. Nothing downstream consumes it. "Complete structure" is correct on the replication axis, and Capability and Loss Function are offered from four points.2. Why I was at this end of the range, and what the real series did
My series are Medicare shared-savings settlements. Each unit is an accountable care organisation, each observation is an annual figure, and I get three or four points per unit. They also trend by construction, because medical costs rise and the benchmark years are built to reflect that.
476 ACOs, per-capita expenditure across three benchmark years plus the performance year:
MR̄ here is largely estimating the slope of medical cost growth rather than variation around a level. That is a property of the estimator at this length rather than anything the library does wrong — but it means the number the library correctly computes is not the number the analyst thinks they asked for.
And a drift test cannot help, for a reason worth stating. With three differences, an independent series is monotone with probability 2/4! = 1/12 ≈ 8.3%. That is above 5%, so the 95th percentile of the null distribution for drift share is 1.000 — the maximum attainable value. No four-point series can exceed it. Simulation at 400,000 replicates gives 8.3%, matching 1/12. The 95th percentile falls to 0.77 at T = 5 and 0.59 at T = 6, so this is specific to the shortest series rather than general.
So the two failure modes point in opposite directions — narrow limits and false alarms on a short stable series, wide limits and silence on a short trending one — and in both cases the design state reads as healthy.
3. A caveat on my own test
Those four points are not four observations of one process. Benchmark years are constructed inputs to a contractual benchmark, risk-adjusted separately and partially restated. A purist would say I fed the package something it should never have been given — and that objection makes the series worse than short, not better.
It was accepted without comment, and a user without the domain background would not know they had done anything wrong.
4. The question
This is the part I'd most value Tom's view on, and it is not rhetorical.
I've assumed the latter and gated my own work accordingly, but that is a judgement, not a result. Wheeler's own guidance is that useful limits can come from 5–6 points and firm up by 15–20, so a reasonable person could argue the estimate should be returned with its 3.94× interval attached and the reader left to discount it. My objection is that a number with a 3.94× interval gets quoted without the interval — which is an argument about institutions, not statistics.
Happy to move this to a Discussion if that is the better venue for a methodology question.
5. A suggestion, offered tentatively
Less "add a concept" than "connect two things already present":
Tis computed and printed in the design report, andSignalConfig.min_observationsalready exists with a default of 20.recommended_chart = Nonewith a stated reason, in the style of the existingads_reasonThe reason for putting it inside the design state rather than beside it: the package's distinguishing claim is that structure determines admissibility before anything is computed. Series length is a structural property, and it currently sits outside the gate everything else passes through.
6. What I liked, said plainly
The
formulate()/execute()separation is the best decision in the API. Making "what does this structure permit" a first-class inspectable object, before computation, is the right decomposition — and it is precisely why the T question stood out, because the architecture is already shaped to hold that kind of constraint.The ODS → ADS distinction is subtle and correct. Detecting incompleteness on raw data and then reporting what survives cleansing is careful work that most tools would collapse into one number.
Validating against Bishop's published reference results on every commit is the standard I would most want other analytical libraries to copy. Most demonstrate internal consistency; checking against an external authority is harder and different.