You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
The PUF tax-detail imputation predicts each major income item from the recipient's own survey value of that same item. The model therefore learns "output ≈ input", and the PUF-clone half of the pool comes out as ASEC income with PUF-style variation around it. This caps the top of the income distribution at ASEC topcodes and carries survey underreporting into the half of the pool that is supposed to carry tax-return income.
Found while diagnosing #958. Read in code at origin/main2b85b7b22 and measured on the 2026-09-12 base-build checkpoints.
Mechanism (read in code)
The predictor set is filing status, tax-unit size, and six income items: wages, self-employment, taxable interest, dividends, short-term gains, long-term gains (us_runtime/puf_support.py:199-208).
The same items are imputation targets (PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS, puf_support.py:210).
On the donor side, each predictor column is a copy of the donor's own PUF amount (_add_predictor_aliases on the donor tax units, puf_support.py:1664 and :1769, via _predictor_source_column). So in training, the wage predictor and the wage target are the same number.
On the recipient side, the predictors are the household's survey amounts (_strict_recipient_predictor_surface, puf_support.py:2693; ASEC wages come from WSAL_VAL, cps_carried.py:172).
Measured: PUF clone vs its own ASEC parent
432,523 PUF-clone persons compared with their clone-0 parent, after 004_qrf_finalization. Before the QRF (002_clone_feature_extraction) every clone value equals its parent exactly; after it, none of the positive values are exact copies, but they stay next to the survey value.
The archived eCPS imputed PUF income onto its CPS clone from demographic predictors only: age, sex, joint filing, dependent count, and tax-unit roles (policyengine-us-data, calibration/puf_impute.py, DEMOGRAPHIC_PREDICTORS), with a stratified PUF subsample that preserves the top percentile.
Microcosm's first PUF commit (63b818730, 2026-06-19) introduced the income predictors. I found no design note in DESIGN.md or docs/ explaining the change.
Status elsewhere
The graph-native path does not change this. The native branches (e.g. 516e6e8fe, b2d7a57c1) reuse the same tuple in full_puf_enrichment.py (PREDICTORS = tuple(support.PUF_TAX_DETAIL_DEFAULT_PREDICTORS); the PUF59 variant keeps the six income predictors). Their native_puf_tail.py ports the Carry PUF high-AGI donors into the own-tail stratum with their full income vector #964 selection, bounded to invented fixtures.
Drop each target from its own predictors while keeping the other income items. Smaller change, but the other items are highly correlated with it, so it may not break the identity by much.
Each option changes the whole PUF half and needs a base build and a scorecard comparison against the current pool. Option 1 is the known design from the eCPS and is the natural first comparison.
Summary
The PUF tax-detail imputation predicts each major income item from the recipient's own survey value of that same item. The model therefore learns "output ≈ input", and the PUF-clone half of the pool comes out as ASEC income with PUF-style variation around it. This caps the top of the income distribution at ASEC topcodes and carries survey underreporting into the half of the pool that is supposed to carry tax-return income.
Found while diagnosing #958. Read in code at
origin/main2b85b7b22and measured on the 2026-09-12 base-build checkpoints.Mechanism (read in code)
us_runtime/puf_support.py:199-208).PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS,puf_support.py:210)._add_predictor_aliaseson the donor tax units,puf_support.py:1664and:1769, via_predictor_source_column). So in training, the wage predictor and the wage target are the same number._strict_recipient_predictor_surface,puf_support.py:2693; ASEC wages come fromWSAL_VAL,cps_carried.py:172).Measured: PUF clone vs its own ASEC parent
432,523 PUF-clone persons compared with their clone-0 parent, after
004_qrf_finalization. Before the QRF (002_clone_feature_extraction) every clone value equals its parent exactly; after it, none of the positive values are exact copies, but they stay next to the survey value.Consequences:
WSAL_VALtopcode) by more than a few percent, while the processed PUF has about 25,000 weighted returns above $10M of income. This is the root cause behind the QRF row of the stage table in US top-income tail is missing above $5M of AGI: Table 1.1 size-of-AGI facts never bind, and the pool's tail stratum carries capital gains only #958 (18 of 231,007 clones at or above $5M).History
policyengine-us-data,calibration/puf_impute.py,DEMOGRAPHIC_PREDICTORS), with a stratified PUF subsample that preserves the top percentile.63b818730, 2026-06-19) introduced the income predictors. I found no design note inDESIGN.mdordocs/explaining the change.Status elsewhere
516e6e8fe,b2d7a57c1) reuse the same tuple infull_puf_enrichment.py(PREDICTORS = tuple(support.PUF_TAX_DETAIL_DEFAULT_PREDICTORS); the PUF59 variant keeps the six income predictors). Theirnative_puf_tail.pyports the Carry PUF high-AGI donors into the own-tail stratum with their full income vector #964 selection, bounded to invented fixtures.Fix options, for decision
Each option changes the whole PUF half and needs a base build and a scorecard comparison against the current pool. Option 1 is the known design from the eCPS and is the natural first comparison.