Skip to content

US PUF imputation predicts each income item from the recipient's own survey value, so the PUF clone half reproduces ASEC income #982

Description

@MaxGhenis

Summary

The PUF tax-detail imputation predicts each major income item from the recipient's own survey value of that same item. The model therefore learns "output ≈ input", and the PUF-clone half of the pool comes out as ASEC income with PUF-style variation around it. This caps the top of the income distribution at ASEC topcodes and carries survey underreporting into the half of the pool that is supposed to carry tax-return income.

Found while diagnosing #958. Read in code at origin/main 2b85b7b22 and measured on the 2026-09-12 base-build checkpoints.

Mechanism (read in code)

  • The predictor set is filing status, tax-unit size, and six income items: wages, self-employment, taxable interest, dividends, short-term gains, long-term gains (us_runtime/puf_support.py:199-208).
  • The same items are imputation targets (PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS, puf_support.py:210).
  • On the donor side, each predictor column is a copy of the donor's own PUF amount (_add_predictor_aliases on the donor tax units, puf_support.py:1664 and :1769, via _predictor_source_column). So in training, the wage predictor and the wage target are the same number.
  • On the recipient side, the predictors are the household's survey amounts (_strict_recipient_predictor_surface, puf_support.py:2693; ASEC wages come from WSAL_VAL, cps_carried.py:172).

Measured: PUF clone vs its own ASEC parent

432,523 PUF-clone persons compared with their clone-0 parent, after 004_qrf_finalization. Before the QRF (002_clone_feature_extraction) every clone value equals its parent exactly; after it, none of the positive values are exact copies, but they stay next to the survey value.

Item Zero/nonzero status agrees with ASEC Rank correlation with ASEC Within 10% of the ASEC value, among ASEC > 0
Wages 100.0% 1.000 100.0%
Self-employment 99.8% 0.979 97.6%
Taxable interest 100.0% 0.999 99.9%
Dividends 99.0% 0.961 82.1%
Long-term gains 100.0% 1.000 90.7%
Short-term gains 100.0% 1.000 78.8%

Consequences:

  1. The top is capped at survey topcodes. No clone's imputed wages exceed the ASEC wage maximum ($2,099,999, the WSAL_VAL topcode) by more than a few percent, while the processed PUF has about 25,000 weighted returns above $10M of income. This is the root cause behind the QRF row of the stage table in US top-income tail is missing above $5M of AGI: Table 1.1 size-of-AGI facts never bind, and the pool's tail stratum carries capital gains only #958 (18 of 231,007 clones at or above $5M).
  2. Survey underreporting carries into the PUF half. A household that reports no dividends, interest or capital gains in the ASEC gets none in its PUF clone 99-100% of the time. Calibration then has to hit SOI income totals with records whose incomes keep the survey's shape. The over-weighting of $200k-$2M returns measured in US top-income tail is missing above $5M of AGI: Table 1.1 size-of-AGI facts never bind, and the pool's tail stratum carries capital gains only #958 is consistent with this, but that link is a hypothesis, not measured.

History

  • The archived eCPS imputed PUF income onto its CPS clone from demographic predictors only: age, sex, joint filing, dependent count, and tax-unit roles (policyengine-us-data, calibration/puf_impute.py, DEMOGRAPHIC_PREDICTORS), with a stratified PUF subsample that preserves the top percentile.
  • Microcosm's first PUF commit (63b818730, 2026-06-19) introduced the income predictors. I found no design note in DESIGN.md or docs/ explaining the change.

Status elsewhere

Fix options, for decision

  1. Demographic predictors, as in the eCPS. The PUF half then carries the PUF's own joint income distribution attached to survey demographics.
  2. Rank bridge. Condition on the household's survey income rank instead of income levels, and draw the PUF value at the matching rank, so top-ranked survey households can receive PUF-scale values (compare the UK rank-preserving approach in UK spine: composition-faithful SPI income draw — bridge predictors and rank-preserving replacement (with-children UC residual) #840).
  3. Drop each target from its own predictors while keeping the other income items. Smaller change, but the other items are highly correlated with it, so it may not break the identity by much.

Each option changes the whole PUF half and needs a base build and a scorecard comparison against the current pool. Option 1 is the known design from the eCPS and is the natural first comparison.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions