Skip to content

Step 2: the value-added block for 2017-2024, all five rows (#538) - #740

Merged
WesIngwersen merged 9 commits into
nowcastfrom
step2_va_rebuilt
Aug 27, 2026
Merged

Step 2: the value-added block for 2017-2024, all five rows (#538)#740
WesIngwersen merged 9 commits into
nowcastfrom
step2_va_rebuilt

Conversation

@WesIngwersen

Copy link
Copy Markdown
Member

Replaces #735. Same Step 2 work, rebuilt onto current nowcast so it carries only the value-added block.

Do not delete #735's branch — see the last section.

Why rebuild instead of resolving #735's conflicts

#735 was never a stack; it sat on two of them. Its history runs through WIP merges of both nipa_va_othertax_538 (#693) and step3_underlying_ii_va:

825389d0  WIP: merge nipa_va_othertax_538 …
ce819e0c  WIP: merge step3_underlying_ii_va …

#693 has since been squash-merged, so most of #735's conflicts were git re-presenting content nowcast already has. Resolving them would have been busywork, and the result would still have dragged the step3 half into nowcast — the opposite of "leaving just the tall stack of intermediate use".

What this is

Six commits cherry-picked onto nowcast in topological order:

commit
3461fe63 Re-evaluate Step 2's 2018-2024 approach: value added is now an observation
6911d588 Record the two Step 2 decisions, and measure the residual's headroom
7f5f016e Grade the QCEW movement series on the 2012→2017 holdout
fa727781 Build V00100 for 2018-2024 on QCEW movement
2201027c Finish the Step 2 VA block for 2018-2024
c679b602 Build T00TOP and T00SUB for 2017-2024

40 files / +9,639, down from #735's 64 files / +16,910. Verified to touch none of the contested step3 files:

nowcast_intermediate.py   unchanged
inputs_structure.py       unchanged
nowcast_sut_gras.py       unchanged
nowcast_targets.py        unchanged

Conflicts resolved

test_nowcast_product_taxes.py#735 independently diagnosed the same #734 regression that #737/#738 fixed here, and reached the same 2024 share (0.073). Its tobacco figure was -26,282 against this lineage's CI-verified -26,278; both sit inside the ±100 band. Kept -26,278 (verified green on this exact lineage) and the grouped asserts, and folded in #735's diagnostic that the 2020 numbers did not move at all — the 2022-vintage leaves reach no earlier FBA, which is the control proving it was the Crosswalk and not the method.

sections.py — took #735's note, which describes what this PR actually delivers: the five rows are not five claims of one kind (V00100 an estimate, T00OTOP a level plus lookups, V00300 a seed, T00TOP/T00SUB converted from the Supply columns).

README.mdnowcast already had two of the three commands from #693; added only the new value_added_timeseries entry.

Two commits deliberately left out

  • 12232f28 "Format sections.py after the two-stack merge" — an artifact of the tangle. Re-ran black instead.
  • 461563e0 "Pin VAPRO per industry as T18, so income-side error stops reaching A" — this touches nowcast_intermediate.py, nowcast_sut_gras.py and nowcast_targets.py. It consumes value added as a GRAS target rather than building the block, and it lands in the middle of the files the step3 stack is still rewriting. It belongs with that stack, not here.

⚠️ 461563e0 exists only on origin/step2_va_timeseries. Closing #735 is fine; deleting its branch would lose that commit. It needs replaying onto the step3/GRAS work before the branch goes anywhere.

WesIngwersen and others added 9 commits August 27, 2026 08:39
…ation

The plan for the value-added time series was 2017 detail shares carried on QCEW
wage growth, renormalised inside T60200D's 69 groups and rescaled to a NIPA
control. Two sources landed after it was written and they change the estimand,
not just an input: UVA205-A gives VAPRO as a column total per detail industry
annually (#712), and TVA113 splits that same VAPRO three ways at 71 summary
industries annually (#538).

value_added_timeseries.py measures the four things that decision turns on.

They nest. The yaml warns the two release archives are different vintages, so
this was checked: rolled to summary, the detail VAPRO panel and TVA113 agree to
$9M in the worst of 923 industry-years. Both margins of a 3 x n block are
therefore observed, and Step 2 estimates a cross-structure rather than a level --
the same reframing Step 3 got from #497. Rolling up needs the INDUSTRY summary
map; the commodity one drops 331314, S00101, S00201 and S00202 and leaves state
and local government enterprises 15% short in every year.

Most of it is already determined. 20 of the 71 summary industries have one
detail child (19.4% of VA, no allocator at all), and 74 of the 138 underlying
leaf lines ARE a single detail industry (63.9% of VA, BEA's own annual number).
Only 26.9% sits in the 17 groups where a within-group allocator earns its keep.

The column control is worth 0.45% of VA in 2018 rising to 2.32% -- $678bn -- in
2024, graded on BEA's own leaf lines so no allocation model enters the
measurement; 3.75% of the $18.1T where the movement is measurable at all, 81%
of it in ten groups led by other retail at 11.7%. A lower bound, and it matters
past its size because Step 5 imposes GO = T005 + VAPRO hard, so value-added
error lands in the intermediate block.

T00TOP's industry axis turns out to be published. #536's rejection of the
market-share conversion (r=0.202) stands, but its conclusion does not: TVA113's
V00200 is T00OTOP + T00TOP - T00SUB by industry, max $1M across all 71, so with
T00OTOP built the product-tax row is an annual summary residual with no operator.

QCEW is still needed for the ~80% of compensation in multi-child groups, but it
no longer carries the level, and it can now be graded on the 2012->2017 detail
holdout instead of BEA's own later years -- which needs QCEW 2012, undeclared
until #728.

Left open for review: whether TVA113 becomes an input, spending summary
V001/V003 as a test. Two arguments that the trade is smaller than #538 thought
are recorded in the plan.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
(cherry picked from commit 3461fe6)
TVA113 is NOT consumed as an input. The row controls stay NIPA's, so summary
V001/V002/V003 keep their ability to grade the build. That withdraws two claims
the first pass made: T00TOP does NOT get a build route from TVA113's V00200
identity (#536's conclusion stands unchanged, the split stays free for Step 5),
and T60200D's 69 groups are NOT superseded. What the identity buys instead is a
grader -- it says where the product-tax money should have landed, so Step 5's
answer can be scored rather than trusted.

The VAPRO column control is a separate decision and it is kept. UVA205-A is the
sibling of UGO305-A, which is already T1.

The slack goes to V00300, never to T005. #710 checked that T1 pins the column's
sum and not the split, and that T5 is deliberately unimposed -- so today
income-side error lands in the intermediate column total, which is the scale of
a column of A, and propagates through L into every N. Pinning VAPRO makes
T005 = GO - VAPRO, both observed, and moves the slack into value added.

residual_headroom() prices that, and states the counterargument rather than
burying it: in percentage terms T005 is the MORE forgiving absorber, because it
is the bigger number -- a 1% compensation error is a median 0.50% of an
industry's T005 against 1.55% of its V00300. It loses because V00300 is
terminal. Nothing reads gross operating surplus: not A, not L, not any emission
factor. A 7% error in a number nothing reads beats a 1.4% error in a column
scale that multiplies through the Leontief inverse.

22 industries cannot absorb much and need a sign guard rather than trust: a 1%
compensation error moves their surplus more than 10%, worst 336414 at 121% on
an $81M surplus under $9.8B of compensation. Only one industry has a published
negative V00300 (S00201, -36,919), so a residual manufacturing new negatives is
visibly wrong and cheap to detect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
(cherry picked from commit 6911d58)
…e carve-out

The test the compensation plan could never run. Two things unblocked it:
SUPPLY-USE_2026-08-24.zip carries V00100 at BEA detail for 2007/2012/2017 on one
2017 code basis (#704), and QCEW 2012 now exists (#728). Scored the way #704
requires -- on the observed benchmark span, never against BEA's carried-forward
2018-2024 -- as the share of a group's compensation dollars on the wrong detail
industry, with the group total given to every candidate so this measures shape
and not level.

                   misplaced $M   % scored   vs frozen
  frozen                 487,348      5.84         --
  qcew                   517,328      6.20      +6.2%
  qcew_covered           462,044      5.54      -5.2%
  qcew_resolvable        438,534      5.26     -10.0%

Applied everywhere QCEW makes the block WORSE. Applied where the concordance can
resolve it, it is a clear go. The damage is essentially one group: GSLG goes
from 4,748 misplaced to 71,694, on its own more than twice the net degradation.

The carve-out is derived, not fitted. The crosswalk puts 47 NAICS codes under
more than one BEA detail industry -- 23 construction, 24 government -- reaching
18 detail industries in exactly five summary groups: 23, GFE, GFGN, GSLE, GSLG.
Decidable before any score is computed, and it is the carve-out Phase 4 argued
for structurally: BEA splits construction by type of structure and NAICS by
trade, so a plumbing contractor's payroll belongs to no single BEA construction
industry.

The coverage ratio is REJECTED as the selector. Trusting QCEW where payroll
covers >=75% of compensation scores -5.2%, but the floor is non-monotonic:
-9.7% at 0.70, -5.2% at 0.75, -3.5% at 0.80, +0.0% at 0.90. Between 0.70 and
0.75 it drops the five groups where QCEW helps most. Coverage and predictive
value are different quantities.

Two vintage traps had to be cleared and each faked a verdict. QCEW 2012 is on
NAICS 2012 against a NAICS 2017 crosswalk -- 28 codes in 2012 not 2017, 20 the
other way -- which put 541700 at 20.9x growth and made raw QCEW score +92%
before the filter. And the ambiguous concordance would not have raised; it would
have produced an even split and called it an answer, exactly as #536 warned.
Neither is visible in a total.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
(cherry picked from commit 7f5f016)
The plan's design was "2017 detail share x QCEW growth, renormalised
in-parent, then the NIPA control", blocked on FBS_outside_flowsa not
working as an attribution source. It is bypassed rather than cleared: a
weight vector that is a rescaling of an FBA already in the method needs
no new source, so a clean_fba socket on the existing BEA_Detail_Use_SUT
attribution source scales the 2017 benchmark weights by QCEW growth
before sector mapping. The blocker is diagnosed to two lines on #731,
where T00OTOP's housing and farm lookups still want it.

All eight years build. 2017 still reproduces the published benchmark row
to $1M, which is the check that wiring the movement in did not disturb
the identity case; 206 industries move in 2018 and $66.8bn changes hands
against frozen shares, 0.61% of the row.

Pointing proportional straight at the QCEW FBA would have been the
trivial build, and it is rejected on measurement: QCEW level shares score
+21.9% worse than frozen at their best and +270.9% raw, against -10.0%
for the movement form. QCEW's share of an industry's compensation is not
that industry's share of compensation, but its change carries signal.

Four corrections, each of which changes the answer, and none of which is
visible in a total:

- The concordance cannot resolve construction or government. 47 NAICS
  codes sit under more than one BEA detail industry and reach exactly
  five summary groups. Applied everywhere QCEW is +6.2%; carved out it is
  -10.0%. Derived from the crosswalk, not fitted.
- QCEW changes NAICS vintage at data year 2022, not 2023. Unbridged it
  costs 31 detail industries their coverage and doubles unmapped payroll
  from 10.4% to 20.7%. The vintage is now detected from the concordance
  and the two years are paired through it, so composition is identical
  across the ratio by construction.
- 482000 rail is outside QCEW's universe -- railroad employees fall under
  the Railroad Retirement Board, not state UI -- so QCEW sees 0.11% of
  its compensation and rail "grows" 4.45x by 2024. A 1% coverage guard
  catches it. The holdout is neutral on that guard (-10.016%, identical),
  because the 2012->2017 span does not contain the failure; it is kept on
  an argument about the source rather than the score, and tuning it
  upward makes the block worse.
- A published zero is a suppression. QCEW reports exactly 0.0 for both
  NAICS under 334610 in 2021 with normal payroll either side; read as an
  observation it zeroes the weight and deletes the industry. One
  occurrence in seven years, fatal each time, so --check sweeps all of
  them.

Also generates the BEA_NIPA FBAs for 2018-2023, which existed without
T60200D and so produced no activity sets at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
(cherry picked from commit fa72778)
…kups

V00100 shipped on QCEW movement last commit; T00OTOP and V00300 were still
2017-only, so derive_initial_value_added raised for every other year. Both
now have per-year files and the guard is lifted: the block builds for
2017-2024, VABAS 18.92tn to 28.29tn.

The three rows are three different claims and the docstring now says so
rather than calling the block "value added, nowcast":

- T00OTOP is a level plus two lookups. The T30500 control is read per year
  (+40.5%), and 43.3% of the row is no longer a frozen share.
- V00300 is a seed and only a seed. Level from the eight-line assembly per
  year (+57.4%), shares frozen at 2017 where drift reaches 12.51% by 2022 -
  six times T00OTOP's. That is acceptable only because T18 makes it the
  residual the balance overwrites, and TVA113 must not "fix" it: that is the
  grader.

T00OTOP's housing and farm lookups are in, and they should have been in the
first time. T70405 B1031C is the 531HSO+531HST pair and T70305 B1017C is the
ten farm codes, both exact, both published annually, together 43.3% of the
row. They were left out because folding a computed weight vector in was read
as needing an FBS_outside_flowsa attribution source, which does not work.
That hatch is still broken and was never on the path: a weight vector that is
a rescaling of an FBA already in the method is a clean_fba socket, which is
how V00100 carries QCEW. Both rows were blocked on a hatch neither needed.
Worth 1.92% to 1.68% of row error in 2024, 0.12-0.39pp across the span -
real, bounded, and smaller than the control's own 2.9% vintage error in 2021.

Two things measured on the way that are not about these rows:

- The summary SUT is a stale grader for exactly 2019-2022. Its own VAPRO
  total sits 0.09-1.21% below current-vintage UVA205-A in those four years
  and matches to the dollar in 2017, 2018, 2023 and 2024. So V00300's
  apparent 2.64% assembly error in 2022 is the workbook being behind, not the
  assembly being wrong - and V00300 shows it at twice VAPRO's rate because it
  is the row a NIPA revision lands in.
- Selecting NIPA controls by line number is only safe because the lines do
  not move. Verified for all eight years and now asserted under --check:
  T30500 line 37 is LA000365 and line 17 is LA000237 in every one, and the
  eight V00300 lines are equally stable. A restructuring would point these at
  different series and no total would look wrong.

Also withdraws stale text in five places saying the T00TOP/T00SUB industry
split is "an output of Step 5's balance, not an input to it".
tax_axis_conversion.py takes that apart: rejecting the Make matrix (r = 0.202)
is not the same as having no operator, and Step 4c's producer/trade level
split plus a few named routings reaches r = 0.948. The split has no industry
target and is still seeded, so it is an input to the balance as well as an
output. The seed itself is not built here - it reads commodity output and
margins, both mid-update.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
(cherry picked from commit 2201027)
…ded block

The last two value-added rows, and the only two that are not estimated: they
are the same money as the Supply table's TOP/MDTY/SUB columns, carried on the
industry axis. So `nowcast_va_taxes` converts rather than re-estimates, and its
whole content is the operator. Both rows are stacked into
`derive_initial_value_added`, which now returns all five rows for 2017-2024.

The two rows needed two different operators, and only one of them is a seed.

T00TOP - the level split, at r = 0.947 and 27.9% error. Market shares fail this
badly (r = 0.202) and not by being noisy: a tax on a product is remitted by
whoever sells it, so they send the whole petroleum tax from wholesalers to
refineries. `top_by_level` already draws the producer-versus-seller line Step 4c
needed for its own reasons. Two of the three legs are exact rather than
estimated - duties are a lookup onto 4200ID, and the ten government columns are
zero by an accounting rule, which the market-share leg was violating by 10,513.
Still a seed; the residual is 20 named trade industries Step 5 moves.

T00SUB - not a seed. A subsidy is paid to an operator, so it stays on its own
code, and code identity's entire residual is two pairs of cells rather than a
smear: S00203 public housing authorities and S00102 federal insurance
enterprises, both government enterprises that produce a subsidised commodity.
Routing those two by name closes 2017 to shape agreement of 7e-17. What is
frozen is one number and it is named - S00203's 54.4% of the housing line, which
NIPA T31300 does not split.

Three things found while building it, all now asserted:

- The routings must not fire in 2020-21, where `sub_decomposition` replaces the
  `other` type with PPP. Routing 5241XX then would put pandemic support on a
  federal enterprise. PPP is already industry-shaped, so identity is at its best
  in those years rather than its worst.
- S00102 is over-subsidised from 2022 (36.6bn against 6.3bn in 2017). The cause
  is upstream: NIPA's `other` line still carries pandemic-era programmes, and
  the same 36.6bn sits on commodity 5241XX with or without this conversion.
  Documented as the `other` line's residue, not as a measurement.
- The T00SUB level carries a 1 $M gap - NIPA's 59,875 against the workbook's
  59,876 - spread proportionally across all twelve subsidised cells. `check`
  tests shape and level separately, because scoring raw levels reports BEA's own
  rounding as error and could hide a real one behind it.

`use_va_detail_sut` now grades five rows instead of three, with the published
T00SUB flipped to the balance's sign convention. The 48 partial column totals
are all T00TOP inside the trade block, which is the documented state of the
seed.

Also updates two constants in test_nowcast_product_taxes that moved when the
Trade FBSs were rebuilt on the NAICS-2022 goods Crosswalk (#734): TOP's residual
now rides a purchaser-price base carrying MCIF, so 2024's tobacco gap went
-26,153 to -26,282 and its share 7.7% to 7.3%. 2020's two numbers did not move
at all, which is the check that the shift is the Crosswalk and not the method.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
(cherry picked from commit c679b60)
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The run on this branch was killed at 15m16s against timeout-minutes: 15, which
GitHub reports as "cancelled" rather than a failure. It was not hung: the log
shows it still building NIPA_VA_surplus FBS methods when the job was cancelled.

Step 2 adds three NIPA FBS methods for each of 2017-2024, and each builds real
source data, so the suite legitimately crossed a ceiling set when it ran 12.5
minutes.

This is a stopgap, not the fix. The time goes into building sources rather than
into assertions, so what actually pays is caching the FBA/FBS parquets between
runs. Raising the ceiling only stops the gate failing for the wrong reason.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
test_compensation_movement failed on CI with FileNotFoundError for 2018 and
2024, while passing locally. The cause was not missing data: qcew_national_payroll
globbed bedrock/extract/output_data/BLS_QCEW_<year>_*.parquet directly, so it
could only ever see a cache generated on a developer machine. CI has no such
cache and no way to obtain one through that path.

getFlowByActivity is what every other FBA in this suite loads through, and it
fetches from the remote before falling back to generating. Routing QCEW through
it makes the module work anywhere rather than only where someone had already
built the files by hand.

The three column filters are load-bearing - the docstring explains that dropping
the Class filter alone inflates 482000's coverage ratio from 0.001 to 1,447 - so
this checks the loader returned all four columns rather than assuming it.

⚠️ compensation_movement_holdout.py globs the same way. It is a manually-run
analysis script with no CI tests, so it is left alone rather than changed
blind; it will need the same treatment if it is ever put under the gate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@WesIngwersen
WesIngwersen merged commit 494d10a into nowcast Aug 27, 2026
5 checks passed
@WesIngwersen
WesIngwersen deleted the step2_va_rebuilt branch August 27, 2026 13:48
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant