Re-evaluate service expense sourcing, and re-rank the plan on impact (#564) - #717
Draft
WesIngwersen wants to merge 21 commits into
Draft
Re-evaluate service expense sourcing, and re-rank the plan on impact (#564)#717WesIngwersen wants to merge 21 commits into
WesIngwersen wants to merge 21 commits into
Conversation
Step 3 freezes every industry's input structure at 2017 and carries it on a price index. For manufacturing the part of that column the annual surveys cannot refresh is the materials bill: AIES publishes all materials, parts and supplies as one cell, 82.5% of the column, so only 8.3% of manufacturing's intermediate is commodity-mappable annually (#564). The commodity breakout is quinquennial Economic Census, and 2022 is a second observation of it sitting close to the middle of the 2018-2025 span. #564 called that a consolation prize. Measured, it is the main prize. Census_EC_MatFuel pulls ecnmatfuel for both vintages -- 4,624 rows in 2017 and 4,399 in 2022, every industry at NAICS-6 and every material an 8-digit code. The orientation is the opposite of Census_EC_PxI on purpose: PxI asks what an industry sells, this asks what it buys, so the industry goes in ActivityConsumedBy and the material in ActivityProducedBy, which is the Use table's own orientation. materials_structure.py answers the two questions that decide whether the source is worth having. Coverage: 66.2% of the 2017 materials bill and 69.1% of 2022 is placeable on a BEA commodity, against 8.3% annually -- 52.9% and 54.0% resolving 1:1 by NAICS prefix, the rest onto a BEA group that needs a within-group split on 2017 Use shares. A third is residual buckets Census could not place, and that is the ceiling. Movement: the mix moved 0.153 between the two censuses against 0.173 for the whole Use column over 2012-2017, with 264 of 345 industries moving more than 10 points. So the largest and least-observed part of the manufacturing column moves as fast as the rest of it, and freezing 2017 out to 2025 discards a reallocation this source can see. Two traps are documented rather than worked around. 00772000 "Total Materials" is the industry total and the named codes sum to it exactly, so summing the FBA unfiltered doubles the table; it is kept because it is the control a suppression recovery subtracts published children from, the same role NAICS 00 plays for PxI. And the vintages sit on different NAICS bases, sharing 345 industries and 291 materials carrying 90% of each year's cost -- 336411 aircraft reallocating 59% of its materials bill is almost certainly a code reassignment, not economics, and is flagged as suspect. Not built yet, and named in the plan: suppression recovery against the 00772000 control, the group-tier within-group split, the vintage code diff, and the interpolation itself. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ted (#698) Census withholds 412 of 4,624 cells in 2017 and 330 of 4,399 in 2022. They are not zero: they sit inside each industry's published 00772000 total, and leaving them there biases a materials mix toward whatever happens to be publishable, which is systematically the large materials. The control turns out to be exact, and that is measured rather than assumed -- for every industry with nothing withheld the named materials sum to 00772000 to within 0.1%, 238 of 238 in 2017 and 247 of 247 in 2022, with fuels carrying their own exact control in 00772002. After recovery all 406 and 386 industry-by-kind controls close to within 0.1% and no negative cell is created; the one negative in the output is published Census data. The prior is chosen by holdout rather than by argument. An economy-wide prior was tried first and is visibly wrong: it hands an idiosyncratic industry the economy's shopping list, and put $8.6bn of motor vehicle seating into aircraft manufacturing while cutting its aircraft engines from $15.5bn to $0.5bn. Masking published cells and recovering them scores every peer-prefix length; NAICS-3 wins on average, at WAPE 0.602 and 0.718 against economy-wide's 0.640 and 1.033. Cross-vintage priors are deliberately excluded even though 2017 is the best predictor of a withheld 2022 cell, because filling 2022 from 2017 biases the movement measurement toward zero -- a recovery must not manufacture the answer the analysis is testing. That WAPE is the finding that matters, and it corrects the last commit. The mass a recovery places is exact, so all 0.6-0.7 of that error is allocation across materials within the column -- which is exactly what a mix score measures. Restricting to the 193 industries with nothing withheld in either year, the materials mix moved 0.1330, not the 0.153 reported before, and the count of columns over 0.25 collapses from 66 to 17. Most of the extremes were the fill, not the economy. 336411 aircraft at 0.592 was the loudest of them and chasing it is what found the defect: its 2022 column is mostly withheld. The argument survives, weaker and better supported: materials mix moves 0.133 over five years against 0.173 for the entire Use column over 2012-2017, with 133 of 193 clean industries moving more than 10 points. So the largest and least-observed part of the manufacturing column moves substantially -- somewhat less than the column as a whole, not more, which is what the previous commit claimed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…lation (#698) The last three #698 items are built. Two dissolved a problem the plan expected to fight; the third overturned the interpolation form the plan had already chosen. materials_structure.py becomes inputs_structure.py, because what these sources reach is no longer just materials. The group-tier split is settled on evidence rather than on which prior sounds more principled. A group cell is divided over the BEA commodities its NAICS could be, on the purchasing industry's own 2017 Use row; scored by demoting every direct cell one prefix and comparing against the commodity Census actually named, the column prior puts 72.0% and 72.9% of the money on the right commodity against 46.9% and 49.5% for an economy-wide one. Accuracy falls off with group breadth -- 79.8% at 2-4 commodities down to 51.9% at 10-29 -- but 73% of group-tier dollars sit in groups of nine or fewer. The bare 33 prefix, 136 commodities, should be read as barely better than residual. The split lifts the placeable bill from $2,097B to $2,681B and the commodities reached from 137 to 204. The vintage code diff turns out not to be a diff. The plan expected to lose 10% of each year's cost to the 2017-2022 revision and it loses none: the material axis shares 289 of 289 and 290 MATFUEL codes, and every off-frame dollar on the industry axis is NAICS 2022 merging pairs of 2017 codes. Connected components of the year concordance put 100% of both vintages on one 365-industry basis with no split assumption. This does not rescue 336411 -- aircraft was on the shared frame all along, so its score is the suppression fill and the plan's guess about why it mattered was wrong. Linear interpolation was called "the obvious first form". It is obvious and it is wrong, and seeing that meant stopping treating 2018-2025 as unobserved. Manufacturing's materials bill is published every year the census misses, so Census_ASM_Expenses and Census_AIES_Expenses now pull it -- ASM through 2021, AIES for 2023. A straight line overstates 2020 by 28.8% because it cannot bend around a pandemic, and past 2022 it gets the sign wrong: the bill fell 6.8% into 2023 and the line says it rose 5.0%. That is the span the nowcast leans on hardest. Scope has to be matched or a definition reads as a growth rate -- against CSTMTOT the census-to-ASM step is a median 1.181, on materials-plus- fuels 1.063, a year of inflation. Those extractors also buy more than a control. ASM and AIES publish electricity, contract work, resales and 9-12 named purchased services at NAICS-6, 11.7% of the manufacturing column, each mapping onto a BEA service commodity -- 91.0% of the column reachable in total, and a separate cheaper task with no suppression recovery or group split to do. AIES's expense block is manufacturing-only, though: sectors 21, 22, 23 and 51-81 publish nothing at any NAICS level, so it confirms #564 on the service drifters rather than overturning it, and it does not cover the mining that the census does. The headline moves down again, and this time because the frame was wrong. 0.133 is a MATFUEL-code score and the 0.173 it was compared against is a BEA detail commodity one. On the same frame the clean subsample gives 0.0941, so the materials block moves roughly half the column's rate rather than "somewhat less". Aggregating 289 materials onto ~200 commodities nets off within-commodity substitution and the split holds 2017 structure fixed inside each group; both are properties of the seed, not corrections to 0.133. Finally, BEA's 2022 and 2023 tables are still annual-survey updates carried over the 2017 benchmark, so differencing a census-seeded block against them is not a check. That cuts both ways, and the second way is the argument for this work: it is information BEA has not yet incorporated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…, #664) S3b was scoped as the cheap half of Step 3 -- ten named cells, no suppression recovery, no group split. It was cheap, but not for the stated reason: the work turned out to be a scope measurement rather than a mapping exercise, and two of the claims it rested on were wrong. The first is the coverage. 11.7% and 91.0% were survey-side dollars, and two of the largest entries are not purchases of a commodity at all -- resales are goods bought and sold on untransformed, which the Use table handles through trade margins, and contract work is manufacturing services whose commodity is the buyer's own industry rather than any fixed row. With the survey's own residual they are reached as expense but cannot be placed. The honest figures are 6.4% seedable and 85.8% reachable, and the seedable block is $228.6B over ten cells. The second is that these cells could be read as levels. They cannot. Census_EC_ Expenses is new and pulls ecnbasic for 2017 and 2022, which publishes the cells under the same variable names ASM uses -- so census and survey form one panel with no crosswalk, and the splice is continuous: 2017 electricity $47.5B against ASM's $51.0B in 2018, repair $52.7B against $55.2B. That 2017 observation is the year the benchmark Use table is built on, and against it the survey and BEA disagree about what the same cells contain by factors of 0.40 to 8.01. The disagreements are structural, not noise: expensed software and computers are operating expense to Census and mostly investment to BEA, repair is one Census question against four BEA rows carrying parts BEA books elsewhere, and professional services runs the other way, one question against BEA's legal, accounting, engineering, consulting and R&D rows together. So the seed moves BEA's cell rather than replacing it -- Use2017 times survey(t)/survey(2017) -- which cancels every one of those because a constant scope factor divides out, and preserves BEA's own level and its own split across the commodities of a multi-row kind. 2017 reproduces the benchmark exactly, which is the check that the form is right. The block carries the same pandemic signature the materials bill does, -1.7% in 2020, and a frozen 2017 understates it by 23.8% by 2023. AIES publishes no telephony and no expensed software -- both variables exist in the 2023 table and both are zero in all 883 rows -- so those two are held at the 2022 census and marked rather than seeded as a collapse. 2024 and 2025 raise rather than quietly extrapolating. Chasing scrap through the same sources then found a defect in the materials placement. Every MATFUEL scrap code begins 33, which is not a NAICS that maps to any single commodity, so the prefix walk was filing purchased metal scrap into the bare 33 group of 136 commodities and smearing it across most of manufacturing. That group was the weak end of the group tier, flagged in the last commit as barely better than residual -- and it turns out to have been entirely scrap. BEA carries S00401 for exactly this concept, $49.1B into manufacturing in 2017, and Census's "excluding home scrap" is the same thing: bought in rather than generated on site. The five codes now map straight onto S00401 ahead of the prefix walk, the 30+ band collapses from $32.1B/$54.9B to $2.5B/$0.6B, and the clean commodity mix score moves 0.0941 to 0.0949. The direct+group frame reaches 200 commodities rather than 204, and the four lost were reached only through the smear. Scrap is metal and only metal in this source -- no wastepaper, no cullet, no plastic regrind, no textile rags. Census reproduces BEA's scrap concentration independently and more finely, separating secondary aluminium at 0.628 from secondary nonferrous at 0.529, and iron and steel mills move 0.328 to 0.440 between the vintages. The fuller picture, including the output side where ecnpxi does carry paper and plastics as wholesale recyclable sales, is written up in cornerstone-data/methods#59. Census_EC_Inventories is extraction only, for #664. ecnbasic carries all three stages of fabrication plus the totals, beginning and end of year, at NAICS-6 for both vintages -- which is the industry x stage cell BEA does not publish anywhere, and the level Hill's rules operate at. Verified additive: the stages sum to the published total at a median ratio of exactly 1.0000. These are stock levels and differencing them imports the holding gains CIPI excludes, so they carry the stage shares and U50705BU1 keeps the level. The ASM annual equivalent is deliberately not pulled yet. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The plan carried 2024 and 2025 as years the estimate had to reach and could not source, which framed an unobservable extrapolation as a risk this work owned. It does not. AIES 2024 still returns 204 No Content, ASM ends at 2021 and the census is quinquennial, so the observed panel runs 2017-2023 and the estimate stops there. Extending it to 2024 is #707, Phase 2 work tied to producing a 2024 table, and 2025 is out of scope entirely. So the open question in S3 is now only which interpolation form to fit, scored on observed years, rather than which form to fit plus how far to extrapolate it past the data. S3b needs no extrapolation at all: its span is fully covered, and nonmaterial_seed() already raises for later years rather than inventing one. unobserved_years() keeps reporting 2024 and 2025 because that is still the fact a caller wants to assert on -- what changed is that they are outside the span rather than gaps inside it. Left alone deliberately: the references to BEA's published summary panel running to 2024, gross output extracted for 2017-2024, the theta fitted on 2022-2024, and FIWS covering to 2025. Those are statements about which data exists, not about how far this estimate reaches, and the column-scaling argument in particular is quantified on a 2024 seed and would lose its point without them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
#705 asked, per drifting column, what could source it. Every one of them already had a source - BEA named one for each at the 2017 benchmark - so the question was whether a later vintage exists on the same basis and beats holding BEA's answer. Four no, one marginally yes. Census_SAS_Expenses splices SAS Table 5 across the two vintages that carry it, sas-17 (2013-2017) and sas-22 (2020-2022). #564 recorded Table 5 as "2020-2022 only"; that is the latest workbook's display window, not the series. 63 industries at 2- to 4-digit NAICS, ~19 mappable items. The two vintages sit on different Economic Census benchmarks, so every row carries the benchmark it was built on in Description. service_expense_seed.py indexes BEA's 2017 531ORE column on it and scores against BEA's current summary Use: +4.4 / +3.8 / +4.5% at 2020-2022, positive at every endpoint on a test biased against the seed, but at the inflation carry's bar rather than over it. --reachable says why: 18.07pp of ORE's movement sits on rows no survey item names against 13.51pp that a seed can touch, and the largest single mover has no counterpart question at all. Two corrections to earlier work in this branch: - An exploratory cut of this score reported +18.9 to +24.7% by applying each item's index to whole summary rows. Temporary staff maps to 561300, $8.6B of the column; at summary it multiplied all of 561, $97.8B, mostly 561700 services to buildings. A coarse commodity mapping inflates a result rather than blurring it. Build at detail, aggregate to score. - _load_usa_summary_sut pins the workbook by year, so --drift's series changes basis at 2023. The same 2022 read from both vintages differs by a dollar-weighted 0.0557 against a measured drift of 0.0986, and by 0.0976 for ORE alone. Near-misses elsewhere on the page are inside that noise. The seed scores on one vintage; fixing the shared diagnostic is follow-on work. Also records the negative results: the trade Business Expenses Supplement exists for 2017 and 2022 on one benchmark but loses every item of 4A0 to suppression; construction's ecnbasic pair is clean but 51% of the column is one undifferentiated materials cell; and ecnpurmode, ecnpurelec and ecnpurgas publish concepts BEA has no cell for. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ives (#705) intermediate_structure_drift read summary Use through io_2017._load_usa_summary_sut, which pins the workbook by year: 2017-2022 from the 2017-2022 release, 2023-2024 from the 1997-2024 one. That is right for FBA consumers, whose published values must not move under BEA's revisions, and wrong for a module that differences years against each other -- it put a vintage seam between 2022 and 2023 in the middle of --drift's series. Read every summary year from the current workbook instead, via this module's own summary_use / summary_intermediate(year, workbook). io_2017's year-pinning is left alone. service_expense_seed.summary_intermediate_current, which had made a local copy of exactly this fix, is deleted in favour of the shared function; its scores are unchanged, confirming the two were equivalent. Add --revision, the same year read from both vintages. The revision table in intermediate_estimation_plan.md was not reproducible by any committed code, which broke that page's norm; the flag reproduces it exactly. The question this was for: does #705's top-drifter ranking survive one basis? It does, bit-for-bit. The seam sits between 2022 and 2023, but 2017 and 2018 are identical across the two vintages (one cell, $3M) and 2024 was only ever read from the current workbook, so the 2017-against-2024 comparison never crossed it. ORE is still 0.141 and the candidate list -- ORE, GSLG/GFGD, 42, 5412OP, 81 -- needs no revisiting. The feared shift, that GFGD and 521CI revise by more than ORE does, could not bite: the revision only starts at 2019. What did move is the middle of the series, understated by 0.005-0.011: 2019 0.0452 -> 0.0504, 2020 0.0743 -> 0.0833, 2021 0.0711 -> 0.0823, 2022 0.0838 -> 0.0859. 2018, 2023 and 2024 are unchanged. The 2022 ranking does not survive -- four columns in, four out of the top ten, ORE 0.084 -> 0.158 -- but nothing on the page ranks at 2022. Two numbers corrected rather than restated. The plan doc and the seed docstring both quoted a 2022 drift of 0.0986 that no code reproduces on either basis; replaced with the measured 0.0859. The seed docstring said the gain was 4.4% at 2022 where the doc and the code both say 4.5%. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Margins.2 specified the margin rate as (TRADE + TRANS) / T013 -- margins over *basic* value. BEA gross output is valued at producers' prices (confirmed), so the price ratio #497 carries already contains the product-tax layer and the rate must be taken over producer value: T014 / (T013 + T015). One factor, not two. Also record the valuation chain the section had been using implicitly. T014 is the margins alone, not a running subtotal: T016 = T013 + T014 + T015, verified to $1M at detail, and independently stated in margins_estimation_plan.md. Restated on the right denominator. The level error from using basic is a median 3.3% and lands hardest on the commodities this section quotes: 315AL apparel 1.372 not 1.793 (+31%), 324 petroleum 0.246 not 0.289, 313TT textiles 0.779 not 0.858, 311FT 0.489 not 0.533. The panel is 26 commodities with a rate above 1%, not 36; median absolute change 3.6pp and p90 12.7pp, not 2.8 and 12.1. Most of the level error divides out of a ratio, so the correction to the carry factor itself is second-order -- median 0.35pp across the 26 receiving commodities above $20B, p90 1.4pp, at most 2.1pp among named ones. Worth having, not decisive; the section now says so rather than implying the fix is large. New guard, which the section needed and did not have. For margin *suppliers* T014 is large and negative -- the margin is allocated away from the trade or transport commodity onto the goods it carries, which is why the columns net to zero -- so mu is -0.94 for 42 wholesale, -0.99 for 486 pipeline, -0.88 for 482 rail. 1 + mu is then 0.06, 0.01 and 0.12 and the factor is a ratio of two near-zero numbers. Set it to 1 wherever mu <= 0. Nearly free: those rows carry almost no dollars in the purchaser-priced intermediate block, for the same reason their mu is negative. Finally, name the experiment. A missing deflator term and substitution under relative-price dispersion produce the same symptom -- a low theta on the summary panel against 1.00 on the detail one -- so theta must be fitted with and without this factor. If adding it pulls summary theta toward 1 the gap was the deflator; if not, the substitution reading stands. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…egative Two results, and the second was not the one being looked for. Add the margin leg of the purchaser deflator -- summary_supply, summary_margin_rate and summary_margin_factor -- on the denominator fixed in 4fd1122: mu_c = T014_c / (T013_c + T015_c), margins over producer value, because BEA gross output is at producers' prices. Margin suppliers are held at 1.0, since their T014 is large and negative (42 is -0.94, 486 is -0.99) and 1 + mu is then a near-zero denominator. Add --theta, which fits the exponent with and without that leg. It was a genuine question: Margins.2 hypothesised that the summary panel's low theta was a missing-deflator artefact rather than substitution, and the two readings predict the same symptom. The hypothesis is wrong. Theta is unmoved in six years of seven and moves away from 1.00 in the seventh; the score differs by under 0.001 either way. Not a null test -- 26 of 73 commodities move, up to 10%, on the largest goods rows (325 chemicals x1.049 on a $705B row, 3361MV x1.047). It fails *because* of that: 22 of the 26 have a factor above 1, 17 of those lost intermediate share, and the touched set lost 3.26pp of the block. Both legs inflate nominal goods shares during real substitution away from goods. Keep the term as the correct deflator, worth a median 0.35pp on the carry factor; drop the claim that it explains theta. Then the larger result. THETA_GRID started at 0.0, which censored the panel: 2023 and 2024 both pinned to the floor and were read as "the carry contributes nothing". They fit -0.25 and -0.50. The frozen structure scores better when commodity shares are moved *against* their own price movement, so #497's theta = 1 is not merely too strong for the target years, it is the wrong sign. The grid now runs -1.0 to 1.5 and carries a warning about the floor. Detail 2012->2017 still fits 1.00 -- its optimum was interior, so nothing measured there moves -- and the +1.00 at 2020 to -0.50 at 2024 span is monotone in the price regime. Consequence for S1: theta must be a parameter, and its default for the recent span is negative rather than 0 or 1. A build that hardcodes 1 applies a correction pointing away from the answer on the block's largest rows. Guard added with the negative exponents: a zero ratio would now raise a fit to infinity rather than harmlessly to zero. No year has one today. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…704) S1 and S0b of intermediate_estimation_plan.md. `nowcast_intermediate.py` is #497 as scoped: seed the published 2017 detail Use SUT interior (402x402, purchaser, before redefinitions), carry each column's shares on the commodity price ratio at theta, renormalise, and scale the column to `GO_producer - VAPRO_seed`. `nowcast.derive_initial_U_intermediate` is the entry point. Runs 2017-2024, bounded by `BEA_Detail_GrossOutput_IO_<year>` rather than by the price index. S0b came with it rather than ahead of it: `use_intermediate_detail_sut` went into sections.py runnable, not `candidate=None`. The 2017 section run passes on every cell - 44,281 populated cells, 366 row totals and 400 column totals inside tolerance, 100% coverage and accuracy. That is the plumbing, not the movement: at 2017 every carry factor is 1.0. The 2017 rescale is not the identity, and the residual is BEA's own rounding. Published T005 is one rounded number; the interior sums 402 separately rounded cells to a different one - $350M on $14.9T, at most $13M on a column. A small column wears that as a large fraction, so `atol` carries those cells rather than `rtol`: 334610 is $482M of intermediates and is rescaled 1.05%, the largest relative error, against a largest absolute error of $6.0M on a $19.2B cell. `reproduction_check` reports both, because either alone reads as the wrong kind of error. The seven published negative cells survive in all eight years and are not clipped. The two structurally empty columns - 4200ID and 814000 - stay empty, and a control that puts real dollars on one raises rather than dropping them. So does a seed column whose nonzero cells cancel: an empty column has no structure to normalise, a cancelling one has structure that cannot be written as shares of its own total, and collapsing both to all-zero would lose that. theta is an argument, defaulting to #497's 1.0. It fits negative at 2023 and 2024; choosing it stays #699. The column control is well levelled and badly allocated. Step 2 is unbuilt, so `vapro_seed` freezes 2017's VA share of gross output, which collapses the control to `GO(t) x T005(2017)/GO(2017)` - gross-output movement and nothing else. Scored against the published summary T005 (`--control`): within 2.3% economy-wide in every year 2018-2024, but weighted MAE by industry runs 2.5% to 8.0% between 2018 and 2022, with GSLG 18.3% low at 2022. Two corrections to claims this branch had been repeating. GSLG topping that list is the column *level* and is not #578. #578 is the commodity *mix* inside the G* columns, sourced from govslocalfin's function x object split and sequenced at S5 behind a go/no-go. Separately, NIPA T31005 matches the published government T005 at $0/$0/$2M in 2017 and is referenced in no source file - so the worst cell of the control table is the one with an exact annual source already identified and never wired. And "Step 5 imposes both margins hard, so Step 3 estimates a shape not levels" holds only once Step 2 exists. Checked in nowcast_targets.py: T1 is hard and real for 2017-2024 but pins the column's sum, not the split. T4 and T6 are both soft and still PLACEHOLDER, and `va_row_targets` does `del year` and reads its values off `published_2017_panel`. T5 - T00OTOP and V00300 - is deliberately not imposed, entering as seed only so the income side stays out-of-sample evidence; V00300 is $7.873T. The balance cannot re-derive VABAS, and there is no VA seed for any year but 2017, so it cannot run on 2024 at all until Step 2 lands. The shape is insulated from this - it is renormalised before the control is applied, so a wrong control rescales a column without moving a share inside it - but the level is not, and the module and plan doc now say so. Tests are structural and run on a toy panel: the arithmetic (`carry_shares`, `apply_column_control`) is split from the data wiring, so they need neither GCS nor the gross-output parquet. The year-by-year numbers are CLI flags on the drift diagnostic (`--seed`, `--control`) rather than tests, so everything quoted above is reproducible from committed code. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…control
BEA publishes intermediate inputs (UII205-A) and value added (UVA205-A)
annually 1997-2024 on its 191-row "underlying" industry frame, in two
workbooks sitting beside the GrossOutput.xlsx this repo already reads.
Nothing read either one. This extracts both, allocates them to the 402
detail industries, and replaces Step 3's frozen-2017 VAPRO seed.
The 191->402 mapping is derived, not hand-written. UGO205-A and UGO305-A
order industries the same way, so the 138 leaves of the 205-A hierarchy
partition the 414 rows of 305-A into contiguous runs; the runs are closed
by matching gross output in all 28 years. All 138 leaves match, exactly
414 of 414 detail rows are consumed, and the 402 codes cover the model
schema once each. derive_underlying_line_mapping reproduces the
checked-in constant exactly, and --mapping on the new diagnostic is how
to re-check it when the BEA vintage moves.
Value added is allocated and intermediate inputs are taken as the
residual GO - VAPRO. That is what makes GO = T005 + VAPRO hold per
industry, which is the form T1 imposes; allocating both independently
broke it by up to $15.2B a cell at 2024. It costs nothing against BEA:
- VAPRO summed back to the 138 lines reproduces UVA205-A exactly
(0.0 on 3,864 cells); T005 reproduces UII205-A to $9M, which is
BEA's own GO=II+VA rounding.
- At 2017 the derived columns reproduce the published detail Use SUT
margins to $0.89M (VAPRO) and $1.48M (T005) per industry.
- Economy-wide VA matches published GDP to at most $9M on $29T in
every year 1997-2024.
- GO - T005 - VAPRO at detail is 2.9e-11.
It also removes the suppression problem. BEA suppresses intermediate
inputs in every year on lines 83 (Customs duties) and 176 (Private
households); both are single-industry lines whose published 2017 T005 is
zero, and the residual recovers that zero rather than needing a special
case.
Step 3's column control is now GO - VAPRO with both sides observed.
Aggregated to summary and scored against the published summary T005 it
is within 0.00007% economy-wide and 0.00023% weighted MAE by industry in
every year 2018-2024, worst summary industry 0.003%. The superseded
frozen-ratio seed scored 0.2-2.3% and 2.5-8.0%, with GSLG 18.3% low at
2022. That is a consistency check and not an independent validation --
UII205-A and the summary Use SUT's T005 are the same BEA estimate
published two ways -- so what it establishes is that the allocation adds
back correctly and the control now *is* BEA's T005. vapro_seed becomes
vapro; the now-dead _row helper and published_gross_output import go.
Signs are carried, not clipped. S00201 has a published 2017 VAPRO of
-$10,069M and stays negative in all 28 years. The residual T005 goes
negative in 11 cells, all in 5191A0 and all in 2002-2015, so the
2017-2024 nowcast span is clear; the other residual direction was worse
(13 spurious negatives across three industries, reaching -$13,117M).
This is not Step 2. VAPRO is the column total; Step 2 owes the split
across the five value-added rows, and T4/T6 are still soft placeholders.
What changes is that there is now a VA level for every year rather than
for 2017 alone, so the balance can run on 2024. It also arrives off the
P1-gated path -- a straight Excel read on the shape of the working
UGO305-A loader, so map_fbs_sectors_to_model_schema never enters it.
Repo-wide: black, ruff, mypy clean (bar the four known Windows-only
settings.py errors); 779 passed, 2 skipped, 1 xfailed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
allocate_underlying_to_detail is public and takes an arbitrary mapping. A code appearing under two lines would be indexed twice by .loc[children] and written twice into the output, so the group totals would silently stop adding up rather than raising. Same for a duplicated line in group_values, where .loc[line] returns a frame instead of a row. The checked-in mapping has neither -- 402 codes, all distinct -- so this guards the function, not the current data. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
BEA publishes 2007, 2012 and 2017 detail Supply and Use SUT as one zip of per-table workbooks with a sheet per year, all three already on the 2017 code basis in one frame. It was a local drop that only the Step 3 drift diagnostic read, through an ad-hoc zip reader. io_2017 now carries `_load_benchmark_detail_supply_use_usa(matrix, year)`, GCS-backed like every other table, with `load_benchmark_detail_U_intermediate_usa` and `load_benchmark_detail_supply_usa` as the typed 402 x 402 accessors. The panel is a second and third observation of every structural question in the build -- Step 3's input mix, Step 4a's commodity mix, the margin rates, the FD splits -- not just this one. The panel's 2017 sheets are the single-year workbooks cell for cell, on both matrices: 0 differing cells on 413 x 424 and 405 x 415. `assert_benchmark_panel_matches_2017()` is the check. The two loaders are kept separate anyway, because `_load_2017_detail_supply_use_usa` is what `bea_parse` emits as the BEA_Detail_Use_SUT / BEA_Detail_Supply FBAs and those stay pinned to their published workbook; the FBAs are not extended to 2007 and 2012 here. BEA's subsidy sign convention holds on all three years, so `_assert_bea_subsidy_signs` runs on every year the panel loader returns. `--holdout` reproduces unchanged, including the 2012->2017 best theta of 1.00. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…699) S2's experiment, and the answer is not the one the plan expected. Every theta on the plan's table starts at 2017, so elapsed years, cumulative inflation, price dispersion and structural drift all move with the calendar and none can be told from the others. The summary Use SUT publishes 1997-2024 and the price index reaches 2012, so 78 non-nested spans are free -- different bases, different lengths, different inflation. On those (`--regime`): crosses the 2021-22 surge R2 0.613 cumulative price level R2 0.525 elapsed years R2 0.142 relative-price dispersion R2 0.014 So the dispersion candidate the plan named is dead -- 1.4% of the variance, and the coefficient points the wrong way -- and so is elapsed time. Holding span length fixed, spans that cross the surge fit theta 0.0-0.5 and spans that do not fit 0.7-0.9, at every length from one to nine years. `default_theta` ships that: 0.75 off the surge, 0.0 across it (fitted 0.755 and 0.141). Rounding up to zero rather than to the target spans' own -0.25/-0.50 is deliberate -- a negative theta says nominal shares move against their own price, and it buys 0.6%. The median gain of the best theta over a frozen A is 5.44% of the score off the surge and 0.59% across it, so in the regime this build targets the carry is worth well under one percent however theta is set. What #497's theta = 1 cost was the 12.6% it gave away by pointing the wrong way. The margin-rate leg is built too, so the deflator is the purchaser one. The non-obvious part is reaching 402 rows annually: detail Supply is published only for benchmark years, so the rate's level is detail-observed at 2017 and only its movement is borrowed from the summary parent. S0a made that testable -- against the observed 2012 detail factor, weighted by 2017 intermediate dollars, the shipped rule is 0.756pp off, the parent's factor taken down unchanged is 1.010pp, and no factor at all is 1.818pp. A year with no published Supply table is refused rather than carried: MARGIN_YEARS (1997-2024, BEA's vintage) is a separate constraint from INTERMEDIATE_YEARS (gross output). They agree at 2024 today, so nothing is blocked; a 2025 build would reach a year with one and not the other, and a silent factor of 1.0 there would read as "margins did not move". The margin leg is inert in exactly the years the build targets -- at theta = 0 every factor is raised to the zero power. It moves 0.21% of the 2019 block and 0.54% of the 2021 block and 0.000% of the 2024 one. Kept because it is the correct deflator for a purchaser-valued cell, not because it changes the current answer. Also restates the plan's Step 3 sections that PR #712 left describing `vapro_seed`: the level table (the block total is theta-independent, so it is a check on the control), the column control table, and the GSLG/T31005 paragraphs. 2017 reproduction is unchanged -- $6.02M max absolute, 1.05% max relative, seven negatives, and the section still scores 100% on 44,281 cells. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`intermediate_estimation_plan.md` covers the theta findings but never states the mechanics, and the name invites reading it as a price ratio or a valuation bridge. It is neither: it is a scalar exponent, one per span, applied identically to all 402 commodity rows and all 402 industry columns. `About_the_price_carry.md` is the reference for that: the carry formula, the two legs of the commodity deflator and their sources, the BAS -> PRO -> PUR chain the margin leg sits in (and why its denominator is producer and not basic value), the theta = 1 - sigma CES reading, what is commodity-specific and what is not, the shipped values, and the approximations the carry rests on. Linked from the nowcasting README, from the plan's Inflation section where a reader first meets theta, and from the module docstring. The README's intermediate_structure_drift entry also gains the --regime flag added in 7761934. Every claim in it re-checked against the code: grid -1.0 to 1.5 in 0.25 steps, theta 0.75 for 2018-2021 and 0.0 for 2022-2024, MARGIN_YEARS 1997-2024 against INTERMEDIATE_YEARS 2017-2024, the four unpriced commodities, and mu of -0.944 for `42` and -0.989 for `486`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The plan sequenced S5's government leg behind a go/no-go on whether Census `govslocalfin`'s function x object split bridges to a commodity mix. It has been run and the answer is no. The model that premise implies has an exact ceiling: with within-function mixes held fixed, commodity-mix movement is bounded by the function-mix movement itself, since each within-function mix sums to 1 over commodities. So the question is answerable without building the bridge. - the function mix moves 0.0464 over five years, against a 0.201 commodity drift in the same columns; - government functions overlap by 36% (mean pairwise dissimilarity 0.639 on BEA's own government columns), so the realistic ceiling is 0.030; - realising the bound on BEA's three general state-and-local columns -- a real function split with observed weights and observed within-function mixes -- buys +2.4%. 97.6% of the drift is *within* function. A finer function list does not help: BEA's 3-function split moves 0.0453 where govslocalfin's 33 functions move 0.0464. Two independent defects would also have sunk it. `Current Operations` by function is published only for 2022-2024, so there is no 2017 anchor at the seed year -- the 137->232 row jump the source note flagged is exactly this. And the only continuous function series, `Total Expenditure`, is a 0.062 proxy for it, larger than the 0.0385 of movement it would carry. `govslocalfin` is also state-and-local only, and federal holds 43.1% of the block's misplaced dollars including S00500, the worst column in the table. This closes govslocalfin as the route to the `G*` mix. It does not close #578: the columns still drift 230-258 $B and the drift is within-function. Corrects two claims in annual_survey_expense_sources.md that are now known wrong -- that the `G*` columns need a total rather than a mix, and that Current Operations by function runs for every year in the span. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The form left open by S3 is settled, and the candidate the plan named is not the answer. Separate the two things called "interpolation" first. The materials LEVEL is observed every year (ASM to 2021, AIES from 2023), so it never needed interpolating and section 4 rules out doing so. The commodity MIX is observed only at 2017 and 2022, and nothing observes its interior -- the fuels share of the census universe stays between 0.98% and 1.14% -- so the open question was only ever about the mix. The price-carried path is rejected. Carrying the 2017 census mix to 2022 on the commodity price ratio fits theta = 0.00 on the unsuppressed frame and -0.25 (+0.4%) on the full one. The level moves with price; the mix does not. The published summary panel cannot arbitrate the form and its answer must not be used: BEA carries the last benchmark forward, so every interior year of a summary span is itself an interpolation, and since BEA has not taken up the 2022 Economic Census its 2022-2024 tables are still 2017-benchmark carries -- so spans ending in the nowcast horizon are contaminated at the endpoint too. The benchmark detail panel can arbitrate it, because 2007, 2012 and 2017 are three independent Economic-Census-anchored observations. Interpolating 2007 -> 2017 and scoring at the observed 2012, manufacturing: frozen 0.0889, linear 0.0764, geometric 0.0710, endpoint 0.1216. So interpolating beats freezing by 20.1%, geometric beats linear by 7.1%, and adopting the newer observation early is worse than freezing. But do not extend the trend past the last observation: reaching 2017 from 2007 and 2012, holding the 2012 mix scores 0.1232 against 0.1569 linear and 0.1513 geometric -- 27.4% worse on manufacturing, 41.7% on the whole table. S3 therefore ships an asymmetry: geometric interpolation between the two censuses, and the 2022 mix held flat for 2023 and 2024. materials_seed() is that seed, an index on the share with the column total held, moving the block 0.0979 off frozen 2017 by 2022 -- alongside the independently computed 0.0949 in the mix score section. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…564) Section S4 rejected SAS Table 5 for services and generalised that into "#564's negative result generalises". It does not survive, for four measurable reasons. The weighting was wrong. Every ranking in the plan -- which columns drift, which sources are worth building, what "a small prize" means -- is dollar-weighted. Cornerstone is an EEIO model, so the quantity that matters is kg CO2e, read here from the shipped v0.3 B_USA_non_finetuned snapshot characterised to CO2e. On the 2017 detail block: manufacturing 31-33 $4,699B 31.6% of dollars 28.0% of impact utilities 22 $350B 2.4% 26.1% agriculture 111/112 $413B 2.8% 23.3% mining 21 $510B 3.4% 11.5% all other (services) $8,260B 55.6% 7.2% Utilities and agriculture are 5.2% of the dollars and 49.4% of the impact. "All other" -- every column S4 ranked -- is 7.2%. The top 10 commodity rows carry 65.3% of direct impact and 221100 electricity alone carries 25.2%. Two consequences: agriculture (#577) is not "a small prize", and purchased electricity in the service columns is the highest-value cell in the step. The AIES finding was an artefact of the wrong endpoint. Census_AIES_Expenses queries timeseries/aies/basic, where every service row is a well-formed zero. timeseries/aies/exp02 publishes all 41 expense variables for 13 service sectors in 2023. Both source notes are corrected. The resource is 97 BEA detail industries holding $6,269B (42.2% of the block), not the one column the built seed uses. Dollar-weighted reach is 41.4%; impact-weighted reach is 63.8%, and per column the gap runs to 65pp -- 481000 reaches 53.8% of dollars and 95.6% of impact. The benchmark seam is not the blocker. On matched three-year spans it adds 3.1% to level movement over an equal span inside a vintage, while also containing 2020; on energy items the seam excess on shares is -7.4%. Shares are 30-54% quieter than levels within a vintage, so a share-based index is right on its own merits and cancels a proportional rebenchmark for free. Also records that wholesale and retail are absent from exp02 entirely and that ecnbasic carries expenses for sectors 21, 23 and 31-33 only, in both 2017 and 2022 -- so the trade counterpart remains the AWTS/ARTS supplement, whose suppression rejection is untouched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The materials interpolation commit d722010 left three constructs black 24.10 rewrites: a parenthesised return annotation on _census_mix_frames, a split %-format key in interpolation_forms' assign, and a nested conditional in the census-mix treatment label. Formatting only -- no behaviour changes. This is the format gate failing on PR #715 and, inherited, on #717 and the service seed branch stacked above it. Fixed here so one change clears all three rather than each branch carrying its own. Gates: black, ruff, mypy clean; pytest 795 passed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
WesIngwersen
added a commit
that referenced
this pull request
Aug 25, 2026
Wes, 2026-08-25: rank by total kg CO2e per dollar -- direct plus indirect --
not by the direct slice. A Use cell is an entry in A, so an error in it
propagates through the whole Leontief inverse; what a row is worth getting
right is its total embodied emissions. Some sectors have much higher N than D
because their indirect is large, and their inputs are still very important.
total_impact_intensity() computes N = C B L with L = (I - (Adom + Aimp))^-1,
matching the inverse cornerstone_disagg_pipeline builds N from. impact_intensity
now takes kind= and defaults to 'total'; direct_impact_intensity keeps D so the
contrast stays legible.
The ranking moves a long way. On the 2017 detail block, by commodity row:
dollars D N
manufacturing 31-33 31.6% 28.0% 44.7%
agriculture 111/112 2.8% 23.7% 16.4%
all other (services) 55.6% 7.2% 14.3%
utilities 22 2.4% 26.1% 13.7%
mining 21 3.4% 11.5% 8.0%
Manufacturing nearly doubles, because a manufactured good carries a long
upstream chain. Utilities halves, because electricity's emissions are almost
entirely direct and D therefore flatters it. Services double. Individual rows
move harder still: 31161A meat processing goes 0.10% -> 2.61%, a 26x change
that is all upstream livestock, and 531ORE goes 1.48% -> 3.97%. The block is
less concentrated on N -- top 10 rows are 48.5% against 65.3% on D.
The seed's score, re-run on N:
2020 +0.5% dollar +10.2% impact (22/33 columns win)
2021 +1.4% dollar +10.8% impact (21/33)
2022 +2.3% dollar +9.2% impact (17/33)
On D these read +22.6/+25.1/+22.6%, so D was overstating the seed by more than
double -- it over-weights electricity, which the survey names well.
It also flips a verdict. Re-scoring S4's per-column no-goes on N: 622
(+11.4/+12.3/+27.5%) and 722 (+7.5/+11.2/+10.6%) overturn, 5412OP still loses
(-9.0/-7.7/-2.0%), and 81 does NOT overturn (-0.8/+2.9/+0.6%) although it
looked like a consistent win on D. More columns win on N than on D, because N
spreads weight off the few electricity-heavy rows.
Every ranking recorded before this is a D ranking, including the table PR #717
re-sorted the plan on; the resource module's docstring is restated accordingly.
Gates: black, ruff, mypy clean; targeted Step 3 tests pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
WesIngwersen
added a commit
that referenced
this pull request
Aug 26, 2026
The materials interpolation commit d722010 left three constructs black 24.10 rewrites: a parenthesised return annotation on _census_mix_frames, a split %-format key in interpolation_forms' assign, and a nested conditional in the census-mix treatment label. Formatting only -- no behaviour changes. This is the format gate failing on PR #715 and, inherited, on #717 and the service seed branch stacked above it. Fixed here so one change clears all three rather than each branch carrying its own. Gates: black, ruff, mypy clean; pytest 795 passed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
WesIngwersen
force-pushed
the
step3_materials_interpolation
branch
from
August 26, 2026 18:30
ec27a63 to
454ddb4
Compare
WesIngwersen
added a commit
that referenced
this pull request
Aug 26, 2026
The materials interpolation commit d722010 left three constructs black 24.10 rewrites: a parenthesised return annotation on _census_mix_frames, a split %-format key in interpolation_forms' assign, and a nested conditional in the census-mix treatment label. Formatting only -- no behaviour changes. This is the format gate failing on PR #715 and, inherited, on #717 and the service seed branch stacked above it. Fixed here so one change clears all three rather than each branch carrying its own. Gates: black, ruff, mypy clean; pytest 795 passed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
WesIngwersen
force-pushed
the
step3_materials_interpolation
branch
from
August 26, 2026 18:31
454ddb4 to
c077b45
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #715.
§S4 rejected SAS Table 5 for services and generalised it to "#564's negative
result generalises". ❌ It does not survive — for four measurable reasons.
Every ranking in the plan is dollar-weighted. Cornerstone is an EEIO model,
so the quantity that matters is kg CO2e — read here straight from the shipped
model, the v0.3
B_USA_non_finetunedsnapshot characterised to CO2e.Utilities + agriculture: 5.2% of dollars, 49.4% of impact. "All other" — every
column §S4 ranked — is 7.2%. The top 10 rows carry 65.3% of direct impact;
221100electricity alone carries 25.2%.Two consequences, neither of them what the plan says:
is 89–91% mappable, annual to 2025, already wired into bedrock.
this step.
Census_AIES_Expensesqueriestimeseries/aies/basic, where every servicerow is a well-formed zero.
timeseries/aies/exp02(AIES00EXP02)publishes all 41 expense variables for 13 service sectors in 2023 — sector 22
electricity $50.3B, sector 54 prof/tech $103.6B. Both source notes corrected.
$6,269B, 42.2% of the intermediate block. The built seed uses one of them.
Per column the gap runs to 65pp:
481000reaches 53.8% of dollars and 95.6% ofimpact;
72100025.5% → 81.8%.On matched three-year spans (annualising biases the test), the seam adds
+3.1% to level movement over an equal span inside a vintage — while also
containing 2020. On energy items the seam excess on shares is −7.4%. Shares
are 30–54% quieter than levels within a vintage, so a share-based index is right
on its own merits and cancels a proportional rebenchmark for free.
No 42/44/45 rows in
exp02at all, andecnbasiccarries expenses for sectors21, 23 and 31-33 only in both 2017 and 2022. So the trade counterpart remains
the AWTS/ARTS supplement, whose suppression rejection is untouched. This is an
API check and does not rule out a data.census.gov table or a later AIES release.
What this does not do
❌ It does not overturn §S4's per-column scores —
5412OP,81,722,622didlose on a whole-column dollar metric. The claim is the metric asked the wrong
question and the seam was the wrong reason. Nothing is built here; this is the
evaluation.
Gates:
ruff check .clean,mypy bedrockclean (bar 4 pre-existing Windows-onlyresourceerrors),pytest795 passed.🤖 Generated with Claude Code