Conversation
NestedSequenceProcessor, NestedFloatsProcessor, DeepNestedSequenceProcessor and DeepNestedFloatsProcessor pad each visit to the longest visit seen in fit() (plus `padding`), but never truncated a longer one. When processors are fitted on a training split, a validation/test visit can be longer, so samples got different widths and batching failed (the default collate pads only the first dimension); the deep processors failed already while building the sample. Longer visits are now truncated to the fitted width, keeping the first codes/values, and each processor logs one warning that names the counts and points to `padding`. The visits-per-group dimension of the deep processors is unchanged: existing tests expect it to grow past the fitted size. - base_processor: shared one-time warning helper. - tests/core/test_nested_processor_truncation.py: all four processors, stacking, one warning per processor, and `padding` keeping more. - docs: "Output width" section on the NestedSequenceProcessor page. - examples/nested_sequence_fit_on_train.py: fit on train, apply to a test patient with a longer visit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
NestedSequenceProcessor,NestedFloatsProcessor,DeepNestedSequenceProcessorandDeepNestedFloatsProcessorpad each visit to the longest visit seen infit(), pluspadding. They never truncate a longer one.When processors are fitted on a training split, a validation or test visit can be longer than any training visit. That sample then has a different width from the others, and batching fails, because
collate_fn_dict_with_paddingpads only the first dimension. The deep processors fail even earlier, while building the sample.On master: fit on visits of 2 codes, process a visit of 4, and you get width 4. Stacking it with a normal sample raises.
Change
paddingas the way to keep more.paddingargument still widens the fitted size, so longer visits fit without truncation.base_processor.py. It tolerates processors pickled before this change.Deliberately unchanged: visits per group in the deep processors
The deep processors also pad the visits-per-group dimension to the size seen in
fit(), and longer groups would break batching in the same way. However, existing tests (test_deep_nested_sequence_processors.py, e.g.test_all_none_values) process groups with more visits than were fitted and expect them to be kept. I left that behaviour alone rather than change those expectations. Happy to truncate that dimension too in a follow-up if that's the intended contract.Tests, docs, example
New
tests/core/test_nested_processor_truncation.pycovers:forward_fillsettings;paddingkeeping longer visits.It fails on master and passes here.
Existing nested and deep-nested processor tests pass (60 in total).
The
NestedSequenceProcessordocs page gets an "Output width" section, and the docstrings now describe truncation.New
examples/nested_sequence_fit_on_train.py: fit on training samples, apply to a test patient with a longer visit, and show thepaddingalternative.Full core suite:
Ran 1388 tests … OK (skipped=76).tools/check_pr_rules.pypasses.Reported by a downstream EHR project that hit
expected sequence of length …errors the first time it fitted processors on the training split only.🤖 Generated with Claude Code