Skip to content

Size tensor embeddings without re-processing samples - #1262

Open
solarsys wants to merge 1 commit into
sunlabuiuc:masterfrom
solarsys:fix/embedding-width-from-processor
Open

solarsys wants to merge 1 commit into
sunlabuiuc:masterfrom
solarsys:fix/embedding-width-from-processor

Conversation

@solarsys

@solarsys solarsys commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator

Problem

For TensorProcessor features, EmbeddingModel.__init__ inferred the input width by calling processor.process(sample[field]) on the first dataset sample (pyhealth/models/embedding.py). But samples from a SampleDataset are already processed, so process() ran a second time.

That's harmless for the plain TensorProcessor. It's wrong for any processor that isn't idempotent: one that selects columns, scales or imputes. Re-processing either raises, or silently produces the wrong width. On master, a column-selecting subclass raises inside MLP(dataset=...).

Changes

  • EmbeddingModel reads the width from the first already-processed sample, which is exactly what the model will receive. It falls back to processor.size() when no sample is available. process() is never called.
  • TensorProcessor.fit() records the feature width (the last dimension, or 1 for scalars), and size() returns it instead of None. Nothing in the library relied on None (checked by grep). Old pickled processors without the attribute still return None.

The processed sample takes precedence over size() on purpose. A subclass that changes the width but doesn't override size() would otherwise report the raw width.

Tests, docs, example

  • New tests/core/test_embedding_tensor_width.py:

    • size() after fit(), before fit, and for scalars;
    • a column-selecting processor whose process() raises if re-run builds an MLP with in_features == 2;
    • plain "tensor" features keep their width.

    It fails on master and passes here.

  • docs/api/processors/pyhealth.processors.TensorProcessor.rst gains a "Custom tensor processors" section. The EmbeddingModel and TensorProcessor docstrings get >>> examples, verified as doctests.

  • New examples/custom_tensor_processor.py: a column-selecting processor with an MLP on synthetic data.

  • Full core suite: Ran 1386 tests … OK (skipped=76). tools/check_pr_rules.py passes.

Reported by a downstream EHR project that uses a column-selecting, standardising tensor processor.

🤖 Generated with Claude Code

EmbeddingModel inferred the input width of a TensorProcessor feature by
calling processor.process() on the first dataset sample. Samples read
from a SampleDataset are already processed, so this ran process() a
second time: wrong for any processor that is not idempotent (selecting
columns, scaling, imputing), and a column-selecting processor raised.

- EmbeddingModel reads the width from the first already-processed sample,
  falling back to processor.size(); process() is never called.
- TensorProcessor.fit() records the feature width (last dimension, 1 for
  scalars) and size() returns it instead of None.
- tests/core/test_embedding_tensor_width.py: size() after fit, a
  column-selecting processor that raises if re-processed builds an MLP
  with the right width, and plain "tensor" features.
- docs: "Custom tensor processors" on the TensorProcessor page.
- examples/custom_tensor_processor.py: a column-selecting processor with
  an MLP.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@solarsys
solarsys requested a review from jhnwu3 October 2, 2026 12:54

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant