Skip to content

proposal: batched inference for TransformerDetector (#23) - #114

Open
YassinNouh21 wants to merge 1 commit into
KRLabsOrg:mainfrom
YassinNouh21:proposal/batched-inference
Open

YassinNouh21 wants to merge 1 commit into
KRLabsOrg:mainfrom
YassinNouh21:proposal/batched-inference

Conversation

@YassinNouh21

@YassinNouh21 YassinNouh21 commented Sep 18, 2026 •

Copy link
Copy Markdown

Summary

Design proposal for #23, no code in this PR. One file, docs/proposals/0023-batched-inference.md: what is wrong today with line references, the batching design, and a test plan. Implementation can follow as a separate PR once the approach is agreed.

The short version: batch only the forward pass. Tokenize pairs together with padding, take each row's real length from attention_mask (the current answer_start formula uses the padded width, hallucination_dataset.py:178), run one forward per micro-batch, then decode every row with one shared _decode_row that the single path also uses.

flowchart LR
    A["validate"] --> B["sort by length,<br/>micro-batch"] --> C["tokenize together,<br/>answer_start per row"] --> D["one forward"] --> E["_decode_row<br/>per row"] --> F["restore order"]
    S["predict_prompt"] --> E
Loading

Also fixes the silent zip truncation at transformer.py:452 and llm.py:612,614 with a shared length check in BaseDetector.

Worked example from the doc, three pairs padded to [3, 13] with the test-suite tokenizer:

row real_len answer_len today L - answer_len - 1 proposed mask.sum() - answer_len - 1
0 9 1 11, points at [PAD] 7, points at paris
1 7 2 10, points at [PAD] 4, points at paris
2 13 3 9, short 9, short

Only the longest row is right today.

Related issue

Refs #23

Type of change

  • Documentation

Testing

  • Other: mkdocs build

Checklist

  • I kept the PR focused on one change.
  • I added or updated tests/docs when needed.
  • I checked that no secrets, API keys, or credentials are included.

Rights & sign-off (required)

  • I certify that I have the right to submit this code and that it may be
    distributed under the repository's MIT license
    (see CONTRIBUTING).

@YassinNouh21
YassinNouh21 force-pushed the proposal/batched-inference branch from aadc3e0 to 1e63a98 Compare September 18, 2026 07:54
@YassinNouh21
YassinNouh21 force-pushed the proposal/batched-inference branch from 1e63a98 to aa78786 Compare September 18, 2026 07:55

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant