Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
72 commits
Select commit Hold shift + click to select a range
eac1b77
Add AMP gradscalar.
Gsunshine Sep 23, 2024
48e3f47
Update README.
Gsunshine Sep 23, 2024
a223882
fix: restore AMP CLI and metrics-none behavior
Jul 14, 2026
4e33194
merge: sync Day 1 bootstrap with main
Jul 14, 2026
43130c7
Merge pull request #1 from hjjjs4vbmv-netizen/leader/day1-bootstrap
buergute Jul 15, 2026
36dcae6
chore: integrate reproducible engineering workflow
Jul 15, 2026
7c638a4
Merge pull request #6 from hjjjs4vbmv-netizen/leader/day2-local-integ…
buergute Jul 15, 2026
57cef9c
Validate Role A clean-container workflow
w10800 Jul 15, 2026
892e497
Add reproducible fixed-seed evaluation workflow
edwards365 Jul 15, 2026
9ebc6c3
Make fixed-seed sampling batch invariant
edwards365 Jul 15, 2026
6576c16
Add fixed-seed sampling artifacts
edwards365 Jul 15, 2026
956c6f4
Pin Hugging Face Hub for clean setup
w10800 Jul 16, 2026
0167aca
docs: refresh clean-container validation evidence
w10800 Jul 16, 2026
2e36590
Close fixed-seed evaluation acceptance gaps
edwards365 Jul 16, 2026
59d39a5
Add official EDM fixed-seed smoke artifacts
edwards365 Jul 16, 2026
abd587a
Merge pull request #8 from w10800/role-a/clean-container-validation
buergute Jul 16, 2026
9b3db63
Refine deterministic evaluation protocol schema
edwards365 Jul 16, 2026
521eae9
Merge remote-tracking branch 'origin/main' into role-d/evaluation-v1
edwards365 Jul 16, 2026
d95094b
chore: validate evaluation workflow on latest main
edwards365 Jul 16, 2026
93a1ffc
Merge pull request #9 from role-d/evaluation-v1
buergute Jul 16, 2026
cb91b4f
Add paired fixed/adaptive training and evidence infrastructure
kangw24 Jul 17, 2026
5344a5c
Add loss-EMA adaptive t-to-r schedule and activation evidence
Alicia24012867 Jul 21, 2026
ef4aa31
Add canonical paired A100 activation and stability evidence
kangw24 Jul 21, 2026
4b80a4f
Add final paired 16 kimg fixed-seed evaluation
w10800 Jul 21, 2026
3279bf8
Add multibudget paired performance evaluation
w10800 Jul 22, 2026
03d3dba
Harden paired analysis provenance checks
Alicia24012867 Jul 22, 2026
c9412dd
Add Role D visual evaluation handoff
edwards365 Jul 22, 2026
ae5a23f
Archive paired 1024 kimg seed0 evaluation
buergute Jul 22, 2026
274786f
Merge pull request #22 from hjjjs4vbmv-netizen/agent/publish-seed0-10…
buergute Jul 22, 2026
3a0d603
Add factorized global-local gap controller experiments (#23)
w10800 Jul 28, 2026
6d4bc7d
Add confirmatory fixed-vs-global training protocol and resume fix (#24)
kangw24 Jul 29, 2026
ab03f9e
Add exploratory q=128 screening and reproducibility materials (#25)
edwards365 Jul 29, 2026
55bb525
docs: freeze staged checkpoint evaluation protocol
Alicia24012867 Jul 29, 2026
f1ee750
feat: add staged evaluation runner and result collector
Alicia24012867 Jul 29, 2026
db7ad33
test: add fixed-seed determinism artifact verifier
Alicia24012867 Jul 29, 2026
f299528
Role C: complete seeds 4/5 confirmatory 256k paired runs
kangw24 Jul 29, 2026
220e2ad
fix: support current scipy in FID evaluation
Alicia24012867 Jul 29, 2026
bb7c689
docs: record quick smoke and formal readiness
Alicia24012867 Jul 29, 2026
0266150
feat: add training integrity receipt checker
Alicia24012867 Jul 29, 2026
b5fe229
docs: clarify seeds 4/5 provenance and checkpoint handoff
kangw24 Jul 30, 2026
1f66a36
docs: record original PR handoff commit explicitly
kangw24 Jul 30, 2026
125ece6
fix: load project modules during integrity checks
Alicia24012867 Jul 30, 2026
9ed0568
feat: add paired staged evaluation statistics
Alicia24012867 Jul 30, 2026
62d205f
docs: record staged formal evaluation results
Alicia24012867 Jul 30, 2026
3c698d8
feat: freeze q256 confirmatory checkpoint matrix
Alicia24012867 Jul 30, 2026
6d2b1a3
runtime-manifest
Alicia24012867 Jul 30, 2026
05e66af
Strengthen the training-integrity receipt
Alicia24012867 Jul 30, 2026
454df27
Implement paired analysis for the fixed/global comparison
Alicia24012867 Jul 30, 2026
82503df
Prevent result-driven promotion from quick to formal
Alicia24012867 Jul 30, 2026
c43f2d7
Make NFE=1 command semantics match the manifest
Alicia24012867 Jul 30, 2026
6d777cd
Record the formal evaluator environment
Alicia24012867 Jul 30, 2026
64fbb7e
docs: pin q256 v2 integrity receipts
Alicia24012867 Jul 30, 2026
8375d46
fix: skip paired analysis for smoke runs
Alicia24012867 Jul 30, 2026
44f795c
docs: record q256 formal evaluation results
Alicia24012867 Jul 30, 2026
e1967d5
docs: correct Role E source provenance classification
edwards365 Jul 30, 2026
d42ef93
docs: qualify Role E executed-source equivalence
edwards365 Jul 30, 2026
6e4038d
Merge pull request #26 from hjjjs4vbmv-netizen/role-c/seeds-4-5-confi…
buergute Jul 30, 2026
6f0cead
Merge pull request #28 from hjjjs4vbmv-netizen/fix/role-e-source-prov…
buergute Jul 30, 2026
6b26c04
Merge pull request #27 from hjjjs4vbmv-netizen/role-d/v1
buergute Jul 30, 2026
fe016f0
docs: package portable q=256 confirmatory formal results
Alicia24012867 Jul 30, 2026
1751642
protocol: freeze q256 budget and fresh q128 evaluation matrices
Alicia24012867 Jul 30, 2026
68ab758
protocol: deliver frozen evaluation matrices and capacity plan
Alicia24012867 Jul 30, 2026
28b04c0
Merge pull request #29 from hjjjs4vbmv-netizen/role-d/v2
buergute Jul 31, 2026
172bb85
Add paper-ready q256 artifacts and budget-curve analysis
Alicia24012867 Aug 1, 2026
42036ef
Add fresh q128 training handoff
edwards365 Aug 3, 2026
f53910b
Add q128 1024k formal evaluation results
Alicia24012867 Aug 3, 2026
495c810
theory: add linear-Gaussian toy model, novelty statement, and handoff…
kangw24 Aug 4, 2026
7ee027e
theory: address PR #33 review — fix stop-grad, add LR-matched control…
kangw24 Aug 4, 2026
c738312
theory: separation as exact-recursion counterexample (PR #33 review r…
kangw24 Aug 4, 2026
eae8de2
Merge pull request #33 from hjjjs4vbmv-netizen/theory/toy-gap-calibra…
buergute Aug 4, 2026
875e0b4
experiment: add q128 gap-response screen
edwards365 Aug 4, 2026
85e5a85
experiment: add q128 gap-screen status
edwards365 Aug 4, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
20 changes: 16 additions & 4 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -161,10 +161,22 @@ cython_debug/


# Project-related
datasets

ct-runs
ct-evals
.DS_Store
/paired_16k_5344a5c_canonical/
datasets/
runs/
checkpoints/
pretrained/
ct-runs/
ct-evals/
wandb/
*.pkl
*.pt
*.pth
*.ckpt
*.safetensors
*.zip
*.pyc

slurm*
debug.sh
Expand Down
127 changes: 127 additions & 0 deletions CONFIRMATORY_COMMANDS.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,127 @@
#!/usr/bin/env bash
# =============================================================================
# Role C — Confirmatory fixed-vs-global (g=1.10) training commands.
# -----------------------------------------------------------------------------
# Provenance:
# training_code_sha : 3a0d603da97dd93ddbb6c7ce49e4a7351d54bb43 (seed-3 was
# trained on this commit; recorded in commit_sha.txt)
# pr_head_sha : 79143c685e5588948972c17457b1c51c7a77bb49 (this PR head;
# only adds docs + the resume fix, not a training baseline)
# Branch : role-c/confirmatory-gap-g110
# Env : conda env `myconda` (python 3.13.5, torch 2.8.0+cu128, A100)
# Dataset : /mnt/ect_project/datasets/cifar10-32x32.zip
# Pretrained (transfer): /mnt/ect_project/pretrained/edm-cifar10-32x32-uncond-vp.pkl
#
# Method definitions (ONLY these two are compared):
# Fixed : --mapping=sigmoid --global-gap-scale=1.0
# (official ECT sigmoid; global_gap_scale==1.0 short-circuits
# to bitwise parity with the official formula)
# Global-only : --mapping=global_sigmoid --global-gap-scale=1.10
# (same official sigmoid gap scaled by a single fixed g=1.10)
#
# INVARIANT: fixed and global differ ONLY by {mapping, global_gap_scale, outdir}.
# No local controller is enabled in either arm.
# =============================================================================
set -Eeuo pipefail

export ECT_BRANCH="role-c/confirmatory-gap-g110"
export ECT_COMMIT="3a0d603da97dd93ddbb6c7ce49e4a7351d54bb43" # training_code_sha
DATA="/mnt/ect_project/datasets/cifar10-32x32.zip"
TRANSFER="/mnt/ect_project/pretrained/edm-cifar10-32x32-uncond-vp.pkl"
PYTHON="${PYTHON:-python}"

# ---- shared parameters (fixed == global except mapping + global-gap-scale) ----
COMMON=(
--data="$DATA"
--cond=False --arch=ddpmpp --precond=ect
--batch=128 --batch-gpu=16 --optim=RAdam --lr=0.0001 --dropout=0.2 --augment=0
-q 256 -k 8 -b 1 -c 0 --double=10000 --ema_beta=0.9993
--fp16=True --enable_amp=True --metrics=none
--transfer="$TRANSFER" --nosubdir
)

# COMMON_RESUME = COMMON without --transfer (resume replaces transfer)
COMMON_RESUME=(
--data="$DATA"
--cond=False --arch=ddpmpp --precond=ect
--batch=128 --batch-gpu=16 --optim=RAdam --lr=0.0001 --dropout=0.2 --augment=0
-q 256 -k 8 -b 1 -c 0 --double=10000 --ema_beta=0.9993
--fp16=True --enable_amp=True --metrics=none --nosubdir
)

# =============================================================================
# Usage:
# MODE=smoke bash CONFIRMATORY_COMMANDS.sh # 2-kimg smoke (seed 3)
# MODE=formal RUN_SEEDS=3 bash CONFIRMATORY_COMMANDS.sh # formal seed 3 (default)
# MODE=formal RUN_SEEDS="4 5" bash CONFIRMATORY_COMMANDS.sh # formal seeds 4+5
#
# Defaults: MODE=formal, RUN_SEEDS=3.
# Running `bash CONFIRMATORY_COMMANDS.sh` with no env vars will ONLY run seed 3.
# Seeds 4 & 5 are never started unless explicitly requested via RUN_SEEDS.
# =============================================================================
MODE="${MODE:-formal}"
RUN_SEEDS="${RUN_SEEDS:-3}"

FORMAL_DURATION=0.256 # 256 kimg
FORMAL_RUN_ARGS=(--tick=10 --snap=0 --dump=0 --ckpt=10 --sample_every=26 --duration=$FORMAL_DURATION)

# =============================================================================
# 1) SMOKE TESTS (seed 3, 2 kimg) -- run BEFORE the formal run
# --duration=0.002 (=2 kimg), --tick=1, --ckpt=1 so a checkpoint is saved.
# =============================================================================
if [ "$MODE" = "smoke" ]; then
SMOKE_DURATION=0.002 # 2 kimg
SMOKE_RUN_ARGS=(--tick=1 --snap=0 --dump=0 --ckpt=1 --seed=3 --duration=$SMOKE_DURATION)

# Smoke A — fixed sigmoid
$PYTHON ct_train.py "${COMMON[@]}" --mapping=sigmoid --global-gap-scale=1.0 \
"${SMOKE_RUN_ARGS[@]}" --outdir=/root/ect_runs/smoke/seed3_fixed

# Smoke B — global-only g=1.10
$PYTHON ct_train.py "${COMMON[@]}" --mapping=global_sigmoid --global-gap-scale=1.10 \
"${SMOKE_RUN_ARGS[@]}" --outdir=/root/ect_runs/smoke/seed3_global110

exit 0
fi

# =============================================================================
# 2) FORMAL PAIRED RUN (256 kimg)
# --duration=0.256 (=256 kimg), --tick=10, --ckpt=10, --sample_every=26
# Default: ONLY seed 3. To run seeds 4 & 5: RUN_SEEDS="4 5" ...
# =============================================================================
for seed in $RUN_SEEDS; do
$PYTHON ct_train.py "${COMMON[@]}" --mapping=sigmoid --global-gap-scale=1.0 --seed="$seed" \
"${FORMAL_RUN_ARGS[@]}" --outdir="/root/ect_runs/confirmatory_256k/seed${seed}_fixed"

$PYTHON ct_train.py "${COMMON[@]}" --mapping=global_sigmoid --global-gap-scale=1.10 --seed="$seed" \
"${FORMAL_RUN_ARGS[@]}" --outdir="/root/ect_runs/confirmatory_256k/seed${seed}_global110"
done

# =============================================================================
# 3) RESUME (after interruption)
# Two modes — both examples are COMMENTED OUT. Uncomment and edit as needed.
#
# a) Verification resume (writes to a NEW outdir, source untouched):
# Use this to test that resume works correctly without risking the
# authoritative run handed off to Role D.
#
# b) Actual interruption resume (writes to the SAME outdir):
# Use this only when the original run was interrupted and must continue
# in-place. Do NOT use this mode for testing.
#
# --resume replaces --transfer; --global-gap-scale is re-stated for safety.
# The training-state also carries the schedule/gap state; verify g==1.10
# after resume.
# =============================================================================

# --- a) Verification resume (new outdir, source untouched) ---
#$PYTHON ct_train.py "${COMMON_RESUME[@]}" --mapping=global_sigmoid --global-gap-scale=1.10 --seed=3 \
# "${FORMAL_RUN_ARGS[@]}" \
# --resume=/root/ect_runs/confirmatory_256k/seed3_global110/training-state-latest.pt \
# --outdir=/root/ect_runs/resume_checks/seed3_global110

# --- b) Actual interruption resume (same outdir, continues the run) ---
#$PYTHON ct_train.py "${COMMON_RESUME[@]}" --mapping=global_sigmoid --global-gap-scale=1.10 --seed=3 \
# "${FORMAL_RUN_ARGS[@]}" \
# --resume=/root/ect_runs/confirmatory_256k/seed3_global110/training-state-latest.pt \
# --outdir=/root/ect_runs/confirmatory_256k/seed3_global110
37 changes: 37 additions & 0 deletions D_HANDOFF.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,37 @@
# Role C → Role D Handoff: Confirmatory 256k (seed 3)

**Provenance:**
- training_code_sha: `3a0d603da97dd93ddbb6c7ce49e4a7351d54bb43` (the commit seed-3 was trained on; recorded in each run dir `commit_sha.txt`)
- pr_head_sha: `79143c685e5588948972c17457b1c51c7a77bb49` (this PR head; only adds docs + the resume fix, not a training baseline)

Frozen training commit: `3a0d603da97dd93ddbb6c7ce49e4a7351d54bb43` (branch role-c/confirmatory-gap-g110, PR #23 merged into main)
All checkpoints below passed integrity: loadable snapshot + training-state, kimg=256 reached, finite loss history, expected schedule, gap scale held.

| Method | Seed | Checkpoint | kimg | Commit | Integrity |
| --- | ---: | --- | ---: | --- | --- |
| Fixed | 3 | /root/ect_runs/confirmatory_256k/seed3_fixed/network-snapshot-latest.pkl | 256 | 3a0d603da97dd93ddbb6c7ce49e4a7351d54bb43 | Passed |
| Global 1.10 | 3 | /root/ect_runs/confirmatory_256k/seed3_global110/network-snapshot-latest.pkl | 256 | 3a0d603da97dd93ddbb6c7ce49e4a7351d54bb43 | Passed |

## Per-run detail

### Fixed (sigmoid, g=1.0)
- run dir: `/root/ect_runs/confirmatory_256k/seed3_fixed`
- snapshot sha256: `09a41e1e7c03dcdf5ffb93bb68687390278b4b190183dfff92bacc1bf79738d9`
- training-state: `training-state-latest.pt` (cur_nimg=256000, cur_tick=27, optimizer+scaler state)
- schedule: sigmoid; gap_over_sigmoid_gap_mean held at 1.000000 throughout
- loss: 2000 finite rows, min 13.193 / max 33.129 / final 17.063 (no NaN/Inf)
- config: `training_options.json` (adj=sigmoid, global_gap_scale=1.0)

### Global 1.10 (global_sigmoid, g=1.10)
- run dir: `/root/ect_runs/confirmatory_256k/seed3_global110`
- snapshot sha256: `24875430eea4679a416ae921c3e9ae16142f6416d2a0edf970764384ef964bed`
- training-state: `training-state-latest.pt` (cur_nimg=256000, cur_tick=27, optimizer+scaler state)
- schedule: global_sigmoid; gap_over_sigmoid_gap_mean held at 1.100000 throughout
- loss: 2000 finite rows, min 13.277 / max 31.419 / final 16.448 (no NaN/Inf)
- config: `training_options.json` (adj=global_sigmoid, global_gap_scale=1.1)

## Recommended evaluation checkpoints
Both `network-snapshot-latest.pkl` are the authoritative 256-kimg endpoints and
are recommended for Role D evaluation. The two arms differ ONLY by
`mapping`/`global_gap_scale`; all other training settings are identical
(verified by resolved-config diff). EMA weights are inside the snapshot (`ema` key).
75 changes: 75 additions & 0 deletions D_HANDOFF_SEEDS45.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,75 @@
# Role C -> Role D Handoff: Confirmatory 256k seeds 4 & 5

## Provenance

- **Executed training source commit:** `ab03f9e03b7b82425282abc3bf661067ca45875a`
- **Training-code baseline:** `6d4bc7d`
- **Original PR head / handoff-document commit:** `f299528dfb9cdb8f5be92576673242ccb0c57464`\n- **Documentation-correction commit:** recorded by the Git commit containing this revised document.

The four training jobs were launched from `ab03f9e03b7b82425282abc3bf661067ca45875a`. Its training code is equivalent to baseline `6d4bc7d`; the PR/handoff-document commit is distinct and must not be described as the executed training source.

## Training health

Each run had one initialization-time AMP-skipped optimizer step, but no NaN values were recorded in the 2000-row loss history. Finite losses were recorded from the first completed training iteration onward.

This is training-health evidence only. Loss is not an evaluation metric and must not be used to infer KID or FID; Role D's independent generation-quality evaluation remains authoritative.

## Checkpoint transfer and verification

The `/root/ect_runs/...` paths below are node-local provenance references, not a shared Role D handoff location. Role D does **not** evaluate by assuming access to the training node or its filesystem.

Transfer archive (staged outside the training node):

- archive: `D:\\seeds45_ckpts_full.tar.gz`
- size: `3,291,447,686` bytes
- archive SHA256: `1bdb147e535fe0b4f069f4106e28a7a6b065b317ca9df38cfbe9643772937608`
- extract: `tar xzf D:\\seeds45_ckpts_full.tar.gz` -> `seeds45_package_full/`

The archive contains all four authoritative `network-snapshot-latest.pkl` files (including EMA weights), their configurations, logs, loss histories, and training states.

**Required Role D acceptance check before evaluation:** obtain the archive through the agreed transfer channel, recompute its SHA256, extract it, recompute the SHA256 of each `network-snapshot-latest.pkl`, and reply on the PR confirming all five values match. Until that confirmation is posted, the archive's accessibility to Role D is not assumed.

| Method | Seed | Run | Training-node checkpoint reference | Checkpoint SHA256 | kimg | Executed source | Integrity |
| --- | ---: | --- | --- | --- | ---: | --- | --- |
| Fixed | 4 | seed4_fixed | `/root/ect_runs/confirmatory_256k/seed4_fixed/network-snapshot-latest.pkl` | `ac94e7b07e5b7628e6b14b26155fb3de09e42373497183d39aba4fe9863663c9` | 256 | `ab03f9e` | Passed on training node |
| Global 1.10 | 4 | seed4_global110 | `/root/ect_runs/confirmatory_256k/seed4_global110/network-snapshot-latest.pkl` | `62a6122a7be523aeb12875d96e96312e9c90efde9eafb75d730c75ceea0e8862` | 256 | `ab03f9e` | Passed on training node |
| Fixed | 5 | seed5_fixed | `/root/ect_runs/confirmatory_256k/seed5_fixed/network-snapshot-latest.pkl` | `21fab0e501bb27032c0e49a553b05a2800ea0fbe20a2a1d94a6bbf5276f2b72a` | 256 | `ab03f9e` | Passed on training node |
| Global 1.10 | 5 | seed5_global110 | `/root/ect_runs/confirmatory_256k/seed5_global110/network-snapshot-latest.pkl` | `491dc887990e6d9f6fde70b5d12775aaf4bfc6155b731682926b02061c253e9b` | 256 | `ab03f9e` | Passed on training node |

## Per-run evidence

### seed4_fixed (seed 4)

- method: Fixed (`mapping=sigmoid`, `global_gap_scale=1.0`)
- training-state: `/root/ect_runs/confirmatory_256k/seed4_fixed/training-state-latest.pt` (`cur_nimg=256000`)
- config/log: `training_options.json` / `log.txt`
- loss: 2000 rows; no recorded NaN; last=16.566, min=13.389, max=30.562
- `gap_over_sigmoid_gap_mean`: 1 -> 1 (target 1.0)

### seed4_global110 (seed 4)

- method: Global 1.10 (`mapping=global_sigmoid`, `global_gap_scale=1.10`)
- training-state: `/root/ect_runs/confirmatory_256k/seed4_global110/training-state-latest.pt` (`cur_nimg=256000`)
- config/log: `training_options.json` / `log.txt`
- loss: 2000 rows; no recorded NaN; last=16.555, min=13.345, max=30.100
- `gap_over_sigmoid_gap_mean`: 1.10000001913 -> 1.10000038269 (target 1.10)

### seed5_fixed (seed 5)

- method: Fixed (`mapping=sigmoid`, `global_gap_scale=1.0`)
- training-state: `/root/ect_runs/confirmatory_256k/seed5_fixed/training-state-latest.pt` (`cur_nimg=256000`)
- config/log: `training_options.json` / `log.txt`
- loss: 2000 rows; no recorded NaN; last=15.188, min=13.347, max=31.666
- `gap_over_sigmoid_gap_mean`: 1 -> 1 (target 1.0)

### seed5_global110 (seed 5)

- method: Global 1.10 (`mapping=global_sigmoid`, `global_gap_scale=1.10`)
- training-state: `/root/ect_runs/confirmatory_256k/seed5_global110/training-state-latest.pt` (`cur_nimg=256000`)
- config/log: `training_options.json` / `log.txt`
- loss: 2000 rows; no recorded NaN; last=15.215, min=13.409, max=25.291
- `gap_over_sigmoid_gap_mean`: 1.1000003469 -> 1.10000058239 (target 1.10)

## Recommended evaluation

After the required transfer verification, Role D evaluates NFE=1, NFE=2, KID-5k, and FID-5k. Fixed and global110 differ only by mapping/gap scale; the remaining settings are identical per seed.
Empty file added HANDOFF_20260804.md
Empty file.
38 changes: 33 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

Pytorch implementation for [Easy Consistency Tuning (ECT)](https://www.notion.so/gsunshine/Consistency-Models-Made-Easy-954205c0b4a24c009f78719f43b419cc).

ECT unlocks state-of-the-art (SoTA) few-step generative abilities through a simple yet principled approach.
ECT unlocks state-of-the-art (SoTA) few-step generative abilities through a simple yet principled approach.
With minimal tuning costs, ECT demonstrates promising early results and scales with training FLOPs and model sizes.

Try your own [Consistency Models](https://arxiv.org/abs/2303.01469)! You only need to fine-tune a bit. :D
Expand All @@ -13,7 +13,7 @@ Try your own [Consistency Models](https://arxiv.org/abs/2303.01469)! You only ne

## Introduction

This repository is organized in a multi-branch structure, with each branch offering a minimal implementation for a specific purpose.
This repository is organized in a multi-branch structure, with each branch offering a minimal implementation for a specific purpose.
The current branches support the following training protocols:

- `main`: ECT on CIFAR-10. Best for understanding CMs and fast prototyping.
Expand Down Expand Up @@ -44,7 +44,7 @@ Prepare the dataset in the EDM's format. See a reference [here](https://github.c

## Training

Run the following command to tune your SoTA 2-step ECM and match Consistency Distillation (CD) within 1 A100 GPU hour.
Run the following command to tune your SoTA 2-step ECM and match Consistency Distillation (CD) within 1 A100 GPU hour.

```bash
bash run_ecm_1hour.sh 1 <PORT> --desc bs128.1hour
Expand All @@ -67,7 +67,7 @@ To enable fp16 and GradScaler, add the following arguments to your script:
bash run_ecm_1hour.sh 1 <PORT> --desc bs128.1hour --fp16=True --enable_amp=True
```

For more information, please refer to this [PR](https://github.com/locuslab/ect/pull/13).
For more information, please refer to this [PR](https://github.com/locuslab/ect/pull/13).
Full support for Automatic Mixed Precision (AMP) will be added later.

## Evaluation
Expand All @@ -78,6 +78,35 @@ Run the following command to calculate FID of a pretrained checkpoint.
bash eval_ecm.sh <NGPUs> <PORT> --resume <CKPT_PATH>
```

### Fixed-seed evaluation

Role D sampling uses seeds 0-63 and verifies that work-group sizes 8 and 16 produce pixel-identical results for both NFE=1 and NFE=2. It also repeats each configuration to verify deterministic output:

See [`docs/EVALUATION_PROTOCOL.md`](docs/EVALUATION_PROTOCOL.md) for the complete protocol, metadata requirements, checkpoint-isolated output layout, and metric boundary.

```bash
bash scripts/sample_checkpoint.sh <CKPT_PATH> \
--outdir /mnt/ect_project/evaluations \
--seeds 0-63 --nfe 1 2 --mid-t 0.821 \
--work-group-size 8 --verify-work-group-size 16 \
--precision fp32
```

The output is isolated under `<outdir>/<checkpoint-stem>-<sha256-prefix>/` and contains one 8x8 grid per NFE, `metadata.json`, `sha256_manifest.txt`, and individual seed images. Keep the individual PNG files under `/mnt`; only commit the grids, metadata, manifest, scripts, and tests.

The unified metric entry point supports explicit one-step or two-step evaluation through `--nfe=1` or `--nfe=2`:

```bash
bash scripts/evaluate_checkpoint.sh 1 <PORT> <CKPT_PATH> \
--outdir ct-evals --data datasets/cifar10-32x32.zip \
--nfe=2 --mid_t=0.821 --metrics=fid50k_full
```

The frozen three-training-seed final comparison uses explicit per-sample seeds,
KID-5k as the primary proxy, FID-5k as an auxiliary proxy, and a method-blinded
A/B ballot. See [`docs/FINAL_PERFORMANCE_EVALUATION.md`](docs/FINAL_PERFORMANCE_EVALUATION.md).
These 5k-sample results are not standard FID-50k benchmarks.

## Generative Performance

### FID Evaluation
Expand Down Expand Up @@ -139,4 +168,3 @@ Feel free to drop me an email at zhengyanggeng@gmail.com if you have additional
year={2024}
}
```

18 changes: 18 additions & 0 deletions RUN_STATUS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
# Confirmatory 256k — Run Status (seed 3)

Training commit: `3a0d603da97dd93ddbb6c7ce49e4a7351d54bb43` (training_code_sha; recorded in each run dir `commit_sha.txt`)
PR head: `79143c685e5588948972c17457b1c51c7a77bb49` (pr_head_sha; docs + resume fix only, not a training baseline)
Output root: `/root/ect_runs/confirmatory_256k/`
GPU: 1x NVIDIA A100-PCIE-40GB (two runs share the GPU, ~7GB total)

| Method | Seed | Outdir | PID | Port | Start (UTC) | Status | Latest kimg |
| --- | ---: | --- | ---: | ---: | --- | --- | ---: |
| Fixed | 3 | /root/ect_runs/confirmatory_256k/seed3_fixed | (see pid.txt) | 29501 | (see start_utc.txt) | COMPLETED | 256 |
| Global 1.10 | 3 | /root/ect_runs/confirmatory_256k/seed3_global110 | (see pid.txt) | 29502 | (see start_utc.txt) | COMPLETED | 256 |

Notes:
- Both launched from identical COMMON args; differ ONLY by mapping + global-gap-scale.
- Checkpoints saved every 10 ticks (`--ckpt=10`); `network-snapshot-latest.pkl` +
`training-state-latest.pt` are the authoritative final artifacts.
- Resume command (if interrupted) is in CONFIRMATORY_COMMANDS.sh section 3.
- Patch applied: `training/ct_training_loop.py:627` `weights_only=False` (PyTorch 2.8 resume compat).
Loading