-
Notifications
You must be signed in to change notification settings - Fork 643
docs(skills): add structured DPA4 workflow support #5961
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Open
SchrodingersCattt
wants to merge
25
commits into
deepmodeling:master
Choose a base branch
from
SchrodingersCattt:docs/add-deepmd-dpa4-skill
base: master
Could not load branches
Branch not found: {{ refName }}
Loading
Could not load tags
Nothing to show
Loading
Are you sure you want to change the base?
Some commits from the old base branch may be removed from the timeline,
and old review comments may become outdated.
Open
Changes from all commits
Commits
Show all changes
25 commits
Select commit
Hold shift + click to select a range
d03cf0b
docs(skills): add DPA4 workflows
4eb0665
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] 979ba46
Refine description for deepmd-finetune-dpa4 skill
SchrodingersCattt a801ee6
docs(skills): leave DPA3 skill unchanged
6e27dc2
docs(skills): clarify pt2 inference limits
dc8ec95
docs(skills): preserve selected DPA4 heads
92476d3
docs(skills): qualify DPA4 pt2 deployment
0e24272
docs(skills): pin online LAMMPS runtime
c41e68a
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] adc8616
docs(skills): qualify pt2 atomic outputs
263c340
docs(skills): make LAMMPS help noninteractive
588023c
docs(skills): clarify DPA4 checkpoint and LoRA selection
SchrodingersCattt d5e2d6c
docs(skills): streamline DPA4 finetuning guidance
SchrodingersCattt f42fdde
docs(skills): add MatMaster DPA4 workflows
weiqichen77 bbe64e0
docs(skills): address DPA4 workflow validation gaps
SchrodingersCattt a9f1279
Revert "docs(skills): add MatMaster DPA4 workflows"
d7c988a
Merge remote-tracking branch 'upstream/master' into docs/add-deepmd-d…
06bc4d9
docs(skills): gate LAMMPS runtime by capability
2541ae0
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] 4ccbd91
docs(skills): define complete held-out evaluation
0807ca6
docs(skills): harden DPA4 LAMMPS deployment
fe8168c
docs(skills): fix DPA4 evaluation and export contracts
dd3fb6a
test(skills): cover DPA4 review contracts
3a871cd
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] 84923b8
docs(skills): reserve held-out detail roots
File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,202 @@ | ||
| --- | ||
| name: deepmd-finetune-dpa4 | ||
| description: Fine-tune a DPA4 model in DeePMD-kit. Use for standard or LoRA fine-tuning from a DPA4/SeZM .pt checkpoint, validation and .pt2 export. | ||
| compatibility: Requires deepmd-kit with the PyTorch backend. DPA4/SeZM training is GPU-oriented. | ||
| license: LGPL-3.0-or-later | ||
| metadata: | ||
| author: SchrodingersCattt | ||
| version: '1.0' | ||
| repository: https://github.com/deepmodeling/deepmd-kit | ||
| --- | ||
|
|
||
| # DeePMD-kit Fine-tuning: DPA4 | ||
|
|
||
| Fine-tune a DPA4/SeZM checkpoint on downstream DeePMD data. This skill covers | ||
| single-task standard and LoRA fine-tuning. Do not infer the model family from a | ||
| `.pt` suffix or filename: DPA3 and DPA4 checkpoints use the same suffix. | ||
|
|
||
| ## Route the checkpoint | ||
|
|
||
| If the user has not already established the model family, inspect the stored | ||
| configuration: | ||
|
|
||
| ```bash | ||
| dp --pt show pretrained.pt descriptor fitting-net type-map | ||
| ``` | ||
|
|
||
| Use this skill only when the descriptor/model configuration identifies DPA4 or | ||
| SeZM. If the checkpoint is multi-task, inspect its branches before selecting a | ||
| head: | ||
|
|
||
| ```bash | ||
| dp --pt show pretrained.pt model-branch descriptor type-map | ||
| ``` | ||
|
|
||
| Do not guess a branch. Use `deepmd-finetune-dpa3` instead when the descriptor is | ||
| DPA3, and stop when the family cannot be established. | ||
|
|
||
| ## Obtain a pretrained checkpoint | ||
|
|
||
| Fine-tuning requires a DPA4/SeZM training checkpoint (`.pt`), not a `.pt2` | ||
| deployment archive. Check whether the installed version provides one: | ||
|
|
||
| ```bash | ||
| dp pretrained download -h | ||
| ``` | ||
|
|
||
| Use only a listed model or a checkpoint supplied by the user or its publisher. | ||
| Record its source and DeePMD-kit version, then verify its descriptor, branch, | ||
| architecture, and `type_map` before use. | ||
|
|
||
| ## Before fine-tuning | ||
|
|
||
| 1. Confirm the checkpoint exists and can be inspected. | ||
| 1. Confirm training, validation, and held-out systems, labels, and element type maps. | ||
| 1. Split correlated frames by independent system, trajectory, or source family; | ||
| do not create a nominal held-out set by randomly splitting adjacent frames. | ||
| 1. Validate each DeePMD system before training: `natoms` is the number of tokens | ||
| in `type.raw`, coordinate and force widths are `3 * natoms`, and every used | ||
| label is finite and frame-aligned. | ||
| 1. Start from the exact checkpoint architecture. Introducing new element types, | ||
| changing architecture, or combining specialized spin/property/multi-task | ||
| configurations requires separate compatibility validation. | ||
| 1. Choose standard fine-tuning or LoRA. Do not assume a built-in DPA4 model name; | ||
| check `dp pretrained download -h` for the installed version. | ||
|
|
||
| ## Decide whether to use LoRA | ||
|
|
||
| Use standard fine-tuning by default. Use LoRA only for a single-task target when | ||
| parameter-efficient adaptation is wanted and the exact base architecture is | ||
| known. LoRA is enabled by a non-null `model.lora` block in the new input; the | ||
| pretrained checkpoint does not need to contain LoRA. Multi-task LoRA targets are | ||
| unsupported. | ||
|
|
||
| Periodic LoRA checkpoints retain adapters and can resume training. Best | ||
| checkpoints may merge the adapters into ordinary DPA4 weights, so absence of | ||
| LoRA metadata does not prove LoRA was never used. | ||
|
|
||
| ## Standard fine-tuning | ||
|
|
||
| The model section in `input.json` must match the checkpoint unless the standard | ||
| pretrained-script mechanism is deliberately used: | ||
|
|
||
| ```bash | ||
| dp --pt train input.json --finetune pretrained.pt | ||
| ``` | ||
|
|
||
| When fine-tuning a single-task target from a multi-task checkpoint and the | ||
| intent is to preserve a particular pretrained fitting head, pass the branch | ||
| selected above: | ||
|
|
||
| ```bash | ||
| dp --pt train input.json --finetune pretrained.pt --model-branch SELECTED_BRANCH | ||
| ``` | ||
|
|
||
| If `--model-branch` is omitted, the fitting net may be initialized from the | ||
| `RANDOM` branch instead. A multi-task target uses `finetune_head` in each target | ||
| branch rather than the command-line option. | ||
|
|
||
| `--use-pretrain-script` replaces only the target `model.descriptor` and | ||
| `model.fitting_net` from the checkpoint. It does not restore the complete model | ||
| configuration. Inspect and reproduce all other required model-level fields in | ||
| the target input, including `type`, spin settings, bridging method/radii, and | ||
| task-specific options: | ||
|
|
||
| ```bash | ||
| dp --pt train input.json --finetune pretrained.pt --use-pretrain-script | ||
| ``` | ||
|
|
||
| Inspect the resulting configuration and run a bounded initial segment before a | ||
| long training job. Do not combine model-specific additions with | ||
| `--use-pretrain-script` unless that combination has been validated. | ||
|
|
||
| ## LoRA fine-tuning | ||
|
|
||
| DPA4/SeZM supports LoRA adapters for single-task fine-tuning. Copy the exact base | ||
| architecture into `lora_ft.json`, then add: | ||
|
|
||
| ```json | ||
| { | ||
| "model": { | ||
| "type": "dpa4", | ||
| "lora": { | ||
| "rank": 16, | ||
| "alpha": 16.0 | ||
| } | ||
| } | ||
| } | ||
| ``` | ||
|
|
||
| Run: | ||
|
|
||
| ```bash | ||
| dp --pt train lora_ft.json --finetune pretrained.pt | ||
|
SchrodingersCattt marked this conversation as resolved.
|
||
| ``` | ||
|
|
||
| When `pretrained.pt` is multi-task, preserve the selected fitting head: | ||
|
|
||
| ```bash | ||
| dp --pt train lora_ft.json --finetune pretrained.pt \ | ||
| --model-branch SELECTED_BRANCH | ||
| ``` | ||
|
|
||
| Use the shorter command only for a single-task source checkpoint. | ||
|
|
||
| The JSON fragment above is not a complete training input. Adapt the full public | ||
| example at `../../examples/water/dpa4/lora_ft.json`, but copy the exact | ||
| architecture from the source checkpoint before adding `model.lora`. Do not add | ||
| `--use-pretrain-script` unless a targeted test confirms that LoRA is retained. | ||
|
|
||
| ## Monitor and validate | ||
|
|
||
| Monitor `lcurve.out` for non-finite values and train/validation divergence. | ||
| Select a checkpoint using validation data, then follow the | ||
| [complete held-out evaluation](../deepmd-python-inference/references/held-out-evaluation.md) | ||
| with that exact native checkpoint and every held-out system. Export for deployment | ||
| only after the complete evaluation meets the task's declared thresholds. | ||
|
|
||
| ## Export and test | ||
|
|
||
| Before export, read | ||
| `../deepmd-python-inference/references/dpa4-freeze-policy.md` and record the | ||
| selected freeze-time inference environment. | ||
|
|
||
| DPA4/SeZM uses the `.pt2` AOTInductor export path rather than the conventional | ||
| PyTorch `.pth` freeze path: | ||
|
|
||
| ```bash | ||
| dp --pt freeze -c ckpt/model.ckpt.pt -o finetuned_model | ||
| dp test -m finetuned_model.pt2 -s /path/to/test_system -n 30 | ||
| ``` | ||
|
|
||
| The freeze command detects DPA4/SeZM and writes `finetuned_model.pt2`. Validate | ||
| the exported archive in the target environment before deployment. | ||
|
|
||
| For a multi-task checkpoint, freeze the selected head explicitly: | ||
|
|
||
| ```bash | ||
| dp --pt freeze -c ckpt/model.ckpt.pt -o finetuned_model --head SELECTED_BRANCH | ||
| ``` | ||
|
|
||
| The resulting `.pt2` contains the selected single head; do not pass a branch | ||
| again when loading that archive. | ||
|
|
||
| ## Checklist | ||
|
|
||
| - [ ] The stored descriptor identifies DPA4/SeZM; the `.pt` suffix was not used as proof. | ||
| - [ ] The checkpoint source, DeePMD-kit revision, architecture, and type map are recorded. | ||
| - [ ] The intended branch is explicit for a multi-task checkpoint. | ||
| - [ ] Training, validation, and held-out systems are independent by source family. | ||
| - [ ] Every admitted system has consistent atom counts, shapes, labels, and type maps. | ||
| - [ ] The input architecture is compatible with the checkpoint. | ||
| - [ ] Standard fine-tuning versus LoRA was selected from the task layout and domain shift. | ||
| - [ ] A resumable LoRA checkpoint is distinguished from a merged best checkpoint. | ||
| - [ ] LoRA uses a complete base configuration and is not silently overwritten. | ||
| - [ ] Complete held-out metrics, sample counts, and reference-label scales are reported. | ||
| - [ ] The selected `.pt` checkpoint was exported to and tested as `.pt2`. | ||
|
|
||
| ## References | ||
|
|
||
| - [DPA4 model and LoRA documentation](https://docs.deepmodeling.com/projects/deepmd/en/latest/model/dpa4.html) | ||
| - [Fine-tuning documentation](https://docs.deepmodeling.com/projects/deepmd/en/latest/train/finetuning.html) | ||
| - [Show model information](https://docs.deepmodeling.com/projects/deepmd/en/latest/model/show-model-info.html) | ||
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
24 changes: 24 additions & 0 deletions
24
skills/deepmd-python-inference/references/dpa4-freeze-policy.md
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,24 @@ | ||
| # DPA4 freeze-time inference policy | ||
|
|
||
| Read this reference before exporting a DPA4 `.pt2`. The generated archive may | ||
| embed inference choices, so do not inherit unknown values of | ||
| `DP_TRITON_INFER`, `DP_TF32_INFER`, or `DP_AMP_INFER` from the shell. | ||
|
|
||
| Choose one explicit policy for the target node. A conservative production | ||
| baseline that avoids the slow default while retaining full-precision | ||
| accumulation is: | ||
|
|
||
| ```bash | ||
| export DP_TRITON_INFER=1 | ||
| export DP_TF32_INFER=0 | ||
| export DP_AMP_INFER=0 | ||
| dp --pt freeze -c model.ckpt.pt -o frozen_model | ||
| ``` | ||
|
|
||
| `DP_TRITON_INFER=2` autotunes for the current hardware and therefore reinforces | ||
| the requirement to freeze and run on the same physical node. Levels 1 and 2 keep | ||
| FP32 accumulation. Level 3, TF32, and AMP change the numerical policy and require | ||
| task-specific accuracy and stability validation before production. | ||
|
|
||
| Record all three values, device/runtime identity, input and output hashes, freeze | ||
| command, log, and true exit code with the artifact. |
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
Uh oh!
There was an error while loading. Please reload this page.