Skip to content

Commit d5e2d6c

Browse files
docs(skills): streamline DPA4 finetuning guidance
Remove smoke-model-specific discussion and detailed checkpoint inspection scripts, retaining concise general requirements for checkpoint selection and LoRA usage.
1 parent 588023c commit d5e2d6c

1 file changed

Lines changed: 18 additions & 111 deletions

File tree

skills/deepmd-finetune-dpa4/SKILL.md

Lines changed: 18 additions & 111 deletions
Original file line numberDiff line numberDiff line change
@@ -37,38 +37,16 @@ DPA3, and stop when the family cannot be established.
3737

3838
## Obtain a pretrained checkpoint
3939

40-
Fine-tuning requires a DPA4/SeZM **training checkpoint** (`.pt`), not a compiled
41-
`.pt2` deployment archive. First check the registry exposed by the installed
42-
version:
40+
Fine-tuning requires a DPA4/SeZM training checkpoint (`.pt`), not a `.pt2`
41+
deployment archive. Check whether the installed version provides one:
4342

4443
```bash
4544
dp pretrained download -h
4645
```
4746

48-
Use a built-in model only when an exact DPA4/SeZM name is listed there. Do not
49-
guess `dp pretrained download DPA4`: some releases have no registered DPA4
50-
checkpoint.
51-
52-
For a workflow smoke test, a DeePMD-kit source checkout contains:
53-
54-
```text
55-
examples/water/dpa4/lmp/pretrained.pt
56-
```
57-
58-
Use that file directly:
59-
60-
```bash
61-
cp examples/water/dpa4/lmp/pretrained.pt ./pretrained.pt
62-
dp --pt show pretrained.pt descriptor fitting-net type-map
63-
```
64-
65-
It is a compact O/H smoke-test model, not a general-purpose pretrained
66-
potential. For scientific fine-tuning, obtain a checkpoint from its documented
67-
publisher or train one with the same DeePMD-kit revision that will perform the
68-
fine-tuning. Pin and record the model source, checksum, DeePMD-kit revision,
69-
architecture, task branch, and `type_map`. Reject a checkpoint whose descriptor,
70-
element set/order, or architecture does not match the intended target. Prefer a
71-
bounded one-step compatibility run before a long job.
47+
Use only a listed model or a checkpoint supplied by the user or its publisher.
48+
Record its source and DeePMD-kit version, then verify its descriptor, branch,
49+
architecture, and `type_map` before use.
7250

7351
## Before fine-tuning
7452

@@ -83,83 +61,15 @@ bounded one-step compatibility run before a long job.
8361

8462
## Decide whether to use LoRA
8563

86-
Use **standard fine-tuning** by default when the target is multi-task, the domain
87-
shift is large, all parameters should adapt, or the workflow combines untested
88-
spin/property/denoising/ZBL changes. Use **LoRA** when the target is single-task,
89-
parameter-efficient adaptation is desired, the downstream domain is reasonably
90-
close to pretraining, and the exact base architecture is known. DPA4 LoRA is
91-
supported by `dp --pt`; do not use it with the exportable training backend.
92-
93-
The pretrained checkpoint does not need to contain LoRA. LoRA is enabled by the
94-
new fine-tuning input through a non-null `model.lora` block. Check an input with:
95-
96-
```bash
97-
python - input.json <<'PY'
98-
import json
99-
import sys
100-
101-
with open(sys.argv[1], encoding="utf-8") as stream:
102-
config = json.load(stream)
103-
104-
model = config.get("model", {})
105-
branches = model.get("model_dict")
106-
lora = model.get("lora")
107-
print("multi_task =", isinstance(branches, dict))
108-
print("lora =", lora)
109-
print("use_lora =", lora is not None)
110-
if isinstance(branches, dict) and lora is not None:
111-
raise SystemExit("DPA4 LoRA targets must be single-task")
112-
PY
113-
```
114-
115-
Interpretation:
116-
117-
- absent or `null` `model.lora` means standard fine-tuning;
118-
- a `model.lora` mapping means LoRA fine-tuning;
119-
- a multi-task target plus LoRA is unsupported.
120-
121-
To diagnose whether a `.pt` file still contains **active, unmerged** LoRA state,
122-
inspect its saved model parameters and adapter tensors without executing pickled
123-
code:
124-
125-
```bash
126-
python - pretrained.pt <<'PY'
127-
import sys
128-
import torch
129-
130-
raw = torch.load(sys.argv[1], map_location="cpu", weights_only=True)
131-
state = raw["model"] if isinstance(raw, dict) and "model" in raw else raw
132-
if not isinstance(state, dict):
133-
raise TypeError(f"Unsupported checkpoint payload: {type(state).__name__}")
134-
135-
extra = state.get("_extra_state", {})
136-
params = extra.get("model_params", {}) if isinstance(extra, dict) else {}
137-
configured = params.get("lora") if isinstance(params, dict) else None
138-
markers = (".A_by_l", ".B_by_l", ".A_m0", ".B_m0", ".A_m.", ".B_m.", ".lora_scaling")
139-
adapter_keys = sorted(
140-
key for key in state if any(marker in key for marker in markers)
141-
)
142-
143-
print("configured_lora =", configured)
144-
print("adapter_tensor_count =", len(adapter_keys))
145-
for key in adapter_keys[:20]:
146-
print("adapter_tensor =", key)
147-
148-
if configured is not None and adapter_keys:
149-
print("classification = active/unmerged LoRA checkpoint")
150-
elif configured is not None or adapter_keys:
151-
print("classification = inconsistent or transitional; inspect before use")
152-
else:
153-
print("classification = plain checkpoint or merged LoRA checkpoint")
154-
PY
155-
```
64+
Use standard fine-tuning by default. Use LoRA only for a single-task target when
65+
parameter-efficient adaptation is wanted and the exact base architecture is
66+
known. LoRA is enabled by a non-null `model.lora` block in the new input; the
67+
pretrained checkpoint does not need to contain LoRA. Multi-task LoRA targets are
68+
unsupported.
15669

157-
Do not infer training history from the last classification. Validation-selected
158-
best LoRA checkpoints fold adapter deltas into ordinary DPA4 weights and remove
159-
LoRA metadata/tensors, so they intentionally look plain. Such merged checkpoints
160-
are suitable for evaluation, new fine-tuning, and `.pt2` export, but do not carry
161-
the optimizer/EMA state needed to resume the same LoRA run. Periodic/final LoRA
162-
checkpoints retain active adapters and are the resumable form.
70+
Periodic LoRA checkpoints retain adapters and can resume training. Best
71+
checkpoints may merge the adapters into ordinary DPA4 weights, so absence of
72+
LoRA metadata does not prove LoRA was never used.
16373

16474
## Standard fine-tuning
16575

@@ -217,12 +127,9 @@ dp --pt train lora_ft.json --finetune pretrained.pt
217127
```
218128

219129
The JSON fragment above is not a complete training input. Adapt the full public
220-
example at `../../examples/water/dpa4/lora_ft.json`, but copy the architecture
221-
from the actual source checkpoint before adding `model.lora`. The compact
222-
`examples/water/dpa4/lmp/pretrained.pt` does not match the larger architecture
223-
in the maintained `lora_ft.json` and must not be paired with it unchanged. Do
224-
not add `--use-pretrain-script` to this LoRA command unless a targeted test
225-
confirms that the intended LoRA configuration is retained.
130+
example at `../../examples/water/dpa4/lora_ft.json`, but copy the exact
131+
architecture from the source checkpoint before adding `model.lora`. Do not add
132+
`--use-pretrain-script` unless a targeted test confirms that LoRA is retained.
226133

227134
## Monitor and validate
228135

@@ -256,12 +163,12 @@ again when loading that archive.
256163
## Checklist
257164

258165
- [ ] The stored descriptor identifies DPA4/SeZM; the `.pt` suffix was not used as proof.
259-
- [ ] The checkpoint source, checksum, DeePMD-kit revision, architecture, and type map are recorded.
166+
- [ ] The checkpoint source, DeePMD-kit revision, architecture, and type map are recorded.
260167
- [ ] The intended branch is explicit for a multi-task checkpoint.
261168
- [ ] Training, validation, and held-out test systems are separate.
262169
- [ ] The input architecture is compatible with the checkpoint.
263170
- [ ] Standard fine-tuning versus LoRA was selected from the task layout and domain shift.
264-
- [ ] Active LoRA state versus a merged best checkpoint is interpreted correctly.
171+
- [ ] A resumable LoRA checkpoint is distinguished from a merged best checkpoint.
265172
- [ ] LoRA uses a complete base configuration and is not silently overwritten.
266173
- [ ] Training and held-out metrics are finite and reported with units.
267174
- [ ] The selected `.pt` checkpoint was exported to and tested as `.pt2`.

0 commit comments

Comments
 (0)