-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathrun_200.err
More file actions
86 lines (81 loc) · 6.51 KB
/
Copy pathrun_200.err
File metadata and controls
86 lines (81 loc) · 6.51 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
[rank: 0] Seed set to 42
0%| | 0/1620 [00:00<?, ?it/s] 3%|▎ | 54/1620 [00:00<00:03, 520.36it/s] 7%|▋ | 107/1620 [00:00<00:02, 521.40it/s] 10%|▉ | 160/1620 [00:00<00:02, 514.51it/s] 13%|█▎ | 212/1620 [00:00<00:02, 512.29it/s] 16%|█▋ | 267/1620 [00:00<00:02, 520.97it/s] 20%|█▉ | 320/1620 [00:00<00:02, 472.53it/s] 23%|██▎ | 379/1620 [00:00<00:02, 507.70it/s] 28%|██▊ | 446/1620 [00:00<00:02, 548.11it/s] 31%|███ | 502/1620 [00:00<00:02, 519.44it/s] 34%|███▍ | 555/1620 [00:01<00:02, 494.06it/s] 37%|███▋ | 605/1620 [00:01<00:02, 453.73it/s] 40%|████ | 652/1620 [00:01<00:02, 415.67it/s] 43%|████▎ | 695/1620 [00:01<00:02, 407.69it/s] 45%|████▌ | 737/1620 [00:01<00:02, 400.41it/s] 48%|████▊ | 779/1620 [00:01<00:02, 405.06it/s] 51%|█████ | 824/1620 [00:01<00:01, 412.05it/s] 53%|█████▎ | 866/1620 [00:01<00:01, 411.41it/s] 56%|█████▌ | 908/1620 [00:02<00:01, 392.95it/s] 59%|█████▉ | 960/1620 [00:02<00:01, 422.46it/s] 62%|██████▏ | 1003/1620 [00:02<00:01, 387.76it/s] 64%|██████▍ | 1043/1620 [00:02<00:01, 355.81it/s] 67%|██████▋ | 1080/1620 [00:02<00:01, 340.66it/s] 69%|██████▉ | 1117/1620 [00:02<00:01, 343.29it/s] 72%|███████▏ | 1170/1620 [00:02<00:01, 392.98it/s] 76%|███████▌ | 1224/1620 [00:02<00:00, 430.37it/s] 78%|███████▊ | 1268/1620 [00:02<00:00, 402.18it/s] 81%|████████ | 1310/1620 [00:03<00:00, 404.17it/s] 83%|████████▎ | 1352/1620 [00:03<00:00, 380.24it/s] 86%|████████▌ | 1392/1620 [00:03<00:00, 384.86it/s] 88%|████████▊ | 1431/1620 [00:03<00:00, 349.84it/s] 91%|█████████▏| 1480/1620 [00:03<00:00, 384.44it/s] 94%|█████████▍| 1520/1620 [00:03<00:00, 381.89it/s] 97%|█████████▋| 1565/1620 [00:03<00:00, 396.27it/s] 99%|█████████▉| 1606/1620 [00:03<00:00, 396.25it/s] /csghome/sm330/emb_env/lib/python3.13/site-packages/torch/utils/data/dataloader.py:627: UserWarning: This DataLoader will create 32 worker processes in total. Our suggested max number of worker in current system is 2, which is smaller than what this DataLoader is going to create. Please be aware that excessive worker creation might get DataLoader running slow or even freeze, lower the worker number to avoid potential slowness/freeze if necessary.
warnings.warn(
Using bfloat16 Automatic Mixed Precision (AMP)
GPU available: True (cuda), used: True
TPU available: False, using: 0 TPU cores
HPU available: False, using: 0 HPUs
/csghome/sm330/emb_env/lib/python3.13/site-packages/pytorch_lightning/trainer/connectors/logger_connector/logger_connector.py:76: Starting from v1.9.0, `tensorboardX` has been removed as a dependency of the `pytorch_lightning` package, due to potential conflicts with other packages in the ML ecosystem. For this reason, `logger=True` will use `CSVLogger` as the default logger, unless the `tensorboard` or `tensorboardX` packages are found. Please `pip install lightning[extra]` or one of them to enable TensorBoard support by default
You are using a CUDA device ('NVIDIA A30') that has Tensor Cores. To properly utilize them, you should set `torch.set_float32_matmul_precision('medium' | 'high')` which will trade-off precision for performance. For more details, read https://pytorch.org/docs/stable/generated/torch.set_float32_matmul_precision.html#torch.set_float32_matmul_precision
0%| | 0/1620 [00:00<?, ?it/s] 10%|▉ | 154/1620 [00:00<00:00, 1531.24it/s] 19%|█▉ | 309/1620 [00:00<00:00, 1539.49it/s] 29%|██▉ | 468/1620 [00:00<00:00, 1559.99it/s] 39%|███▉ | 636/1620 [00:00<00:00, 1605.23it/s] 50%|████▉ | 805/1620 [00:00<00:00, 1633.51it/s] 60%|██████ | 972/1620 [00:00<00:00, 1644.58it/s] 70%|███████ | 1142/1620 [00:00<00:00, 1661.57it/s] 81%|████████ | 1313/1620 [00:00<00:00, 1674.53it/s] 91%|█████████▏| 1482/1620 [00:00<00:00, 1678.01it/s] LOCAL_RANK: 0 - CUDA_VISIBLE_DEVICES: [0]
/csghome/sm330/emb_env/lib/python3.13/site-packages/pytorch_lightning/utilities/model_summary/model_summary.py:231: Precision bf16-mixed is not supported by the model summary. Estimated model size in MB will not be accurate. Using 32 bits instead.
| Name | Type | Params | Mode
-----------------------------------------------------
0 | model | FusionModel | 5.5 M | train
1 | loss_fn | Embed2HeightLoss | 0 | train
-----------------------------------------------------
5.5 M Trainable params
0 Non-trainable params
5.5 M Total params
22.056 Total estimated model params size (MB)
230 Modules in train mode
0 Modules in eval mode
SLURM auto-requeueing enabled. Setting signal handlers.
/csghome/sm330/emb_env/lib/python3.13/site-packages/pytorch_lightning/utilities/data.py:79: Trying to infer the `batch_size` from an ambiguous collection. The batch size we found is 16. To avoid any miscalculations, use `self.log(..., batch_size=batch_size)`.
/csghome/sm330/emb_env/lib/python3.13/site-packages/pytorch_lightning/utilities/data.py:79: Trying to infer the `batch_size` from an ambiguous collection. The batch size we found is 4. To avoid any miscalculations, use `self.log(..., batch_size=batch_size)`.
Metric val_total_loss improved. New best score: 1.608
Metric val_total_loss improved by 0.581 >= min_delta = 0.0. New best score: 1.028
Metric val_total_loss improved by 0.155 >= min_delta = 0.0. New best score: 0.873
Metric val_total_loss improved by 0.177 >= min_delta = 0.0. New best score: 0.696
Metric val_total_loss improved by 0.056 >= min_delta = 0.0. New best score: 0.640
Metric val_total_loss improved by 0.016 >= min_delta = 0.0. New best score: 0.624
Metric val_total_loss improved by 0.019 >= min_delta = 0.0. New best score: 0.604
Metric val_total_loss improved by 0.002 >= min_delta = 0.0. New best score: 0.602
Metric val_total_loss improved by 0.020 >= min_delta = 0.0. New best score: 0.582
Metric val_total_loss improved by 0.002 >= min_delta = 0.0. New best score: 0.580
Metric val_total_loss improved by 0.002 >= min_delta = 0.0. New best score: 0.578