Skip to content

Commit 96b2ce9

Browse files
committed
[0.14.1] 8B rewriter runs the sft
1 parent 8b3830d commit 96b2ce9

13 files changed

Lines changed: 657 additions & 145 deletions

File tree

CHANGELOG.md

Lines changed: 56 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -6,6 +6,62 @@ The version in `pyproject.toml`, the git tag and the release on GitHub always sa
66
the same thing; the release workflow refuses a tag that disagrees with
77
`pyproject.toml`, or one that neither changelog has a section for.
88

9+
## 0.14.1 - 2026-08-21
10+
11+
### Added
12+
13+
- **The 8B rewriter runs the safetensors build as well as the GGUF one.** The
14+
adapter is published as `adapter_model.safetensors` on
15+
`lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA-8B`, trained on the official
16+
`Qwen/Qwen3-VL-8B-Instruct` folder, and until now the node could only reach
17+
the GGUF conversion of both. Now the base model list offers either shape and
18+
the node picks the engine from it.
19+
20+
What the route buys is residency. A GGUF base with frames runs through
21+
`llama-mtmd-cli`, a fresh process each time, so `keep_model_loaded` could only
22+
ever work for T2VA; safetensors loads in ComfyUI's own process through
23+
Transformers and PEFT and stays there for every task. What it costs is the
24+
download: 17.5 GB of base and 2.8 GB of adapter against 4.7 + 0.7 + 0.7.
25+
26+
- **`quantization` on the 8B rewriter**, the same widget the 27B has and for the
27+
same reason: `nf4` needs about 8 GB of VRAM, `int8` about 13, `bfloat16` about
28+
20. Ignored for a GGUF base, which carries its own.
29+
30+
- **The "Open model list" button on two nodes that were missing it** - the 8B
31+
rewriter and Multi Reference Caption. Both pick their model from `models.json`
32+
like every other node, so both had every reason to offer the button and no
33+
reason not to.
34+
35+
### Changed
36+
37+
- **The base-model check knows which adapter is about to be applied.** It
38+
compared four numbers from `config.json` against constants that always
39+
described Qwen3.6-27B, so it could only ever answer for the 27B; the four now
40+
travel together as a `Shape` and the caller says which one it means. A
41+
Qwen3-VL-4B is refused for the 8B adapter by name and number -
42+
`hidden_size is 2560, the adapter needs 4096` - rather than after the
43+
download. Nothing changes for the 27B, which passes its own shape.
44+
45+
- The transformers engine splits generation into a step that prepares inputs and
46+
a step that runs them, so the multimodal path could be added without a second
47+
copy of the sampling, streaming and interrupt handling. A checkpoint's
48+
processor is loaded and cached beside its model, and a text-only checkpoint
49+
simply has none.
50+
51+
- **The adapter sections reach your `models.json` at last.** `adapters` and
52+
`adapters_8b` are dicts, and the merge that keeps a live list current is set
53+
algebra over named entries in a *list*, so it had never walked them: no
54+
adapter entry has ever been written into anybody's copy. Reading still worked,
55+
because an unconfigured entry falls back to the packaged value, but there was
56+
no line to point at a conversion of your own - which is the whole reason the
57+
file is yours to edit. They are now merged one format at a time and recorded
58+
in `seed_offered` under the section name, so a format published later arrives
59+
and one you delete on purpose stays deleted.
60+
61+
- A chat template found on the tokenizer rather than on the processor is used
62+
rather than refused. Qwen3-VL-4B is one such checkpoint, and its tokenizer's
63+
template writes the same image placeholders.
64+
965
## 0.14.0 - 2026-08-21
1066

1167
### Added

CHANGELOG_RU.md

Lines changed: 54 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -6,6 +6,60 @@
66
workflow релиза отклонит тег, который расходится с `pyproject.toml`, и тег, для
77
которого нет раздела ни в одном из двух changelog'ов.
88

9+
## 0.14.1 - 2026-08-21
10+
11+
### Добавлено
12+
13+
- **Ререйтер 8B работает и на сборке safetensors, а не только на GGUF.** Адаптер
14+
издан как `adapter_model.safetensors` в
15+
`lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA-8B` и обучен на официальной папке
16+
`Qwen/Qwen3-VL-8B-Instruct`, а нода до сих пор умела дотянуться только до
17+
GGUF-конвертации того и другого. Теперь список базовых моделей предлагает обе
18+
формы, и нода выбирает движок по ней.
19+
20+
Что этот путь даёт - резидентность. База в GGUF с кадрами идёт через
21+
`llama-mtmd-cli`, каждый раз новым процессом, поэтому `keep_model_loaded` мог
22+
работать только для T2VA; safetensors грузится в сам процесс ComfyUI через
23+
Transformers и PEFT и остаётся там для любой задачи. Чего стоит - загрузки:
24+
17.5 ГБ базы и 2.8 ГБ адаптера против 4.7 + 0.7 + 0.7.
25+
26+
- **`quantization` на ноде 8B** - тот же виджет, что у 27B, и по той же причине:
27+
`nf4` требует около 8 ГБ VRAM, `int8` около 13, `bfloat16` около 20. Для базы
28+
в GGUF игнорируется: там своя квантизация.
29+
30+
- **Кнопка "Open model list" на двух нодах, где её не было** - на ререйтере 8B и
31+
на Multi Reference Caption. Обе берут модель из `models.json`, как и все
32+
остальные, так что причин её не показывать не было никаких.
33+
34+
### Изменено
35+
36+
- **Проверка базовой модели знает, какой адаптер сейчас будут применять.** Она
37+
сравнивала четыре числа из `config.json` с константами, которые всегда
38+
описывали Qwen3.6-27B, и потому могла отвечать только за 27B. Теперь эти
39+
четыре числа ездят вместе как `Shape`, и вызывающий говорит, о каком адаптере
40+
речь. Qwen3-VL-4B отклоняется для адаптера 8B по имени и числам -
41+
`hidden_size is 2560, the adapter needs 4096`, - а не после загрузки. Для 27B
42+
не меняется ничего: он передаёт свою форму.
43+
44+
- Движок на transformers разделён на подготовку входа и его прогон, чтобы
45+
мультимодальный путь появился без второй копии сэмплинга, стриминга и
46+
обработки прерывания. Процессор чекпоинта грузится и кэшируется рядом с
47+
моделью, а у текстового чекпоинта его просто нет.
48+
49+
- **Секции с адаптерами наконец доходят до вашего `models.json`.** `adapters` и
50+
`adapters_8b` — словари, а слияние, которое держит списки актуальными, есть
51+
теоретико-множественная операция над именованными записями *списка*, так что до них оно
52+
никогда не доходило: ни одна запись об адаптере никогда не попадала ни в чей файл.
53+
Чтение при этом работало — ненастроенная запись откатывается к пакетному значению, —
54+
но строки, которую можно нацелить на свою конвертацию, не было, а ведь ради этого
55+
файл и отдаётся вам. Теперь они сливаются поформатно и записываются в
56+
`seed_offered` под именем секции: формат, изданный позже, приходит, а удалённый
57+
вами намеренно — остаётся удалённым.
58+
59+
- Шаблон чата, лежащий на токенизаторе, а не на процессоре, теперь используется,
60+
а не приводит к отказу. Qwen3-VL-4B - как раз такой чекпоинт, и шаблон его
61+
токенизатора расставляет те же плейсхолдеры под картинки.
62+
963
## 0.14.0 - 2026-08-21
1064

1165
### Добавлено

README.md

Lines changed: 24 additions & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -141,21 +141,29 @@ interchangeable downstream.
141141
**Inputs**
142142

143143
- `prompt`, `resolution`, `duration`, `greedy`, `seed` — as above.
144-
- `model` — a Qwen3-VL base and its projector, which are two files from the same
145-
conversion. Entries prefixed `on disk:` are pairs already in your model
146-
folders. **Only the 8B fits the adapter**; a Qwen3-VL of another size loads and
147-
then runs as a plain model with no rewriter, and the node says so before
148-
downloading anything.
144+
- `model` — a Qwen3-VL-8B base, in either shape the adapter is published for.
145+
A **GGUF** entry is two files from the same conversion, the model and its
146+
projector. A **safetensors** entry is the official
147+
[Qwen3-VL-8B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-8B-Instruct)
148+
folder, which is what the adapter was trained on and what it is published as.
149+
Entries prefixed `on disk:` are already in your model folders.
150+
**Only the 8B fits the adapter**; a Qwen3-VL of another size is refused by name
151+
and number before anything is downloaded.
152+
- `quantization` — how to load a **safetensors** base: `nf4` needs about 8 GB of
153+
VRAM, `int8` about 13, `bfloat16` about 20. Ignored for GGUF, which carries its
154+
own.
149155
- `task``T2VA`, `I2VA`, `FL2VA`, `L2VA`. The model's own name for these is
150156
T2AV, I2AV, FL2AV and L2AV; they are the same four tasks.
151157
- `first_frame` / `last_frame` — optional IMAGE inputs. `I2VA` reads
152158
`first_frame`, `L2VA` reads `last_frame`, `FL2VA` reads both, `T2VA` reads
153159
neither. Connect the wrong one and the node says which is missing before
154160
anything loads — which end of the clip a picture belongs to is part of what
155161
the model is told.
156-
- `keep_model_loaded`**only `T2VA` can honour it.** The three tasks with
157-
frames run through `llama-mtmd-cli`, a fresh process each time, which takes
158-
the model with it when it exits.
162+
- `keep_model_loaded` — on a **safetensors** base every task honours it: the
163+
model is loaded in ComfyUI's own process and stays there. On a **GGUF** base
164+
only `T2VA` can, because the three tasks with frames run through
165+
`llama-mtmd-cli`, a fresh process each time that takes the model with it when
166+
it exits. The node says which it did rather than ignoring the switch.
159167
- `options` — the same options node as everything else. Its `adapter` dropdown
160168
lists both LoRAs; the first entry picks whichever one matches the base model
161169
you chose, so it needs no attention.
@@ -166,6 +174,14 @@ interchangeable downstream.
166174
|---|---|---|
167175
| Q4_K_M base + projector + Q8_0 adapter | 4.7 + 0.7 + 0.7 GB | ~9 GB |
168176
| Q8_0 base + projector + F16 adapter | 8.1 + 0.7 + 1.3 GB | ~13 GB |
177+
| safetensors base + adapter, `nf4` | 17.5 + 2.8 GB | ~8 GB |
178+
| safetensors base + adapter, `bfloat16` | 17.5 + 2.8 GB | ~20 GB |
179+
180+
The GGUF route is much the smaller download and needs nothing installed. The
181+
safetensors route is the shape the adapter was published in, keeps the model
182+
resident for every task rather than only for `T2VA`, and is the one to reach for
183+
if you already have the checkpoint — but it needs `transformers` and `peft`,
184+
which the pack lists as dependencies.
169185

170186
**What to expect of it.** All four tasks produce the trained shape: `T2VA` starts
171187
straight in on the three fields, and the other three open with the alignment

README_RU.md

Lines changed: 22 additions & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -140,20 +140,27 @@ python_embeded\python.exe -m pip install -r ComfyUI\custom_nodes\MiniMax-H3-Prom
140140
**Входы**
141141

142142
- `prompt`, `resolution`, `duration`, `greedy`, `seed` — как выше.
143-
- `model` — база Qwen3-VL и её проектор: два файла из одной конвертации. Записи с
144-
префиксом `on disk:` — это пары, уже лежащие в ваших папках моделей. **К
145-
адаптеру подходит только 8B**; Qwen3-VL другого размера загрузится и будет
146-
работать как обычная модель без ререйтера, и нода скажет об этом до того, как
147-
что-то скачает.
143+
- `model` — база Qwen3-VL-8B в любой из двух форм, в которых издан адаптер.
144+
**GGUF** — это два файла из одной конвертации: модель и её проектор.
145+
**safetensors** — официальная папка
146+
[Qwen3-VL-8B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-8B-Instruct), то
147+
самое, на чём адаптер обучался и в чём он опубликован. Записи с
148+
префиксом `on disk:` уже лежат в ваших папках моделей. **К адаптеру
149+
подходит только 8B**; Qwen3-VL другого размера отклоняется по имени и
150+
числам до того, как что-то скачается.
151+
- `quantization` — как грузить базу в **safetensors**: `nf4` — около 8 ГБ VRAM,
152+
`int8` — около 13, `bfloat16` — около 20. Для GGUF игнорируется: там своя.
148153
- `task``T2VA`, `I2VA`, `FL2VA`, `L2VA`. Сама модель называет их T2AV, I2AV,
149154
FL2AV и L2AV; это те же четыре задачи.
150155
- `first_frame` / `last_frame` — опциональные входы IMAGE. `I2VA` читает
151156
`first_frame`, `L2VA``last_frame`, `FL2VA` — оба, `T2VA` — ни одного. Если
152157
подключить не тот, нода назовёт недостающий до загрузки: к какому концу клипа
153158
относится картинка — часть того, что модели сообщают.
154-
- `keep_model_loaded`**работает только для `T2VA`.** Три задачи с кадрами идут
155-
через `llama-mtmd-cli`, каждый раз новым процессом, и модель уходит вместе с
156-
ним.
159+
- `keep_model_loaded` — на базе в **safetensors** работает для любой задачи:
160+
модель грузится в сам процесс ComfyUI и там же остаётся. На **GGUF** — только
161+
для `T2VA`: три задачи с кадрами идут через `llama-mtmd-cli`, каждый раз новым
162+
процессом, и модель уходит вместе с ним. Нода сообщает, как поступила, а не
163+
молча игнорирует переключатель.
157164
- `options` — та же нода настроек, что и везде. В её списке `adapter` обе LoRA, а
158165
первая запись подставляет ту, что подходит выбранной базовой модели, — трогать
159166
её не нужно.
@@ -164,6 +171,13 @@ python_embeded\python.exe -m pip install -r ComfyUI\custom_nodes\MiniMax-H3-Prom
164171
|---|---|---|
165172
| База Q4_K_M + проектор + адаптер Q8_0 | 4.7 + 0.7 + 0.7 ГБ | ~9 ГБ |
166173
| База Q8_0 + проектор + адаптер F16 | 8.1 + 0.7 + 1.3 ГБ | ~13 ГБ |
174+
| База safetensors + адаптер, `nf4` | 17.5 + 2.8 ГБ | ~8 ГБ |
175+
| База safetensors + адаптер, `bfloat16` | 17.5 + 2.8 ГБ | ~20 ГБ |
176+
177+
Путь через GGUF — заметно меньшая загрузка, и ставить ничего не надо. Путь через
178+
safetensors — та форма, в которой адаптер издан; модель остаётся в памяти на любой
179+
задаче, а не только на `T2VA`, и это то, что стоит брать, если чекпоинт уже есть.
180+
Взамен ему нужны `transformers` и `peft` — они у пака в зависимостях.
167181

168182
**Чего от неё ждать.** Все четыре задачи дают обученную форму: `T2VA` сразу
169183
начинает с трёх полей, остальные три открываются той самой строкой выравнивания,

minimax_h3_rewriter/catalog.py

Lines changed: 39 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -194,6 +194,43 @@ def _names(entries) -> list[str]:
194194
return [str(raw["name"]) for raw in entries if isinstance(raw, dict) and raw.get("name")]
195195

196196

197+
def _merge_adapters(merged: dict, seed: dict, offered: dict, changes: list[str]) -> None:
198+
"""Fold new adapter entries in, one format at a time.
199+
200+
An adapter section is a dict, not a list, so the set algebra that keeps the
201+
lists current never reached it: a section or a format published after
202+
somebody's copy was made simply never appeared in their file. Reading still
203+
worked -- ``adapter`` falls back to the packaged value for anything
204+
unconfigured -- but there was no line for them to point at a conversion of
205+
their own, which is the whole reason the file is theirs to edit.
206+
207+
Tracked in ``seed_offered`` under the section name, the same way and for the
208+
same reason as the lists: a format somebody deleted on purpose stays
209+
deleted, and only something genuinely new arrives.
210+
"""
211+
for section in ADAPTER_SECTIONS:
212+
available = seed.get(section)
213+
if not isinstance(available, dict) or not available:
214+
continue
215+
216+
current = merged.get(section)
217+
current = dict(current) if isinstance(current, dict) else {}
218+
seen = set(offered.get(section) or [])
219+
fresh = [
220+
fmt for fmt, entry in available.items()
221+
if isinstance(entry, dict) and fmt not in current and fmt not in seen
222+
]
223+
if fresh:
224+
for fmt in fresh:
225+
current[fmt] = json.loads(json.dumps(available[fmt]))
226+
merged[section] = current
227+
changes.append(f"{section}: added {', '.join(fresh)}")
228+
elif not isinstance(merged.get(section), dict) and current:
229+
merged[section] = current
230+
231+
offered[section] = sorted(seen | set(available))
232+
233+
197234
def merge(live: dict, seed: dict) -> tuple[dict, list[str]]:
198235
"""Fold new seed entries into a live list. Returns ``(merged, what changed)``.
199236
@@ -228,6 +265,8 @@ def merge(live: dict, seed: dict) -> tuple[dict, list[str]]:
228265

229266
offered[section] = sorted(set(offered.get(section) or []) | set(_names(available)))
230267

268+
_merge_adapters(merged, seed, offered, changes)
269+
231270
if offered == (live.get(OFFERED_KEY) or {}) and not changes:
232271
return merged, changes
233272

0 commit comments

Comments
 (0)