Skip to content

Commit de51791

Browse files
committed
[0.16.5] reference caption issue (#7)
1 parent df518e6 commit de51791

12 files changed

Lines changed: 328 additions & 35 deletions

File tree

.gitignore

Lines changed: 4 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -223,4 +223,7 @@ uv.exe
223223
# Local development: junction helper and the ComfyUI path it remembers
224224
dev-junction.ps1
225225
dev-junction.bat
226-
comfyui-path.txt
226+
comfyui-path.txt
227+
# Seeded when the package is imported outside ComfyUI, where folder_paths is
228+
# absent and the user directory falls back into the package itself
229+
minimax_h3_rewriter/_user/

CHANGELOG.md

Lines changed: 44 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -6,6 +6,50 @@ The version in `pyproject.toml`, the git tag and the release on GitHub always sa
66
the same thing; the release workflow refuses a tag that disagrees with
77
`pyproject.toml`, or one that neither changelog has a section for.
88

9+
## 0.16.5 - 2026-08-24
10+
11+
### Fixed
12+
13+
- **A captioner no longer reserves the whole context the model was trained on.**
14+
`--ctx-size 0` reads as "let llama.cpp decide" and means "the context this
15+
model was trained for", and the packaged captioners hid the difference:
16+
Qwen2.5-Omni asks for 32k and fits anywhere. Qwen3-VL asks for 262144 tokens,
17+
and its cache is 36 layers of 8 KV heads at 128 dimensions, K and V, in f16 -
18+
144 KiB a token, so 36 GiB of KV cache allocated before a single pixel has
19+
been read. On a 32 GB card the run died at `failed to allocate buffer for kv
20+
cache`, having never looked at the picture, and llama.cpp's own auto-fit could
21+
not rescue it because the node pins `--n-gpu-layers`.
22+
23+
0 now means "what this run needs, and never more than the card can hold". The
24+
context length and the shape of the cache are read from the GGUF header the
25+
model scan already parses, and sized against the number of references and the
26+
memory of the device the run is bound for. A number typed into `context_size`
27+
is still honoured exactly as typed, and a header too thin to size against
28+
falls back to letting llama.cpp decide, as before.
29+
30+
Reported by [@808charlie](https://github.com/808charlie) in
31+
[#7](https://github.com/pytraveler/MiniMax-H3-Prompt-Rewriter-ComfyUI/issues/7).
32+
33+
### Changed
34+
35+
- **A failed caption says which way it failed.** The note about projector
36+
formats was printed on every non-zero exit from `llama-mtmd-cli`, so an
37+
out-of-memory came back dressed as a model mtmd cannot read - and sent the
38+
reader to study the model while the child had already said `cudaMalloc
39+
failed`. An allocation that did not fit and a context too small for the frames
40+
now each get their own answer, naming the widget that moves them, and the
41+
projector note is left for the exits that are actually about the projector.
42+
43+
- **The captioner scan stops narrating.** A model sharing a folder with an
44+
mmproj it could not be paired with printed a line naming every projector in
45+
that folder, at INFO. A flat `models/LLM` holding a dozen unrelated quants and
46+
five projectors therefore announced twelve models by five names each - four
47+
times over, once for every node that offers a captioner dropdown, every time
48+
ComfyUI asked for the node definitions. One line per folder now, at DEBUG.
49+
50+
Reported by [@808charlie](https://github.com/808charlie) in
51+
[#7](https://github.com/pytraveler/MiniMax-H3-Prompt-Rewriter-ComfyUI/issues/7).
52+
953
## 0.16.4 - 2026-08-24
1054

1155
### Changed

CHANGELOG_RU.md

Lines changed: 45 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -6,6 +6,51 @@
66
workflow релиза отклонит тег, который расходится с `pyproject.toml`, и тег, для
77
которого нет раздела ни в одном из двух changelog'ов.
88

9+
## 0.16.5 - 2026-08-24
10+
11+
### Исправлено
12+
13+
- **Капшенер больше не резервирует весь контекст, на котором модель обучали.**
14+
`--ctx-size 0` читается как "пусть llama.cpp решит сама", а означает "контекст,
15+
на который модель обучена", и упакованные капшенеры эту разницу скрывали:
16+
Qwen2.5-Omni просит 32k и помещается где угодно. Qwen3-VL просит 262144
17+
токена, а её кэш - это 36 слоёв по 8 KV-голов размерности 128, K и V, в f16:
18+
144 КиБ на токен, то есть 36 ГиБ KV-кэша, выделяемых до того, как прочитан
19+
первый пиксель. На карте в 32 ГБ запуск падал на `failed to allocate buffer
20+
for kv cache`, так и не посмотрев на картинку, а собственный auto-fit
21+
llama.cpp его не спасал, потому что нода жёстко задаёт `--n-gpu-layers`.
22+
23+
Теперь 0 означает "сколько нужно этому запуску и не больше, чем поместится в
24+
карту". Длина контекста и форма кэша читаются из заголовка GGUF, который скан
25+
моделей и так разбирает, и соотносятся с числом референсов и памятью
26+
устройства, куда запуск направлен. Число, вписанное в `context_size`,
27+
по-прежнему выполняется буквально, а заголовок, из которого посчитать нечего,
28+
как и раньше оставляет решение за llama.cpp.
29+
30+
Сообщил [@808charlie](https://github.com/808charlie) в
31+
[#7](https://github.com/pytraveler/MiniMax-H3-Prompt-Rewriter-ComfyUI/issues/7).
32+
33+
### Изменено
34+
35+
- **Упавший капшенер говорит, в какую сторону он упал.** Заметка про форматы
36+
проекторов печаталась при любом ненулевом коде выхода `llama-mtmd-cli`, и
37+
нехватка памяти возвращалась в костюме модели, которую mtmd не умеет читать, -
38+
отправляя читателя изучать модель в тот момент, когда потомок уже сказал
39+
`cudaMalloc failed`. У невлезшего выделения и у слишком тесного под кадры
40+
контекста теперь свои ответы, с именем виджета, который их двигает, а заметка
41+
про проектор осталась тем выходам, которые действительно о проекторе.
42+
43+
- **Скан капшенеров перестал вести репортаж.** Модель, лежащая в одной папке с
44+
mmproj, к которому её не удалось привязать, печатала строку с перечислением
45+
всех проекторов этой папки, уровнем INFO. Плоский `models/LLM` с дюжиной
46+
посторонних квантов и пятью проекторами объявлял двенадцать моделей по пять
47+
имён каждая - и четырежды, по разу на каждую ноду с выпадающим списком
48+
капшенеров, при каждом запросе описаний нод от ComfyUI. Теперь одна строка на
49+
папку, уровнем DEBUG.
50+
51+
Сообщил [@808charlie](https://github.com/808charlie) в
52+
[#7](https://github.com/pytraveler/MiniMax-H3-Prompt-Rewriter-ComfyUI/issues/7).
53+
954
## 0.16.4 - 2026-08-24
1055

1156
### Изменено

README.md

Lines changed: 8 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -658,10 +658,14 @@ Other inputs:
658658
two seconds at 25 fps is 56 images through the vision tower, and thirty seconds
659659
is 750. This is what makes the cost of describing a clip independent of its
660660
length.
661-
- `context_size``0` means the model's own, which is what its projector was
662-
sized against; one 1024×1024 frame is already twenty-odd media chunks. Lower it
663-
only to cut the KV cache, and know that too small a value fails the run instead
664-
of truncating it.
661+
- `context_size``0` sizes the context from the references and from the card,
662+
which is not the same as the model's own: llama.cpp reserves the whole KV cache
663+
before it reads a pixel, and a model trained for 262144 tokens asks tens of
664+
gigabytes for it — Qwen3-VL-8B asks 36 GB, on a card that would have captioned
665+
the picture in three seconds. One 1024×1024 frame is already twenty-odd media
666+
chunks, so the count of references is half of the answer and the memory of the
667+
device is the other half. Type a number to say it yourself; too small a value
668+
fails the run instead of truncating it.
665669

666670
> **Not every multimodal GGUF works here**, and the ones that do not fail loudly.
667671
> llama.cpp's `mtmd` has to understand the projector format: Gemma 4's aborts the

README_RU.md

Lines changed: 8 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -649,10 +649,14 @@ delivery, speaking in Russian».*
649649
Все сразу переполнили бы и контекст, и терпение: две секунды при 25 fps — это
650650
56 картинок через визуальную башню, а тридцать секунд — 750. Именно это и
651651
делает стоимость описания клипа не зависящей от его длины.
652-
- `context_size``0` означает «контекст самой модели», под который её проектор
653-
и рассчитан: один кадр 1024×1024 — это уже больше двадцати медиа-чанков.
654-
Уменьшайте только чтобы срезать KV-кэш, и учтите, что слишком малое значение
655-
роняет запуск, а не обрезает его.
652+
- `context_size``0` подбирает контекст под референсы и под карту, а это не то
653+
же самое, что «контекст самой модели»: llama.cpp резервирует весь KV-кэш до
654+
того, как прочтёт первый пиксель, и модель, обученная на 262144 токена,
655+
требует под него десятки гигабайт — Qwen3-VL-8B ровно 36 ГБ, на карте, которая
656+
описала бы картинку за три секунды. Один кадр 1024×1024 — это уже больше двадцати
657+
медиа-чанков, поэтому число референсов — половина ответа, а память устройства —
658+
вторая. Вписанное число выполняется буквально; слишком малое роняет запуск,
659+
а не обрезает его.
656660

657661
> **Не всякий мультимодальный GGUF здесь работает**, и те, что не работают,
658662
> падают громко. `mtmd` в llama.cpp должен понимать формат проектора: у Gemma 4

minimax_h3_rewriter/devices.py

Lines changed: 22 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -70,6 +70,28 @@ def describe(index: int) -> str:
7070
return ""
7171

7272

73+
def vram_bytes(spec: str = AUTO) -> int:
74+
"""Total memory of the card a run is bound for, or 0 when it cannot be read.
75+
76+
Total rather than free, deliberately. What is free at the moment of asking
77+
is mostly a statement about the diffusion model that is about to be evicted
78+
anyway, so it would size a context differently depending on what happened to
79+
be loaded when the node ran -- and the same graph would then fail only on
80+
the second pass. The headroom that covers the difference is the caller's.
81+
"""
82+
if is_cpu(spec):
83+
return 0
84+
ordinal = index(spec) or 0
85+
try:
86+
torch = _torch()
87+
if not torch.cuda.is_available():
88+
return 0
89+
return int(torch.cuda.get_device_properties(ordinal).total_memory)
90+
except Exception:
91+
log.debug("[minimax_h3_rewriter.devices.vram_bytes] %s unreadable", spec, exc_info=True)
92+
return 0
93+
94+
7395
def choices() -> list[str]:
7496
"""The values offered by the options node, in a stable spelling."""
7597
return [AUTO] + [f"{PREFIX}{index}" for index in range(count())] + [CPU]

minimax_h3_rewriter/discovery.py

Lines changed: 63 additions & 7 deletions
Original file line numberDiff line numberDiff line change
@@ -393,9 +393,54 @@ def consider(directory: str, name: str) -> None:
393393
return found
394394

395395

396+
ARCH_KEYS = (
397+
"block_count",
398+
"embedding_length",
399+
"context_length",
400+
"attention.head_count",
401+
"attention.head_count_kv",
402+
"attention.key_length",
403+
"attention.value_length",
404+
)
405+
406+
KV_ELEMENT_BYTES = 2
407+
396408
_GGUF_HEADER_CACHE: dict[tuple[str, int, int], dict] = {}
397409

398410

411+
def _header_int(value) -> int | None:
412+
"""One integer out of a header value that is sometimes a per-layer array."""
413+
if isinstance(value, (list, tuple)):
414+
value = max(value) if value else None
415+
try:
416+
return int(value)
417+
except (TypeError, ValueError):
418+
return None
419+
420+
421+
def _kv_per_token(value, arch: str, blocks: int | None, width: int | None) -> int | None:
422+
"""Bytes one token costs in the KV cache, or ``None`` from a header too thin.
423+
424+
``key_length`` and ``value_length`` are optional and often absent; the head
425+
dimension is then ``embedding_length / head_count``, which is the same
426+
fallback llama.cpp makes. ``head_count_kv`` missing means the model has no
427+
grouped-query attention and every head keeps its own cache.
428+
"""
429+
if not arch or not blocks:
430+
return None
431+
heads = _header_int(value(f"{arch}.attention.head_count"))
432+
kv_heads = _header_int(value(f"{arch}.attention.head_count_kv")) or heads
433+
if not kv_heads:
434+
return None
435+
key_length = _header_int(value(f"{arch}.attention.key_length"))
436+
if key_length is None and heads and width:
437+
key_length = width // heads
438+
value_length = _header_int(value(f"{arch}.attention.value_length")) or key_length
439+
if not key_length or not value_length:
440+
return None
441+
return blocks * kv_heads * (key_length + value_length) * KV_ELEMENT_BYTES
442+
443+
399444
def gguf_header(path: str) -> dict:
400445
"""Architecture, type and shape of a GGUF, read from its header only.
401446
@@ -407,7 +452,10 @@ def gguf_header(path: str) -> dict:
407452
Cached per file identity, so a folder of large quants costs no more than a
408453
stat each after the first pass.
409454
"""
410-
empty = {"arch": "", "kind": "", "blocks": None, "width": None, "vision": False, "audio": False}
455+
empty = {
456+
"arch": "", "kind": "", "blocks": None, "width": None,
457+
"vision": False, "audio": False, "context": None, "kv_per_token": None,
458+
}
411459
try:
412460
stat = os.stat(path)
413461
except OSError:
@@ -423,7 +471,7 @@ def gguf_header(path: str) -> dict:
423471

424472
def also(found: dict) -> tuple[str, ...]:
425473
arch = found.get("general.architecture")
426-
return (f"{arch}.block_count", f"{arch}.embedding_length") if arch else ()
474+
return tuple(f"{arch}.{name}" for name in ARCH_KEYS) if arch else ()
427475

428476
value = gguf_meta.keys(path, HEADER_KEYS, probe=also, verify=True).get
429477

@@ -438,6 +486,10 @@ def also(found: dict) -> tuple[str, ...]:
438486
width = value(f"{header['arch']}.embedding_length")
439487
header["blocks"] = int(blocks) if blocks is not None else None
440488
header["width"] = int(width) if width is not None else None
489+
header["context"] = _header_int(value(f"{header['arch']}.context_length"))
490+
header["kv_per_token"] = _kv_per_token(
491+
value, header["arch"], header["blocks"], header["width"]
492+
)
441493
except Exception:
442494
log.debug("[minimax_h3_rewriter.gguf_header] %s unreadable", path, exc_info=True)
443495

@@ -653,20 +705,18 @@ def scan_captioner_gguf(arch: str | None = None) -> list[tuple[str, str, str]]:
653705
elif header["kind"] == "model":
654706
models.append(path)
655707

656-
for _directory, (models, projectors) in sorted(by_directory.items()):
708+
for directory, (models, projectors) in sorted(by_directory.items()):
657709
if not projectors:
658710
continue
711+
unpaired: list[str] = []
659712
for model in sorted(models):
660713
# After counting the models, not before: how many there are is what
661714
# decides whether a lone projector in the folder can be trusted.
662715
if arch is not None and gguf_header(model)["arch"] != arch:
663716
continue
664717
projector = _pair_mmproj(model, projectors, len(models))
665718
if not projector:
666-
log.info(
667-
"[minimax_h3_rewriter.scan_captioner_gguf] no obvious projector for %s among %s",
668-
model, [os.path.basename(p) for p in projectors],
669-
)
719+
unpaired.append(os.path.basename(model))
670720
continue
671721
header = gguf_header(projector)
672722
modalities = ", ".join(
@@ -678,5 +728,11 @@ def scan_captioner_gguf(arch: str | None = None) -> list[tuple[str, str, str]]:
678728
size = 0.0
679729
label = f"{os.path.basename(model)} [+mmproj, {modalities}, {size:.1f} GB]"
680730
found.append((label, model, projector))
731+
if unpaired:
732+
log.debug(
733+
"[minimax_h3_rewriter.scan_captioner_gguf] %s: %d model(s) with no projector "
734+
"of their own among %s: %s",
735+
directory, len(unpaired), [os.path.basename(p) for p in projectors], unpaired,
736+
)
681737

682738
return found

0 commit comments

Comments
 (0)