Skip to content

Best prompt nodepack! Thanks. Ideas: amended instruction for 'MiniMax-H3 Reference Caption' (and instruction field in 'Multi Reference Caption') #8

Description

@808charlie

Now I've got things running, this really is great. Very configurable and flexible. Thank you!

I observed some models caption a picture well (albeit with an empty at the start of the description).
Other models at times return a ton of thinking (particularly Qwen 3.5, particularly variants of original model). They get to a great description, but think first.

The prompt you use for the Reference Caption node is output to the cli, so I'd already amended the prompt using a text input to the 'instruction' field in MiniMax-H3 Reference Caption. I then found adding "return only the prompt not your thinking" fixed the problem with 'thinking' in every case! I daresay other similar instructions to 'return only...' would have a similar effect.

The 'instruction' field for custom instructions is therefore great.
Request 1: any chance of an instruction field in the Multi Reference Caption node? Appreciate it might need one of pic/aud/vid as the prompt is not the same.

Request 2: Even more usefully, is it possible the 'MiniMax-H3 Prompt Writer' could include an 'instruction' field. You provide the 'system prompt' (instruction) exactly through the 'minimx-H3 guide promt (and LLM) node. Users could then edit it if needed, but importantly edit it for other models.

Why so much interest?

Best node pack I've found.

  • Your nodes have significant advantage of others like 'ComfyUI-QwenVL' or 'ComfyUI_Qwen3VL-instruct' nodes, which require editing json files and changes to the node stop them trying to download models and to specify personally selected and pre-downloaded models. They use safetensors.
  • MiniCPM node has above problems, but was one which used GGUF, but it had other issues as I've mentioned before.
  • pixaroma has some prompt nodes, but whilst some of his are great, I have real issues with the slick functionality and JS (Claude coded) and how non-native they are and fragile the interaction feels - it feels like the difference between 'windows' and 'linux' :D
  • Your nodes are faster than the native 'text-generate' that use safetensors and is much more fussy on models and much slower. Indeed, I've used safetensors models with yours, as you know, via clip-load, and for some reason it also takes WAY longer than loading a Q8 GGUF. Way longer. I suspect all the pinning and over-management of memory to be related.

Whilst I'd not normally be a fan of GGUF within ComfyUI (on my hardware), it really hits the spot for quick calls to LLMs doesn't it. Really impressed so far thanks.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions