Skip to content

Add support for Hybrid Models in extract hidden states#159

Draft
rahul-tuli wants to merge 6 commits into
mainfrom
extract-hidden-states-hma
Draft

Add support for Hybrid Models in extract hidden states#159
rahul-tuli wants to merge 6 commits into
mainfrom
extract-hidden-states-hma

Conversation

@rahul-tuli

Copy link
Copy Markdown
Member
  • Add CacheOnlySpec(MLAAttentionSpec) to kv_cache_interface.py so it duck-types through all existing AttentionSpec code paths
  • Pre-filter CacheOnlySpec in get_kv_cache_groups() before type- unification routing to prevent crashes with mixed spec types
  • Joint budget calculation in get_kv_cache_config_from_groups() via extra_bytes_per_block parameter on get_num_blocks()
  • Gate HMA disable in config with supports_hma() check so hybrid models keep their per-group block allocators
  • Add SupportsHMA to ExampleHiddenStatesConnector with correct cache_group_idx for block_ids
  • Resolve CacheOnly slot_mapping from per-layer mappings in the proposer instead of using main group's common_attn_metadata

- Add CacheOnlySpec(MLAAttentionSpec) to kv_cache_interface.py so it
    duck-types through all existing AttentionSpec code paths
- Pre-filter CacheOnlySpec in get_kv_cache_groups() before type-
    unification routing to prevent crashes with mixed spec types
- Joint budget calculation in get_kv_cache_config_from_groups() via
    extra_bytes_per_block parameter on get_num_blocks()
- Gate HMA disable in config with supports_hma() check so hybrid
    models keep their per-group block allocators
- Add SupportsHMA to ExampleHiddenStatesConnector with correct
    cache_group_idx for block_ids
- Resolve CacheOnly slot_mapping from per-layer mappings in the
    proposer instead of using main group's common_attn_metadata

Signed-off-by: Rahul-Tuli <rtuli@redhat.com>
…V cache groups

CacheOnly layers (used for hidden-state extraction) previously created a
separate KV cache group, which caused issues with hybrid model support:
scattered isinstance checks, memory accounting hacks (extra_bytes_per_block),
and potential 3-way hybrid coordinator failures.

This commit makes CacheOnly layers "supplementary" -- they piggyback on
group 0's block table and slot mappings instead of being full group
participants. They still get their own allocated tensors but are invisible
to the KV cache coordinator. This eliminates the need for separate block
management while keeping the memory accounting clean.

Key changes:
- Add supplementary_specs field to KVCacheConfig
- Strip CacheOnly from groups in get_kv_cache_groups() via split_supplementary_specs()
- Account for supplementary memory in get_kv_cache_config_from_groups() budget
- Handle supplementary layers in attn_utils (init_attn_backend, _allocate_kv_cache, _reshape_kv_cache)
- Revert proposer to use group 0's slot_mapping directly
- Remove _cache_group_idx from ExampleHiddenStatesConnector
- Remove CacheOnlySpec from single_type_kv_cache_manager spec_manager_map

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

Signed-off-by: Rahul-Tuli <rtuli@redhat.com>
…eshape

The model runner has its own _allocate_kv_cache_tensors and
_reshape_kv_cache_tensors methods (separate from attn_utils.py).
These were missing supplementary layer support:

1. _allocate_kv_cache_tensors: assertion didn't include supplementary
   layer names, causing "Some layers are not correctly initialized"
2. _reshape_kv_cache_tensors: supplementary layers were never reshaped
   or included in kv_caches, so they wouldn't be bound

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

Signed-off-by: Rahul-Tuli <rtuli@redhat.com>
CacheOnlyAttentionLayer captured cache_config.block_size at __init__
time (16), but for hybrid models with Mamba layers the block_size is
later adjusted upward (e.g. to 528) to match the Mamba page size.

Since CacheOnly layers share slot_mapping with group 0, slot indices
are computed as block_id * 528 + offset, but the CacheOnly KV cache
was shaped with block_size=16 — causing CUDA "index out of bounds".

Fix: read vllm_config.cache_config.block_size at get_kv_cache_spec()
time instead of using the stale self.block_size from init.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

Signed-off-by: Rahul-Tuli <rtuli@redhat.com>
Use the GPU slot_mapping from the attention metadata when extracting
hidden states from the KV cache, rather than recomputing it from
scheduler block IDs. On hybrid models (e.g. Qwen3.5), kernel block
splitting causes the two to diverge, producing all-zeros extraction.

Also map supplementary CacheOnly layers to group 0's slot_mapping in
both build_slot_mappings_by_layer and _get_slot_mappings, since they
share group 0's block table.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

Signed-off-by: Rahul-Tuli <rtuli@redhat.com>
Replace the inlined 37-line reshape block in gpu_model_runner's
_reshape_kv_cache_tensors with a call to the shared _reshape_one_layer
helper already factored out in attn_utils.

Include supplementary specs (e.g. CacheOnly) in the memory budget
used by _max_memory_usage_bytes_from_groups, _estimate_max_model_len,
and _auto_fit_max_model_len so that the "enough memory" check and
auto-fit context length estimation account for supplementary tensors.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

Signed-off-by: Rahul-Tuli <rtuli@redhat.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant