Hanoona Rasheed1 · Haania Siddiqui1 · Ming-Hsuan Yang2 · Fahad Shahbaz Khan1,3 · Salman Khan1,3
1 Mohamed bin Zayed University of Artificial Intelligence · 2 University of California, Merced · 3 Apertix
Spatio-temporal video grounding requires identifying when a queried event occurs and localizing the referred entity throughout that interval. We introduce Parallel Tube Decoding (PTD), which predicts the temporal interval first and then decodes all time-conditioned spatial blocks in parallel. PTD reduces tube generation to two decoding rounds, independent of tube length, while improving both localization accuracy and inference efficiency.
Autoregressive localization vs. Parallel Tube Decoding. Given a video and a referring expression, STVG predicts when the event occurs and the bounding box of the referred entity throughout that interval. PTD generates all time-conditioned spatial blocks in parallel after temporal localization, reducing the sequential decoding depth to
- Parallel tube generation. PTD removes token-level and trajectory-level dependencies by generating the temporal block in one round and all spatial blocks in a second round.
- Decoupled Block Attention. Each spatial block uses the shared video-query context and predicted temporal interval without depending on other generated boxes.
- Localization-aware optimization. Complementary temporal and spatial rewards improve event boundaries, bounding-box geometry, and target consistency.
-
Efficiency and accuracy. PTD achieves
$79\times$ lower Tube Completion Latency and$92\times$ higher spatial decoding throughput than standard autoregressive decoding while improving grounding performance.
| Component | Location |
|---|---|
| Load the released PTD checkpoint | MBZUAI/ParallelTubeDecoding-Qwen3-VL-4B on Hugging Face |
| Install the PTD environment | Installation guide |
| Prepare training and evaluation annotations | Data preparation guide, including SFT preparation for VidSTG/HC-STVG and evaluation preparation for all four released benchmarks |
| Run SFT, merge the adapter, and run GRPO | Training guide, SFT launcher, merge script, and GRPO launcher |
| Evaluate VidSTG, HC-STVG, Charades-STA, and ActivityNet | Evaluation guide and lmms-eval task definitions |
| Measure Quantized/PTD TCL and BPS | Efficiency runner, reporter, and the evaluation guide |
After installing the environment, launch the Gradio demo from the
repository root. It loads the released merged PTD checkpoint from Hugging Face
by default; --disable_flash_attention uses the broadly supported SDPA backend.
PYTHONPATH=src python src/serve/app.py --disable_flash_attentionUpload a video and use the grounding prompt below, replacing the text in angle brackets with the target event:
Given the query: '<description of the target event>' Localize the described object throughout the video. Use object reference tokens, time tokens, and box tokens. Return the object reference, event time segment, and per-time bbox coordinates.
Pass --model-path /path/to/merged-ptd-checkpoint to use a local merged
checkpoint. For benchmark inference with the paper's preprocessing and
evaluation settings, follow the evaluation guide.
Attention masks for Sequential Block Decoding and PTD. Sequential Block Decoding retains causal attention across spatial blocks. PTD replaces this cross-box dependency with Decoupled Block Attention: every spatial block accesses the shared multimodal prefix and temporal block while all other spatial blocks remain masked.
Table 1: Comparison of decoding strategies on VidSTG. We report temporal and video IoU for declarative and interrogative queries, together with Tube Completion Latency (TCL) and Boxes Per Second (BPS). Parallel Tube Decoding achieves the strongest grounding performance, lowest latency, and highest throughput.
Analysis of decoding efficiency and trajectory-level dependency. (a) Tube completion latency as the number of grounded frames increases. PTD maintains nearly constant latency, while token-based and block decoding scale with tube length. (b) Attention distribution across tube decoding for Sequential Block Decoding (dotted) and PTD (solid). Sequential decoding progressively shifts attention from the video toward prior text. (c) History-correction analysis for Sequential Block Decoding. Replacing an erroneous box
Table 2: Results on VidSTG. Comparison with backbone baselines and prior multimodal large language models for declarative and interrogative spatio-temporal grounding. Our compact Qwen3-VL-4B model with GRPO and PTD delivers the strongest overall results.
Table 3: Results on HC-STVG. Comparison with backbone baselines and prior methods on HC-STVG v1 and v2. PTD achieves strong temporal localization and the best video IoU across both benchmark versions.
Table 4: Zero-shot temporal grounding. Results on Charades-STA and ActivityNet Captions without task-specific training. Our model improves over the strongest prior zero-shot methods, including under stricter temporal-overlap thresholds.
Table 6: Ablation of localization-aware rewards. Temporal and spatial rewards provide complementary improvements for declarative and interrogative grounding. Combining both produces the strongest overall spatio-temporal localization performance.
Comparison with prior STVG methods. PTD directly generates both the temporal interval and complete spatial tube. The examples highlight fine-grained target identification among visually similar distractors, large changes in object scale, and event-specific temporal localization. Bounding boxes show spatial predictions, while the horizontal lines indicate the predicted temporal intervals.
Effect of the temporal reward. The first example shows how the reward recovers the complete interaction when SFT localizes only its most salient portion. The second shows how it prevents the prediction from extending beyond the queried event while the target remains visible.
Effect of the spatial reward. The examples highlight tighter localization under rapid motion and more consistent target identity despite substantial scale changes and interference from nearby entities.
Sequential Block Decoding vs. PTD. Abrupt scale changes and partial occlusion expose how Sequential Block Decoding propagates localization errors across subsequent boxes. PTD maintains tighter target localization by removing cross-box dependencies.
Additional qualitative results. PTD accurately localizes the queried entity under visually similar distractors, target motion and scale changes, partial occlusion, progressive object reveal, and multi-instance interactions.
Representative failure cases. Temporally subtle state changes can produce ambiguous event boundaries, while spatial localization becomes difficult for small, rapidly moving, or occluded targets.
This codebase is built on Qwen-VL-Series-Finetune. We thank its authors for releasing their fine-tuning framework. We also thank the Qwen team for releasing Qwen3-VL and acknowledge Locate Anything for making its work and implementation publicly available.
@misc{rasheed2026locatevideosrethinkingefficient,
title = {Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding},
author = {Rasheed, Hanoona and Siddiqui, Haania and Yang, Ming-Hsuan and Khan, Fahad Shahbaz and Khan, Salman},
year = {2026},
eprint = {2608.28192},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2608.28192}
}
















