Dynamic Resolution Routing for Efficient Egocentric Grounding

NeurIPS 2026

Huixin Sun1Wangbo Zhao2Fanyue Wei1Qiuxia Lin3Pengzhan Sun1Angela Yao1
1 National University of Singapore2 The Hong Kong University of Science and Technology3 Nanyang Technological University

Overview

Egocentric visual grounding requires high-resolution inputs to localize small objects. However, scaling Multimodal Large Language Models to this domain is constrained by the excessive cost of visual token processing. We identify that current efficient strategies based on token reduction are unreliable for selecting object-centric spatial evidence. SmartRes performs efficiency optimization in pixel space via dynamic resolution routing. A low-resolution view provides global context, while a lightweight router activates high-resolution patches in object-centric regions and constructs an order-preserving visual sequence. A margin-regularized routing objective improves foreground recall under severe foreground–background imbalance. Experiments on Ego4D and EgoIntention demonstrate a favorable trade-off between fine-grained grounding and inference efficiency, including strong performance on small objects.

Pixel-space routing preserves relevant object evidence. Token pruning can instead retain background regions, while high-resolution vision encoding remains a major computational cost.

Method

A lightweight routing head operates on low-resolution features to select object-centric patches for high-resolution encoding. The selected features are combined with global context in raster-scan order. A margin-regularized routing objective improves foreground recall under foreground–background imbalance.

The SmartRes framework. Routing reduces visual encoding work at its source, and the shorter visual sequence also lowers language-model computation.

Experimental Results

We evaluate SmartRes on Ego4D and the Context and Uncommon subsets of EgoIntention, using Qwen2.5-VL-3B-Instruct as the backbone. Grounding accuracy is measured with P@0.3, P@0.5, and mean IoU. Evaluation on RefCOCO, RefCOCO+, and RefCOCOg further examines generalization to standard referring expression comprehension.

Comparisons include token-pruning methods FastV and Dyn-LLaVA, alongside uniform down-scaling. The table compares grounding accuracy and computational efficiency across representative baselines and the two SmartRes configurations. See Section 4 of the paper for the full comparisons and evaluation protocol.

Egocentric grounding results with Qwen2.5-VL-3B-Instruct. Dataset columns and Overall report P@0.5 (%).
MethodTokens ↓Ego4D ↑EgoInt-C ↑EgoInt-U ↑Overall ↑Retention ↑FLOPs (T) ↓Latency (ms) ↓
Full resolution100%63.2260.0154.7759.33100.0%11.32553.6
Down-scaling32%32.7631.8323.1029.2352.3%3.94240.8
Down-scaling50%52.1750.2542.7048.3781.8%5.89323.6
FastV50%45.5247.3235.7442.8673.6%7.15410.2
FastV70%51.6550.3343.6148.5383.3%8.92495.5
Dyn-LLaVA50%52.4850.7639.1947.4881.4%7.41419.3
Dyn-LLaVA70%54.4252.1244.5450.3685.1%9.16506.7
SmartRes-Lite33%54.5552.1546.6351.1186.4%4.05252.6
SmartRes-Pro55%55.4553.8450.0453.1189.9%6.58365.8

EgoInt-C and EgoInt-U are the Context and Uncommon subsets of EgoIntention. Overall is the mean P@0.5 across the three datasets. Retention is the mean performance retention across all nine metrics (P@0.3, P@0.5, and mIoU on each dataset), relative to full resolution. Tokens are reported relative to full resolution. Values are from Table 1 of the paper; ↑ / ↓ indicate higher / lower is better.

Qualitative grounding results. Selected regions preserve fine-grained evidence while retaining global scene context.

Citation

@inproceedings{sun2026smartres,
  title   = {Dynamic Resolution Routing for Efficient Egocentric Grounding},
  author  = {Sun, Huixin and Zhao, Wangbo and Wei, Fanyue and Lin, Qiuxia
             and Sun, Pengzhan and Yao, Angela},
  booktitle = {Advances in Neural Information Processing Systems},
  year    = {2026}
}