Dynamic Resolution Routing for Efficient Egocentric Grounding
NeurIPS 2026
Method
A lightweight routing head operates on low-resolution features to select object-centric patches for high-resolution encoding. The selected features are combined with global context in raster-scan order. A margin-regularized routing objective improves foreground recall under foreground–background imbalance.
Experimental Results
We evaluate SmartRes on Ego4D and the Context and Uncommon subsets of EgoIntention, using Qwen2.5-VL-3B-Instruct as the backbone. Grounding accuracy is measured with P@0.3, P@0.5, and mean IoU. Evaluation on RefCOCO, RefCOCO+, and RefCOCOg further examines generalization to standard referring expression comprehension.
Comparisons include token-pruning methods FastV and Dyn-LLaVA, alongside uniform down-scaling. The table compares grounding accuracy and computational efficiency across representative baselines and the two SmartRes configurations. See Section 4 of the paper for the full comparisons and evaluation protocol.
| Method | Tokens ↓ | Ego4D ↑ | EgoInt-C ↑ | EgoInt-U ↑ | Overall ↑ | Retention ↑ | FLOPs (T) ↓ | Latency (ms) ↓ |
|---|---|---|---|---|---|---|---|---|
| Full resolution | 100% | 63.22 | 60.01 | 54.77 | 59.33 | 100.0% | 11.32 | 553.6 |
| Down-scaling | 32% | 32.76 | 31.83 | 23.10 | 29.23 | 52.3% | 3.94 | 240.8 |
| Down-scaling | 50% | 52.17 | 50.25 | 42.70 | 48.37 | 81.8% | 5.89 | 323.6 |
| FastV | 50% | 45.52 | 47.32 | 35.74 | 42.86 | 73.6% | 7.15 | 410.2 |
| FastV | 70% | 51.65 | 50.33 | 43.61 | 48.53 | 83.3% | 8.92 | 495.5 |
| Dyn-LLaVA | 50% | 52.48 | 50.76 | 39.19 | 47.48 | 81.4% | 7.41 | 419.3 |
| Dyn-LLaVA | 70% | 54.42 | 52.12 | 44.54 | 50.36 | 85.1% | 9.16 | 506.7 |
| SmartRes-Lite | 33% | 54.55 | 52.15 | 46.63 | 51.11 | 86.4% | 4.05 | 252.6 |
| SmartRes-Pro | 55% | 55.45 | 53.84 | 50.04 | 53.11 | 89.9% | 6.58 | 365.8 |
EgoInt-C and EgoInt-U are the Context and Uncommon subsets of EgoIntention. Overall is the mean P@0.5 across the three datasets. Retention is the mean performance retention across all nine metrics (P@0.3, P@0.5, and mIoU on each dataset), relative to full resolution. Tokens are reported relative to full resolution. Values are from Table 1 of the paper; ↑ / ↓ indicate higher / lower is better.
Citation
@inproceedings{sun2026smartres,
title = {Dynamic Resolution Routing for Efficient Egocentric Grounding},
author = {Sun, Huixin and Zhao, Wangbo and Wei, Fanyue and Lin, Qiuxia
and Sun, Pengzhan and Yao, Angela},
booktitle = {Advances in Neural Information Processing Systems},
year = {2026}
}