Follow
Yuhao Dong
Yuhao Dong
Tsinghua University, Nanyang Technological University, Moonshot AI
Verified email at mails.tsinghua.edu.cn
Title
Cited by
Cited by
Year
Kimi k2. 5: Visual agentic intelligence
K Team, T Bai, Y Bai, Y Bao, SH Cai, Y Cao, Z Chai, Y Charles, HS Che, ...
arXiv preprint arXiv:2602.02276, 2026
801*2026
Lmms-eval: Accelerating the development of large multimodal models
B Li, P Zhang, K Zhang, F Pu, X Du, Y Dong, H Liu, Y Zhang, G Zhang, ...
https://github.com/EvolvingLMMs-Lab/lmms-eval., 2024
556*2024
Kimi-vl technical report
K Team, A Du, B Yin, B Xing, B Qu, B Wang, C Chen, C Zhang, C Du, ...
arXiv preprint arXiv:2504.07491, 2025
494*2025
Insight-v: Exploring long-chain visual reasoning with multimodal large language models
Y Dong, Z Liu, HL Sun, J Yang, W Hu, Y Rao, Z Liu
2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR …, 2025
2242025
Oryx mllm: On-demand spatial-temporal understanding at arbitrary resolution
Z Liu*, Y Dong*, Z Liu, W Hu, J Lu, Y Rao
arXiv preprint arXiv:2409.12961, 2024
1942024
Egolife: Towards egocentric life assistant
J Yang*, S Liu*, H Guo*, Y Dong*, X Zhang, S Zhang, P Wang, Z Zhou, ...
Proceedings of the Computer Vision and Pattern Recognition Conference, 28885 …, 2025
1932025
Are vlms ready for autonomous driving? an empirical study from the reliability, data, and metric perspectives
S Xie, L Kong, Y Dong, C Sima, W Zhang, QA Chen, Z Liu, L Pan
2025 IEEE/CVF International Conference on Computer Vision (ICCV), 6285-6297, 2025
188*2025
Kimi k3: Open frontier intelligence
K Team, T Bai, Y Bai, Y Bao, J Cai, X Cai, P Cao, Y Cao, Z Chai, ...
arXiv preprint arXiv:2607.24653, 2026
181*2026
Kimi k2. 6
M AI
Hugging Face model card, 2026
163*2026
3dtopia-xl: Scaling high-quality 3d asset generation via primitive diffusion
Z Chen, J Tang, Y Dong, Z Cao, F Hong, Y Lan, T Wang, H Xie, T Wu, ...
2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR …, 2025
1512025
Octopus: Embodied vision-language programmer from environmental feedback
J Yang*, Y Dong*, S Liu*, B Li*, Z Wang, H Tan, C Jiang, J Kang, Y Zhang, ...
European conference on computer vision, 20-38, 2024
1302024
Mme-survey: A comprehensive survey on evaluation of multimodal llms
C Fu, YF Zhang, S Yin, B Li, X Fang, S Zhao, H Duan, X Sun, Z Liu, ...
arXiv preprint arXiv:2411.15296, 2024
1212024
Ola: Pushing the frontiers of omni-modal language model
Z Liu*, Y Dong*, J Wang, Z Liu, W Hu, J Lu, Y Rao
arXiv preprint arXiv:2502.04328, 2025
1102025
Chain-of-spot: Interactive reasoning improves large vision-language models
Z Liu*, Y Dong*, Y Rao, J Zhou, J Lu
arXiv preprint arXiv:2403.12966, 2024
922024
Ego-r1: Chain-of-tool-thought for ultra-long egocentric video reasoning
S Tian, R Wang, H Guo, P Wu, Y Dong, X Wang, J Yang, H Zhang, H Zhu, ...
arXiv preprint arXiv:2506.13654, 2025
78*2025
Coarse correspondence elicit 3d spacetime understanding in multimodal language model
B Liu*, Y Dong*, Y Wang*, Y Rao, Y Tang, WC Ma, R Krishna
arXiv e-prints, arXiv: 2408.00754, 2024
58*2024
Shotbench: Expert-level cinematic understanding in vision-language models
H Liu, J He, Y Jin, D Zheng, Y Dong, F Zhang, Z Huang, Y He, W Chen, ...
Advances in Neural Information Processing Systems 38, 129987-130019, 2026
532026
Complete-to-partial 4D distillation for self-supervised point cloud sequence representation learning
Z Zhang*, Y Dong*, Y Liu, L Yi
Proceedings of the IEEE/CVF conference on computer vision and pattern …, 2023
522023
Efficient inference of vision instruction-following models with elastic cache
Z Liu, B Liu, J Wang, Y Dong, G Chen, Y Rao, R Krishna, J Lu
European Conference on Computer Vision, 54-69, 2024
472024
RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark
Y Shi*, Y Dong*†, Y Ding, Y Wang, X Zhu, S Zhou, W Liu, H Tian, R Wang, ...
arXiv preprint arXiv:2509.24897, 2025
37*2025
The system can't perform the operation now. Try again later.
Articles 1–20