Follow
Zhengyuan Yang
Zhengyuan Yang
Microsoft AI Superintelligence
Verified email at microsoft.com - Homepage
Title
Cited by
Cited by
Year
Mm-vet: Evaluating large multimodal models for integrated capabilities
W Yu*, Z Yang*, L Li, J Wang, K Lin, Z Liu, X Wang, L Wang
The 41st International Conference on Machine Learning (ICML), 2024
15982024
The dawn of lmms: Preliminary explorations with gpt-4v (ision)
Z Yang, L Li, K Lin, J Wang, CC Lin, Z Liu, L Wang
arXiv preprint arXiv:2309.17421, 2023
12842023
Git: A generative image-to-text transformer for vision and language
J Wang, Z Yang, X Hu, L Li, K Lin, Z Gan, Z Liu, C Liu, L Wang
Transactions on Machine Learning Research (TMLR), 2022
10532022
Mm-react: Prompting chatgpt for multimodal reasoning and action
Z Yang, L Li, J Wang, K Lin, E Azarnasab, F Ahmed, Z Liu, C Liu, M Zeng, ...
arXiv preprint arXiv:2303.11381, 2023
7252023
An empirical study of gpt-3 for few-shot knowledge-based vqa
Z Yang, Z Gan, J Wang, X Hu, Y Lu, Z Liu, L Wang
Proceedings of the AAAI conference on artificial intelligence 36 (3), 3081-3089, 2022
6872022
TransVG: End-to-End Visual Grounding with Transformers
J Deng, Z Yang, T Chen, W Zhou, H Li
IEEE International Conference on Computer Vision (ICCV), 2021
6652021
A fast and accurate one-stage approach to visual grounding
Z Yang, B Gong, L Wang, W Huang, D Yu, J Luo
IEEE International Conference on Computer Vision (ICCV), 4683-4693, 2019
5832019
Multimodal foundation models: From specialists to general-purpose assistants
C Li*, Z Gan*, Z Yang*, J Yang*, L Li*, L Wang, J Gao
Foundations and Trends® in Computer Graphics and Vision 16 (1-2), 1-214, 2024
5362024
Scaling up vision-language pretraining for image captioning
X Hu, Z Gan, J Wang, Z Yang, Z Liu, Y Lu, L Wang
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR …, 2022
4972022
Prompting gpt-3 to be reliable
C Si, Z Gan, Z Yang, S Wang, J Wang, J Boyd-Graber, L Wang
International Conference on Learning Representations (ICLR 23), 2022
4712022
Improving One-stage Visual Grounding by Recursive Sub-query Construction
Z Yang, T Chen, L Wang, J Luo
European Conference on Computer Vision (ECCV), 2020
3952020
Showui: One vision-language-action model for generalist gui agent
KQ Lin, L Li, D Gao, Z Yang, Z Bai, W Lei, L Wang, MZ Shou
NeurIPS 2024 Workshop on Open-World Agents, 2024
336*2024
Disco: Disentangled control for referring human dance generation in real world
T Wang, L Li, K Lin, CC Lin, Z Yang, H Zhang, Z Liu, L Wang
CVPR, 2024
324*2024
Ragen: Understanding self-evolution in llm agents via multi-turn reinforcement learning
Z Wang, K Wang, Q Wang, P Zhang, L Li, Z Yang, X Jin, K Yu, MN Nguyen, ...
arXiv preprint arXiv:2504.20073, 2025
3192025
Promptcap: Prompt-guided image captioning for vqa with gpt-3
Y Hu, H Hua, Z Yang, W Shi, NA Smith, J Luo
Proceedings of the IEEE/CVF International Conference on Computer Vision …, 2023
316*2023
Design2code: Benchmarking multimodal code generation for automated front-end engineering
C Si, Y Zhang, R Li, Z Yang, R Liu, D Yang
Proceedings of the 2025 Conference of the Nations of the Americas Chapter of …, 2025
2992025
ReCo: Region-Controlled Text-to-Image Generation
Z Yang, J Wang, Z Gan, L Li, K Lin, C Wu, N Duan, Z Liu, C Liu, M Zeng, ...
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023
2832023
Attentive relational networks for mapping images to scene graphs
M Qi, W Li, Z Yang, Y Wang, J Luo
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 3957-3966, 2019
2372019
End-to-end multi-modal multi-task vehicle control for self-driving cars with visual perceptions
Z Yang, Y Zhang, J Yu, J Cai, J Luo
2018 24th international conference on pattern recognition (ICPR), 2289-2294, 2018
2322018
UniTAB: Unifying Text and Box Outputs for Grounded Vision-Language Modeling
Z Yang, Z Gan, J Wang, X Hu, F Ahmed, Z Liu, Y Lu, L Wang
European Conference on Computer Vision (ECCV), 521--539, 2022
228*2022
The system can't perform the operation now. Try again later.
Articles 1–20