Follow
Shengyi Qian
Shengyi Qian
Research Scientist, Meta FAIR
Verified email at meta.com - Homepage
Title
Cited by
Cited by
Year
The llama 4 herd: The beginning of a new era of natively multimodal ai innovation
AI Meta
6072025
LLM-Grounder: Open-vocabulary 3d visual grounding with large language model as an agent
J Yang, X Chen, S Qian, N Madaan, M Iyengar, DF Fouhey, J Chai
2024 IEEE International Conference on Robotics and Automation (ICRA), 7694-7701, 2024
2372024
Affordancellm: Grounding affordance from vision language models
S Qian, W Chen, M Bai, X Zhou, Z Tu, LE Li
2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition …, 2024
1442024
Multi-object hallucination in vision language models
X Chen, Z Ma, X Zhang, S Xu, J Yang, DF Fouhey, J Chai, S Qian
Advances in Neural Information Processing Systems 37, 44393-44418, 2024
1372024
OASIS: A Large-Scale Dataset for Single Image 3D in the Wild
W Chen, S Qian, D Fan, N Kojima, M Hamilton, J Deng
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020
1022020
Learning single-image depth from videos using quality assessment networks
W Chen, S Qian, J Deng
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019
852019
Planar Surface Reconstruction from Sparse Views
L Jin, S Qian, A Owens, DF Fouhey
International Conference on Computer Vision (ICCV), 2021
602021
3d-grand: A million-scale dataset for 3d-llms with better grounding and less hallucination
J Yang, X Chen, N Madaan, M Iyengar, S Qian, DF Fouhey, J Chai
2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR …, 2025
582025
3d-mvp: 3d multiview pretraining for manipulation
S Qian, K Mo, V Blukis, DF Fouhey, D Fox, A Goyal
2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR …, 2025
502025
Understanding 3D Object Interaction from a Single Image
S Qian, DF Fouhey
International Conference on Computer Vision (ICCV), 2023
472023
Understanding 3D Object Articulation in Internet Videos
S Qian, L Jin, C Rockwell, S Chen, DF Fouhey
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022
472022
Pitfalls in link prediction with graph neural networks: Understanding the impact of target-link inclusion & better practices
J Zhu, Y Zhou, VN Ioannidis, S Qian, W Ai, X Song, D Koutra
Proceedings of the 17th ACM International Conference on Web Search and Data …, 2024
372024
The llama 4 herd: The beginning of a new era of natively multimodal ai innovation
M AI
https://ai. meta. com/blog/llama-4-multimodal-intelligence/, 2025
342025
Mosaic of modalities: A comprehensive benchmark for multimodal graph learning
J Zhu, Y Zhou, S Qian, Z He, T Zhao, N Shah, D Koutra
2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR …, 2025
332025
Beyond language modeling: An exploration of multimodal pretraining
S Tong, D Fan, J Nguyen, E Brown, G Zhou, S Qian, B Zheng, T Vallaeys, ...
arXiv preprint arXiv:2603.03276, 2026
302026
Associative3D: Volumetric Reconstruction from Sparse Views
S Qian, L Jin, D Fouhey
European Conference on Computer Vision (ECCV), 2020
262020
Sound Localization from Motion: Jointly Learning Sound Direction and Camera Rotation
Z Chen, S Qian, A Owens
International Conference on Computer Vision (ICCV), 2023
222023
Multimodal graph benchmark
J Zhu, Y Zhou, S Qian, Z He, T Zhao, N Shah, D Koutra
arXiv preprint arXiv:2406.16321 10, 2024
192024
Linkgpt: Teaching large language models to predict missing links
Z He, J Zhu, S Qian, J Chai, D Koutra
arXiv preprint arXiv:2406.04640, 2024
162024
Disco balances the scales: Adaptive domain-and difficulty-aware reinforcement learning on imbalanced data
Y Zhou, J Zhu, S Qian, Z Zhao, X Wang, X Liu, M Li, P Xu, W Ai, F Huang
arXiv preprint arXiv:2505.15074, 2025
132025
The system can't perform the operation now. Try again later.
Articles 1–20