MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
MMSI-Bench evaluates whether MLLMs can integrate evidence across multiple images to solve grounded, real-world spatial reasoning problems.
ICLR, 2026
Researcher · Multimodal Learning & Evaluation
ByteDance Seed Singapore
欢迎围绕多模态学习、大模型评测与视频理解开展学术合作,也欢迎对开源评测工具和基准感兴趣的研究者与工程师联系我。
I am open to academic collaborations on multimodal learning, LLM/LMM evaluation, and video understanding. Please feel free to reach out by email.
My research focuses on multimodal learning, LLM/LMM evaluation, and video understanding. I build open-source evaluation infrastructure and benchmarks that make model capabilities easier to measure, compare, and reproduce.
Three papers are listed in the NeurIPS 2026 program: Castle-in-the-Air, RISE-Video, and Training Long-Context Vision-Language Models Effectively.
LEGO-Puzzles is published at ECCV 2026.
Seed2.1 is released, advancing multimodal reasoning and agentic productivity across the Seed model family.
Five papers are published at CVPR 2026, covering GUI-agent evaluation, spatial self-supervised RL, agentic reward modeling, visual reasoning, and poster intelligence.
Seed2.0 is officially launched, strengthening multimodal understanding, complex instruction execution, and real-world agent capabilities.
Three papers are published at ICLR 2026: MMSI-Bench, VisualPRM, and MM-HELIX.
Seed1.8 is officially released as a generalized agentic model with multimodal, search, coding, and GUI capabilities.
I joined ByteDance Seed in Singapore, after two years at Shanghai AI Laboratory working on OpenCompass and large-model evaluation.
RISEBench is accepted by NeurIPS 2025 Datasets & Benchmarks as an Oral presentation.
MMSI-Bench evaluates whether MLLMs can integrate evidence across multiple images to solve grounded, real-world spatial reasoning problems.
ICLR, 2026
LEGO-Puzzles probes spatial understanding and multi-step planning through interpretable LEGO assembly tasks that remain difficult for frontier MLLMs.
ECCV, 2026
Visual-RFT brings reinforcement fine-tuning with verifiable rewards to visual perception tasks, improving data-efficient adaptation with limited examples.
IEEE/CVF International Conference on Computer Vision (ICCV), 2025
RISEBench measures whether image editors can follow reasoning-heavy temporal, causal, spatial, and logical instructions while preserving visual quality.
NeurIPS D&B Oral, 2025
OmniAlign-V improves multimodal assistants with diverse preference-alignment data and introduces MM-AlignBench for open-ended alignment evaluation.
ACL, 2025
This work quantifies dimension, instance, and cross-benchmark redundancy and turns the findings into practical principles for efficient MLLM evaluation.
ACL, 2025
Condor synthesizes and refines knowledge-intensive instruction data to strengthen language-model alignment without relying on a fixed seed dataset.
ACL, 2025
MMBench provides a bilingual, ability-structured benchmark and CircularEval protocol for robustly diagnosing the capabilities of multimodal models.
European Conference on Computer Vision (ECCV), 2024
MMStar curates genuinely vision-dependent questions to expose data leakage and inflated gains in large vision-language model evaluation.
Conference on Neural Information Processing Systems (NeurIPS), 2024
VLMEvalKit unifies model inference, benchmark access, and standardized scoring into a widely adopted toolkit for reproducible multimodal evaluation.
ACM International Conference on Multimedia (MM), 2024
ShareGPT4Video scales detailed video captioning to improve both video-language understanding and text-to-video generation.
Conference on Neural Information Processing Systems (NeurIPS), Datasets & Benchmarks Track, 2024
ProSA measures prompt sensitivity at the instance level and analyzes why semantically equivalent instructions can produce unstable LLM performance.
Empirical Methods in Natural Language Processing (EMNLP), Findings, 2024
MMBench-Video evaluates long-form, multi-shot video understanding with free-form questions and a fine-grained taxonomy of temporal capabilities.
Conference on Neural Information Processing Systems (NeurIPS), Datasets & Benchmarks Track, 2024
BotChat evaluates multi-turn conversational ability through scalable bot-to-bot interaction and fine-grained dialogue-quality assessment.
Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), 2024
Prism decouples perception and reasoning to diagnose VLM capability bottlenecks and combine compact specialist components more effectively.
Conference on Neural Information Processing Systems (NeurIPS), 2024
Ada-LEval adapts task length to model context windows, enabling controlled evaluation of retrieval, ordering, and full-text comprehension at scale.
Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), 2024
JourneyDB pairs millions of generated images with prompts and annotations to benchmark captioning, prompt inversion, retrieval, and visual question answering.
Conference on Neural Information Processing Systems (NeurIPS), Datasets & Benchmarks Track, 2023
SkeleTR models both intra-person motion and inter-person relationships for skeleton action recognition in crowded, unconstrained scenes.
IEEE/CVF International Conference on Computer Vision (ICCV), 2023
PoseC3D represents skeletons as 3D heatmap volumes, yielding a simple, robust, and scalable alternative to graph-based action recognition.
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022
PYSKL consolidates diverse skeleton-action algorithms, training practices, and benchmarks into one reproducible research toolbox.
ACM International Conference on Multimedia (MM), 2022
DG-STGCN learns dynamic spatial graphs and multi-scale temporal structure instead of relying on a fixed skeleton topology.
arXiv preprint arXiv:2210.05895, 2022
TransRank reframes transformation recognition as relative ranking to learn stronger self-supervised video representations from temporal and spatial transformations.
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022
OmniSource unifies images, short clips, and untrimmed web videos to make large-scale web supervision more data-efficient for video recognition.
European Conference on Computer Vision (ECCV), 2020
TRB combines skeleton and contour keypoints to represent human pose and shape compactly for estimation, editing, and conditional generation.
IEEE/CVF International Conference on Computer Vision (ICCV), 2019
An all-in-one evaluation toolkit for large vision-language and multimodal models. It connects model inference, benchmark execution, scoring, and result comparison in one reproducible workflow used by both researchers and practitioners.
A comprehensive platform for evaluating large language and multimodal models. Its modular configs, task scheduling, evaluators, and reporting tools make broad and reproducible model comparisons easier to run and extend.
A unified toolbox for skeleton-based action recognition, with strong baselines spanning GCN- and CNN-based methods. It packages training recipes, pretrained models, and benchmark results so new methods can be compared on consistent footing.
OpenMMLab's modular video-understanding toolbox for action recognition, localization, skeleton-based recognition, and related tasks. It offers reusable components, extensive model coverage, and standardized training and inference pipelines.