Skip to content
View yanwei-li's full-sized avatar
🎯
Focusing
🎯
Focusing

Highlights

  • Pro

Block or report yanwei-li

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results
Lean 8,867 853 Updated Oct 6, 2026

[NeurIPS 2026] Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models

Python 139 7 Updated Sep 30, 2026

DeepSeek Harness: Everything is a Plugin.

TypeScript 245,131 29,364 Updated Oct 3, 2026

From Foundation to Application

Python 1,000 117 Updated Sep 22, 2026

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

Python 964 47 Updated Sep 17, 2026

Building General-Purpose Robots Based on Embodied Foundation Model

Python 1,284 99 Updated Sep 18, 2026

Cambrian-P: Pose-Grounded Video Understanding

Python 116 3 Updated Jul 28, 2026

Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond

TypeScript 325 11 Updated Sep 30, 2026

Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding

Python 362 3 Updated Aug 5, 2026

JoyAI-Image is the unified multimodal foundation model for image understanding, text-to-image generation, and instruction-guided image editing.

Python 2,163 164 Updated Aug 5, 2026

[ICLR 2026] Efficient Reasoning with Balanced Thinking

Python 338 8 Updated May 30, 2026

GigaBrain-0: A World Model-Powered Vision-Language-Action Model

Python 2,666 208 Updated Sep 9, 2026

Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence

Python 403 20 Updated Jun 20, 2026

[RSS 2026] Causal video-action world model for generalist robot control

Python 1,925 186 Updated Jul 9, 2026

Qwen-Image-Layered: Layered Decomposition for Inherent Editablity

Python 2,127 170 Updated Dec 31, 2025

[NeurIPS 2025] Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance

Python 659 37 Updated Jan 5, 2026

SAM 3D Objects

Python 7,489 888 Updated Jun 2, 2026

Depth Anything 3

Python 6,438 712 Updated Jul 27, 2026

[ECCV2026] Visual Spatial Tuning

Jupyter Notebook 212 9 Updated Mar 25, 2026

Cambrian-S: Towards Spatial Supersensing in Video

Python 571 20 Updated Apr 3, 2026

Native Multimodal Models are World Learners

Python 1,561 69 Updated Dec 30, 2025

[NeurIPS'25] Official repository of Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations

Python 540 28 Updated Aug 18, 2026

The official implement of VITA, VITA15, LongVITA, VITA-Audio, VITA-VLA, and VITA-E.

Python 165 4 Updated Oct 28, 2025

This is a project on visual spatial reasoning tasks-SIBench

Python 28 1 Updated Jan 12, 2026

Fully Open Framework for Democratized Multimodal Training

Python 1,216 75 Updated Oct 7, 2026

Official Code for "Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search"

Python 427 19 Updated Jan 29, 2026

MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech

Python 202 11 Updated Jun 20, 2026

verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework

Python 23,773 4,712 Updated Oct 7, 2026

Code and dataset link for "DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World"

129 2 Updated Oct 2, 2025

The implementation of Extreme Viewpoint 4D Video Generation

Python 270 18 Updated Sep 6, 2025
Next