-
MMLab, CUHK
- Hong Kong, China
- https://caraj7.github.io/
Stars
Agent-driven creation of editable, tastefully crafted visual artifacts.
Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System
OpenCoF: Learning to Reason Through Video Generation
Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1
Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepowe…
AI-powered slide workspace for creating, editing, versioning, and presenting beautiful reveal.js decks from prompts and source files.
[ECCV26]CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation
Implement search image generation similar to Nano-banana-pro / Seedream / FLUX. [SIGGRAPH Asia 2026]
[CVPR 2026] The official implementation of The paper "Are We Ready for RL in Text-to-3D Generation? A Progressive Investigation"
[ECCV 2026 Oral] RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards.
Offical Repository for Paper: DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
The first Interleaved framework for textual reasoning within the visual generation process
ULMEvalKit: One-Stop Eval ToolKit for Image Generation
Qwen-Image text to image lora trainer
A curated collection of fun and creative examples generated with Nano Banana & Nano Banana Pro🍌, Gemini-2.5-flash-image based model. We also release Nano-consistent-150K openly to support the commu…
Echo-4o: Harnessing Proprietary Models’ Synthetic Images for Improved Image Generation
A curated gallery and toolkit designed to provide inspiration for scientific illustrations, project sites, and visual storytelling in research.
CLIP+MLP Aesthetic Score Predictor
[CVPR 2025 (Oral)] Open implementation of "RandAR"
Official codebase for "Self Forcing: Bridging Training and Inference in Autoregressive Video Diffusion" (NeurIPS 2025 Spotlight)
[NeurIPS 2025] MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual information for complex reasoning, planning, and generation.
Jodi: Unification of Visual Generation and Understanding via Joint Modeling
CrossLMM: Decoupling Long Video Sequences from LMMs via Dual Cross-Attention Mechanisms
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning