Skip to content
View CaraJ7's full-sized avatar

Organizations

@MME-Benchmarks

Block or report CaraJ7

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Agent-driven creation of editable, tastefully crafted visual artifacts.

Python 1,149 80 Updated Sep 29, 2026

Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

Python 115 3 Updated Aug 6, 2026

OpenCoF: Learning to Reason Through Video Generation

Python 77 Updated Oct 3, 2026

Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1

Python 78,114 12,526 Updated Sep 28, 2026

GenClaw: Code-Driven Agentic Image Generation

Python 315 11 Updated Aug 28, 2026
Python 175 8 Updated Mar 18, 2026

Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepowe…

TeX 13,332 941 Updated Jun 16, 2026

AI-powered slide workspace for creating, editing, versioning, and presenting beautiful reveal.js decks from prompts and source files.

TypeScript 17 Updated Apr 14, 2026

[ECCV26]CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation

Python 56 1 Updated Aug 7, 2026

Implement search image generation similar to Nano-banana-pro / Seedream / FLUX. [SIGGRAPH Asia 2026]

Python 98 9 Updated Mar 10, 2026

[CVPR 2026] The official implementation of The paper "Are We Ready for RL in Text-to-3D Generation? A Progressive Investigation"

Python 122 2 Updated Feb 28, 2026

[ECCV 2026 Oral] RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards.

Python 316 18 Updated Jul 27, 2026

Offical Repository for Paper: DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation

19 Updated Dec 7, 2025

The first Interleaved framework for textual reasoning within the visual generation process

Python 167 2 Updated Mar 16, 2026

Are Video Models Ready as Zero-shot Reasoners?

Python 88 4 Updated Nov 24, 2025

ULMEvalKit: One-Stop Eval ToolKit for Image Generation

Python 56 2 Updated Dec 17, 2025

A huge collection of SVG logos

SVG 6,838 756 Updated Sep 28, 2026

Qwen-Image text to image lora trainer

Python 768 71 Updated Dec 16, 2025

A curated collection of fun and creative examples generated with Nano Banana & Nano Banana Pro🍌, Gemini-2.5-flash-image based model. We also release Nano-consistent-150K openly to support the commu…

23,842 2,394 Updated Sep 3, 2026

Echo-4o: Harnessing Proprietary Models’ Synthetic Images for Improved Image Generation

Jupyter Notebook 506 28 Updated Dec 9, 2025

A curated gallery and toolkit designed to provide inspiration for scientific illustrations, project sites, and visual storytelling in research.

1,059 30 Updated Mar 28, 2026

CLIP+MLP Aesthetic Score Predictor

Python 1,349 115 Updated Jul 1, 2024

Open-source unified multimodal model

Python 6,189 544 Updated May 4, 2026

[CVPR 2025 (Oral)] Open implementation of "RandAR"

Python 209 10 Updated Jul 14, 2025

Official codebase for "Self Forcing: Bridging Training and Inference in Autoregressive Video Diffusion" (NeurIPS 2025 Spotlight)

Python 3,534 285 Updated Sep 12, 2025

[NeurIPS 2025] MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning

Python 108 5 Updated Sep 19, 2025

Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual information for complex reasoning, planning, and generation.

1,511 46 Updated Mar 9, 2026

Jodi: Unification of Visual Generation and Understanding via Joint Modeling

Python 94 2 Updated Mar 6, 2026

CrossLMM: Decoupling Long Video Sequences from LMMs via Dual Cross-Attention Mechanisms

25 Updated Dec 21, 2025

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning

Python 239 11 Updated May 30, 2025
Next