Skip to content
View Aleafy's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report Aleafy

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Just a VLM Agent Can Play Robots

Python 519 32 Updated Oct 6, 2026

GPT-6 Astra for embodied AI and robotics.

1,099 26 Updated Oct 7, 2026

In-Context Robot Learning with VLM Agents

Python 302 5 Updated Oct 3, 2026

HY-SOAR:Self-Correction for Optimal Alignment and Refinement in Diffusion Models

Python 747 65 Updated Apr 21, 2026

A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.

Python 1,350 93 Updated Jul 14, 2026

Project Lyra: Open Generative 3D World Models

Python 2,644 259 Updated Jul 20, 2026

A unified inference and post-training framework for accelerated video generation.

Python 4,564 487 Updated Oct 7, 2026

Unified Codebase for Advanced World Models.

Python 879 56 Updated Sep 23, 2026

An in-the-wild benchmark for AI agents in the production harness.

Python 529 66 Updated Sep 18, 2026

EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL

Python 5,183 398 Updated Sep 19, 2026

ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework, no lock-in — works…

Python 17,098 1,437 Updated Oct 7, 2026

🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning

Python 27,985 5,825 Updated Oct 7, 2026

VIGA: Vision-as-Inverse-Graphics Agent

Python 1,315 126 Updated May 6, 2026

A unified framework for easy fine-tuning in Flow-Matching models

Python 716 59 Updated Oct 3, 2026

This repository provides FlashPortrait custom nodes for ComfyUI.

Python 25 2 Updated Dec 29, 2025

A Unified Visual Generator with Interleaved OmniModal Context

Python 235 3 Updated Mar 5, 2026

[ICLR 26 Oral] Stable Video Infinity: Infinite-Length Video Generation with Error Recycling

Python 2,582 232 Updated Sep 17, 2026

Official code for StoryMem: Multi-shot Long Video Storytelling with Memory

Python 772 74 Updated Jul 22, 2026

The official implementation of InfiniteVGGT

Python 391 21 Updated Apr 19, 2026

Mixture-of-Groups Attention for End-to-End Long Video Generation

100 Updated Oct 22, 2025

The repository provides code for running inference with the Meta Segment Anything Audio Model (SAM-Audio), links for downloading the trained model checkpoints, and example notebooks that show how t…

Python 3,651 332 Updated May 26, 2026

Qwen-Image-Layered: Layered Decomposition for Inherent Editablity

Python 2,127 170 Updated Dec 31, 2025

Official Implementation of "MemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video Narratives"

Python 220 8 Updated Dec 29, 2025

[NeurIPS 24] The implementation and dataset of LiveScene: Language Embedding Interactive Radiance Fields for Physical Scene Rendering and Control

Python 60 2 Updated Mar 31, 2025

[Siggraph Asia 25] SS4D: Native 4D Generative Model via Structured Spacetime Latents

Python 36 3 Updated Dec 17, 2025

HY-World 1.5: A Systematic Framework for Interactive World Modeling with Real-Time Latency and Geometric Consistency

Python 1,610 146 Updated Jun 10, 2026

[CVPR2026]We present FlashPortrait, an end-to-end video diffusion transformer capable of synthesizing ID-preserving, infinite-length videos while achieving up to 6$\times$ acceleration in inference…

Python 479 37 Updated Feb 21, 2026

[CVPR 2026] V-RGBX: Video Editing with Accurate Controls over Intrinsic Properties

Python 112 7 Updated Jan 17, 2026

[CVPR 2026] Official Code for "ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning"

Python 198 12 Updated Feb 13, 2026

ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation

118 5 Updated Dec 11, 2025
Next