-
MMLab, CUHK
- Hong Kong, China
- https://caraj7.github.io/
Stars
Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1
《代码随想录》LeetCode 刷题攻略:200道经典题目刷题顺序,共60w字的详细图解,视频难点剖析,50余张思维导图,支持C++,Java,Python,Go,JavaScript等多语言版本,从此算法学习不再迷茫!🔥🔥 来看看,你会发现相见恨晚!🚀
Build and share delightful machine learning apps, all in Python. 🌟 Star to support our work!
An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.
[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
A curated collection of fun and creative examples generated with Nano Banana & Nano Banana Pro🍌, Gemini-2.5-flash-image based model. We also release Nano-consistent-150K openly to support the commu…
深度学习入门教程, 优秀文章, Deep Learning Tutorial
Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepowe…
🏄 Scalable embedding, reasoning, ranking for images and sentences with CLIP
LAVIS - A One-stop Library for Language-Vision Intelligence
Refine high-quality datasets and visual AI models
Accessible large language models via k-bit quantization for PyTorch.
BertViz: Visualize Attention in Transformer Models
Script Commands let you tailor Raycast to your needs. Think of them as little productivity boosts throughout your day.
[ICLR 2024] Fine-tuning LLaMA to follow Instructions within 1 Hour and 1.2M Parameters
The unofficial python package that returns response of Google Bard through cookie value.
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
Witness the aha moment of VLM with less than $3.
OpenDILab Decision AI Engine. The Most Comprehensive Reinforcement Learning Framework B.P.
Official codebase for "Self Forcing: Bridging Training and Inference in Autoregressive Video Diffusion" (NeurIPS 2025 Spotlight)
🦦 Otter, a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo), trained on MIMIC-IT and showcasing improved instruction-following and in-context learning ability.
SSD: Single Shot MultiBox Detector | a PyTorch Tutorial to Object Detection
PPO x Family DRL Tutorial Course(决策智能入门级公开课:8节课帮你盘清算法理论,理顺代码逻辑,玩转决策AI应用实践 )