https://adongwanai.github.io/AgentGuide | AI Agent开发指南 | LangGraph实战 | 高级RAG | 转行大模型 | 大模型面试 | 算法工程师 | 面试题库 | 强化学习|数据合成
-
Updated
Sep 8, 2026 - MDX
https://adongwanai.github.io/AgentGuide | AI Agent开发指南 | LangGraph实战 | 高级RAG | 转行大模型 | 大模型面试 | 算法工程师 | 面试题库 | 强化学习|数据合成
Your AI intranet: network the computers you already own for inference and training.
Procedural data generators for verifiable reasoning, synthetic pretraining, post-training, evaluation, and RL.
Multimodal RL training framework for diffusion & omni models
Pi extension for resumable Bash-specialization workflows with verified teacher data, SSH GPU orchestration, SFT, and GRPO.
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).
A lightweight, scalable RL post-training framework for agentic environments.
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
Build decoder-only Transformers from scratch in PyTorch, covering tokenization, pretraining, supervised fine-tuning, and alignment using core tensor operations.
Automate Microsoft Account OAuth registration and token imports with this Chrome extension for the Microsoft Account Manager API.
🚀 Accelerate large language model training and fine-tuning with Surogate’s high-performance, mixed-precision framework in C++ and Python.
🎤 Fine-tune the Llasa TTS model with GRPO using Hugging Face tools to enhance performance and evaluate rewards with Whisper ASR and WER metrics.
GRPO integration for multi-turn retail agents with Tau2 environments, VERL AgentLoop, tool rollouts, and terminal rewards.
Post-training pipeline for math reasoning: DeepMath data curation, SFT, and GRPO with composite rewards on Qwen2.5-3B
Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training. Reinforcement learning for Qwen3.6, GPT-OSS, Llama, and more!
Grammar-aided Constrained Decondig for self-aligned LLM with SFT and GRPO
To associate your repository with the grpo topic, visit your repo's landing page and select "manage topics."