OpenClaw-RL: Train any agent simply by talking
-
Updated
May 23, 2026 - Python
OpenClaw-RL: Train any agent simply by talking
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
Awesome List for On-Policy Distillation
A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of large language models
A curated collection of papers and resources on On-Policy Distillation for Large Language Models.
🔥 DanceOPD: On-Policy Generative Field Distillation
🔥 On-Policy Self-Distillation in Diffusion Models
Run a documented subset of verl-style OPD on one consumer GPU—typed config, Parquet prompts, and PEFT scale-out artifacts.
A user-friendly & efficient knowledge distillation framework for LLMs, supporting off-policy, on-policy (OPD), cross-tokenizer, multimodal, and on-policy self-distillation.
[ICML 2026] [OnlineSPEC] When Drafts Evolve: Speculative Decoding Meets Online Learning
Source code of paper "RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation"
Weak-to-Strong Generalization via Direct On-Policy Distillation
Official implementation of On-Policy Delta Distillation (OPD2)
Tiny-R2: A hybrid architecture integrating SWA, CSA, HCA, mHC, and DSMoE under the DeepSeek V4 design paradigm, enabling single-GPU OPD post-training.
Official code for "Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate" (arXiv:2605.01347).
[Preprint 2026] StreamOPD: A Post-Training Recipe with Spatio-Temporal Cue Gating for Streaming Video Understanding
A curated paper list on On-Policy Distillation (OPD) & On-Policy Self-Distillation (OPSD) for LLMs — student rollouts + teacher feedback, covering white-box/black-box methods, OPD-RL hybrids, agentic & multimodal applications, and industrial multi-teacher recipes.
The local inference server for coding agents. Pure Rust, one binary, Apple Silicon + NVIDIA. Anthropic + OpenAI APIs; the KV cache survives across turns, so turn 20 starts as fast as turn 2. OPD training on the same runtime.
CaOPD: Calibration-Aware On-Policy Distillation
To associate your repository with the on-policy-distillation topic, visit your repo's landing page and select "manage topics."