arxiv:2605.15726
Chanuk Lee
tally0818
AI & ML interests
LLM post-training
Recent Activity
upvoted a paper about 10 hours ago
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning upvoted a paper about 10 hours ago
DAPD: Dual-Anchored Policy Distillation upvoted a paper about 11 hours ago
On-Policy Self-Distillation without Any SupervisionOrganizations
None yet