Chanuk Lee
tally0818
AI & ML interests
LLM post-training
Recent Activity
upvoted a paper about 13 hours ago
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning upvoted a paper about 13 hours ago
DAPD: Dual-Anchored Policy Distillation upvoted a paper about 13 hours ago
On-Policy Self-Distillation without Any SupervisionOrganizations
None yet