Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Paper • 2607.17524 • Published 7 days ago • 6
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Paper • 2607.17524 • Published 7 days ago • 6
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Paper • 2607.17524 • Published 7 days ago • 6
Value-Aware Stochastic KV Cache Eviction for Reasoning Models Paper • 2606.03928 • Published Jun 2 • 8
Value-Aware Stochastic KV Cache Eviction for Reasoning Models Paper • 2606.03928 • Published Jun 2 • 8
Value-Aware Stochastic KV Cache Eviction for Reasoning Models Paper • 2606.03928 • Published Jun 2 • 8
deqing/convergent-llama-300M-muon-6digit-addition_6digit_llama Text Generation • 0.3B • Updated Jun 3 • 149 • 1
deqing/convergent-llama-300M-muon-6digit-addition_6digit_custom3 Text Generation • 0.2B • Updated Jun 2 • 49 • 1
deqing/convergent-llama-300M-muon-base15-addition_base15 Text Generation • 0.2B • Updated May 31 • 28
deqing/convergent-llama-300M-muon-6digit-addition_6digit_llama Text Generation • 0.3B • Updated Jun 3 • 149 • 1
deqing/convergent-llama-300M-muon-base12-addition_base12 Text Generation • 0.2B • Updated May 30 • 77
deqing/convergent-llama-300M-muon-4digit-addition_4digit_custom3_right2left 0.2B • Updated May 30 • 5