stefanocarrera/sqlautophagycode_D_test_Qwen3-8B_t1.25_g5_run0_metrics Viewer • Updated about 8 hours ago • 579 • 1
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Paper • 2607.14952 • Published 5 days ago • 182
EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos Paper • 2607.09701 • Published 30 days ago • 17
timaeus/rl-lm-pythia1b-sentiment-neg-alpha0-grpo-nostd-gs4-tp1-tk0-pt80000-lr1e-6-bs600-seed1 Updated 4 days ago • 1
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Paper • 2607.08964 • Published 12 days ago • 74
Rank-Then-Act: Reward-Free Control from Frame-Order Progress Paper • 2607.01897 • Published 19 days ago • 7