Auxiliary datasets for tmax.
Hamish Ivison
hamishivi
AI & ML interests
NLP :)
Recent Activity
updated a model 2 days ago
hamishivi/vip_glm52_hypers_1507_qwen3_4b_math published a model 3 days ago
hamishivi/vip_glm52_hypers_1507_qwen3_4b_math updated a dataset 24 days ago
allenai/tmax-15k-open-instructOrganizations
RLVE
Models for "RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments" - https://arxiv.org/abs/2511.07317
-
RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments
Paper • 2511.07317 • Published • 18 -
hamishivi/OpenThinker3-1.5B-RLVE
Text Generation • 2B • Updated • 36 • • 2 -
hamishivi/Nemotron-Research-Reasoning-Qwen-1.5B-v2-RLVE
Text Generation • 2B • Updated • 29 • • 3
Tmax Extras
Auxiliary datasets for tmax.
RLVE
Models for "RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments" - https://arxiv.org/abs/2511.07317
-
RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments
Paper • 2511.07317 • Published • 18 -
hamishivi/OpenThinker3-1.5B-RLVE
Text Generation • 2B • Updated • 36 • • 2 -
hamishivi/Nemotron-Research-Reasoning-Qwen-1.5B-v2-RLVE
Text Generation • 2B • Updated • 29 • • 3
models 336
hamishivi/vip_glm52_hypers_1507_qwen3_4b_math
Updated • 10
hamishivi/swerl_qwen35_27b_fp32lm_dppo_tmax__42_step300
2.65M • Updated • 11
hamishivi/swerl_qwen35_27b_fp32lm_dppo_tmax_step240
2.65M • Updated • 6
hamishivi/swerl_qwen35_27b_fp32lm_dppo_tmax_step200
2.65M • Updated • 5
hamishivi/swerl_qwen35_27b_fp32lm_dppo_tmax_step160
2.65M • Updated • 5
hamishivi/swerl_qwen35_4b_fp32lm_dppo
Updated • 9
hamishivi/swerl_qwen35_2b_fp32lm_dppo
Updated • 4
hamishivi/swerl_qwen35_9b_fp32lm_dppo_swesmith
Updated • 6
hamishivi/qwen35_9b_tmax_skill_tax_no_tool_call_sft
9B • Updated • 6
hamishivi/swerl_qwen35_27b_fp32lm_dppo_tmax_step100
2.65M • Updated • 4
datasets 228
hamishivi/qwen35-4b-drpo-vs0f49th-trainer-logprobs
Preview • Updated • 1.42k
hamishivi/tmax-sft-big
Viewer • Updated • 327k • 107
hamishivi/agent-task-endless-terminals
Viewer • Updated • 2.49k • 16
hamishivi/sft_ablations_scientific_minimax_v1_sanitized
Viewer • Updated • 2.23k • 19
hamishivi/sft_ablations_bc_only_v1_sanitized
Viewer • Updated • 5.59k • 32
hamishivi/agent-task-cli-gym
Viewer • Updated • 1.55k • 18
hamishivi/agent-task-swe-smith
Viewer • Updated • 59.1k • 20
hamishivi/agent-task-r2e-gym
Viewer • Updated • 4.58k • 8
hamishivi/agent-task-combined
Viewer • Updated • 27k • 379
hamishivi/sft_ablations_redsearcher_sft_sanitized
Viewer • Updated • 9.81k • 10