AI & ML interests
None defined yet.
Recent Activity
View all activity
Papers
ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
datasets 167
benchflow/frontierphysics-traj
Updated • 6.99k • 1
benchflow/frontierphysics-pr638-evidence
Updated • 6
benchflow/frontierphysics-pr608-evidence
Updated • 3
benchflow/frontierphysics-pr660-evidence
Updated • 10
benchflow/frontierphysics-pr615-evidence
Updated • 5
benchflow/frontierphysics-pr651-evidence
Updated • 7
benchflow/frontierphysics-pr634-evidence
Updated • 8
benchflow/frontierphysics-pr623-evidence
Updated • 8
benchflow/frontierphysics-pr633-evidence
Updated • 8
benchflow/frontierphysics-pr631-evidence
Updated • 9