Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up
JackMa's picture

JackMa

JacckMa
6
·

AI & ML interests

None yet

Recent Activity

upvoted a paper 1 day ago
Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs
upvoted a paper 5 months ago
Heterogeneous Agent Collaborative Reinforcement Learning
upvoted a paper 5 months ago
Does Your Reasoning Model Implicitly Know When to Stop Thinking?
View all activity

Organizations

None yet

upvoted a paper 1 day ago

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs

Paper • 2608.01755 • Published 3 days ago • 137
upvoted 2 papers 5 months ago

Heterogeneous Agent Collaborative Reinforcement Learning

Paper • 2603.02604 • Published Mar 3 • 199

Does Your Reasoning Model Implicitly Know When to Stop Thinking?

Paper • 2602.08354 • Published Feb 9 • 267
upvoted a paper 6 months ago

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger

Paper • 2602.08222 • Published Feb 9 • 290
upvoted a paper 7 months ago

Your Group-Relative Advantage Is Biased

Paper • 2601.08521 • Published Jan 13 • 158
upvoted an article over 1 year ago
view article
Article

Illustrating Reinforcement Learning from Human Feedback (RLHF)

  • +2
natolambert, LouisCastricato, lvwerra, Dahoas
•
Dec 9, 2022
• 422
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs