Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
205516.8
TFLOPS
Leandro von Werra
PRO
lvwerra
910
95
129
Follow
Lemba's profile picture
Snufkin1's profile picture
Griffinsandra222's profile picture
835 followers
·
87 following
https://www.lvwerra.com
lvwerra
lvwerra
lvwerra
AI & ML interests
NLP and RL
Recent Activity
new
activity
4 minutes ago
rl-llm-wiki/knowledge-base:
algorithms/learning-from-feedback: new node (hindsight/language/reflective feedback) + taxonomy; de-orphans 7 sources
new
activity
5 minutes ago
rl-llm-wiki/knowledge-base:
Add expert sub-article: pluralistic preference optimization (multi-objective, personalized, robust)
new
activity
27 minutes ago
rl-llm-wiki/knowledge-base:
alignment-tax: runnable check for the HMA/model-averaging mechanism (SS4.1)
View all activity
Organizations
lvwerra
's papers
17
arxiv:
2510.08697
arxiv:
2506.20920
arxiv:
2504.05299
arxiv:
2502.02737
arxiv:
2501.08365
arxiv:
2410.24198
arxiv:
2406.17557
arxiv:
2405.18392
arxiv:
2402.19173
arxiv:
2310.16944
arxiv:
2308.07124
arxiv:
2305.06161
arxiv:
2303.03915
arxiv:
2301.03988
arxiv:
2211.15533
arxiv:
2211.05100
arxiv:
2210.01970