Subscribe
Sign in
Home
Archive
About
Positive Gradients, Negative Gradients
+ the Importance of Pre-Training Priors
Dec 19, 2025
•
Alex Nikulkov
1
1
Most Popular
View all
Bandits vs Reinforcement Learning from Human Feedback
Apr 30, 2024
•
Alex Nikulkov
5
2
Reward Modeling for RLHF
Jan 10, 2024
•
Alex Nikulkov
2
Reward Model Overoptimization: Root Causes and Mitigations
Apr 7, 2024
•
Alex Nikulkov
1
Bandits vs Reinforcement Learning from Human Feedback
Are single-step or multi-step models better suited for RLHF? Find out the main differences between them in this post
Apr 30, 2024
•
Alex Nikulkov
5
2
Reward Model Overoptimization: Root Causes and Mitigations
When I first ran an RLHF training job, I was surprised at how easily the reward model scores increased during the training process.
Apr 7, 2024
•
Alex Nikulkov
1
Reward Modeling for RLHF
An introduction to reward models
Jan 10, 2024
•
Alex Nikulkov
2
Hello World
Welcome to Reinforced!
Jan 7, 2024
•
Alex Nikulkov
Reinforced
A newsletter/blog about Reinforcement Learning and Generative AI
Subscribe
Recommendations
Interconnects AI
Nathan Lambert
Machine Learning Frontiers
Samuel Flender
RL News
Ryan Lee
Reinforced
Subscribe
About
Archive
Recommendations
Sitemap
This site requires JavaScript to run correctly. Please
turn on JavaScript
or unblock scripts