Back to home
RLHF
4 articles tagged with this topic
OpenAIAnthropic
AI's Invisible Workforce: Millions of Annotators Power LLMs, Make Zero Headlines
A Lobsters thread resurfaces a forgotten fact: millions of Global South data annotators power today's LLMs at under $2/hour. Why only now?
6d ago2 min read
LLMRLHF
AI Spouts 'Minted' and 'Escape Hatch': A Cure for Silicon Valley-Speak
Reddit's r/LocalLLaMA flagged AI quirks like 'minted' replacing 'created'. We investigate why LLMs learned to posture and share practical remedies.
Aug 222 min read
DeepSeekQwen
US-China LLMs Are Copying One Training Pipeline — Pretraining Isn't the Secret
US-China LLMs are converging on one training pipeline. Pretraining takes 90% of compute, but mid-training (5%) is what actually builds capability.
Aug 142 min read
PPORLHF
Why LLMs Obey Without Crashing: The PPO Algorithm Behind ChatGPT Explained
PPO is the core algorithm letting LLMs learn human preferences without crashing. Like a cautious coach limiting steps, it ensures safe AI deployment,
May 12 min read