Back to NewsSFT Memorizes, RL GeneralizesFebruary 12, 2025ResearchSFT Memorizes, RL Generalizes. DeepSeek has shown the power of Reinforcement Learning (RL) without Supervised Fine-Tuning (SFT). What does RL learn differently than SFT? Well, as the title, SFT memorizes, RL generalizes.PrevCerebras — Fastest Inference PlatformNextAs AIs Get Smarter, They Develop Coherent Value Systems