DeepSeek unveiled DeepSeek-R1 and DeepSeek-R1-Zero, fully open-weights reasoning models that match OpenAI's o1 performance on AIME, MATH-500, and Codeforces benchmarks using large-scale reinforcement learning without prior supervised fine-tuning.
Democratizes state-of-the-art chain-of-thought test-time compute reasoning for researchers, enterprises, and local hardware builders worldwide with MIT licensing.