deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B
Parameters1.8B
Context128K
Input → outputtext → text
Licensemit
PrecisionBF16
ArchitectureQwen2ForCausalLM
ReleasedJan 20, 2025
Downloads1.2M
Other sizes and formats in the family
| Model | Params | Precision | Downloads |
|---|---|---|---|
| deepseek-ai/DeepSeek-R1-Zerofp8 | 685B | F8_E4M3 | 8K |
| deepseek-ai/DeepSeek-R1fp8 | 684B | F8_E4M3 | 1.1M |
| deepseek-ai/DeepSeek-R1-Distill-Llama-70B | 70.6B | BF16 | 56.2K |
| deepseek-ai/DeepSeek-R1-Distill-Qwen-32B | 32.8B | BF16 | 409K |
| deepseek-ai/DeepSeek-R1-Distill-Qwen-14B | 14.8B | BF16 | 328.5K |
| deepseek-ai/DeepSeek-R1-Distill-Llama-8B | 8B | BF16 | 164.2K |
| deepseek-ai/DeepSeek-R1-Distill-Qwen-7B | 7.6B | BF16 | 322.6K |
This release in the news
No coverage of this exact release yet.
Recent news about Qwen
- Oct 11, 2026 · Opinion / analysisReverse Engineering w/ Local?
- Oct 11, 2026 · Opinion / analysisUPDATE: Qwen 3.8 27B 140 tok/s on single RTX 3090 Megakernel: KL divergence 0.0009 vs llama.cpp
- Oct 11, 2026 · Opinion / analysis[P] Pecision models that score every allowed label from the logits: Jebadiah v2.1 (27B, 9B), open weights and self-run benchmark results [P]
- Oct 11, 2026 · Opinion / analysisPSA: DeepSeek V4.1 Flash habitually exfiltrates API keys. It is dangerously misaligned and may be hazardous to use
- Oct 11, 2026 · Opinion / analysisQwen 3.8 27B Q5 vs Qwen 3.8 Next Q3_S for document analysis
- Oct 11, 2026 · Open-source releaseConverting dense models into Mixture-of-Experts
- Oct 11, 2026 · Tutorial / explainerRunning Next Flash IQ3_XXS at ~70 tok/s with 100k context or 2 instances of Qwen 3.6 35B A3B IQ4 at ~145 tok/s with 256k all on $500 of ex mining BC-250 boards
- Oct 11, 2026 · Opinion / analysisQwen3.8 Flash Next fixed my GNOME extension