Published event
ArtificialIntelligence Research 1 source(s)

Apriel-H1: The Surprising Key to Distilling Efficient Reasoning Models

Updated September 26, 2026 · 2:45 PM · source date November 19, 2025

Summary

Apriel-H1: The Surprising Key to Distilling Efficient Reasoning Models Apriel-H1: The Surprising Key to Distilling Efficient Reasoning Models Enterprise Article Published November 19, 2025 Upvote 35 Torsten Scholak tscholak ServiceNow-AI Oleksiy Ostapenko ostapeno ServiceNow-AI Raymond Li RaymondLi ServiceNow-AI Luke Kumar nitsanluke ServiceNow-AI Joel Lamy-Poirier jlamypoirier ServiceNow-AI We converted our 15B reasoning model to a Mamba hybrid achieving 2.1x throughput with minimal quality loss. A non-obvious insight about what data to distill on, and why intuition fails here.

Why it matters

This Research is relevant to the technology intelligence record because it involves NVIDIA, Mistral AI, Apple, Hugging Face. The source article should remain the factual reference for follow-up coverage.

Key facts
  • Apriel-H1: The Surprising Key to Distilling Efficient Reasoning Models Enterprise Article Published November 19, 2025 Upvote 35 Torsten Scholak tscholak ServiceNow-AI Oleksiy Ostapenko ostapeno ServiceNow-AI Raymond Li RaymondLi ServiceNow-AI Luke Kumar nitsanluke ServiceNow-AI Joel Lamy-Poirier jlamypoirier ServiceNow-AI We converted our 15B reasoning model to a Mamba hybrid achieving 2.1x throughput with minimal quality loss.
  • A non-obvious insight about what data to distill on, and why intuition fails here.
  • When MiniMax published their M2 post-mortem in October explaining why they abandoned efficient attention at 230B scale, the narrative briefly became "efficient attention is dead." Within days, Kimi Linear proved otherwise.
  • The real lesson: it depends on your constraints.
  • Our constraint was simple: we had a strong 15B reasoning model and needed to make it efficient without starting over.
  • No infinite compute for 20T-token pretraining.
Entities in this story
Related events