Published event
ArtificialIntelligence
ModelRelease
1 source(s)
Release v5.9.0
Summary
Release v5.9.0 huggingface / transformers Public Notifications You must be signed in to change notification settings Fork 34.7k Star 167k Release v5.9.0 Cyrilvallez released this 20 May 14:12 · 1222 commits to main since this release v5.9.0 0a2757d This commit was signed with the committer’s verified signature . Cyrilvallez Cyril Vallez SSH Key Fingerprint: OjK+mdCRyLrQ4vIh2+8FffCzuDZs1WByt+LD1Z6WNw4 Verified Learn about vigilant mode .
Why it matters
This ModelRelease is relevant to the technology intelligence record because it involves Cohere, GitHub, Qwen. The source article should remain the factual reference for follow-up coverage.
Key facts
- huggingface / transformers Public Notifications You must be signed in to change notification settings Fork 34.7k Star 167k Release v5.9.0 Cyrilvallez released this 20 May 14:12 · 1222 commits to main since this release v5.9.0 0a2757d This commit was signed with the committer’s verified signature .
- Cyrilvallez Cyril Vallez SSH Key Fingerprint: OjK+mdCRyLrQ4vIh2+8FffCzuDZs1WByt+LD1Z6WNw4 Verified Learn about vigilant mode .
- Release v5.9.0 New Model additions Cohere2Moe Command A+ is a Mixture-of-Experts (MoE) language model from Cohere that features a hybrid attention pattern combining sliding window and full attention layers.
- The model incorporates both shared and routed experts and supports a very large context window for processing extensive text sequences.
- Links: Documentation Add new cohere2_moe model ( #46115 ) by @Cyrilvallez in #46115 Parakeet tdt ( #44171 ) Parakeet tdt ( #44171 ) by @lmaksym HRM-Text HRM-Text is an improved autoregressive language-modeling variant of the Hierarchical Reasoning Model (HRM) that uses a hierarchical recurrent forward pass with two transformer stacks - one for slow, abstract planning (H) and one for fast, detailed computation (L) - reused inside a nested recurrence.
- It features PrefixLM attention where instruction tokens attend bidirectionally while response tokens attend causally, per-head sigmoid output gates, and parameterless RMSNorm.
Entities in this story
AI models
qwen→Related events