Published event
ArtificialIntelligence
ProductUpdate
1 source(s)
v0.31.1
Summary
v0.31.1 ollama / ollama Public Notifications You must be signed in to change notification settings Fork 18k Star 182k v0.31.1 github-actions released this 30 Jun 22:10 · 296 commits to main since this release v0.31.1 710292f This commit was created on GitHub.com and signed with GitHub’s verified signature . GPG key ID: B5690EEEBB952194 Verified Learn about vigilant mode .
Why it matters
This ProductUpdate is relevant to the technology intelligence record because it involves Apple, GitHub, llama. The source article should remain the factual reference for follow-up coverage.
Key facts
- ollama / ollama Public Notifications You must be signed in to change notification settings Fork 18k Star 182k v0.31.1 github-actions released this 30 Jun 22:10 · 296 commits to main since this release v0.31.1 710292f This commit was created on GitHub.com and signed with GitHub’s verified signature .
- GPG key ID: B5690EEEBB952194 Verified Learn about vigilant mode .
- Faster Gemma 4 on Apple Silicon Gemma 4 is now significantly faster in Ollama on Apple Silicon, generating tokens nearly 90% faster on average across a coding-agent benchmark by leveraging multi-token prediction (MTP).
- Ollama auto-tunes how many tokens to draft as it runs, so the speedup is on by default, requires no configuration, and does not change the model's output.
- What's Changed Tightened Gemma 4 MoE model loading in the MLX engine Updated the MLX engine to the latest version, including a new small-batch matmul kernel Updated the underlying llama.cpp engine to build 9840 Improved Gemma 4 multi-token prediction (MTP) performance Full Changelog : v0.30.12...v0.31.1 Assets 19 Loading Uh oh!
- There was an error while loading.
Entities in this story
Related events