ollama/ollama v0.34.1
GGUF model creation now requires using llama.cpp tooling for safetensor conversion and quantization.
Key points
- MLX safetensors ollama create no longer experimental.
- Runaway repeat token detection now requires 100 repeat tokens for reduced false positives (e.g. OCR)
- /api/tags is much faster on large model libraries (3.1 s → 294 ms cold in testing), and model capabilities are now reported consistently.
- Deprecated typicalp: it can no longer be set when creating new models, existing GGUF models retain support.
Sources (1)
- [1]ollama/ollama v0.34.1GitHub: ollama/ollama · Sep 14, 10:14 PM
GGUF model creation now requires using llama.cpp tooling for safetensor conversion and quantization.
* MLX safetensors `ollama create` no longer experimental.
Extractive summary: sentences quoted from the sources.