ollama/ollama v0.32.15
New desktop onboarding flow on first launch
Key points
- Caches resolved model metadata between requests, cutting time-to-first-token by roughly half (TTFT dropped from ~995 ms to ~524 ms in benchmarks)
- Fixes a bug where chat and generate could wedge after a mid-stream parser error
- Qwen 3.8 system messages are now normalized so non-leading system messages are handled consistently
Sources (1)
- [1]ollama/ollama v0.32.15GitHub: ollama/ollama · Aug 19, 05:25 PM
* New desktop onboarding flow on first launch
* Caches resolved model metadata between requests, cutting time-to-first-token by roughly half (TTFT dropped from ~995 ms to ~524 ms in benchmarks)
Extractive summary: sentences quoted from the sources.