← All tools

Local / Open Source

Ollama

Updates

  1. Feature3d ago· GitHub Releases

    v0.35.1

    Ollama v0.35.1 allows up to 10 web searches per response and adds explicit model capability declarations during model creation. It also updates its MLX and llama.cpp dependencies.

  2. Feature4d ago· GitHub Releases

    v0.35.0

    Ollama 0.35.0 adds decision models through the /v1/systemone endpoint, returning choices, probabilities, and scores instead of generated text. Initial supported models are Bespoke Labs’ Nimble and Together AI’s Tev1.

  3. Feature7d ago· GitHub Releases

    v0.40.0

    Ollama v0.40.0 enables its MLX runner by default for supported model architectures on Apple Silicon. Additional models will be tested and enabled during the release candidate period.

  4. Feature9d ago· GitHub Releases

    v0.34.4

    Ollama v0.34.4 fixes intermittent model-not-found errors and improves structured outputs for reasoning models. It also updates llama.cpp and MLX, adds dynamic Gemma 4 image resolution, and accelerates Qwen 3.8 prompt processing on MLX.

  5. Feature13d ago· GitHub Releases

    v0.34.3

    Ollama v0.34.3 adds model thinking controls and defaults to the GET /api/show response. It also supports Nemotron H vision models on Apple Silicon through MLX and fixes unwanted window reopening in the macOS app.

  6. Feature17d ago· GitHub Releases

    v0.34.2

    Ollama 0.34.2 updates its llama.cpp dependency. The release notes provide no additional details.

  7. Feature18d ago· GitHub Releases

    v0.34.1

    Ollama v0.34.1 improves MLX model memory management, cache eviction, and model-loading behavior. It also updates llama.cpp and MLX, raises the token repeat limit, fixes incomplete-result handling, and refreshes parts of the app interface.

  8. Integration27d ago· GitHub Releases

    v0.34.0

    Ollama v0.34.0 lets macOS users run Ollama-hosted open models directly in ChatGPT Desktop. It also improves structured outputs on Apple Silicon and adds OpenAI-compatible tool search, response compaction, and image handling in compacted responses.

  9. FeatureSep 2· GitHub Releases

    v0.33.3

    Ollama v0.33.3 adds reporting for cached prompt tokens and honors default parameters defined in GGUF models. It also updates MLX, MLX-C, and llama.cpp dependencies.

  10. FeatureAug 27· GitHub Releases

    v0.33.2

    Ollama 0.33.2 restores system dark mode, keeps proxy requests running when the model catalog changes, and synchronizes macOS app handoff.

  11. FeatureAug 26· GitHub Releases

    v0.33.1

    Ollama 0.33.1 adds Qwen3.8 Flash Next and structured-output support to its MLX runner. It also updates MLX and llama.cpp and mitigates Metal GPU timeouts when loading models from slow storage.

  12. IntegrationAug 21· GitHub Releases

    v0.33.0

    Ollama 0.33.0 adds Claude Desktop integration for selecting local models and managing app connections. It also fixes cancelled or resumed prefill caching, Claude Code prompt behavior, and DeepSeek Harness startup fallback.

  13. FeatureAug 19· GitHub Releases

    v0.32.15

    Ollama v0.32.15 adds a model metadata cache to reduce per-request overhead.

  14. FeatureAug 15· GitHub Releases

    v0.32.14

    Ollama v0.32.14 adds WebP image transcoding for llama-server and updates Qwen renderers to accept system messages that are not first in the conversation.

  15. FeatureAug 14· GitHub Releases

    v0.32.13

    Ollama v0.32.13 adds support for developer instructions when running Qwen3.8 models.

  16. FeatureAug 14· GitHub Releases

    v0.32.12

    Ollama v0.32.12 adds support for the Qwen 3.8 27B model. The release includes Apple Silicon optimizations aimed at repeated tasks and coding-agent workloads.

  17. FeatureAug 14· GitHub Releases

    v0.32.11

    Ollama 0.32.11 adds launch support for DeepSeek Harness and Meta's Muse Code agentic coding CLI. Its OpenAI-compatible Responses API now supports web search.

  18. FeatureAug 12· GitHub Releases

    v0.32.10

    Ollama 0.32.10 changes the default repeat penalty from 1.1 to 1.0, improving speculative decoding compatibility and speed. It also accelerates NVFP4 MLX prefill and fixes OCI blob verification.

  19. FeatureAug 11· GitHub Releases

    v0.32.9

    Ollama 0.32.9 adds the Nemotron 3 architecture and support for NVIDIA’s open Nemotron 3.5 Lightning MoE model. It also fixes a boundary condition in Muse Glimmer function-call parsing.

  20. NewsAug 11· Changelog

    Ollama changelog updated

    :root { --tab-size-preference: 4; } pre, code { tab-size: var(--tab-size-preference); } Releases · ollama/ollama · GitHub Skip to content Navigation Menu Sign in Appearance settings Platform AI CODE CREATION GitHub Copilot Write better code with AI GitHub Copilot app Direct agents from issue to merge MCP Registry Integ

  21. NewsAug 10· GitHub Releases

    v0.32.8

    ## What's Changed * Add Muse Glimmer support for NVIDIA, AMD, and additional platforms **Full Changelog**: https://github.com/ollama/ollama/compare/v0.32.7...v0.32.8-rc0

  22. NewsAug 10· GitHub Releases

    v0.32.7

    ## Muse Glimmer > Note: Muse Glimmer is currently available via initial support via Ollama's MLX engine on Apple Silicon. Additional support and optimizations for Apple Silicon, NVIDIA, AMD, and other platforms will be available in the coming days. **Muse Glimmer**, Meta's newest open model and the first released

  23. NewsAug 4· GitHub Releases

    v0.32.6

    ## What's Changed - Qwen3.5 is faster on Apple GPUs: the MLX engine now uses the model's MTP head for speculative decoding automatically - `/v1/chat/completions` streaming now matches OpenAI's wire format: `role` only on the first chunk, `finish_reason` on its own chunk, and usage in a separate chunk with `str