The Ollama game logo

Ollama

PatchBot will keep your Discord channel up-to-date on all the latest Ollama patch notes.

Servers
756
Gamers
150,627

Game details

  • Description This AI tool simplifies the process of running machine learning models locally, enabling developers to easily deploy and interact with various models without extensive setup. Its user-friendly interface streamlines the integration of AI into applications.
  • Release Date Jul 8, 2023

Latest updates

v0.35.1

Clef decision models Ollama now supports Clef and Clef Flash, Cloudflare's new open-source decision models, through /v1/systemone. Clef (27B) and Clef Flash (9B) are multimodal: requests can now include images alongside the text state, shared by all questions and scored jointly with it.


curl http://localhost:11434/v1/systemone -d '{
  "model": "clef-flash",
  "state": "The user took this screenshot.",
...
Git tag

v0.35.1

v0.35.0

Decision models Ollama now supports decision models through /v1/systemone, based on TypeSafe’s Jev API. Decision models return choices, probabilities, and scores instead of text. Use them for tasks such as ticket triage, model routing, and content classification. Available models:

  • Nimble from Bespoke Labs
  • Tev1 from Together AI
    ollama pull nimble

    ...

Git tag

v0.35.0

v0.34.4

What's Changed

  • Structured outputs on thinking models now apply in a single pass, making them faster and more reliable.
  • Fixed intermittent "model not found" errors with a large local library
  • Fixed the macOS app becoming unresponsive when checking if ChatGPT or Codex is running.
  • Qwen 3.8 prompt processing is faster on Apple Silicon.
  • Gemma 4 on Apple Silicon now picks the best image resolution per image, keeping more detail in high-resolution images. ...
Git tag

v0.34.4

v0.34.3

What's Changed GET /api/show now advertises each model's thinking controls and default: Available in the CLI with:

ollama show gemma4
    thinking
        levels     false, true
        default    true

Available in the API with:

curl http://localhost:11434/api/show -d '{"model": "glm-5.3-flash:cloud"}'
{
  "thinking": {
    "values": ["low", "high", "max"],
    "default": "max"
  }
}

Also available on ollama.com directly for cloud models. ...

Git tag

v0.34.3

v0.34.2

What's Changed

  • Added first-run setup when running ollama, with options to sign in or continue locally. Setup completion is shared with the desktop app on macOS and Windows.
  • Added ollama://apps to open the desktop app’s Apps page directly on macOS and Windows.
  • Fixed excessive memory growth during long generations with MLX speculative decoding.
  • Updated llama.cpp.
Git tag

v0.34.2

v0.34.1

What's Changed

  • MLX safetensors ollama create no longer experimental. GGUF model creation now requires using llama.cpp tooling for safetensor conversion and quantization.
  • Improved MLX memory handling on Apple Silicon
  • Runaway repeat token detection now requires 100 repeat tokens for reduced false positives (e.g. OCR)
  • /api/tags is much faster on large model libraries (3.1 s → 294 ms cold in testing), and model capabilities are now reported consistently. ...
Git tag

v0.34.1