DescriptionThis AI tool simplifies the process of running machine learning models locally, enabling developers to easily deploy and interact with various models without extensive setup. Its user-friendly interface streamlines the integration of AI into applications.
Clef decision models
Ollama now supports Clef and Clef Flash, Cloudflare's new open-source decision models, through /v1/systemone.
Clef (27B) and Clef Flash (9B) are multimodal: requests can now include images alongside the text state, shared by all questions and scored jointly with it.
curl http://localhost:11434/v1/systemone -d '{
"model": "clef-flash",
"state": "The user took this screenshot.",
...
Decision models
Ollama now supports decision models through /v1/systemone, based on TypeSafe’s Jev API.
Decision models return choices, probabilities, and scores instead of text. Use them for tasks such as ticket triage, model routing, and content classification.
Available models:
Added first-run setup when running ollama, with options to sign in or continue locally. Setup completion is shared with the desktop app on macOS and Windows.
Added ollama://apps to open the desktop app’s Apps page directly on macOS and Windows.
Fixed excessive memory growth during long generations with MLX speculative decoding.
MLX safetensors ollama create no longer experimental. GGUF model creation now requires using llama.cpp tooling for safetensor conversion and quantization.
Improved MLX memory handling on Apple Silicon
Runaway repeat token detection now requires 100 repeat tokens for reduced false positives (e.g. OCR)
/api/tags is much faster on large model libraries (3.1 s → 294 ms cold in testing), and model capabilities are now reported consistently.
...
What's ChangedModels run on MLX on Apple Silicon by default
In this release, on Apple Silicon devices, model architectures supported by the MLX runtime will automatically run on MLX.
What's ChangedModels run on MLX on Apple Silicon by default
In this release, on Apple Silicon devices, model architectures supported by the MLX runtime will automatically run on MLX.
Decision models
Ollama now supports decision models through /v1/systemone, based on TypeSafe’s Jev API.
Decision models return choices, probabilities, and scores instead of text. Use them for tasks such as ticket triage, model routing, and content classification.
Available models: