Ollama announced on September 29, 2026 that it now supports decision models based on TypeSafe's Jev API[1]. You send a piece of text along with a set of named questions, and a model running on your own machine answers all of them in a single request. Ollama says there are no additional costs and that running locally lowers latency. The feature requires Ollama 0.35 or later.

A New API for Decisions, Not Text Generation

These models do not generate long passages. They answer named questions about a given state. Ollama points to tasks that need quick decisions, such as ticket triage, model routing, and content or safety moderation[1].

To use it, send text to the new /v1/systemone endpoint added in Ollama 0.35. You can include several questions at once and receive every answer in one response[1].

91 ms per Decision When Run Locally

Because requests do not travel over a network, latency stays low. According to Ollama, the 9B model nimble averaged 91 ms per decision in a Pac-Man demo running locally on an M5 Max, which the company says is fast enough for tasks like playing a game or processing content in real time[1].

Three Models Available Today

The following models are available now[1].

  • nimble: an open-source 9B-parameter decision model developed by Bespoke Labs
  • tev1: an experimental 4B decision model from Together AI
  • tev1:0.8b: an experimental 0.8B decision model from Together AI

Ollama says more decision models are coming, including models served by Ollama's cloud[1].

How to Get Started

First update Ollama to the latest version, then pull a decision model such as nimble.

ollama pull nimble

You can send requests with curl or through TypeSafe's official Python SDK. Install the SDK with uv or pip and set environment variables such as TYPESAFE_BASE_URL[1].

uv add typesafe-sdk # or: pip install typesafe-sdk
export TYPESAFE_BASE_URL=http://localhost:11434
export TYPESAFE_API_KEY=ollama
export TYPESAFE_DEFAULT_MODEL=nimble

In Ollama's example, a ticket reading "I was charged twice. Please refund the extra payment." is paired with three questions: which team should handle it (choice), whether the customer explicitly asks for a refund (noul), and how urgent it is (score). The answers come back with probabilities and confidence values. Here the team is judged to be billing, and the refund request scores 0.997[1].

What Comes Next

Ollama calls this the first of many releases adding decision model support. Planned updates include faster performance on Apple Silicon powered by MLX and more models specializing in different kinds of decision making[1].

Summary

Instead of asking an LLM to write text, Ollama's decision models make it easy to run classification and routing-style judgments quickly and cheaply on local hardware. For developers who want to keep ticket triage or moderation pre-processing on their own machines, there is now a simple option to try.

出典:https://ollama.com/blog/ollama-now-supports-jev-style-decision-models