Model setup recipes
Run Jev in Ollama: get classifications and probabilities instead of prose
Instead of asking the model to rewrite an answer for every inquiry, first classify it using criteria you have already defined.
To automate choosing the team responsible for an inquiry, you can use a classification path that selects a candidate rather than asking a chat model to write an answer. Available from Ollama 0.35.0, /v1/systemone combines choice, noul and score questions with shared input called state, returning structured answers and candidate probabilities. This guide downloads Nimble, checks a call routing one inquiry, and explains how to evaluate its output without treating probability as an accuracy guarantee.
When you want to choose the team before drafting a reply
When refund questions and payment errors enter one queue, someone has to read each message and choose its team. Asking a chat model to explain which team should handle it produces natural prose, but the automation then has to interpret that prose again. Writing an explanation and selecting one predefined candidate are different jobs.
System One is a request format for the latter. Send the inquiry as shared state, then define candidates and decision instructions for each question. Ollama added `/v1/systemone` in 0.35.0, released September 28, 2026. Check the version rather than expecting the same address to work in 0.34.x or earlier. This example classifies one payment-queue inquiry as billing or technical.
Before starting, check that Ollama is installed and running, with enough disk space and memory for the model. Bespoke Labs lists Nimble 9B's original BF161 weights at about 18GB, but the Ollama nimble tag may use a different format or quantization2. Read its downloaded size in the SIZE column of `ollama list`; that size is not running memory usage. Ollama's 0.35.0 release does not guarantee minimum memory or endpoint latency for a particular Mac or GPU3.
ollama --version
ollama pull nimble
ollama list
Choose between two candidates for one inquiry
Fixing the input and question makes failures easier to isolate. The request below classifies a checkout 500 error that began this morning as either a technical or billing issue. Candidate descriptions tell the model what to distinguish, while instructions specify the decision for this inquiry. Changing a description can change the classification boundary, so use wording consistent with your operating rules.
A choice question accepts 2 to 26 candidates. Its result contains the key with the highest probability and probabilities for every candidate. curl sends one request and prints the completed JSON. This API4 does not stream, so do not expect tokens5 to appear one by one as in a chat response.
Find model, answers.route.choice, answers.route.probabilities, answers.route.confidence and usage in the response. choice is the highest-probability key. Candidate probabilities are normalized within your supplied candidates and sum to 1; this does not recover a missing candidate. With only billing and technical, even a greeting or account-takeover report must fall into one of them. Design separate rules or candidates if your queue needs an exception route.
| Response field | Meaning | What to watch for |
|---|---|---|
| choice | Key of the highest-probability candidate | An answer absent from the candidate list cannot be selected. |
| probabilities | Normalized probability for each candidate key | Compare only within this request's candidate list. |
| confidence | How much the distribution concentrates on one side relative to a uniform distribution | It is not confidence calibrated to correctness. |
| usage | Input and output token usage reported by the server | Token counts alone do not establish latency or quality. |
Each response field supports a different decision.
choice
- Meaning
- Key of the highest-probability candidate
- What to watch for
- An answer absent from the candidate list cannot be selected.
probabilities
- Meaning
- Normalized probability for each candidate key
- What to watch for
- Compare only within this request's candidate list.
confidence
- Meaning
- How much the distribution concentrates on one side relative to a uniform distribution
- What to watch for
- It is not confidence calibrated to correctness.
usage
- Meaning
- Input and output token usage reported by the server
- What to watch for
- Token counts alone do not establish latency or quality.
curl -sS http://127.0.0.1:11434/v1/systemone -H 'Content-Type: application/json' -d '{"model":"nimble","state":"Our checkout has returned 500 errors since this morning.","questions":{"route":{"type":"choice","instructions":"Which team should handle this ticket?","criteria":{"billing":"Payments, charges, and refunds","technical":"Application errors, failed requests, and outages"}}}}'Selection, true/false and ordinal scores ask different questions
Use choice to select one team. A near yes/no decision such as whether an outage is urgent can use noul, whose default candidates are false and true. You can describe those candidates as strings in criteria if needed. The answer is the numerical probability of true, not a boolean, so your application must set a threshold to convert it into true or false.
For ordered priorities, use score and put criteria in ascending order. Its score is the sum of each zero-based candidate index multiplied by its probability, so it can be non-integer. For probabilities 0.2, 0.3 and 0.5 at levels 0, 1 and 2, the score is 0×0.2 + 1×0.3 + 2×0.5 = 1.3. That is not a fourth category; your service needs separate boundaries to assign a grade.
Multiple questions in one request each evaluate the same state independently. The first answer is not automatically passed into the next question. Bundling questions returns several fields in one network call, not a conversation or dependency chain. Before sending, give every question a unique nonempty key and check that candidate counts stay within 2–26.
{
"model": "nimble",
"state": "Customer reports repeated checkout failures and requests a refund.",
"questions": {
"route": {
"type": "choice",
"instructions": "Which team should handle this ticket?",
"criteria": {
"billing": "Payments and refunds",
"technical": "Application errors and outages"
}
},
"urgent": {
"type": "noul",
"instructions": "Is the checkout unavailable for multiple customers?",
"criteria": {
"false": "No evidence of a broad outage",
"true": "Multiple customers cannot complete checkout"
}
}
}
}
Do not turn high confidence into automatic approval
confidence can rise when probability concentrates on one candidate. It measures concentration based on entropy, not a calibrated probability of correctness. An unfamiliar inquiry can be confidently misclassified. Do not adopt a rule such as confidence above 0.9 means correct without validation data.
Before deployment, build a validation set of representative and exceptional past inquiries. Have people confirm each actual team, then rerun with the same state and candidate descriptions. Beyond overall accuracy, separately count billing inquiries sent to technical and technical failures buried in billing. The cost of each error determines automatic assignment, human-review intervals and retry routes.
With 100 validation inquiries, record both the correct count out of 100 and errors within categories, such as how many of 40 actual billing inquiries were misrouted. These are illustrative denominators, not model results. Published accuracy or latency needs your dataset, model tag and revision, hardware, Ollama version and repetitions. Ollama 0.35.0 does not supply a Nimble /v1/systemone performance comparison.
For unexpected output, first check the endpoint and Ollama version, then whether the model tag is a System One-compatible Nimble or Tev. The request type must be choice, noul or score, with supported candidate counts. Start with simple string candidates and criteria arrays. Do not assume Ollama 0.35.0 accepts every complex candidate-description format in the TypeSafe reference API; use JSON actually accepted by Ollama.

Keep the chat API when you need a written answer
System One selects and scores; it does not explain a decision to a customer or write refund instructions. If you need a reply after selecting a team, pass the classification into application logic and make a separate chat request. Keeping these roles separate in code makes candidate errors and writing-style problems easier to diagnose independently.
If existing rules already distinguish every type without someone reading the inquiry, adding a model may be unnecessary. Introduce it only if recurring variations defeat those rules and validation shows classification reduces manual work. A safe start is shadow mode: save the results and compare them with human decisions without automatically forwarding tickets.
For timing, fix the request and separate warm runs6 with the model loaded from cold loading. Wall-clock latency per inquiry matters directly for classification, but output tokens per second from ordinary chat cannot alone compare this endpoint. After warm-up, repeat the same samples at least three times and record median latency plus errors to evaluate model or candidate-description changes.
Official specifications and execution paths
The API paths and field meanings here follow Ollama 0.35.0, the Ollama Python client's System One documentation and the current Ollama OpenAPI definition. Supported models and candidate types can change; check the official documentation for your installed version before applying them.
Terminology notes
BF16 — A 16-bit floating-point format for storing and computing model values. Support depends on the hardware and runtime.
Back to the textQuantization — Representing model values with fewer bits. Memory use, accuracy, or execution speed may change; the effects depend on the format and implementation.
Back to the textGPU — A processor designed to handle many calculations in parallel. It performs model computations during AI inference.
Back to the textAPI — A defined interface that lets other code call a program’s functions. The term API alone does not imply sending data to an external server.
Back to the textToken — A unit into which a model divides input or output for processing. One token does not equal one character or a fixed duration.
Back to the textWarm run — A measurement made after model loading and initialization. It may exclude the loading wait from the first run.
Back to the text