jiwo-0.8b
Version 2 (2026-10-10). jiwo-0.8b v2 scores 29.58 on the public part of the
Decision Index 0.3 (v1: 28.72). v1 is 1st of 27
models under 1B parameters on the official board, with a Full score of 24.32. With JIWO_CUDA_GRAPHS=1, the median
request takes 13.1 ms on one RTX PRO 6000.
jiwo-0.8b is a decision model. It reads a state and one or more typed questions. For each question, it returns a probability for every option in one forward pass. It does not generate text.
The request and response format is the format of TypeSafe's Jev API (POST /v1/systemone). jiwo is an independent
project. TypeSafe does not endorse it.
| Base model | Qwen/Qwen3.5-0.8B |
| Parameters | 0.75B |
| Weights | Full weights: decoder, tokenizer and readout |
| Question types | choice (at most 255 options), score (2 to 10 levels), noul (true or false) |
| Training context length | 8,192 tokens |
| Licence | Apache-2.0 |
Code, server and documentation: github.com/jiwidi/jiwo. The other model of the family is eljiwo/jiwo-4b.
Versions
| Version | Date | What changed | Decision Index 0.3 public | Board Full score | Load it with |
|---|---|---|---|---|---|
v2 (main) |
2026-10-10 | Same recipe, a larger training mix that adds many new domains | 29.58 | pending | revision="v2" |
| v1 | 2026-10-05 | First release | 28.72 | 24.32, 1st of 27 under 1B | revision="v1" |
DecisionModel.from_pretrained("eljiwo/jiwo-0.8b", revision="v1") # the first release
Evaluation (v2)
| Benchmark | v2 | v1 |
|---|---|---|
| Decision Index 0.3, public index | 29.58 | 28.72 (board) |
| Decision Index 0.2.1 | 28.72 | 28.00 |
| Area skill ร100 (0.3): knowledge / language / retrieval and classification / tools / arts | 10.9 / 36.1 / 44.2 / 40.3 / 12.2 | 10.3 / 32.3 / 44.5 / 38.0 / 11.2 (0.2.1) |
| Three public datasets never used in training (accuracy) | 61.3% | 54.6% |
| Seven more public datasets from new domains, never used in training (accuracy) | 59.6% | |
| Untrained Qwen3.5-0.8B on Decision Index 0.2.1, same server and calibration | 6.99: +21.7 points for v2 | |
Median request latency on one RTX PRO 6000, JIWO_CUDA_GRAPHS=1 (879-request sample) |
13.1 ms (26.6 ms without graphs) |
The scores come from complete runs of the public suite (150,317 requests for 0.2.1, 140,178 for 0.3), served by
jiwo serve with JIWO_MAX_LENGTH=65536. On the board, the public suite is 20% of the Full score. The maintainers
run the private tests (80%) after review. The v1 board scores come from the board of 2026-10-07.
Use the model
pip install "jiwo @ git+https://github.com/jiwidi/jiwo" # on NVIDIA GPUs: "jiwo[cuda] @ git+..."
from jiwo.model import DecisionModel
model = DecisionModel.from_pretrained("eljiwo/jiwo-0.8b")
question = {"type": "noul", "instructions": "The customer is angry."}
response = model.decide("Refund please, the parcel arrived crushed.", {"angry": question})
print(response["answers"]["angry"]["noul"]) # the probability that the statement is true
Or serve it over HTTP. The server listens on 127.0.0.1:8765:
JIWO_CHECKPOINT=eljiwo/jiwo-0.8b JIWO_CUDA_GRAPHS=1 jiwo serve # on a CPU or a Mac, leave out JIWO_CUDA_GRAPHS
curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d '{
"state": "Refund please, the parcel arrived crushed.",
"questions": {
"team": {"type": "choice", "instructions": "Which team handles this?",
"criteria": {"billing": "Payments and refunds", "shipping": "Damaged or lost parcels", "account": null}},
"angry": {"type": "noul", "instructions": "The customer is angry."},
"urgency": {"type": "score", "instructions": "How urgent is it?", "criteria": ["Not urgent", "Soon", "Now"]}}}'
JIWO_CUDA_GRAPHS=1 captures CUDA graphs at start-up (about 2 minutes) and replays them for each request. A small
model spends most of an eager request on kernel launches, so the graphs make it 2.0 times faster on a GPU. 99.7% of
the fields pick the same option as the eager server, and the others are near ties. The jiwo README describes the
response, the errors and all settings.
Request format
A request has a state (a string, a JSON object or a JSON array) and at most 64 named questions.
| Type | Meaning | criteria |
|---|---|---|
choice |
Pick one of N named options | An object that maps option keys to descriptions or null |
score |
Rate the state on 2 to 10 ordered levels | A list of level descriptions, lowest level first |
noul |
Decide whether a statement is true | Optional. Descriptions for the keys true and false |
How the model works
- The decoder of the base model reads one prompt for each (state, question) pair.
- A linear readout with one row per answer code (A, B, ..., Z, AA, ...) gives one logit for each option.
- A softmax with a fitted temperature for each question type gives the probabilities.
Training data
The model was fine-tuned on the train splits of public datasets and on synthetic data. The training data is mostly English, with some data in more than 20 other languages.
Limits
- The training prompts had at most 8,192 tokens. By default, the server refuses a longer prompt with a 400 error.
You can set a larger
JIWO_MAX_LENGTH, but the model did not see longer prompts in training. - A
choicequestion can have at most 255 options. - The model was trained mostly on English. Other languages are evaluated only on a few datasets.
- The model gives probabilities, not explanations. Check its answers before you use them for decisions with a high cost.
Licence
Apache-2.0, the licence of the base model. See LICENSE.
- Downloads last month
- 159