jiwo-0.8b

Version 2 (2026-10-10). jiwo-0.8b v2 scores 29.58 on the public part of the Decision Index 0.3 (v1: 28.72). v1 is 1st of 27 models under 1B parameters on the official board, with a Full score of 24.32. With JIWO_CUDA_GRAPHS=1, the median request takes 13.1 ms on one RTX PRO 6000.

jiwo-0.8b is a decision model. It reads a state and one or more typed questions. For each question, it returns a probability for every option in one forward pass. It does not generate text.

The request and response format is the format of TypeSafe's Jev API (POST /v1/systemone). jiwo is an independent project. TypeSafe does not endorse it.

Base model Qwen/Qwen3.5-0.8B
Parameters 0.75B
Weights Full weights: decoder, tokenizer and readout
Question types choice (at most 255 options), score (2 to 10 levels), noul (true or false)
Training context length 8,192 tokens
Licence Apache-2.0

Code, server and documentation: github.com/jiwidi/jiwo. The other model of the family is eljiwo/jiwo-4b.

Versions

Version Date What changed Decision Index 0.3 public Board Full score Load it with
v2 (main) 2026-10-10 Same recipe, a larger training mix that adds many new domains 29.58 pending revision="v2"
v1 2026-10-05 First release 28.72 24.32, 1st of 27 under 1B revision="v1"
DecisionModel.from_pretrained("eljiwo/jiwo-0.8b", revision="v1")  # the first release

Evaluation (v2)

Benchmark v2 v1
Decision Index 0.3, public index 29.58 28.72 (board)
Decision Index 0.2.1 28.72 28.00
Area skill ร—100 (0.3): knowledge / language / retrieval and classification / tools / arts 10.9 / 36.1 / 44.2 / 40.3 / 12.2 10.3 / 32.3 / 44.5 / 38.0 / 11.2 (0.2.1)
Three public datasets never used in training (accuracy) 61.3% 54.6%
Seven more public datasets from new domains, never used in training (accuracy) 59.6%
Untrained Qwen3.5-0.8B on Decision Index 0.2.1, same server and calibration 6.99: +21.7 points for v2
Median request latency on one RTX PRO 6000, JIWO_CUDA_GRAPHS=1 (879-request sample) 13.1 ms (26.6 ms without graphs)

The scores come from complete runs of the public suite (150,317 requests for 0.2.1, 140,178 for 0.3), served by jiwo serve with JIWO_MAX_LENGTH=65536. On the board, the public suite is 20% of the Full score. The maintainers run the private tests (80%) after review. The v1 board scores come from the board of 2026-10-07.

Use the model

pip install "jiwo @ git+https://github.com/jiwidi/jiwo"    # on NVIDIA GPUs: "jiwo[cuda] @ git+..."
from jiwo.model import DecisionModel

model = DecisionModel.from_pretrained("eljiwo/jiwo-0.8b")
question = {"type": "noul", "instructions": "The customer is angry."}
response = model.decide("Refund please, the parcel arrived crushed.", {"angry": question})
print(response["answers"]["angry"]["noul"])  # the probability that the statement is true

Or serve it over HTTP. The server listens on 127.0.0.1:8765:

JIWO_CHECKPOINT=eljiwo/jiwo-0.8b JIWO_CUDA_GRAPHS=1 jiwo serve    # on a CPU or a Mac, leave out JIWO_CUDA_GRAPHS
curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d '{
  "state": "Refund please, the parcel arrived crushed.",
  "questions": {
    "team":  {"type": "choice", "instructions": "Which team handles this?",
              "criteria": {"billing": "Payments and refunds", "shipping": "Damaged or lost parcels", "account": null}},
    "angry": {"type": "noul", "instructions": "The customer is angry."},
    "urgency": {"type": "score", "instructions": "How urgent is it?", "criteria": ["Not urgent", "Soon", "Now"]}}}'

JIWO_CUDA_GRAPHS=1 captures CUDA graphs at start-up (about 2 minutes) and replays them for each request. A small model spends most of an eager request on kernel launches, so the graphs make it 2.0 times faster on a GPU. 99.7% of the fields pick the same option as the eager server, and the others are near ties. The jiwo README describes the response, the errors and all settings.

Request format

A request has a state (a string, a JSON object or a JSON array) and at most 64 named questions.

Type Meaning criteria
choice Pick one of N named options An object that maps option keys to descriptions or null
score Rate the state on 2 to 10 ordered levels A list of level descriptions, lowest level first
noul Decide whether a statement is true Optional. Descriptions for the keys true and false

How the model works

  • The decoder of the base model reads one prompt for each (state, question) pair.
  • A linear readout with one row per answer code (A, B, ..., Z, AA, ...) gives one logit for each option.
  • A softmax with a fitted temperature for each question type gives the probabilities.

Training data

The model was fine-tuned on the train splits of public datasets and on synthetic data. The training data is mostly English, with some data in more than 20 other languages.

Limits

  • The training prompts had at most 8,192 tokens. By default, the server refuses a longer prompt with a 400 error. You can set a larger JIWO_MAX_LENGTH, but the model did not see longer prompts in training.
  • A choice question can have at most 255 options.
  • The model was trained mostly on English. Other languages are evaluated only on a few datasets.
  • The model gives probabilities, not explanations. Check its answers before you use them for decisions with a high cost.

Licence

Apache-2.0, the licence of the base model. See LICENSE.

Downloads last month
159
Safetensors
Model size
0.8B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for eljiwo/jiwo-0.8b

Finetuned
(494)
this model
Quantizations
1 model

Spaces using eljiwo/jiwo-0.8b 2