jev-omni.js weights
A browser-ready export of Jev-Omni, akhilaaa3's Gemma 4 12B decision classifier: a state, a question and 2-256 options in, one probability per option out, from a single forward pass. It runs in Chrome on WebGPU through jev-omni.js, for text and images. Live demo.
This repo only holds converted weights. The model, its training, its decision head and its prompt format are Jev-Omni's. Jev-Omni states that it is independent of TypeSafe AI's Jev, and so is this repo.
import * as ort from "onnxruntime-web/webgpu"; // 1.30.0: see Requirements
import { loadJevOmni } from "@ai-ecoverse/jev-omni.js";
const jev = await loadJevOmni("https://huggingface.co/ai-ecoverse/jev-omni.js/resolve/main/jev-omni", { ort });
const res = await jev.predict({
state: "The meeting starts at 10 AM. It is now 9 AM.",
question: "Has the meeting started?",
options: ["Yes", "No"],
});
res.probabilities; // { Yes: ..., No: ... }
Contents
| Path | What | Size |
|---|---|---|
jev-omni/manifest.json |
file list with sizes and SHA-256, graph inputs, parity numbers; revision 1544c2b7ed5277a0 |
|
jev-omni/r-c050d51/q8f32/ |
decoder: int8 weights (MatMulNBits, block 32), fp32 activations, fp16 attention in the 8 global layers; 437 files | 13.34 GB |
jev-omni/r-c050d51/vision/ |
Gemma 4's vision embedder and projection, fp32 | 0.20 GB |
jev-omni/r-c050d51/head.safetensors, tokenizer*.json |
the 256-way decision head, the tokenizer | 36 MB |
Total download: 13.58 GB in 445 files of at most 32 MB.
Provenance
- Model: akhilaaa3/Jev-Omni at
c050d51354147985d13286cf4acf90f562f2c631(Apache-2.0),unified/model.safetensors(bf16) andhead.pt. - Base: google/gemma-4-12B-it at
707f0a3b8a3c7ad586ed01e27eafbad8a27dd0f7(Apache-2.0). - Export: a vendored draft of the onnxruntime-genai model builder with Gemma 4 support
(microsoft/onnxruntime-genai#2473,
603c36c), without the LM head, then graph rewrites for WebGPU and the image path. Details are in the jev-omni.js docs.
Parity with the fp32 PyTorch model
Chrome on WebGPU (Apple M4 Max) against Jev-Omni in fp32 PyTorch:
| Set | Accuracy, browser / reference | Answers changed | Mean / max abs. probability difference |
|---|---|---|---|
| DecisionBench medium, 293 text questions | 86.35% / 86.35% | 0 | 0.0019 / 0.131 |
| kev.js vision-v1, 106 image questions | 98.11% / 99.06% | 1 | 0.0014 / 0.065 |
| kev.js vision-v2, 128 image questions | 85.16% / 85.16% | 1 | 0.0037 / 0.065 |
Without quantization, a 6-layer fp32 export matches PyTorch to a relative error of 5.0e-06 at the last position.
Requirements
- Memory: about 16 GB of GPU memory for the WebGPU session. While loading, the tab also holds the 13.6 GB of files in memory until they are on the GPU. On Apple Silicon that means 32 GB of unified memory at the very least, and 64 GB or more to be comfortable. It has been tested on an Apple M4 Max with 128 GB only.
- Download: 13.6 GB on the first load. jev-omni.js keeps it in Cache Storage, so later loads read from disk.
- onnxruntime-web
1.30.0: its attention needs an n×n fp32 buffer per head, so prompts are capped at about 8,000 tokens on GPUs with 4 GB buffers (Apple Silicon) and fewer on GPUs with smaller ones. jev-omni.js rejects longer prompts with a clear error. The parity numbers above were measured with a 1.31 development build, whose sliding-window attention runs prompts up to 9k tokens; on prompts of 1-8k tokens, 1.30 gives the same answers (max. probability difference to 1.31: 0.019). - Latency (M4 Max): about 1.2-1.6 s per image question, about 7 s for a text question under 2k tokens.