jev-omni.js weights

A browser-ready export of Jev-Omni, akhilaaa3's Gemma 4 12B decision classifier: a state, a question and 2-256 options in, one probability per option out, from a single forward pass. It runs in Chrome on WebGPU through jev-omni.js, for text and images. Live demo.

This repo only holds converted weights. The model, its training, its decision head and its prompt format are Jev-Omni's. Jev-Omni states that it is independent of TypeSafe AI's Jev, and so is this repo.

import * as ort from "onnxruntime-web/webgpu";   // 1.30.0: see Requirements
import { loadJevOmni } from "@ai-ecoverse/jev-omni.js";

const jev = await loadJevOmni("https://huggingface.co/ai-ecoverse/jev-omni.js/resolve/main/jev-omni", { ort });
const res = await jev.predict({
  state: "The meeting starts at 10 AM. It is now 9 AM.",
  question: "Has the meeting started?",
  options: ["Yes", "No"],
});
res.probabilities;   // { Yes: ..., No: ... }

Contents

Path What Size
jev-omni/manifest.json file list with sizes and SHA-256, graph inputs, parity numbers; revision 1544c2b7ed5277a0
jev-omni/r-c050d51/q8f32/ decoder: int8 weights (MatMulNBits, block 32), fp32 activations, fp16 attention in the 8 global layers; 437 files 13.34 GB
jev-omni/r-c050d51/vision/ Gemma 4's vision embedder and projection, fp32 0.20 GB
jev-omni/r-c050d51/head.safetensors, tokenizer*.json the 256-way decision head, the tokenizer 36 MB

Total download: 13.58 GB in 445 files of at most 32 MB.

Provenance

  • Model: akhilaaa3/Jev-Omni at c050d51354147985d13286cf4acf90f562f2c631 (Apache-2.0), unified/model.safetensors (bf16) and head.pt.
  • Base: google/gemma-4-12B-it at 707f0a3b8a3c7ad586ed01e27eafbad8a27dd0f7 (Apache-2.0).
  • Export: a vendored draft of the onnxruntime-genai model builder with Gemma 4 support (microsoft/onnxruntime-genai#2473, 603c36c), without the LM head, then graph rewrites for WebGPU and the image path. Details are in the jev-omni.js docs.

Parity with the fp32 PyTorch model

Chrome on WebGPU (Apple M4 Max) against Jev-Omni in fp32 PyTorch:

Set Accuracy, browser / reference Answers changed Mean / max abs. probability difference
DecisionBench medium, 293 text questions 86.35% / 86.35% 0 0.0019 / 0.131
kev.js vision-v1, 106 image questions 98.11% / 99.06% 1 0.0014 / 0.065
kev.js vision-v2, 128 image questions 85.16% / 85.16% 1 0.0037 / 0.065

Without quantization, a 6-layer fp32 export matches PyTorch to a relative error of 5.0e-06 at the last position.

Requirements

  • Memory: about 16 GB of GPU memory for the WebGPU session. While loading, the tab also holds the 13.6 GB of files in memory until they are on the GPU. On Apple Silicon that means 32 GB of unified memory at the very least, and 64 GB or more to be comfortable. It has been tested on an Apple M4 Max with 128 GB only.
  • Download: 13.6 GB on the first load. jev-omni.js keeps it in Cache Storage, so later loads read from disk.
  • onnxruntime-web 1.30.0: its attention needs an n×n fp32 buffer per head, so prompts are capped at about 8,000 tokens on GPUs with 4 GB buffers (Apple Silicon) and fewer on GPUs with smaller ones. jev-omni.js rejects longer prompts with a clear error. The parity numbers above were measured with a 1.31 development build, whose sliding-window attention runs prompts up to 9k tokens; on prompts of 1-8k tokens, 1.30 gives the same answers (max. probability difference to 1.31: 0.019).
  • Latency (M4 Max): about 1.2-1.6 s per image question, about 7 s for a text question under 2k tokens.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ai-ecoverse/jev-omni.js

Quantized
(4)
this model