How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf mindchain/decider-2b-vision-GGUF:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf mindchain/decider-2b-vision-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf mindchain/decider-2b-vision-GGUF:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf mindchain/decider-2b-vision-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf mindchain/decider-2b-vision-GGUF:Q4_K_M
# Run inference directly in the terminal:
./llama-cli -hf mindchain/decider-2b-vision-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf mindchain/decider-2b-vision-GGUF:Q4_K_M
# Run inference directly in the terminal:
./build/bin/llama-cli -hf mindchain/decider-2b-vision-GGUF:Q4_K_M
Use Docker
docker model run hf.co/mindchain/decider-2b-vision-GGUF:Q4_K_M
Quick Links

decider-2b-vision Q4_K_M + mmproj — der Bild-Entscheider der decider-Familie

Typed decisions from an image in one forward pass — Kamera-Bild rein, kalibrierte Wahrscheinlichkeiten raus. Der fehlende Baustein für den Kamera-Pfad (Android, alte GPUs, Edge).

File Größe Zweck
decider-2b-vision.Q4_K_M.gguf 1,27 GB Sprach-/Entscheidungsteil (Qwen3.5-2B-Basis, v5 Text-Gewichte)
mmproj-decider-2b-vision-f16.gguf 668 MB Vision-Projektor (f16, wie bei Qwen-VL/smolVLM üblich)
Revision Commit-gepinnt 863e290863655f1d6b69324d77d09ac972d21609 (Repo hat keine Tags)
Toolchain llama.cpp Commit 9575389: Text-Pfad --no-mtp + bf16 → Q4_K_M; mmproj separat mit --mmproj --outtype f16 exportiert
⚠️ Toolchain-Warnung llama.cpp master (e351231, Sep 2026) baut kaputte qwen3_5_text-GGUFs (falsche Metadaten + verschobene Gewichte, Output ?). Pin euren Commit; Qwen3_5ForConditionalGeneration ist im gepinnten Baum in qwen.py (Text) und qwen3vl.py (VL) registriert.

Usage (llama-server, zwei Files)

llama-server -m decider-2b-vision.Q4_K_M.gguf \
  --mmproj mmproj-decider-2b-vision-f16.gguf \
  -ngl 99 --alias decider-2b-vision

Rolle im JEV-Stack

Kamera-/Bild-Pfad der Tier-0/1-Kette: Bild rein → Entscheidung + Wahrscheinlichkeiten in einem Forward-Pass. Q4 ~1,3 GB läuft auf einer GTX 1070 neben anderen Services; Text-Schwester decider-2b liest bei T=1.03 (Vision-Temperatur: im decider/-Paket des Upstream-Repos — vor Produktionsnutzung gegen den bf16-Original-Readout verifizieren, wie bei allen unseren Quants).

Downloads last month
334
GGUF
Model size
2B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mindchain/decider-2b-vision-GGUF

Quantized
(4)
this model

Space using mindchain/decider-2b-vision-GGUF 1

Collection including mindchain/decider-2b-vision-GGUF