Instructions to use Omartificial-Intelligence-Space/Nun-Vision-30B-Lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Omartificial-Intelligence-Space/Nun-Vision-30B-Lora with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("unsloth/Muse-Glimmer-30B-unsloth-bnb-4bit") model = PeftModel.from_pretrained(base_model, "Omartificial-Intelligence-Space/Nun-Vision-30B-Lora") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Unsloth Desktop
ن Nūn Vision — 30B LoRA
Nūn Vision is the Arabic-calligraphy model behind Nūn · نون, our entry to the AI in the Service of Islamic Content Challenge 2026 (Track 03 — interactive experiences that introduce Islam).
Point a camera at a calligraphy panel in a mosque, a museum or a home, and Nūn Vision returns, in one structured answer:
- the text it reads, line by line;
- where each line is in the photo (bounding boxes);
- the script style — Thuluth, Diwani, Naskh, Kufic, Ruq'ah, Nasta'liq;
- the theme — Quranic, Names of Allah, supplication, hadith, names, non-religious…
It is a LoRA fine-tune of the 30B Muse Glimmer vision-language model (4-bit), trained by the Nūn team on DuwatBench Arabic calligraphy, and released with open weights so that museums, mosques and researchers can build on it.
In Nūn: how the model is used safely
A model that reads Quranic calligraphy must never write the Quran. In Nūn the model's reading is never shown to a visitor. It is used to find the verse, and the verse is then shown verbatim from the approved Mushaf text:
photo ─► known panel? ── yes ─► verified card (geometric match)
│ no
▼
Nūn Vision reads it ─► nearest passage in the whole Quran ─► independent check agrees? ─► card
(text, lines, script, (1–3-ayah windows, fuzzy match; (a second model, blind (Mushaf text,
theme) the Mushaf text replaces the reading) to our reading) translation,
recitation)
otherwise ─► "not a Quran verse" / "not sure", with the script and text regions it saw
Results
Real calligraphy, 195 panels (Naskh 60, Thuluth 59, Diwani 59, Diwani-jelli 11, Kufic 5, Muhaqaq 1), none of them in Nūn's collection. Question: which verse is in this photo? "Hallucination" = a wrong verse is shown.
| system | hallucination ↓ | precision when it answers ↑ |
|---|---|---|
| GPT-5.5 (closed, general) | 35.4% | 62% |
| Claude Opus 5 (closed, general) | 10.3% | 89% |
| Nūn (Nūn Vision + Mushaf search + independent check) | 1.0% | 96% |
General-purpose vision models answer almost every image and confidently name a wrong verse on 10–35% of panels — for Quranic content that is the failure that matters. Nūn shows a verse only when it is sure, and otherwise says so.
What the model itself gets right (same 195 panels):
| Naskh | Thuluth | Diwani | |
|---|---|---|---|
| script style recognised | 85% | 86% | 36% |
| verse found from its reading (strong match) | 83% | 12% | 15% |
DuwatBench held-out subset (50 images): valid structured JSON 100%, script style-set 76%, text-region IoU 0.71, reading chrF2 46.3.
Known limits
- Ornate scripts (Thuluth, Diwani) are hard: the model can recall a well-known verse instead of reading the panel. This is why Nūn never shows a verse from the model alone, and why short famous phrases are treated with extra care.
- Trained on DuwatBench (1,272 images); real-world photos with glare, angle or decoration lower accuracy.
- Not an authoritative transcription system for religious, legal or archival material.
Usage
import torch
from PIL import Image
from peft import PeftModel
from transformers import AutoModelForImageTextToText, AutoProcessor
BASE = "unsloth/Muse-Glimmer-30B-unsloth-bnb-4bit"
ADAPTER = "Omartificial-Intelligence-Space/Nun-Vision-30B-Lora"
processor = AutoProcessor.from_pretrained(ADAPTER)
base = AutoModelForImageTextToText.from_pretrained(BASE, device_map={"": 0}, dtype=torch.bfloat16)
model = PeftModel.from_pretrained(base, ADAPTER).eval()
PROMPT = """حلّل لوحة الخط العربي وحدّد أنواع الخط، والتصنيف الموضوعي، والنصوص العربية ومواقعها.
استخدم bbox_1000_xywh بصيغة [x, y, width, height] وأعداد صحيحة بين 0 و1000. أعد JSON صالحًا فقط."""
image = Image.open("panel.jpg").convert("RGB")
image.thumbnail((896, 896)) # the model was trained at ≤ 448×448 pixels of area, ≤ 896 px side
messages = [{"role": "user", "content": [{"type": "image"}, {"type": "text", "text": PROMPT}]}]
prompt = processor.apply_chat_template(messages, add_generation_prompt=True, tokenize=False) + " to=user<|message|>"
inputs = processor(text=[prompt], images=[[image]], add_special_tokens=False, return_tensors="pt").to("cuda")
out = model.generate(**inputs, max_new_tokens=512, do_sample=False)
print(processor.tokenizer.decode(out[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))
# {"styles": ["Thuluth"], "theme": "quranic", "regions": [{"bbox_1000_xywh": [...], "text": "..."}]}
About 22 GB of GPU memory in 4-bit; ~6–7 s per panel on one RTX 5090. Note the " to=user<|message|>" suffix, which
matches the fine-tuning template.
Training
| Base | Muse Glimmer 30B, 4-bit (Unsloth), Apache-2.0 |
| Method | LoRA, rank 16, α 16 (PEFT 0.19) |
| Data | DuwatBench — 1,272 calligraphy images (567 Quranic, 344 devotional, 97 Names of Allah, 140 other); images are used for training only and are not redistributed |
| Tasks | full reading, script style, theme, structured analysis with text regions |
| Image policy | ≤ 448×448 pixels of area, ≤ 896 px on the long side |
License
other: use is subject to the base model's terms (Apache-2.0) and to the DuwatBench terms. The training images are not
redistributed.
Nūn
Nūn (نون — the letter that opens Sūrat al-Qalam, «نٓ ۚ وَٱلۡقَلَمِ») is an AI museum guide: photograph Arabic calligraphy, see the verse exactly as in the Mushaf with an approved translation and a human recitation, then ask questions in any language and get answers grounded in approved sources, with citations.
Team: Omer Nacar, Saeed Al-Zahrani.
Citation
@misc{nun_vision_2026,
title = {Nūn Vision: a LoRA fine-tuned vision-language model for Arabic calligraphy understanding},
author = {Nacar, Omer and Al-Zahrani, Saeed},
year = {2026},
howpublished = {\url{https://huggingface.co/Omartificial-Intelligence-Space/Nun-Vision-30B-Lora}}
}
- Downloads last month
- 36
Model tree for Omartificial-Intelligence-Space/Nun-Vision-30B-Lora
Base model
meta-models/Muse-Glimmer-30BDataset used to train Omartificial-Intelligence-Space/Nun-Vision-30B-Lora
Evaluation results
- Structured JSON validity on DuwatBench held-out subset (50 images)test set self-reported1.000
- Script style-set accuracy on DuwatBench held-out subset (50 images)test set self-reported0.760
- Text-region IoU (mean, matched) on DuwatBench held-out subset (50 images)test set self-reported0.707
- chrF2 (reading) on DuwatBench held-out subset (50 images)test set self-reported46.270