ن Nūn Vision — 30B LoRA

Nūn Vision is the Arabic-calligraphy model behind Nūn · نون, our entry to the AI in the Service of Islamic Content Challenge 2026 (Track 03 — interactive experiences that introduce Islam).

Point a camera at a calligraphy panel in a mosque, a museum or a home, and Nūn Vision returns, in one structured answer:

  • the text it reads, line by line;
  • where each line is in the photo (bounding boxes);
  • the script style — Thuluth, Diwani, Naskh, Kufic, Ruq'ah, Nasta'liq;
  • the theme — Quranic, Names of Allah, supplication, hadith, names, non-religious…

It is a LoRA fine-tune of the 30B Muse Glimmer vision-language model (4-bit), trained by the Nūn team on DuwatBench Arabic calligraphy, and released with open weights so that museums, mosques and researchers can build on it.

In Nūn: how the model is used safely

A model that reads Quranic calligraphy must never write the Quran. In Nūn the model's reading is never shown to a visitor. It is used to find the verse, and the verse is then shown verbatim from the approved Mushaf text:

photo ─► known panel? ── yes ─► verified card (geometric match)
             │ no
             ▼
   Nūn Vision reads it  ─►  nearest passage in the whole Quran  ─►  independent check agrees?  ─►  card
   (text, lines, script,     (1–3-ayah windows, fuzzy match;          (a second model, blind      (Mushaf text,
    theme)                    the Mushaf text replaces the reading)    to our reading)             translation,
                                                                                                   recitation)
             otherwise ─► "not a Quran verse" / "not sure", with the script and text regions it saw

Results

Real calligraphy, 195 panels (Naskh 60, Thuluth 59, Diwani 59, Diwani-jelli 11, Kufic 5, Muhaqaq 1), none of them in Nūn's collection. Question: which verse is in this photo? "Hallucination" = a wrong verse is shown.

system hallucination ↓ precision when it answers ↑
GPT-5.5 (closed, general) 35.4% 62%
Claude Opus 5 (closed, general) 10.3% 89%
Nūn (Nūn Vision + Mushaf search + independent check) 1.0% 96%

General-purpose vision models answer almost every image and confidently name a wrong verse on 10–35% of panels — for Quranic content that is the failure that matters. Nūn shows a verse only when it is sure, and otherwise says so.

What the model itself gets right (same 195 panels):

Naskh Thuluth Diwani
script style recognised 85% 86% 36%
verse found from its reading (strong match) 83% 12% 15%

DuwatBench held-out subset (50 images): valid structured JSON 100%, script style-set 76%, text-region IoU 0.71, reading chrF2 46.3.

Known limits

  • Ornate scripts (Thuluth, Diwani) are hard: the model can recall a well-known verse instead of reading the panel. This is why Nūn never shows a verse from the model alone, and why short famous phrases are treated with extra care.
  • Trained on DuwatBench (1,272 images); real-world photos with glare, angle or decoration lower accuracy.
  • Not an authoritative transcription system for religious, legal or archival material.

Usage

import torch
from PIL import Image
from peft import PeftModel
from transformers import AutoModelForImageTextToText, AutoProcessor

BASE = "unsloth/Muse-Glimmer-30B-unsloth-bnb-4bit"
ADAPTER = "Omartificial-Intelligence-Space/Nun-Vision-30B-Lora"

processor = AutoProcessor.from_pretrained(ADAPTER)
base = AutoModelForImageTextToText.from_pretrained(BASE, device_map={"": 0}, dtype=torch.bfloat16)
model = PeftModel.from_pretrained(base, ADAPTER).eval()

PROMPT = """حلّل لوحة الخط العربي وحدّد أنواع الخط، والتصنيف الموضوعي، والنصوص العربية ومواقعها.
استخدم bbox_1000_xywh بصيغة [x, y, width, height] وأعداد صحيحة بين 0 و1000. أعد JSON صالحًا فقط."""

image = Image.open("panel.jpg").convert("RGB")
image.thumbnail((896, 896))  # the model was trained at ≤ 448×448 pixels of area, ≤ 896 px side
messages = [{"role": "user", "content": [{"type": "image"}, {"type": "text", "text": PROMPT}]}]
prompt = processor.apply_chat_template(messages, add_generation_prompt=True, tokenize=False) + " to=user<|message|>"
inputs = processor(text=[prompt], images=[[image]], add_special_tokens=False, return_tensors="pt").to("cuda")
out = model.generate(**inputs, max_new_tokens=512, do_sample=False)
print(processor.tokenizer.decode(out[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))
# {"styles": ["Thuluth"], "theme": "quranic", "regions": [{"bbox_1000_xywh": [...], "text": "..."}]}

About 22 GB of GPU memory in 4-bit; ~6–7 s per panel on one RTX 5090. Note the " to=user<|message|>" suffix, which matches the fine-tuning template.

Training

Base Muse Glimmer 30B, 4-bit (Unsloth), Apache-2.0
Method LoRA, rank 16, α 16 (PEFT 0.19)
Data DuwatBench — 1,272 calligraphy images (567 Quranic, 344 devotional, 97 Names of Allah, 140 other); images are used for training only and are not redistributed
Tasks full reading, script style, theme, structured analysis with text regions
Image policy ≤ 448×448 pixels of area, ≤ 896 px on the long side

License

other: use is subject to the base model's terms (Apache-2.0) and to the DuwatBench terms. The training images are not redistributed.

Nūn

Nūn (نون — the letter that opens Sūrat al-Qalam, «نٓ ۚ وَٱلۡقَلَمِ») is an AI museum guide: photograph Arabic calligraphy, see the verse exactly as in the Mushaf with an approved translation and a human recitation, then ask questions in any language and get answers grounded in approved sources, with citations.

Team: Omer Nacar, Saeed Al-Zahrani.

Citation

@misc{nun_vision_2026,
  title  = {Nūn Vision: a LoRA fine-tuned vision-language model for Arabic calligraphy understanding},
  author = {Nacar, Omer and Al-Zahrani, Saeed},
  year   = {2026},
  howpublished = {\url{https://huggingface.co/Omartificial-Intelligence-Space/Nun-Vision-30B-Lora}}
}
Downloads last month
36
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Omartificial-Intelligence-Space/Nun-Vision-30B-Lora

Dataset used to train Omartificial-Intelligence-Space/Nun-Vision-30B-Lora

Evaluation results