⚡ Laya-fa Universal Support (v3.1 Contrastive Edition)

🇮🇷 مطالعه مستندات به زبان فارسی (Persian Documentation)

Laya System-1 Architecture

HF Model Author Persian README Accuracy Disambiguation Latency Domains License

Laya-fa Universal Support is an ultra-fast, sub-140ms System-1 decision and intent-routing engine engineered specifically for the Iranian Telegram bot and e-commerce ecosystem. Powered by an mmBERT / ModernBERT backbone (322M parameters) with contrastive hard-negative training, it delivers calibrated decision probabilities directly to bot inline keyboard controllers—bypassing the multi-second latency, GPU costs, and hallucination risks of traditional generative LLMs.


📑 Table of Contents

  1. Why System-1 Instead of Generative LLMs?
  2. Dataset Statistics & Corpus Breakdown
  3. Comprehensive E2E Benchmark (350 Production Queries)
  4. The 12 Universal E-Commerce Domains
  5. Hard Negative Disambiguation Matrix (v3.1 Breakthrough)
  6. Inference Latency & Temperature Calibration
  7. Repository Structure (src/ & data/)
  8. Step-by-Step Training & Reproduction Guide
  9. Quickstart & Integration Examples
  10. Available Versions & Git Branches
  11. Acknowledgments

⚡ Why System-1 Instead of Generative LLMs?

When an e-commerce Telegram bot receives a customer inquiry, prompt-engineering a 70B parameter generative model (like Llama-3 or GPT-4) creates severe architectural bottlenecks:

  • High Latency: Generative LLMs generate tokens autoregressively, taking 3,000 to 8,000 ms to answer. Laya-fa System-1 classifies the intent in 139 ms on a standard CPU.
  • Hallucinations & Markdown Errors: Generative models often invent policies or fail to output parseable JSON. Laya-fa outputs clean, deterministic probability vectors.
  • Marginal Cost: Operating LLM APIs at 100,000 requests/day costs hundreds of dollars monthly. Laya-fa runs on-premises on modest CPU servers for $0.00 in API fees.

📈 Dataset Statistics & Corpus Breakdown

Laya-fa v3.1 was developed using a multi-tiered dataset engineered to eliminate subword tokenizer collisions, informal Persian typos, and domain polysemy:

Corpus Split / Test Suite Sample Count Format Primary Objective
Final Training Set (train.jsonl) 5,556 JSONL Full-coverage supervision across all 12 commercial domains.
Validation / Dev Set (dev.jsonl) 740 JSONL Stratified cross-validation and temperature calibration ($T=1.20$).
Contrastive Hard Negatives 625 Pairs High-leverage adversarial pairs (miner vs. mine, config vs. cancel, etc.).
Blind Production Benchmark 350 JSONL Completely uncurated real queries collected from live Telegram e-commerce bots.
Edge Disambiguation Suite 34 JSONL Extreme lexical edge cases and boundary collisions.
Total Supervised Persian Queries 6,296 — Cleaned, deduplicated, and normalized with ZWNJ (نیم‌فاصله).

📊 Comprehensive E2E Benchmark (350 Production Queries)

The model was evaluated against 350 real-world, blind production queries collected from active Telegram commercial bots in Iran, tested side-by-side against official baselines and the premier overseas commercial cloud engine (TypeSafe Jev):

Benchmark Comparison

Detailed Comparative Metrics:

Model Architecture Deployment Overall Acc Macro F1 Edge Ambiguity (34) Short Acronyms (15) Finglish (15) Latency (p50) Cost / 1k req
Laya Base (Official) Local CPU 23.4% 18.6% 12.0% 46.7% 53.3% 428 ms $0.00
Tarfandoon FA Support Local CPU 39.1% 38.2% 29.4% 60.0% 26.7% 141 ms $0.00
Laya-fa v1 (Early) Local CPU 60.0% 59.2% 44.1% 100.0% 46.7% 134 ms $0.00
Laya-fa v2 (Universal) Local CPU 82.3% 82.3% 52.9% 26.7% ❌ 60.0% 144 ms $0.00
TypeSafe Jev (Cloud API) Cloud API 88.3% 88.2% 85.3% 100.0% 100.0% 255 ms $$$ (Paid API)
Laya-fa v3 (Universal v3) Local CPU 90.0% 89.9% 76.5% 100.0% 93.3% 139 ms $0.00
Laya-fa v3.1 (Contrastive) Local CPU 88.0% 87.9% 100.0% (34/34) 🔥 100.0% 93.3% 139 ms $0.00 (Self-hosted)

Key Finding: While generic cloud APIs perform well on standard English or formal Persian, they miss colloquial Iranian cultural slangs and local e-commerce terminology. Laya-fa v3.1 achieved 100% precision on edge ambiguity scenarios while responding 1.8x faster with zero internet-dependency.


🌐 The 12 Universal E-Commerce Domains

Laya-fa v3.1 routes all incoming traffic into 12 mutually exclusive commercial categories:

Domain Identifier Business Department Canonical Persian Intent / Slangs
billing_banking Banking & Shetab Transactions واریز کارت به کارت، کسر وجه، شماره شبا، تایید فیش، تمدید اشتراک، کدهای ستاره مربع USSD
stars_telegram Telegram Stars & Channel Gifts خرید استارز، نرخ ستاره، گیفت ۵۰۰ استار، استار، ۵۰ star، ستاره چنل
ton_crypto_web3 TON Blockchain, Gram & Web3 والت تون‌کیپر، ماینر سخت‌افزاری، استخراج ارز، هش‌ریت، مینی‌اپ، فرگمنت (+888)
telegram_premium Telegram Premium Subscriptions اشتراک پریمیوم ۳ ماهه و ۱ ساله، فعال‌سازی با آیدی بدون پسورد، گیفت‌کد
network_vpn_proxy Proxies, VPNs & Censorship Circumvention قطعی همراه اول/ایرانسل، کانفینگ v2ray، لینک ساب در ویتوری، یه فیلتر بده، اینترنت استارلینک
hosting_domain_server VPS Hosting & Domain Registration سرور مجازی vps، وی پی اس آلمان، ساب‌دومین کلودفلر، ارورهای ۵۰۲/۵۰۴، کانفیگ nginx
gaming_minecraft Gaming Servers & Minecraft سرور ماین، مخفف‌های mc و ام سی، لانچر، رنک VIP، پورت بدراک و جاوا، اکانت پرمیوم
account_security Account Security & Access Recovery عدم دریافت OTP، فراموشی رمز ۲ مرحله‌ای، گزارش هک، نشست‌های مشکوک
refund_cancellation Cancellations & Money-Back Guarantees لغو قطعی تمدید خودکار، انصراف آنی، استرداد وجه به کارت، بستن اشتراک
partnership_wholesale B2B Wholesale & API Partnerships خرید عمده (۲۰k تا ۵۰k)، همکاری در فروش، وب‌سرویس API، توکن بات‌فادر تلگرام
onboarding_guide User Onboarding & FAQs راهنمای شروع، کلیپ آموزش کار با بات، پین پیام در تلگرام، لیست فرامین
trolling_chitchat_spam Troll Filtering, Chit-Chat & Noise متن ترول و شعر، امتیاز ۵ ستاره به اپلیکیشن، پین لوکیشن نقشه، واحد وزن بار تن سیمان

🎯 Hard Negative Disambiguation Matrix (v3.1 Breakthrough)

In subword tokenization, phonetically or structurally similar words often trigger severe neural network confusion. Version 3.1 introduces 625 contrastive hard-negative pairs to firmly separate overlapping boundaries:

Disambiguation Matrix

Live Verification Examples on Production Edge Queries:

Query Under Test Legacy v3 Route v3.1 Contrastive Route Confidence Technical Disambiguation Feat
«ماینر روشن نمیشه فنش کار نمیکنه» gaming_minecraft ❌ ton_crypto_web3 ✅ 99.5% Separates crypto mining hardware (ماینر) from Minecraft (ماین)
«کانفینگ جدید بدید قبلی قطع شد» refund_cancellation ❌ network_vpn_proxy ✅ 93.6% Eliminates bias on the colloquial typo «کانفینگ» (with extra noon)
«کانفینگ همراه اول بده» refund_cancellation ❌ network_vpn_proxy ✅ 98.5% Correctly maps network request instead of triggering refund
«لینک ساب رو داخل ویتوری کپی کردم» trolling_chitchat ❌ network_vpn_proxy ✅ 99.9% Identifies V2Ray subscription links vs. DNS subdomains
«روی کلودفلر یه رکورد ساب دامین A زدم» trolling_chitchat ❌ hosting_domain_server ✅ 99.5% Correctly maps Cloudflare DNS routing
«ام سی بالا نمیاد ارور میده» trolling_chitchat ❌ gaming_minecraft ✅ 100.0% Understands phonetic Persian spelling of English MC (Minecraft)
«سرور وی پی اس آلمان با رم ۸ گیگ» network_vpn_proxy ❌ hosting_domain_server ✅ 98.9% Disambiguates Virtual Private Server (VPS) from VPN
«به برنامه‌تون ۵ تا استار دادم عالی هستید» stars_telegram ❌ trolling_chitchat_spam ✅ 98.7% Separates 5-star app feedback from Telegram Stars currency
«اینترنت ماهواره‌ای استارلینک دارید؟» stars_telegram ❌ network_vpn_proxy ✅ 93.5% Separates Starlink satellite equipment from Telegram Stars
«توکن بات‌فادر رو تو کانفیگ بذارم ران میشه؟» network_vpn_proxy ❌ partnership_wholesale ✅ 68.4% Routes Telegram BotFather API token inquiries correctly
«دو تن سیمان بار ماشین کردیم» network_vpn_proxy ❌ trolling_chitchat_spam ✅ 97.9% Separates physical metric ton weight from TON cryptocurrency
«لوکیشن دفتر رو روی گوگل مپ پین کن» onboarding_guide ❌ trolling_chitchat_spam ✅ 99.9% Distinguishes Google Map GPS pins from pinned Telegram messages

⏱ Inference Latency & Calibration

  • Median Latency (p50): 139.3 ms (on standard Intel Xeon CPU).
  • 95th Percentile Latency (p95): 260.6 ms under high concurrency.
  • Calibrated Softmax Temperature: $T = 1.20$.
  • Expected Calibration Error (ECE): 0.0592.
  • Confidence Safety Margin:
    • Mean confidence on correct decisions: 98.5%.
    • Mean confidence on ambiguous / edge queries: 69.2% (giving developers an unambiguous 30% margin to safely trigger bot fallback menus).

📁 Repository Structure (src/ & data/)

This repository is fully reproducible and structured cleanly:

├── assets/                       # High-resolution infographics and charts
├── data/                         # Benchmark & evaluation test suites
│   ├── benchmark_350_suite.jsonl # Complete 350-query blind production test set
│   ├── edge_ambiguity_suite.jsonl# 34 critical lexical boundary test cases
│   └── contrastive_sample.jsonl  # Representative hard-negative training pairs
├── encoder/                      # mmBERT backbone configuration
├── src/                          # Standalone Python training & evaluation pipeline
│   ├── train.py                  # Warm-start fine-tuning & contrastive training script
│   └── evaluate.py               # Standalone E2E benchmark evaluator
├── tokenizer/                    # Tokenizer vocabulary and configs
├── model.safetensors             # Calibrated neural network weights (bfloat16)
├── rl_agent_config.json          # System-1 label map & temperature manifest
├── README.md                     # Official English documentation
├── README_FA.md                  # Comprehensive Persian documentation
└── LICENSE                       # MIT Open-Source License

🛠 Step-by-Step Training & Reproduction Guide

1. Requirements & Setup

pip install torch transformers datasets accelerate safetensors huggingface_hub

2. Reproduce the Production Benchmark in 1-Line:

You can directly evaluate the published model against the 350-query blind suite and 34 edge ambiguity cases:

python3 src/evaluate.py --model mmahdi-sz/Laya-fa-universal-support --device cpu

3. Running Custom Warm-Start Fine-Tuning:

python3 src/train.py \
  --base mmahdi-sz/Laya-fa-universal-support \
  --train ./data/train.jsonl \
  --dev ./data/dev.jsonl \
  --contrastive ./data/contrastive_sample.jsonl \
  --epochs 2 \
  --temperature 1.20 \
  --out ./output_model_v31

4. Hyperparameters Reference:

  • Base Architecture: mmBERT (ModernBERT multilingual, 322M parameters)
  • Precision: bfloat16
  • Optimizer: AdamW (weight_decay=0.01)
  • Learning Rate (Backbone): 3e-5 with linear warmup
  • Learning Rate (Classification Head): 1e-4
  • Batch Size: 16 (per device) with gradient accumulation
  • Contrastive Oversampling: 2x on hard-negative pairs to firmly stamp out boundary collisions

🚀 Quickstart & Integration Examples

Python (Native httpx async client):

import httpx

async def get_bot_route(user_message: str):
    async with httpx.AsyncClient() as client:
        response = await client.post(
            "http://127.0.0.1:8888/v1/systemone",
            headers={"Authorization": "Bearer YOUR_SECRET_API_KEY"},
            json={
                "state": user_message,
                "model": "laya-fa-universal",
                "questions": {
                    "department": {
                        "type": "choice",
                        "instructions": "Route user message to the correct department",
                        "criteria": {
                            "billing_banking": "Banking, receipts, and card transfers",
                            "stars_telegram": "Telegram Stars currency purchases",
                            "ton_crypto_web3": "TON, crypto wallets, and miners",
                            "network_vpn_proxy": "VPNs, proxies, and configs",
                            "gaming_minecraft": "Minecraft servers and gaming ranks",
                            "trolling_chitchat_spam": "Spam, chit-chat, or ratings"
                        }
                    }
                }
            }
        )
        return response.json()

cURL Execution:

curl -X POST http://127.0.0.1:8888/v1/systemone \
  -H "Authorization: Bearer YOUR_SECRET_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "state": "ماینر روشن نمیشه فنش کار نمیکنه",
    "model": "laya-fa-universal",
    "questions": {
      "department": {
        "type": "choice",
        "instructions": "Select appropriate team"
      }
    }
  }'

📦 Available Versions & Git Branches

To ensure backward compatibility, every major version of Laya-fa is preserved as an independent branch:

Version Tag / Branch Architecture Target Scope Key Milestones
main / v3.1 mmBERT-322M 12 Domains Latest Production Release. 100% resolution of slangs, typos (کانفینگ), and subword collisions (ماینر vs ماین).
v3.0 mmBERT-322M 12 Domains Introduced single-word and short acronym optimization (mc, star, stars).
v2.0 mmBERT-322M 12 Domains Initial 12-domain universal expansion across 6,000 training samples.
v1.0 mmBERT-322M 4 Domains Early fine-tuned baseline for Telegram star purchases.

To pin an earlier version in code:

from transformers import AutoModel
# Pinning version 2.0 explicitly:
model = AutoModel.from_pretrained("mmahdi-sz/Laya-fa-universal-support", revision="v2.0")

🙏 Acknowledgments

Special thanks to @NotMmd for the initial inspiration, insightful architectural feedback, discovering nuanced Persian edge cases, and valuable contributions throughout the development and benchmarking of this model.


📜 Citation & License

This project is licensed under the MIT License.

@misc{laya_fa_universal_2026,
  author = {Mahdi (mmahdi-sz)},
  title = {Laya-fa Universal Support: Sub-140ms System-1 Intent Router for Persian E-Commerce},
  year = {2026},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/mmahdi-sz/Laya-fa-universal-support}}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support