Instructions to use mmahdi-sz/Laya-fa-universal-support with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mmahdi-sz/Laya-fa-universal-support with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="mmahdi-sz/Laya-fa-universal-support")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("mmahdi-sz/Laya-fa-universal-support", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- ⚡ Laya-fa Universal Support (v3.1 Contrastive Edition)
- 📑 Table of Contents
- ⚡ Why System-1 Instead of Generative LLMs?
- 📈 Dataset Statistics & Corpus Breakdown
- 📊 Comprehensive E2E Benchmark (350 Production Queries)
- 🌐 The 12 Universal E-Commerce Domains
- 🎯 Hard Negative Disambiguation Matrix (v3.1 Breakthrough)
- ⏱ Inference Latency & Calibration
- 📁 Repository Structure (
src/&data/) - 🛠 Step-by-Step Training & Reproduction Guide
- 🚀 Quickstart & Integration Examples
- 📦 Available Versions & Git Branches
- 🙏 Acknowledgments
- 📜 Citation & License
- 📑 Table of Contents
⚡ Laya-fa Universal Support (v3.1 Contrastive Edition)
🇮🇷 مطالعه مستندات به زبان فارسی (Persian Documentation)
Laya-fa Universal Support is an ultra-fast, sub-140ms System-1 decision and intent-routing engine engineered specifically for the Iranian Telegram bot and e-commerce ecosystem. Powered by an mmBERT / ModernBERT backbone (322M parameters) with contrastive hard-negative training, it delivers calibrated decision probabilities directly to bot inline keyboard controllers—bypassing the multi-second latency, GPU costs, and hallucination risks of traditional generative LLMs.
📑 Table of Contents
- Why System-1 Instead of Generative LLMs?
- Dataset Statistics & Corpus Breakdown
- Comprehensive E2E Benchmark (350 Production Queries)
- The 12 Universal E-Commerce Domains
- Hard Negative Disambiguation Matrix (v3.1 Breakthrough)
- Inference Latency & Temperature Calibration
- Repository Structure (
src/&data/) - Step-by-Step Training & Reproduction Guide
- Quickstart & Integration Examples
- Available Versions & Git Branches
- Acknowledgments
⚡ Why System-1 Instead of Generative LLMs?
When an e-commerce Telegram bot receives a customer inquiry, prompt-engineering a 70B parameter generative model (like Llama-3 or GPT-4) creates severe architectural bottlenecks:
- High Latency: Generative LLMs generate tokens autoregressively, taking 3,000 to 8,000 ms to answer. Laya-fa System-1 classifies the intent in 139 ms on a standard CPU.
- Hallucinations & Markdown Errors: Generative models often invent policies or fail to output parseable JSON. Laya-fa outputs clean, deterministic probability vectors.
- Marginal Cost: Operating LLM APIs at 100,000 requests/day costs hundreds of dollars monthly. Laya-fa runs on-premises on modest CPU servers for $0.00 in API fees.
📈 Dataset Statistics & Corpus Breakdown
Laya-fa v3.1 was developed using a multi-tiered dataset engineered to eliminate subword tokenizer collisions, informal Persian typos, and domain polysemy:
| Corpus Split / Test Suite | Sample Count | Format | Primary Objective |
|---|---|---|---|
Final Training Set (train.jsonl) |
5,556 | JSONL | Full-coverage supervision across all 12 commercial domains. |
Validation / Dev Set (dev.jsonl) |
740 | JSONL | Stratified cross-validation and temperature calibration ($T=1.20$). |
| Contrastive Hard Negatives | 625 | Pairs | High-leverage adversarial pairs (miner vs. mine, config vs. cancel, etc.). |
| Blind Production Benchmark | 350 | JSONL | Completely uncurated real queries collected from live Telegram e-commerce bots. |
| Edge Disambiguation Suite | 34 | JSONL | Extreme lexical edge cases and boundary collisions. |
| Total Supervised Persian Queries | 6,296 | — | Cleaned, deduplicated, and normalized with ZWNJ (نیمفاصله). |
📊 Comprehensive E2E Benchmark (350 Production Queries)
The model was evaluated against 350 real-world, blind production queries collected from active Telegram commercial bots in Iran, tested side-by-side against official baselines and the premier overseas commercial cloud engine (TypeSafe Jev):
Detailed Comparative Metrics:
| Model Architecture | Deployment | Overall Acc | Macro F1 | Edge Ambiguity (34) | Short Acronyms (15) | Finglish (15) | Latency (p50) | Cost / 1k req |
|---|---|---|---|---|---|---|---|---|
| Laya Base (Official) | Local CPU | 23.4% | 18.6% | 12.0% | 46.7% | 53.3% | 428 ms | $0.00 |
| Tarfandoon FA Support | Local CPU | 39.1% | 38.2% | 29.4% | 60.0% | 26.7% | 141 ms | $0.00 |
| Laya-fa v1 (Early) | Local CPU | 60.0% | 59.2% | 44.1% | 100.0% | 46.7% | 134 ms | $0.00 |
| Laya-fa v2 (Universal) | Local CPU | 82.3% | 82.3% | 52.9% | 26.7% ❌ | 60.0% | 144 ms | $0.00 |
| TypeSafe Jev (Cloud API) | Cloud API | 88.3% | 88.2% | 85.3% | 100.0% | 100.0% | 255 ms | $$$ (Paid API) |
| Laya-fa v3 (Universal v3) | Local CPU | 90.0% | 89.9% | 76.5% | 100.0% | 93.3% | 139 ms | $0.00 |
| Laya-fa v3.1 (Contrastive) | Local CPU | 88.0% | 87.9% | 100.0% (34/34) 🔥 | 100.0% | 93.3% | 139 ms | $0.00 (Self-hosted) |
Key Finding: While generic cloud APIs perform well on standard English or formal Persian, they miss colloquial Iranian cultural slangs and local e-commerce terminology. Laya-fa v3.1 achieved 100% precision on edge ambiguity scenarios while responding 1.8x faster with zero internet-dependency.
🌐 The 12 Universal E-Commerce Domains
Laya-fa v3.1 routes all incoming traffic into 12 mutually exclusive commercial categories:
| Domain Identifier | Business Department | Canonical Persian Intent / Slangs |
|---|---|---|
billing_banking |
Banking & Shetab Transactions | واریز کارت به کارت، کسر وجه، شماره شبا، تایید فیش، تمدید اشتراک، کدهای ستاره مربع USSD |
stars_telegram |
Telegram Stars & Channel Gifts | خرید استارز، نرخ ستاره، گیفت ۵۰۰ استار، استار، ۵۰ star، ستاره چنل |
ton_crypto_web3 |
TON Blockchain, Gram & Web3 | والت تونکیپر، ماینر سختافزاری، استخراج ارز، هشریت، مینیاپ، فرگمنت (+888) |
telegram_premium |
Telegram Premium Subscriptions | اشتراک پریمیوم ۳ ماهه و ۱ ساله، فعالسازی با آیدی بدون پسورد، گیفتکد |
network_vpn_proxy |
Proxies, VPNs & Censorship Circumvention | قطعی همراه اول/ایرانسل، کانفینگ v2ray، لینک ساب در ویتوری، یه فیلتر بده، اینترنت استارلینک |
hosting_domain_server |
VPS Hosting & Domain Registration | سرور مجازی vps، وی پی اس آلمان، سابدومین کلودفلر، ارورهای ۵۰۲/۵۰۴، کانفیگ nginx |
gaming_minecraft |
Gaming Servers & Minecraft | سرور ماین، مخففهای mc و ام سی، لانچر، رنک VIP، پورت بدراک و جاوا، اکانت پرمیوم |
account_security |
Account Security & Access Recovery | عدم دریافت OTP، فراموشی رمز ۲ مرحلهای، گزارش هک، نشستهای مشکوک |
refund_cancellation |
Cancellations & Money-Back Guarantees | لغو قطعی تمدید خودکار، انصراف آنی، استرداد وجه به کارت، بستن اشتراک |
partnership_wholesale |
B2B Wholesale & API Partnerships | خرید عمده (۲۰k تا ۵۰k)، همکاری در فروش، وبسرویس API، توکن باتفادر تلگرام |
onboarding_guide |
User Onboarding & FAQs | راهنمای شروع، کلیپ آموزش کار با بات، پین پیام در تلگرام، لیست فرامین |
trolling_chitchat_spam |
Troll Filtering, Chit-Chat & Noise | متن ترول و شعر، امتیاز ۵ ستاره به اپلیکیشن، پین لوکیشن نقشه، واحد وزن بار تن سیمان |
🎯 Hard Negative Disambiguation Matrix (v3.1 Breakthrough)
In subword tokenization, phonetically or structurally similar words often trigger severe neural network confusion. Version 3.1 introduces 625 contrastive hard-negative pairs to firmly separate overlapping boundaries:
Live Verification Examples on Production Edge Queries:
| Query Under Test | Legacy v3 Route | v3.1 Contrastive Route | Confidence | Technical Disambiguation Feat |
|---|---|---|---|---|
| «ماینر روشن نمیشه فنش کار نمیکنه» | gaming_minecraft ❌ |
ton_crypto_web3 ✅ |
99.5% | Separates crypto mining hardware (ماینر) from Minecraft (ماین) |
| «کانفینگ جدید بدید قبلی قطع شد» | refund_cancellation ❌ |
network_vpn_proxy ✅ |
93.6% | Eliminates bias on the colloquial typo «کانفینگ» (with extra noon) |
| «کانفینگ همراه اول بده» | refund_cancellation ❌ |
network_vpn_proxy ✅ |
98.5% | Correctly maps network request instead of triggering refund |
| «لینک ساب رو داخل ویتوری کپی کردم» | trolling_chitchat ❌ |
network_vpn_proxy ✅ |
99.9% | Identifies V2Ray subscription links vs. DNS subdomains |
| «روی کلودفلر یه رکورد ساب دامین A زدم» | trolling_chitchat ❌ |
hosting_domain_server ✅ |
99.5% | Correctly maps Cloudflare DNS routing |
| «ام سی بالا نمیاد ارور میده» | trolling_chitchat ❌ |
gaming_minecraft ✅ |
100.0% | Understands phonetic Persian spelling of English MC (Minecraft) |
| «سرور وی پی اس آلمان با رم ۸ گیگ» | network_vpn_proxy ❌ |
hosting_domain_server ✅ |
98.9% | Disambiguates Virtual Private Server (VPS) from VPN |
| «به برنامهتون ۵ تا استار دادم عالی هستید» | stars_telegram ❌ |
trolling_chitchat_spam ✅ |
98.7% | Separates 5-star app feedback from Telegram Stars currency |
| «اینترنت ماهوارهای استارلینک دارید؟» | stars_telegram ❌ |
network_vpn_proxy ✅ |
93.5% | Separates Starlink satellite equipment from Telegram Stars |
| «توکن باتفادر رو تو کانفیگ بذارم ران میشه؟» | network_vpn_proxy ❌ |
partnership_wholesale ✅ |
68.4% | Routes Telegram BotFather API token inquiries correctly |
| «دو تن سیمان بار ماشین کردیم» | network_vpn_proxy ❌ |
trolling_chitchat_spam ✅ |
97.9% | Separates physical metric ton weight from TON cryptocurrency |
| «لوکیشن دفتر رو روی گوگل مپ پین کن» | onboarding_guide ❌ |
trolling_chitchat_spam ✅ |
99.9% | Distinguishes Google Map GPS pins from pinned Telegram messages |
⏱ Inference Latency & Calibration
- Median Latency (p50):
139.3 ms(on standard Intel Xeon CPU). - 95th Percentile Latency (p95):
260.6 msunder high concurrency. - Calibrated Softmax Temperature: $T = 1.20$.
- Expected Calibration Error (ECE):
0.0592. - Confidence Safety Margin:
- Mean confidence on correct decisions: 98.5%.
- Mean confidence on ambiguous / edge queries: 69.2% (giving developers an unambiguous 30% margin to safely trigger bot fallback menus).
📁 Repository Structure (src/ & data/)
This repository is fully reproducible and structured cleanly:
├── assets/ # High-resolution infographics and charts
├── data/ # Benchmark & evaluation test suites
│ ├── benchmark_350_suite.jsonl # Complete 350-query blind production test set
│ ├── edge_ambiguity_suite.jsonl# 34 critical lexical boundary test cases
│ └── contrastive_sample.jsonl # Representative hard-negative training pairs
├── encoder/ # mmBERT backbone configuration
├── src/ # Standalone Python training & evaluation pipeline
│ ├── train.py # Warm-start fine-tuning & contrastive training script
│ └── evaluate.py # Standalone E2E benchmark evaluator
├── tokenizer/ # Tokenizer vocabulary and configs
├── model.safetensors # Calibrated neural network weights (bfloat16)
├── rl_agent_config.json # System-1 label map & temperature manifest
├── README.md # Official English documentation
├── README_FA.md # Comprehensive Persian documentation
└── LICENSE # MIT Open-Source License
🛠 Step-by-Step Training & Reproduction Guide
1. Requirements & Setup
pip install torch transformers datasets accelerate safetensors huggingface_hub
2. Reproduce the Production Benchmark in 1-Line:
You can directly evaluate the published model against the 350-query blind suite and 34 edge ambiguity cases:
python3 src/evaluate.py --model mmahdi-sz/Laya-fa-universal-support --device cpu
3. Running Custom Warm-Start Fine-Tuning:
python3 src/train.py \
--base mmahdi-sz/Laya-fa-universal-support \
--train ./data/train.jsonl \
--dev ./data/dev.jsonl \
--contrastive ./data/contrastive_sample.jsonl \
--epochs 2 \
--temperature 1.20 \
--out ./output_model_v31
4. Hyperparameters Reference:
- Base Architecture: mmBERT (ModernBERT multilingual, 322M parameters)
- Precision:
bfloat16 - Optimizer: AdamW (
weight_decay=0.01) - Learning Rate (Backbone):
3e-5with linear warmup - Learning Rate (Classification Head):
1e-4 - Batch Size:
16(per device) with gradient accumulation - Contrastive Oversampling: 2x on hard-negative pairs to firmly stamp out boundary collisions
🚀 Quickstart & Integration Examples
Python (Native httpx async client):
import httpx
async def get_bot_route(user_message: str):
async with httpx.AsyncClient() as client:
response = await client.post(
"http://127.0.0.1:8888/v1/systemone",
headers={"Authorization": "Bearer YOUR_SECRET_API_KEY"},
json={
"state": user_message,
"model": "laya-fa-universal",
"questions": {
"department": {
"type": "choice",
"instructions": "Route user message to the correct department",
"criteria": {
"billing_banking": "Banking, receipts, and card transfers",
"stars_telegram": "Telegram Stars currency purchases",
"ton_crypto_web3": "TON, crypto wallets, and miners",
"network_vpn_proxy": "VPNs, proxies, and configs",
"gaming_minecraft": "Minecraft servers and gaming ranks",
"trolling_chitchat_spam": "Spam, chit-chat, or ratings"
}
}
}
}
)
return response.json()
cURL Execution:
curl -X POST http://127.0.0.1:8888/v1/systemone \
-H "Authorization: Bearer YOUR_SECRET_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"state": "ماینر روشن نمیشه فنش کار نمیکنه",
"model": "laya-fa-universal",
"questions": {
"department": {
"type": "choice",
"instructions": "Select appropriate team"
}
}
}'
📦 Available Versions & Git Branches
To ensure backward compatibility, every major version of Laya-fa is preserved as an independent branch:
| Version Tag / Branch | Architecture | Target Scope | Key Milestones |
|---|---|---|---|
main / v3.1 |
mmBERT-322M | 12 Domains | Latest Production Release. 100% resolution of slangs, typos (کانفینگ), and subword collisions (ماینر vs ماین). |
v3.0 |
mmBERT-322M | 12 Domains | Introduced single-word and short acronym optimization (mc, star, stars). |
v2.0 |
mmBERT-322M | 12 Domains | Initial 12-domain universal expansion across 6,000 training samples. |
v1.0 |
mmBERT-322M | 4 Domains | Early fine-tuned baseline for Telegram star purchases. |
To pin an earlier version in code:
from transformers import AutoModel
# Pinning version 2.0 explicitly:
model = AutoModel.from_pretrained("mmahdi-sz/Laya-fa-universal-support", revision="v2.0")
🙏 Acknowledgments
Special thanks to @NotMmd for the initial inspiration, insightful architectural feedback, discovering nuanced Persian edge cases, and valuable contributions throughout the development and benchmarking of this model.
📜 Citation & License
This project is licensed under the MIT License.
@misc{laya_fa_universal_2026,
author = {Mahdi (mmahdi-sz)},
title = {Laya-fa Universal Support: Sub-140ms System-1 Intent Router for Persian E-Commerce},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/mmahdi-sz/Laya-fa-universal-support}}
}