Laya Experts
Laya Experts are fine-tuned versions ("flavours") of Laya. Each one is trained for a single job.
Laya doesn't write text. You give it some input and a fixed list of possible answers, and it returns a probability for each answer in one quick pass. The probabilities are calibrated, so when it says 90% it is right about 90% of the time. Out of the box Laya is a generalist and not very accurate on any one task. Fine-tuning on one task fixes that.
| flavour | what it does | headline | folder |
|---|---|---|---|
| pii | Finds personal data in English text (names, phone numbers, emails, addresses, ID and card numbers, …) so you can redact it | 0.967 entity F1 | pii/ |
| banking77 | Works out what a banking customer is asking about, out of 77 intents | 91.7% accuracy | banking77/ |
🔒 Try the PII flavour: a Gradio demo lives in demo/. Run python demo/app.py after installing demo/requirements.txt.
Each flavour is a full copy of the 421M-parameter model (ModernBERT-large encoder plus Laya's decision head), fine-tuned on one NVIDIA A40 GPU. It runs on a laptop CPU.
Benchmarks
Every number comes from a held-out test set the model never saw during training. "Zero-shot Laya" is the original model before fine-tuning. Jev is TypeSafe's hosted model, which uses the same question-and-answer interface and makes a good reference point.
PII
The eval set is English documents from ai4privacy OpenPII: emails, forms, chat logs and medical notes, with personal data across 19 types. Documents that closely copied a training template were removed.
As a redactor, measured over whole documents:
| model | entity F1 | recall | PII characters missed | documents with any leak | extra text redacted |
|---|---|---|---|---|---|
| Zero-shot Laya | 0.156 | 0.187 | 19.3% | 54.6% | 12.5% |
| Jev (hosted) | 0.736 | 0.845 | 7.6% | 43.0% | 6.1% |
| DeBERTa-v3 tagger (trained on the same data) | 0.964 | 0.969 | 0.41% | 6.0% | 0.13% |
| Laya Experts · pii | 0.967 | 0.978 | 0.41% | 7.0% | 3.1% |
As a classifier, one guess per candidate span (22 answers including "not pii"):
| model | accuracy | macro-F1 | top-3 | calibration error (ECE) |
|---|---|---|---|---|
| Zero-shot Laya | 16.2% | 0.101 | 39.1% | 0.791 |
| Jev (hosted) | 69.7% | 0.561 | 94.3% | 0.040 |
| Laya Experts · pii | 93.9% | 0.863 | 98.5% | 0.045 |
At finding personal data, it is level with a purpose-built DeBERTa tagger. Across two Laya seeds and four DeBERTa seeds, Laya was about half a point ahead every time. It is worse at drawing tight boundaries (below). If you only need to find PII, either works. If you need exact character offsets, the tagger is better.
| model | entity F1, exact boundaries |
|---|---|
| DeBERTa-v3 tagger | 0.916 |
| Laya Experts · pii | 0.695 |
Per-type results (Laya Experts · pii)
| type | precision | recall | F1 |
|---|---|---|---|
| date | 0.989 | 1.000 | 0.995 |
| given name | 0.914 | 0.993 | 0.952 |
| surname | 0.941 | 0.954 | 0.948 |
| email address | 1.000 | 1.000 | 1.000 |
| city | 0.976 | 0.976 | 0.976 |
| title | 0.992 | 1.000 | 0.996 |
| phone number | 0.995 | 0.995 | 0.995 |
| age | 0.967 | 0.962 | 0.964 |
| street | 0.967 | 1.000 | 0.983 |
| building number | 0.921 | 0.953 | 0.937 |
| postal code | 0.969 | 1.000 | 0.984 |
| id card number | 0.951 | 0.893 | 0.921 |
| credit card number | 1.000 | 1.000 | 1.000 |
| driver licence number | 0.855 | 0.926 | 0.889 |
| tax number | 0.892 | 0.952 | 0.921 |
| gender | 0.942 | 0.961 | 0.952 |
| passport number | 0.933 | 0.972 | 0.952 |
| social security number | 0.960 | 0.986 | 0.973 |
| sex | 0.900 | 0.940 | 0.920 |
If you merge the look-alike types (every government ID number as one type, and given name, surname and title as "person name"), every group scores 0.96 F1 or higher.
Banking77
Customer messages from the Banking77 test set, 77 intents.
| model | accuracy | macro-F1 | top-3 | top-5 | calibration error (ECE) |
|---|---|---|---|---|---|
| Zero-shot Laya | 39.4% | 0.376 | 50.0% | 53.9% | 0.472 |
| Jev (hosted) | 79.8% | 0.790 | 91.8% | 93.6% | 0.089 |
| Laya Experts · banking77 | 91.7% | 0.916 | 97.4% | 98.6% | 0.009 |
Speed
Measured on an Apple M3 laptop. The model runs locally and costs nothing per call. Jev times include the network round trip from Singapore to the US.
| Laya Experts, p50 | Laya Experts, p99 | Jev, p50 | Jev, p99 | |
|---|---|---|---|---|
| one PII span | 87 ms | 131 ms | 344 ms | 805 ms |
| one Banking77 message | 208 ms | 382 ms | 309 ms | 760 ms |
A PII document is about 30 candidate spans, so a paragraph takes a few seconds on CPU. Batching
the spans, as pii_demo.py does, and running on a GPU both help a lot.
Recommended hardware
Each flavour is 421M parameters, about 850 MB on disk. You don't need a GPU.
| minimum | recommended | |
|---|---|---|
| memory (RAM) | 4 GB free | 8 GB or more |
| processor | any recent 64-bit CPU | Apple silicon (M1 or later) or a modern x86 CPU |
| GPU | not needed | any NVIDIA GPU with 4 GB+ of memory, or Apple silicon via mps |
| disk | 1 GB per flavour | 2 GB for both flavours |
Measured on an Apple M3 with 16 GB: the model uses about 3.1 GB of memory at peak,
PyTorch included. A 27-span PII document takes 3.8 s on CPU and 2.3 s on the Apple GPU
(device="mps"). On an NVIDIA GPU the model runs in bfloat16 (float16 on cards older than the
Ampere generation) and needs well under 2 GB of GPU memory. That figure is an estimate, not a
measurement.
To fine-tune your own flavour, a GPU with about 24 GB makes training all layers comfortable. That is an estimate. Smaller cards work with a smaller batch and gradient accumulation. These flavours were trained on a 48 GB NVIDIA A40.
How to use
pip install "laya>=0.3.6,!=0.3.7"
Every flavour loads the same way. Pick the folder with subfolder, and only that folder is
downloaded.
Banking77
import json, laya
from huggingface_hub import hf_hub_download
agent = laya.Agent("goku-san/laya-experts", subfolder="banking77")
spec = json.load(open(hf_hub_download("goku-san/laya-experts", "banking77/labels.json")))
# The model was trained on short option names ("card arrival"); map them back to labels.
back = {alias: label for label, alias in spec["aliases"].items()}
question = {"intent": {"type": "choice", "instructions": spec["instructions"],
"criteria": list(spec["aliases"].values())}}
answer = agent.predict("I still haven't received my new card", question)["answers"]["intent"]
print(back[answer["choice"]], answer["confidence"]) # card_arrival 0.88
PII
PII detection has two steps. A rule-based finder suggests every span that could be personal
data, and the model labels each one. The finder has to be the same one used in training, so
it ships as a single file, demo/pii_demo.py
in this repo.
from pii_demo import PiiRedactor
redactor = PiiRedactor() # loads goku-san/laya-experts, subfolder "pii"
result = redactor.detect(
"Hi, I'm Priya Nair. Call me on +44 7700 900123 or email priya.nair@example.com.",
threshold=0.5, # lower catches more and redacts more
region="GB", # US, GB, CA, IN or SG
)
print(result.redacted)
# Hi, I'm [GIVEN NAME] [SURNAME]. Call me on [PHONE NUMBER] or email [EMAIL ADDRESS].
for e in result.entities:
print(e.text, e.label, round(e.score, 3))
To call the model directly on one span, the input is a JSON object laid out exactly as below.
pii/labels.json holds the instruction and the 22 answers.
state = {"document_type": "email", "language": "English", "region": "GB",
"left": "Kind regards,\n", "span": "Oliver", "right": " Bennett\nOperations Lead"}
Apple silicon
The same weights run under laya-mlx on an M-series Mac.
They give the same answers as PyTorch: 100% argmax agreement, with probabilities within 0.005.
Limits
- English only, and the PII model was trained on synthetic documents (ai4privacy). Test it on your own data before trusting it with real records.
- The PII rule-based finder catches 99.9% of the personal data in the eval set. Anything it misses, the model never sees.
- Look-alike numbers sometimes get the wrong type, for example a space-separated card number labelled as a tax number. They are still redacted. Sex vs gender and the different government ID types are the least reliable labels.
- Redacted spans tend to be a little wider than needed: 3.1% of non-personal text gets redacted, against 0.13% for a DeBERTa tagger.
- Banking77 only knows its 77 intents. For a message about anything else it will still pick one of them.
- Input is capped at 512 tokens, including the answer list. The PII flavour sends at most 400 characters of context either side of each span.
Training
| pii | banking77 | |
|---|---|---|
| base | convaiinnovations/laya |
convaiinnovations/laya |
| training data | candidate spans from ai4privacy training documents | Banking77 train split (10% held out for validation) |
| layers trained | all 28 (421M parameters) | all 28 (421M parameters) |
| epochs / learning rate / batch | 3 / 2e-5 / 32 | 4 / 1e-5 / 16 |
| hardware / time | 1× A40, 96 min | 1× A40, 15 min |
Each folder also has the full training record in finetune.json. Training refits the
probability temperature, which is most of the jump in calibration.
License
Apache-2.0, the same as the Laya base model. Training data: ai4privacy OpenPII and Banking77 (CC-BY-4.0). Please follow their terms too.
Model tree for goku-san/laya-experts
Base model
convaiinnovations/layaDatasets used to train goku-san/laya-experts
ai4privacy/pii-masking-openpii-1.5m
Evaluation results
- entity micro-F1 on ai4privacy OpenPII (English, held-out documents)self-reported0.967
- entity macro-F1 on ai4privacy OpenPII (English, held-out documents)self-reported0.961
- span accuracy (22-way) on ai4privacy OpenPII (English, held-out documents)self-reported0.939
- accuracy on Banking77 (test)test set self-reported0.917
- macro-F1 on Banking77 (test)test set self-reported0.916