Laya Experts

Laya Experts are fine-tuned versions ("flavours") of Laya. Each one is trained for a single job.

Laya doesn't write text. You give it some input and a fixed list of possible answers, and it returns a probability for each answer in one quick pass. The probabilities are calibrated, so when it says 90% it is right about 90% of the time. Out of the box Laya is a generalist and not very accurate on any one task. Fine-tuning on one task fixes that.

flavour what it does headline folder
pii Finds personal data in English text (names, phone numbers, emails, addresses, ID and card numbers, …) so you can redact it 0.967 entity F1 pii/
banking77 Works out what a banking customer is asking about, out of 77 intents 91.7% accuracy banking77/

🔒 Try the PII flavour: a Gradio demo lives in demo/. Run python demo/app.py after installing demo/requirements.txt.

Each flavour is a full copy of the 421M-parameter model (ModernBERT-large encoder plus Laya's decision head), fine-tuned on one NVIDIA A40 GPU. It runs on a laptop CPU.


Benchmarks

Every number comes from a held-out test set the model never saw during training. "Zero-shot Laya" is the original model before fine-tuning. Jev is TypeSafe's hosted model, which uses the same question-and-answer interface and makes a good reference point.

PII

The eval set is English documents from ai4privacy OpenPII: emails, forms, chat logs and medical notes, with personal data across 19 types. Documents that closely copied a training template were removed.

As a redactor, measured over whole documents:

model entity F1 recall PII characters missed documents with any leak extra text redacted
Zero-shot Laya 0.156 0.187 19.3% 54.6% 12.5%
Jev (hosted) 0.736 0.845 7.6% 43.0% 6.1%
DeBERTa-v3 tagger (trained on the same data) 0.964 0.969 0.41% 6.0% 0.13%
Laya Experts · pii 0.967 0.978 0.41% 7.0% 3.1%

As a classifier, one guess per candidate span (22 answers including "not pii"):

model accuracy macro-F1 top-3 calibration error (ECE)
Zero-shot Laya 16.2% 0.101 39.1% 0.791
Jev (hosted) 69.7% 0.561 94.3% 0.040
Laya Experts · pii 93.9% 0.863 98.5% 0.045

At finding personal data, it is level with a purpose-built DeBERTa tagger. Across two Laya seeds and four DeBERTa seeds, Laya was about half a point ahead every time. It is worse at drawing tight boundaries (below). If you only need to find PII, either works. If you need exact character offsets, the tagger is better.

model entity F1, exact boundaries
DeBERTa-v3 tagger 0.916
Laya Experts · pii 0.695
Per-type results (Laya Experts · pii)
type precision recall F1
date 0.989 1.000 0.995
given name 0.914 0.993 0.952
surname 0.941 0.954 0.948
email address 1.000 1.000 1.000
city 0.976 0.976 0.976
title 0.992 1.000 0.996
phone number 0.995 0.995 0.995
age 0.967 0.962 0.964
street 0.967 1.000 0.983
building number 0.921 0.953 0.937
postal code 0.969 1.000 0.984
id card number 0.951 0.893 0.921
credit card number 1.000 1.000 1.000
driver licence number 0.855 0.926 0.889
tax number 0.892 0.952 0.921
gender 0.942 0.961 0.952
passport number 0.933 0.972 0.952
social security number 0.960 0.986 0.973
sex 0.900 0.940 0.920

If you merge the look-alike types (every government ID number as one type, and given name, surname and title as "person name"), every group scores 0.96 F1 or higher.

Banking77

Customer messages from the Banking77 test set, 77 intents.

model accuracy macro-F1 top-3 top-5 calibration error (ECE)
Zero-shot Laya 39.4% 0.376 50.0% 53.9% 0.472
Jev (hosted) 79.8% 0.790 91.8% 93.6% 0.089
Laya Experts · banking77 91.7% 0.916 97.4% 98.6% 0.009

Speed

Measured on an Apple M3 laptop. The model runs locally and costs nothing per call. Jev times include the network round trip from Singapore to the US.

Laya Experts, p50 Laya Experts, p99 Jev, p50 Jev, p99
one PII span 87 ms 131 ms 344 ms 805 ms
one Banking77 message 208 ms 382 ms 309 ms 760 ms

A PII document is about 30 candidate spans, so a paragraph takes a few seconds on CPU. Batching the spans, as pii_demo.py does, and running on a GPU both help a lot.

Recommended hardware

Each flavour is 421M parameters, about 850 MB on disk. You don't need a GPU.

minimum recommended
memory (RAM) 4 GB free 8 GB or more
processor any recent 64-bit CPU Apple silicon (M1 or later) or a modern x86 CPU
GPU not needed any NVIDIA GPU with 4 GB+ of memory, or Apple silicon via mps
disk 1 GB per flavour 2 GB for both flavours

Measured on an Apple M3 with 16 GB: the model uses about 3.1 GB of memory at peak, PyTorch included. A 27-span PII document takes 3.8 s on CPU and 2.3 s on the Apple GPU (device="mps"). On an NVIDIA GPU the model runs in bfloat16 (float16 on cards older than the Ampere generation) and needs well under 2 GB of GPU memory. That figure is an estimate, not a measurement.

To fine-tune your own flavour, a GPU with about 24 GB makes training all layers comfortable. That is an estimate. Smaller cards work with a smaller batch and gradient accumulation. These flavours were trained on a 48 GB NVIDIA A40.


How to use

pip install "laya>=0.3.6,!=0.3.7"

Every flavour loads the same way. Pick the folder with subfolder, and only that folder is downloaded.

Banking77

import json, laya
from huggingface_hub import hf_hub_download

agent = laya.Agent("goku-san/laya-experts", subfolder="banking77")
spec = json.load(open(hf_hub_download("goku-san/laya-experts", "banking77/labels.json")))

# The model was trained on short option names ("card arrival"); map them back to labels.
back = {alias: label for label, alias in spec["aliases"].items()}
question = {"intent": {"type": "choice", "instructions": spec["instructions"],
                       "criteria": list(spec["aliases"].values())}}

answer = agent.predict("I still haven't received my new card", question)["answers"]["intent"]
print(back[answer["choice"]], answer["confidence"])   # card_arrival 0.88

PII

PII detection has two steps. A rule-based finder suggests every span that could be personal data, and the model labels each one. The finder has to be the same one used in training, so it ships as a single file, demo/pii_demo.py in this repo.

from pii_demo import PiiRedactor

redactor = PiiRedactor()          # loads goku-san/laya-experts, subfolder "pii"
result = redactor.detect(
    "Hi, I'm Priya Nair. Call me on +44 7700 900123 or email priya.nair@example.com.",
    threshold=0.5,                # lower catches more and redacts more
    region="GB",                  # US, GB, CA, IN or SG
)
print(result.redacted)
# Hi, I'm [GIVEN NAME] [SURNAME]. Call me on [PHONE NUMBER] or email [EMAIL ADDRESS].
for e in result.entities:
    print(e.text, e.label, round(e.score, 3))

To call the model directly on one span, the input is a JSON object laid out exactly as below. pii/labels.json holds the instruction and the 22 answers.

state = {"document_type": "email", "language": "English", "region": "GB",
         "left": "Kind regards,\n", "span": "Oliver", "right": " Bennett\nOperations Lead"}

Apple silicon

The same weights run under laya-mlx on an M-series Mac. They give the same answers as PyTorch: 100% argmax agreement, with probabilities within 0.005.


Limits

  • English only, and the PII model was trained on synthetic documents (ai4privacy). Test it on your own data before trusting it with real records.
  • The PII rule-based finder catches 99.9% of the personal data in the eval set. Anything it misses, the model never sees.
  • Look-alike numbers sometimes get the wrong type, for example a space-separated card number labelled as a tax number. They are still redacted. Sex vs gender and the different government ID types are the least reliable labels.
  • Redacted spans tend to be a little wider than needed: 3.1% of non-personal text gets redacted, against 0.13% for a DeBERTa tagger.
  • Banking77 only knows its 77 intents. For a message about anything else it will still pick one of them.
  • Input is capped at 512 tokens, including the answer list. The PII flavour sends at most 400 characters of context either side of each span.

Training

pii banking77
base convaiinnovations/laya convaiinnovations/laya
training data candidate spans from ai4privacy training documents Banking77 train split (10% held out for validation)
layers trained all 28 (421M parameters) all 28 (421M parameters)
epochs / learning rate / batch 3 / 2e-5 / 32 4 / 1e-5 / 16
hardware / time 1× A40, 96 min 1× A40, 15 min

Each folder also has the full training record in finetune.json. Training refits the probability temperature, which is most of the jump in calibration.

License

Apache-2.0, the same as the Laya base model. Training data: ai4privacy OpenPII and Banking77 (CC-BY-4.0). Please follow their terms too.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for goku-san/laya-experts

Finetuned
(120)
this model

Datasets used to train goku-san/laya-experts

Evaluation results