Instructions to use CraneAILabs/afrimed-pii-afriberta with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use CraneAILabs/afrimed-pii-afriberta with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="CraneAILabs/afrimed-pii-afriberta")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("CraneAILabs/afrimed-pii-afriberta", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- AfriMed-PII-AfriBERTa: clinical PII de-identification for Luganda, Kiswahili, Kinyarwanda and Dholuo
- Model details
- Categories
- Uses
- How to use
- Training details
- Evaluation
- Evaluation data
- Metrics
- Results, held-out translated-clinical test (model only)
- Results, names never seen in training
- Results, tribe, religion and clan (synthetic dev; in-distribution ceiling)
- Out-of-domain
- Comparison with off-the-shelf systems
- Effect of a rule layer and threshold calibration (not part of these weights)
- Evaluation on real clinical records
- Transfer to other languages
- Bias, risks and limitations
- Environmental impact
- Technical specifications
- Citation
- Funding
AfriMed-PII-AfriBERTa: clinical PII de-identification for Luganda, Kiswahili, Kinyarwanda and Dholuo
AfriBERTa (Ogueji, Zhu and Lin, 2021; University of Waterloo, released through the Castorini group) is an encoder pretrained from scratch on under 1 GB of text in 11 African languages, including Kiswahili and Kinyarwanda (as Gahuza) but not Luganda or Dholuo. Of the three backbones released in this series it is the smallest and, model-for-model, has the highest patient-name recall in every language; its precision is the lowest of the three, so it finds the most identifiers and marks the most non-identifier text in doing so. The tables below quantify both.
This release contains four per-language checkpoints fine-tuned from castorini/afriberta_large for
token-level detection of personal information in clinical text. Each checkpoint emits character-offset
spans with one of 14 categories. The model outputs positions and labels; it does not generate text.
Model details
| Developed by | Crane AI Labs, under Gates Foundation grant INV-107957 (Workstream 1) |
| Model type | Transformer encoder with a token-classification head; BIO tagging over 14 categories (29 tags) |
| Base model | castorini/afriberta_large (126M parameters, MIT); Hub revision 231d92411ab8a99add67eda412ee948c74a4b179 |
| Languages | Luganda (lug), Kiswahili (swa), Kinyarwanda (kin), Dholuo (luo) |
| Licence | Apache-2.0 for the fine-tuned weights; the base model's MIT notice is retained |
| Released | 2026-09 |
Checkpoints in this release
| Language | Folder | Run | Base | Learning rate |
|---|---|---|---|---|
| Luganda | lug/ |
afriberta_large__lug__seed42__opt3_age_lr5e-5 |
castorini/afriberta_large (released weights) | 5e-5 |
| Kiswahili | swa/ |
afriberta_large__swa__seed42__opt3_age_lr3e-5 |
castorini/afriberta_large (released weights) | 3e-5 |
| Kinyarwanda | kin/ |
afriberta_large__kin__seed42__opt3_age_lr1e-5 |
castorini/afriberta_large (released weights) | 1e-5 |
| Dholuo | luo/ |
afriberta_large__luo__seed42__opt3_age_lr3e-5 |
castorini/afriberta_large (released weights) | 3e-5 |
Each checkpoint is the seed selected on the development split (see Model selection). The remaining seeds are retained and are available on request; every figure below is a mean over all five.
Dholuo was included as a deliberate stress case. The other three languages are the jurisdiction-aligned core; Dholuo is Nilotic where they are Bantu, has one to two orders of magnitude less monolingual text (about 5 to 6 million tokens of licence-clean text in total), and no backbone in this series was pretrained on it. Its results therefore bound what the approach can do for a genuinely low-resource language of a different family, rather than measuring it under favourable conditions.
Categories
The 14 categories follow the project's medical-PII taxonomy (D3). Identifier categories with a fixed surface form, national ID, phone, date, record number, age, are detectable by pattern; the remaining categories require the model.
| Category | Description |
|---|---|
PERSON |
Patient's name |
PERSON_THIRD_PARTY |
Any other person's name, relative, next of kin, clinician |
ID_NATIONAL |
National identity number (UG NIN, KE ID, RW NID, TZ NIDA) |
ID_RECORD |
Facility or patient record number |
CONTACT_PHONE |
Telephone number |
CONTACT_OTHER |
Email address, handle, or other contact string |
LOCATION_VILLAGE |
Village, parish, district or other administrative place |
LOCATION_FACILITY |
Named health facility |
DATE |
Calendar date |
AGE |
Patient age, including its unit word |
PROFESSION |
Occupation |
TRIBE |
Ethnic or tribal affiliation |
RELIGION |
Religious affiliation |
PERSON_CLAN |
Clan name |
Uses
The model is a component for de-identifying clinical free text in the four languages. Its measured operating envelope is given in Evaluation; the per-category tables there describe what proportion of each kind of identifier it finds and how much non-identifier text it marks in doing so, and are the basis on which to judge fitness for a given workflow.
It performs best on clinical-register text. Its behaviour on other registers was measured on news text and is reported under Out-of-domain.
How to use
from transformers import AutoTokenizer, AutoModelForTokenClassification
repo = "CraneAILabs/afrimed-pii-afriberta"
lang = "lug" # or "swa", "kin", "luo"
tok = AutoTokenizer.from_pretrained(repo, subfolder=lang)
mdl = AutoModelForTokenClassification.from_pretrained(repo, subfolder=lang)
The models were trained over a whitespace-and-punctuation word grid and expect input pre-split into
words (tokenizer(words, is_split_into_words=True)), with predictions aligned back to character
offsets. A reference implementation of that alignment, together with an optional rule layer for the
fixed-format identifiers, is provided in deid.py in this repository. The three calls above were verified
to load and run for all four language checkpoints on transformers 5.13 / torch 2.13 with no
trust_remote_code and no additional files; warm inference on a 219-character record takes ~6 ms on a GB10
GPU and ~100 ms on CPU.
The system on a clinical note
The screenshots below are the deployed system (AfriBERTa checkpoints, rule layer and calibration) on a short synthetic clinical note per language, captured from the live demo. Category labels replace the spans the system found. They are shown unedited, including what the system gets wrong, because that is what a user of these models will see.
In the Luganda note the system finds the patient's name, the age, the national ID, the phone number, the date, both places and the father's name; nothing identifying is left in the text. Across the four notes the rule layer takes the age, the national ID, the phone number and the date in every language, the template layer takes the named facility in Kinyarwanda, Kiswahili and Dholuo, and the model takes the names and places. Two errors are visible and both are the model's: in the Dholuo note the district Nyando is typed as a third-party name (the location/person confusion reported under Bias, risks and limitations), and the patient-name span swallows the preceding word Jatuo (patient). Both still redact the identifier; the cost is a wrong label and one extra word.
Two properties of the rule layer matter when reading these notes. Age is matched in both word orders (46 emyaka and emyaka 46); the over-matches this produces are durations such as after 10 years of follow-up, which are not identifiers and are accepted as the cost of recall 1.000. Facility names in natural phrasing (Kiswahili Hospitali ya, Zahanati ya, Kituo cha Afya; Kinyarwanda Ibitaro bya, Ikigo nderabuzima cya; Dholuo osipital mar; Luganda eddwaliro lya) are labelled by the template layer. The translated test writes facility names in their English form, so these phrases never fire there (0 matches in 707 documents): the system figures in this card do not depend on them, and their recall on natural text is a claim from the naming conventions, not a measurement.
The full-page captures for all four languages, with the exact input and output text, are in figures/demo/.
Training details
Training data
The models were fine-tuned on a synthetic clinical corpus (internal designation D5, version 1): 8,000 training and 1,000 development documents per language, 36,000 in total. Documents are clinical-style records generated from language-specific templates, filled from curated East African vocabularies of given and family names, villages, districts and health facilities, and country-specific identifier formats for Uganda, Kenya, Rwanda and Tanzania. A native-speaker review of 600 documents was conducted across the four languages, and the Kinyarwanda partition was regenerated with reviewer-sourced lexical terms. Every span is labelled with one of the 14 categories above.
For this backbone the corpus was additionally regenerated so that age spans include their unit word (e.g. 46 emyaka, 46 miaka), matching the annotation convention of the evaluation corpus; the original convention labelled the bare number. This is the only change from the corpus used for the other two backbones and it accounts for the age-category results.
No real patient data was used. Evaluation corpora, including MasakhaNER, which is licensed for non-commercial use, were used for evaluation and development only and never for training. The training corpus itself is not part of this release.
Kinyarwanda tribe and clan are out of scope, by ruling. All three backbones score 0.000 on Kinyarwanda TRIBE and PERSON_CLAN on the synthetic evaluation set. This is not a detection failure: the project's taxonomy ruling removed both categories for Kinyarwanda, Rwandan clinical records do not record Hutu/Tutsi/Twa under post-1994 law and policy, and clan has a Luganda-register basis, so the Kinyarwanda training corpus the released weights learned from contains no spans of either. The evaluation edition reported here is an earlier one that still annotates 76 tribe and 35 clan spans; the zeros record that the released Kinyarwanda checkpoint was never trained to emit those labels, which is the intended behaviour. Kinyarwanda RELIGION, which remains in scope, scores 0.86 to 0.93.
Training procedure
Each language is trained separately as a token-classification model: the released base encoder with a randomly initialised linear head over 29 BIO tags, trained with cross-entropy over a whitespace-and-punctuation word grid (first sub-token of each word carries the label). The base weights are the public castorini/afriberta_large checkpoint; no continued pretraining was applied to this backbone, because in a controlled comparison it did not measurably help an encoder already pretrained on African-language text. Five seeds were trained per language.
Hyperparameters
| Setting | Value |
|---|---|
| Learning rate (per language) | Luganda 5e-5, Kiswahili 3e-5, Kinyarwanda 1e-5, Dholuo 3e-5 |
| Epochs | 6, with early stopping on the development split (patience 2) |
| Batch size | 32 (no gradient accumulation) |
| Max sequence length | 256 |
| Warm-up ratio | 0.1 |
| Weight decay | 0.01 |
| Schedule | linear |
| Precision | bf16 |
| Optimiser | AdamW (Transformers defaults) |
| Max gradient norm | 1.0 |
| Seeds | 13, 42, 99, 123, 2026 |
Model selection
The learning rate was swept per language over {1e-5, 2e-5, 3e-5, 5e-5}, and the (learning rate, epoch) pair was chosen on a real held-out development set, the MasakhaNER validation split plus the 30% development slice of the translated-clinical corpus, by per-category macro-recall over categories with at least ten gold spans (age excluded as an annotation-convention artefact at the time). The synthetic development split, on which all configurations saturate near F1 0.99, was deliberately not used for selection because it does not predict performance on real text. The test split was never used for any selection. The released checkpoint per language is the seed selected on that development set; all five seeds are reported.
Compute
Fine-tuning and the learning-rate sweep ran on a single NVIDIA DGX Spark (GB10). The complete lane for this backbone, the learning-rate sweep, two ablated levers, the AGE-convention regeneration and five-seed finals for four languages, with scoring, used 7.0 GPU-hours as logged by the lane; 5-seed training of the four released configurations alone accounts for roughly 1.3 of those.
Evaluation
Evaluation data
Three evaluation sets are reported. Together they cover every category and separate three questions, performance on the clinical register, generalisation to names never seen in training, and behaviour out of domain.
- Held-out translated-clinical test (
teval). The MEDDOCAN Spanish clinical corpus (Marimon et al., 2019) machine-translated into each of the four languages with span offsets re-aligned; split 30% development / 70% test by document index. Test size per language: Luganda 161 docs / 3,351 spans; Kiswahili 163 / 3,387; Kinyarwanda 158 / 3,278; Dholuo 168 / 3,487. The source schema has no tribe, religion or clan annotations, so those categories are evaluated on set 3. - Name-disjoint variant (
teval-novel). The same test documents with every person-name span replaced by a well-formed surrogate drawn from a Wikidata (CC0) name pool whose tokens are disjoint from the training vocabulary, verified at 0.0% overlap in all four languages against nine sources, with document and span counts preserved exactly. A reshuffle control (original names permuted across documents) separates the effect of unfamiliar names from that of merely different names. - Synthetic development set (
d5-dev), an in-distribution ceiling, not a held-out score. 1,000 documents per language from the synthetic generator. These are the documents on which training early-stopped (best epoch selected by F1 on this set), so no training document appears in it but the released weights are, by construction, the epoch that scored best here. Its figures are therefore an upper bound for in-distribution synthetic text. It is the only set carrying gold spans forTRIBE,RELIGIONandPERSON_CLAN, and is reported for those categories only, labelled as such.
Out-of-domain behaviour is measured on the MasakhaNER 1.0 test split (news text; PERSON, LOCATION,
DATE), which is used for evaluation only.
Metrics
Scores are entity-level, strict: a predicted span counts only if its character offsets and category both match the gold span exactly. A relaxed score (any character overlap with matching category) is also reported where it was computed, as a diagnostic of boundary behaviour: the gap between the two is the share of identifiers the model finds but mis-delimits. All figures are micro averages unless labelled per category, and are means over five training seeds with the seed standard deviation.
We report precision, recall and F1 together, because in de-identification each answers a different question, and this follows the convention of the field's shared tasks. The 2014 i2b2/UTHealth track reported entity-level micro precision, recall and F1 with F1 as the primary metric for comparing systems; MEDDOCAN 2019 reported the same and additionally a leak score, the proportion of identifiers missed, as an explicit safety measure, observing that systems tend to favour precision over recall. Accordingly:
- Recall is the safety metric: an identifier the model does not find remains in the text.
- Precision is the utility metric: text wrongly marked as an identifier is removed unnecessarily.
- F1 is the comparison metric between systems.
- Leak = 1 − recall, reported per category. This is MEDDOCAN's safety measure and is included so the tables are read in the same terms as that benchmark.
Strict versus relaxed, and why both appear. Under the strict criterion a prediction counts only if its character offsets and its category both match the gold span exactly; this is the criterion the shared tasks use and it is the headline everywhere in this card. Under the relaxed criterion a gold span counts as found if some prediction of the same category overlaps it by at least one character. Relaxed is not an alternative score and is never compared against strict figures from elsewhere; it is a diagnostic that separates two different failures. When strict is low and relaxed is high, the model found the identifier but drew its boundary differently (an attached determiner, included punctuation, a tokeniser split); when both are low, it did not find it. A relaxed value of 1.000 therefore means every gold span of that category was touched by a correctly-typed prediction, which is common for well-formed categories such as dates and is the reason those values are not rare here. The clearest case in this series is age on the XLM-R checkpoints: strict 0.00, relaxed 0.98 to 0.99, because the model tags the number and the gold span includes the unit word.
Seed-to-seed variation in name recall spans SD 0.008-0.094 across languages; per-language differences smaller than about 0.05 are within that variation and are not interpreted.
Results, held-out translated-clinical test (model only)
These tables describe the released weights alone: no rule layer and no threshold adjustment. Each row gives the number of gold spans, so low-count categories can be weighed accordingly.
Luganda, held-out translated-clinical test, model only, 5 seeds
| Category | Gold spans | Precision | Recall (strict) | Recall (relaxed) | F1 | Leak (1 − R) |
|---|---|---|---|---|---|---|
| Patient name | 324 | 0.550 ±.03 | 0.612 ±.06 | 0.651 ±.06 | 0.578 ±.03 | 0.388 |
| Third-party name | 379 | 0.177 ±.02 | 0.400 ±.02 | 0.760 ±.03 | 0.245 ±.02 | 0.601 |
| National ID | 126 | 0.761 ±.04 | 0.935 ±.06 | 0.935 ±.06 | 0.838 ±.04 | 0.065 |
| Record number | 359 | 0.757 ±.10 | 0.977 ±.01 | 1.000 ±.00 | 0.850 ±.07 | 0.023 |
| Phone | 25 | 0.752 ±.08 | 0.768 ±.02 | 0.880 ±.00 | 0.758 ±.05 | 0.232 |
| Other contact | 154 | 0.529 ±.08 | 0.716 ±.02 | 0.853 ±.00 | 0.605 ±.05 | 0.284 |
| Village / area | 1143 | 0.778 ±.03 | 0.518 ±.04 | 0.597 ±.04 | 0.621 ±.03 | 0.482 |
| Facility | 124 | 0.188 ±.04 | 0.315 ±.05 | 0.805 ±.02 | 0.235 ±.04 | 0.685 |
| Date | 377 | 0.797 ±.12 | 0.887 ±.13 | 0.966 ±.06 | 0.839 ±.12 | 0.113 |
| Age | 336 | 0.541 ±.02 | 0.980 ±.00 | 1.000 ±.00 | 0.697 ±.02 | 0.020 |
| Profession | 4 | 0.001 ±.00 | 0.250 ±.22 | 0.250 ±.22 | 0.003 ±.00, n=3 | 0.750 |
| Micro (all gold-bearing) | 0.443 ±.03 | 0.669 ±.02 | 0.780 ±.02 | 0.533 ±.03 | 0.331 |
Not in this table: tribe, religion, clan. The translated test carries no gold spans for them (the source schema has none); they are scored on the synthetic development set under Results, tribe, religion and clan below.
Kiswahili, held-out translated-clinical test, model only, 5 seeds
| Category | Gold spans | Precision | Recall (strict) | Recall (relaxed) | F1 | Leak (1 − R) |
|---|---|---|---|---|---|---|
| Patient name | 328 | 0.634 ±.06 | 0.706 ±.09 | 0.809 ±.05 | 0.665 ±.06 | 0.294 |
| Third-party name | 374 | 0.355 ±.03 | 0.489 ±.05 | 0.786 ±.09 | 0.410 ±.03 | 0.511 |
| National ID | 127 | 0.976 ±.02 | 0.721 ±.22 | 0.721 ±.22 | 0.812 ±.16 | 0.279 |
| Record number | 361 | 0.693 ±.12 | 0.981 ±.01 | 1.000 ±.00 | 0.807 ±.08 | 0.019 |
| Phone | 22 | 0.901 ±.02 | 0.991 ±.02 | 1.000 ±.00 | 0.944 ±.02 | 0.009 |
| Other contact | 157 | 0.567 ±.04 | 0.850 ±.00 | 1.000 ±.00 | 0.679 ±.03 | 0.150 |
| Village / area | 1152 | 0.700 ±.01 | 0.633 ±.02 | 0.751 ±.03 | 0.664 ±.01 | 0.367 |
| Facility | 142 | 0.114 ±.01 | 0.522 ±.02 | 1.000 ±.00 | 0.187 ±.02 | 0.478 |
| Date | 383 | 0.897 ±.03 | 0.992 ±.01 | 0.999 ±.00 | 0.942 ±.02 | 0.008 |
| Age | 336 | 0.795 ±.03 | 0.992 ±.00 | 0.999 ±.00 | 0.882 ±.02 | 0.008 |
| Profession | 5 | 0.011 ±.01 | 0.480 ±.10 | 0.880 ±.10 | 0.021 ±.01 | 0.520 |
| Micro (all gold-bearing) | 0.555 ±.02 | 0.748 ±.01 | 0.862 ±.02 | 0.637 ±.01 | 0.252 |
Not in this table: tribe, religion, clan. The translated test carries no gold spans for them (the source schema has none); they are scored on the synthetic development set under Results, tribe, religion and clan below.
Kinyarwanda, held-out translated-clinical test, model only, 5 seeds
| Category | Gold spans | Precision | Recall (strict) | Recall (relaxed) | F1 | Leak (1 − R) |
|---|---|---|---|---|---|---|
| Patient name | 318 | 0.451 ±.05 | 0.504 ±.01 | 0.526 ±.01 | 0.474 ±.03 | 0.496 |
| Third-party name | 361 | 0.123 ±.03 | 0.378 ±.03 | 0.859 ±.01 | 0.184 ±.04 | 0.622 |
| National ID | 124 | 0.993 ±.01 | 0.150 ±.13 | 0.150 ±.13 | 0.240 ±.18 | 0.850 |
| Record number | 344 | 0.287 ±.04 | 0.974 ±.01 | 1.000 ±.00 | 0.443 ±.04 | 0.026 |
| Phone | 19 | 0.108 ±.04 | 0.853 ±.04 | 1.000 ±.00 | 0.189 ±.06 | 0.147 |
| Other contact | 159 | 0.131 ±.06 | 0.728 ±.04 | 0.994 ±.00 | 0.215 ±.09 | 0.272 |
| Village / area | 1123 | 0.649 ±.04 | 0.611 ±.03 | 0.638 ±.03 | 0.628 ±.03 | 0.389 |
| Facility | 117 | 0.141 ±.06 | 0.595 ±.03 | 0.992 ±.01 | 0.221 ±.08 | 0.405 |
| Date | 385 | 0.558 ±.07 | 0.813 ±.08 | 0.934 ±.04 | 0.660 ±.07 | 0.187 |
| Age | 325 | 0.469 ±.04 | 0.945 ±.04 | 0.993 ±.01 | 0.626 ±.04 | 0.055 |
| Profession | 3 | 0.001 ±.00 | 0.600 ±.13 | 0.600 ±.13 | 0.002 ±.00 | 0.400 |
| Micro (all gold-bearing) | 0.233 ±.01 | 0.659 ±.02 | 0.773 ±.01 | 0.344 ±.02 | 0.341 |
Not in this table: tribe, religion, clan. The translated test carries no gold spans for them (the source schema has none); they are scored on the synthetic development set under Results, tribe, religion and clan below.
Dholuo, held-out translated-clinical test, model only, 5 seeds
| Category | Gold spans | Precision | Recall (strict) | Recall (relaxed) | F1 | Leak (1 − R) |
|---|---|---|---|---|---|---|
| Patient name | 338 | 0.637 ±.04 | 0.732 ±.06 | 0.783 ±.04 | 0.681 ±.05 | 0.268 |
| Third-party name | 403 | 0.262 ±.02 | 0.570 ±.03 | 0.821 ±.03 | 0.358 ±.02 | 0.430 |
| National ID | 131 | 0.959 ±.03 | 0.458 ±.14 | 0.458 ±.14 | 0.605 ±.15 | 0.542 |
| Record number | 378 | 0.607 ±.06 | 0.992 ±.02 | 1.000 ±.00 | 0.751 ±.05 | 0.008 |
| Phone | 24 | 0.874 ±.03 | 0.917 ±.00 | 1.000 ±.00 | 0.894 ±.01 | 0.083 |
| Other contact | 169 | 0.600 ±.11 | 0.832 ±.04 | 0.982 ±.00 | 0.692 ±.07 | 0.168 |
| Village / area | 1151 | 0.647 ±.03 | 0.574 ±.02 | 0.727 ±.02 | 0.608 ±.01 | 0.426 |
| Facility | 118 | 0.104 ±.03 | 0.624 ±.03 | 1.000 ±.00 | 0.177 ±.05 | 0.376 |
| Date | 426 | 0.780 ±.03 | 0.959 ±.06 | 0.993 ±.01 | 0.859 ±.03 | 0.041 |
| Age | 345 | 0.473 ±.04 | 0.981 ±.00 | 1.000 ±.00 | 0.637 ±.04 | 0.019 |
| Profession | 4 | 0.001 ±.00 | 0.250 ±.00 | 0.650 ±.12 | 0.001 ±.00 | 0.750 |
| Micro (all gold-bearing) | 0.342 ±.05 | 0.733 ±.01 | 0.846 ±.01 | 0.464 ±.04 | 0.267 |
Not in this table: tribe, religion, clan. The translated test carries no gold spans for them (the source schema has none); they are scored on the synthetic development set under Results, tribe, religion and clan below.
Across the three released backbones
The same fine-tuning data, recipe family and evaluation were used for all three backbones in this series, so the figures below are directly comparable. Patient-name recall is the safety-critical category; micro F1 is the balanced comparison metric.
Results, names never seen in training
| Language | Category | Familiar names (teval) | Unseen names (teval-novel) | Δ | Reshuffle control |
|---|---|---|---|---|---|
| Luganda | Patient name | R 0.612 / F1 0.578 | R 0.502 / F1 0.493 | -0.111 | R 0.612 (+0.000) |
| Luganda | Third-party name | R 0.400 / F1 0.245 | R 0.411 / F1 0.243 | +0.012 | R 0.404 (+0.004) |
| Kiswahili | Patient name | R 0.706 / F1 0.665 | R 0.654 / F1 0.663 | -0.052 | R 0.691 (-0.015) |
| Kiswahili | Third-party name | R 0.489 / F1 0.410 | R 0.458 / F1 0.388 | -0.031 | R 0.463 (-0.026) |
| Kinyarwanda | Patient name | R 0.504 / F1 0.474 | R 0.469 / F1 0.453 | -0.036 | R 0.493 (-0.011) |
| Kinyarwanda | Third-party name | R 0.379 / F1 0.184 | R 0.384 / F1 0.187 | +0.006 | R 0.377 (-0.002) |
| Dholuo | Patient name | R 0.732 / F1 0.681 | R 0.531 / F1 0.571 | -0.201 | R 0.673 (-0.059) |
| Dholuo | Third-party name | R 0.570 / F1 0.358 | R 0.501 / F1 0.322 | -0.068 | R 0.570 (+0.001) |
Δ is recall on unseen names minus recall on familiar names. The reshuffle control keeps the original names but permutes them across documents, so it measures context change alone; the part of Δ not explained by the reshuffle column is the effect of name novelty itself. Model only, entity-strict.
The deployed system on the same unseen-name test. The name thresholds were fitted on the unseen-name development slice with a no-recall-regression guard. On the test slice it holds only approximately: the system gives back between 0.006 and 0.010 of patient-name recall on unseen names, against micro F1 gains of 0.20 to 0.42. Held-out unseen-name test, entity-strict, 5-seed mean.
| Language | Patient-name recall, model only | Patient-name recall, with rules + calibration | Δ recall | Micro F1, model only | Micro F1, with rules + calibration |
|---|---|---|---|---|---|
| Luganda | 0.502 | 0.494 | -0.008 | 0.493 | 0.746 |
| Kiswahili | 0.654 | 0.643 | -0.010 | 0.614 | 0.810 |
| Kinyarwanda | 0.469 | 0.462 | -0.006 | 0.315 | 0.738 |
| Dholuo | 0.531 | 0.524 | -0.007 | 0.424 | 0.755 |
Results, tribe, religion and clan (synthetic dev; in-distribution ceiling)
| Language | Tribe | Religion | Clan |
|---|---|---|---|
| Luganda | R 0.984 / F1 0.984 (n=63) | R 0.994 / F1 0.997 (n=72) | R 0.973 / F1 0.976 (n=96) |
| Kiswahili | R 0.953 / F1 0.943 (n=64) | R 1.000 / F1 1.000 (n=76) | R 0.966 / F1 0.983 (n=29) |
| Kinyarwanda | R 0.000 / F1 0.000 (n=76) | R 0.855 / F1 0.592 (n=87) | R 0.000 / F1 0.000 (n=35) |
| Dholuo | R 0.960 / F1 0.956 (n=50) | R 0.998 / F1 0.998 (n=80) | R 0.988 / F1 0.958 (n=32) |
Synthetic development set, the set training early-stopped on, so an in-distribution ceiling rather than a held-out score. Reported because it is the only set with gold for these categories, which the clinical test corpus lacks. Kinyarwanda tribe and clan are out of scope by ruling (see Training data).
All three released backbones on the same synthetic development set; AfriBERTa is a 5-seed mean, XLM-R and AfroXLMR a single checkpoint. Kinyarwanda tribe and clan are out of scope by ruling and carry no bar.
Out-of-domain
| Language | Patient-name recall (news) | Micro F1 (news) |
|---|---|---|
| Luganda | 0.350 ±.02 | 0.281 ±.02 |
| Kiswahili | 0.389 ±.02 | 0.292 ±.01 |
| Kinyarwanda | 0.449 ±.03 | 0.248 ±.03 |
| Dholuo | 0.214 ±.02 | 0.165 ±.01 |
MasakhaNER 1.0 test (news). Names in news text are drawn from a different distribution and register than clinical notes; the gap between this table and the clinical results is the model's domain sensitivity.
Comparison with off-the-shelf systems
Publicly available de-identification and NER systems were run zero-shot on the same held-out test split with the same scorer, so the comparison is like for like. None was trained for these languages or this task; they are the systems a practitioner could pick up today, and the comparison measures what is available without training, not what each architecture could reach if fine-tuned on this data.
GLiNER-multi deserves a specific note, because its zero-shot label prompts can be chosen and its confidence threshold can be set, and a fixed choice understates it. It is shown twice: at its published default (the 13 category names as prompts, threshold 0.3), and at the best operating point reachable by tuning the prompt wording and threshold on the development split, read once on the test split. Tuning lifts it by 0.04 to 0.12 F1, almost entirely from clinical-descriptive prompts and a lower threshold. It remains a general-domain model trained on English-centric span data with no clinical text and none of these four languages, so it is being asked to transfer across domain and language at once. On Kinyarwanda the tuned GLiNER exceeds the model-only score of every released backbone; the released system with its rule layer and calibration (next section) is well above it in every language.
| System | Luganda | Kiswahili | Kinyarwanda | Dholuo |
|---|---|---|---|---|
| presidio_rules | 0.242 | 0.190 | 0.226 | 0.230 |
| piiranha_v1 | 0.149 | 0.295 | 0.266 | 0.125 |
| openmed_pf_ml_v2 | 0.127 | 0.173 | 0.144 | 0.221 |
| openai_privacy_filter | 0.244 | 0.351 | 0.262 | 0.291 |
| gliner_multi, zero-shot | 0.351 | 0.415 | 0.379 | 0.328 |
| gliner_multi, prompts and threshold tuned on dev | 0.391 | 0.515 | 0.501 | 0.379 |
| lfm25_pii_detector | 0.260 | 0.277 | 0.268 | 0.284 |
| masakhane_afroxlmr_ner | 0.271 | 0.470 | 0.287 | 0.294 |
| This model | 0.533 | 0.637 | 0.344 | 0.464 |
Micro F1, entity-strict, on the same held-out translated-clinical test. The off-the-shelf systems are run zero-shot; this model is fine-tuned on synthetic data for the task. The tuned GLiNER row selects, per language, the label-prompt set and confidence threshold with the best micro F1 on the development split only, then reads the test split once (Luganda V2 at 0.25, Kiswahili V3 at 0.15, Kinyarwanda V2 at 0.1, Dholuo V2 at 0.15; V2 is a clinical-descriptive prompt set, V3 adds a generic person label; the full grid is in baselines_gliner_calibrated.json and GLINER_CALIBRATION.md).
Effect of a rule layer and threshold calibration (not part of these weights)
The deployable system built on these weights adds two components that are not part of the released model and are reported separately so the model's own performance is not conflated with them.
Why a rule layer. National ID, phone, date, record number, other contact and age have fixed surface forms. Deterministic detectors handle them, and the model is used for names and places. Measured on the held-out translated-clinical test (70% slice), rule layer alone, entity-strict: recall 1.000 in every category in all four languages; precision 1.000 for national ID, record number and phone, 0.987 to 0.994 for other contact, 0.997 to 0.998 for date, and 0.894 to 0.913 for age. The age precision is the cost of a deliberate choice: the age rule accepts both word orders (46 emyaka and emyaka 46), and in this corpus the unit-first form also appears in durations such as emyaka 3 ("three years", of treatment or follow-up), which the rule redacts and the gold does not mark. That is an over-redaction of a number that is not an identifier, accepted so that no patient age is left in the text. Record numbers that carry a facility's own prefix rather than a register name (MUL/2024/00871, KNH/24/00871) are covered by a structure rule (two to four capitals, a two- or four-digit year, four to six digits): it never fires on the translated test, where the register rule already reaches recall 1.000 on 2,078 gold spans, and the same structure matches 1,109 times there, all inside gold.
Why calibration on unseen names. The model attaches a confidence to each span. Fitting the decision threshold on the ordinary development split, where names are token-familiar and confidence is high, suppresses correct low-confidence predictions on unfamiliar names. The name thresholds are therefore fitted on the name-disjoint development slice, with an asymmetric precision floor and a no-recall-regression guard. The same procedure, unchanged, was applied to all three backbones.
| Language | Model only P / R / F1 | With rules + calibration P / R / F1 | Over-redaction, model → system |
|---|---|---|---|
| Luganda | 0.405 / 0.669 / 0.504 | 0.772 / 0.743 / 0.757 | 0.596 → 0.228 |
| Kiswahili | 0.530 / 0.748 / 0.621 | 0.845 / 0.783 / 0.813 | 0.470 → 0.155 |
| Kinyarwanda | 0.207 / 0.659 / 0.315 | 0.751 / 0.730 / 0.740 | 0.793 → 0.249 |
| Dholuo | 0.320 / 0.733 / 0.445 | 0.796 / 0.760 / 0.777 | 0.680 → 0.204 |
Same test split, entity-strict; supervisor re-scored. The effect has the same shape on all three backbones: the gain is predominantly precision, with recall preserved or improved.
How the calibration is done. The model attaches a confidence to every span it proposes. A single decision threshold is applied per language: spans above it are kept, spans below it are dropped. The threshold is fitted, never hand-set, on development data only, in three steps. First, a rule layer takes over the fixed-format categories (national ID, phone, date, record number, age), where deterministic detectors reach recall 1.000, so the model's threshold only governs names and places. Second, the threshold for the two name categories is fitted on the name-disjoint development slice, the same documents used for the unseen-name results, rather than on the ordinary development split. This matters because on familiar names the model is confident and any threshold looks fine; on unfamiliar names its confidence is lower, and a threshold fitted on familiar names silently discards correct low-confidence predictions. Fitting on unfamiliar names recovered most of that recall (for example, unseen-name patient-name recall on the AfriBERTa checkpoints rose from 0.343 to 0.494 in Luganda and from 0.235 to 0.524 in Dholuo) at a cost of 1 to 2 percentage points of over-redaction. Third, the fit is asymmetric: for names it maximises recall subject to a precision floor, because a missed name is a disclosure while an over-redacted token is only lost text, and a guard rejects any threshold that would lower recall relative to the uncalibrated model. The test split is not read at any point in this procedure. The identical procedure was applied to all three backbones; the shipped thresholds are in router_config.json.
Evaluation on real clinical records
These models have not been evaluated on real clinical records in any of the four languages. No annotated corpus of that kind existed when this work was done; the clinical evaluation set is a translation of a Spanish corpus, and the project's pre-registered methodology reserves deployment-grade claims for physician-labelled spans. The figures in this card are therefore development-grade measurements on a proxy. As annotated East African clinical text becomes available through the project's clinical workstream, the same evaluation harness applies unchanged, and the numbers here should be read as the estimate to be replaced by that measurement rather than as the final word.
Transfer to other languages
Nothing in the recipe is specific to these four languages beyond the training data. Extending it to another language needs three things: a synthetic clinical corpus in that language generated from the same template-and-vocabulary approach (name, place and facility lists plus the country's identifier formats), a small real-text development set for model selection and threshold fitting, and a base encoder with some pretraining coverage of the language. The last matters most. In a controlled comparison across these four languages, gains from richer training data appeared only where the backbone had pretrained on the language, and vocabulary diversification hurt where it had not. For a language none of the available encoders covers, continued pretraining on monolingual text is the first step, and the Dholuo results here are the realistic expectation for that case.
Bias, risks and limitations
Stated as measurements, on the translated-clinical test unless noted.
- Patient names are found at recall 0.50-0.73; the remainder is the principal residual risk.
- Third-party names are the weakest name category, recall 0.38-0.57. Threshold adjustment did not raise it without lowering recall elsewhere, and retraining on richer data did not move it.
- Village / area names, the largest category by span count, are found at recall 0.52-0.63. Some are mis-typed as third-party names; this is representational rather than a threshold effect.
- Names never seen in training are found at lower recall than familiar ones (table above); the gap is the memorisation effect and is the figure to plan against for real patients.
- Out-of-domain text (news) is handled at substantially lower recall (table above); thresholds fitted on clinical text do not transfer to other registers.
- Profession has 3-5 gold spans per language in the clinical test; its scores are low-powered and are reported for completeness.
- Tribe, religion and clan are evaluated only on the synthetic held-out set, which is in-distribution with training; treat those figures as upper bounds.
- Evaluation corpus. The clinical test is machine-translated from Spanish; it is a proxy for native clinical text. Physician-labelled spans in these languages did not exist at release.
- Training names. Person names in the synthetic training data are recombined from public given- and family-name vocabularies; as with any such generator, some recombinations coincide with the names of real individuals. The model emits offsets and labels only and cannot reproduce a name.
Environmental impact
Hardware: NVIDIA DGX Spark (GB10), on-premises. Usage: 7.0 GPU-hours for the full lane (sweep, ablations, finals and scoring). Emissions were not separately metered for this lane and are not estimated here.
Technical specifications
| Architecture | AfriBERTa-large (Ogueji, Zhu and Lin, 2021): a from-scratch multilingual encoder pretrained on 11 African languages (10 layers, hidden 768, 6 heads), with a linear token-classification head |
| Parameters | 126M parameters |
| Max sequence length | 256 tokens (longer documents are windowed) |
| Label scheme | BIO over 14 categories = 29 tags |
| Input | Pre-split words over a `\w+ |
| Precision | bf16 training; fp32 inference |
Citation
@misc{craneai2026afrimedpii,
title = {AfriMed-PII: Clinical de-identification models for Luganda, Kiswahili, Kinyarwanda and Dholuo},
author = {{Crane AI Labs}},
year = {2026},
note = {Gates Foundation INV-107957, Workstream 1. Technical report forthcoming.}
}
Funding
This material is based on research funded by the Gates Foundation. The findings and conclusions contained within are those of the authors and do not necessarily reflect positions or policies of the Gates Foundation.
Model tree for CraneAILabs/afrimed-pii-afriberta
Base model
castorini/afriberta_large







