Update README.md
Browse files
README.md
CHANGED
|
@@ -10,4 +10,36 @@ pipeline_tag: text-classification
|
|
| 10 |
tags:
|
| 11 |
- dialect
|
| 12 |
- low-resource languages
|
| 13 |
-
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 10 |
tags:
|
| 11 |
- dialect
|
| 12 |
- low-resource languages
|
| 13 |
+
---
|
| 14 |
+
|
| 15 |
+
### Model Description
|
| 16 |
+
|
| 17 |
+
Meertje is intended as a dialect classifier, developed at the Meertens Institute to distinguish between dialect material and Standard Dutch.
|
| 18 |
+
The model was trained on the Dialect Novel Corpus, using a subcorpus of linguistic material from Drenthe.
|
| 19 |
+
|
| 20 |
+
|
| 21 |
+
### Intended Use
|
| 22 |
+
|
| 23 |
+
Isolating dialect material from Dutch texts containing both dialect and Standard Dutch.
|
| 24 |
+
|
| 25 |
+
|
| 26 |
+
### Training Data
|
| 27 |
+
|
| 28 |
+
Sentences containing Drents vs. Standard Dutch sentences. Balanced train/dev/test at 2122/730/730.
|
| 29 |
+
|
| 30 |
+
### Evaluation
|
| 31 |
+
|
| 32 |
+
| Material | F1 (weighted avg) | support |
|
| 33 |
+
| ----------------- | ----------------- | ------- |
|
| 34 |
+
| Test set (Drents) | 0.95 | 730 |
|
| 35 |
+
| Drents | 0.95 | 7362 |
|
| 36 |
+
| Gronings | 0.94 | 605 |
|
| 37 |
+
| Twents | 0.98 | 1496 |
|
| 38 |
+
| Zeeuws-Vlaams | 0.90 | 3231 |
|
| 39 |
+
|
| 40 |
+
|
| 41 |
+
### Further Resources
|
| 42 |
+
|
| 43 |
+
Background article on the Meertje-project ([NL](https://meertens.knaw.nl/2026/03/05/kan-een-computer-dialect-herkennen/) or [Eng](https://www.the-low-countries.com/article/can-a-computer-recognise-dialect/))
|
| 44 |
+
[Finetuning script](https://colab.research.google.com/drive/1XYuAeNQQrCZMvq24mdMRPNeay-MNhPrz?usp=sharing) (Colab Notebook)
|
| 45 |
+
[Usage script](https://colab.research.google.com/drive/16LyvCtORY4NBo4PpN-Xwo_lP0zH2hwOI?usp=sharing) (Colab Notebook)
|