nikkibyr commited on
Commit
09b58d3
·
verified ·
1 Parent(s): aa64bef

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +33 -1
README.md CHANGED
@@ -10,4 +10,36 @@ pipeline_tag: text-classification
10
  tags:
11
  - dialect
12
  - low-resource languages
13
- ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
10
  tags:
11
  - dialect
12
  - low-resource languages
13
+ ---
14
+
15
+ ### Model Description
16
+
17
+ Meertje is intended as a dialect classifier, developed at the Meertens Institute to distinguish between dialect material and Standard Dutch.
18
+ The model was trained on the Dialect Novel Corpus, using a subcorpus of linguistic material from Drenthe.
19
+
20
+
21
+ ### Intended Use
22
+
23
+ Isolating dialect material from Dutch texts containing both dialect and Standard Dutch.
24
+
25
+
26
+ ### Training Data
27
+
28
+ Sentences containing Drents vs. Standard Dutch sentences. Balanced train/dev/test at 2122/730/730.
29
+
30
+ ### Evaluation
31
+
32
+ | Material | F1 (weighted avg) | support |
33
+ | ----------------- | ----------------- | ------- |
34
+ | Test set (Drents) | 0.95 | 730 |
35
+ | Drents | 0.95 | 7362 |
36
+ | Gronings | 0.94 | 605 |
37
+ | Twents | 0.98 | 1496 |
38
+ | Zeeuws-Vlaams | 0.90 | 3231 |
39
+
40
+
41
+ ### Further Resources
42
+
43
+ Background article on the Meertje-project ([NL](https://meertens.knaw.nl/2026/03/05/kan-een-computer-dialect-herkennen/) or [Eng](https://www.the-low-countries.com/article/can-a-computer-recognise-dialect/))
44
+ [Finetuning script](https://colab.research.google.com/drive/1XYuAeNQQrCZMvq24mdMRPNeay-MNhPrz?usp=sharing) (Colab Notebook)
45
+ [Usage script](https://colab.research.google.com/drive/16LyvCtORY4NBo4PpN-Xwo_lP0zH2hwOI?usp=sharing) (Colab Notebook)