microsoft
/

deberta-v3-base

Model card Files Files and versions

Fix typo in README.md

#10

by nelsonauner - opened May 11, 2024

base: refs/heads/main

←

from: refs/pr/10

Discussion Files changed

Files changed (1) hide show

README.md +1 -1

README.md CHANGED Viewed

@@ -16,7 +16,7 @@ In [DeBERTa V3](https://arxiv.org/abs/2111.09543), we further improved the effic
 Please check the [official repository](https://github.com/microsoft/DeBERTa) for more implementation details and updates.
-The DeBERTa V3 base model comes with 12 layers and a hidden size of 768. It has only 86M backbone parameters  with a vocabulary containing 128K tokens which introduces 98M parameters in the Embedding layer.  This model was trained using the 160GB data as DeBERTa V2.
 #### Fine-tuning on NLU tasks

 Please check the [official repository](https://github.com/microsoft/DeBERTa) for more implementation details and updates.
+The DeBERTa V3 base model comes with 12 layers and a hidden size of 768. It has only 86M backbone parameters  with a vocabulary containing 128K tokens which introduces 98M parameters in the Embedding layer.  This model was trained using the same 160GB data as DeBERTa V2.
 #### Fine-tuning on NLU tasks