bge-m3-italian

Italian retrieval, fine-tuned from BAAI/bge-m3. Drop-in replacement: same 1024 dims, same 8192 context, all three heads (dense + sparse + ColBERT) trained and shipped.

Modello di retrieval italiano: sostituisce bge-m3 senza cambiare una riga di codice.

Results

Use it

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("albertobarnabo/bge-m3-italian")
emb = model.encode(["quanto dura la validitΓ  della carta d'identitΓ  elettronica?"])

For hybrid search (dense + sparse), which is what this model is best at:

from FlagEmbedding import BGEM3FlagModel

model = BGEM3FlagModel("albertobarnabo/bge-m3-italian", use_fp16=True)
out = model.encode(docs, return_dense=True, return_sparse=True)
# score = dense_similarity + 0.3 * lexical_matching_score

ONNX (dense head), for Text Embeddings Inference, fastembed, or the sentence-transformers ONNX backend. onnx/model.onnx is fp32; onnx/model_qint8_avx512_vnni.onnx is dynamically quantized int8 (the same file serves ARM CPUs):

model = SentenceTransformer("albertobarnabo/bge-m3-italian", backend="onnx",
                            model_kwargs={"file_name": "onnx/model_qint8_avx512_vnni.onnx"})

Both were checked against the PyTorch weights on STS22 (Italian, Spearman): PyTorch 0.7995, ONNX fp32 0.7995 with cosine 1.0000 to the original embeddings, int8 0.8019 on AVX-512 VNNI hardware. Dense embeddings only: the sparse and ColBERT heads stay in sparse_linear.pt / colbert_linear.pt.

fastembed (the library behind Qdrant's client), with the int8 file β€” CPU-only, 0.57 GB, cosine 0.994 to the PyTorch embeddings:

from fastembed import TextEmbedding
from fastembed.common.model_description import PoolingType, ModelSource

TextEmbedding.add_custom_model(
    model="albertobarnabo/bge-m3-italian",
    pooling=PoolingType.CLS,
    normalization=True,
    sources=ModelSource(hf="albertobarnabo/bge-m3-italian"),
    dim=1024,
    model_file="onnx/model_qint8_avx512_vnni.onnx",
)
model = TextEmbedding(model_name="albertobarnabo/bge-m3-italian")
emb = list(model.embed(["quanto dura la carta d'identitΓ  elettronica?"]))

Results

Measured against the base model with identical code; 10,000-sample paired bootstrap.

Italian retrieval, nDCG@10 BAAI/bge-m3 ours v1.0 ours v1.1 v1.1 vs base
Short passages, dense 0.751 0.777 0.777 +0.027 Β· CI [.021, .032]
Short passages, hybrid 0.752 0.787 0.783 +0.031 Β· CI [.026, .037]
Long documents, sparse 0.740 0.736 0.781 +0.041 Β· CI [.005, .080]
Long documents, hybrid 0.693 0.692 0.748 +0.055 Β· CI [.025, .087]
Long documents, dense 0.627 0.595 0.609 βˆ’0.019 Β· not significant
Short passages, sparse 0.474 0.482 0.450 βˆ’0.023 Β· see limitations

In hybrid mode β€” the recommended one β€” this model beats bge-m3 on both short passages and long documents.

Short passages: mMARCO-it dev, 6,304 queries over 206k passages. Long documents: MLDR-it test, 200 queries over 10k documents at 8192 tokens.

In practice, on short passages: the right answer is ranked #1 for 66% of questions (up from 62.5%), and the share of questions where it isn't found in the top 10 at all drops from 11.4% to 8.6%.

Pair it with a reranker for the best results:

pipeline (2,000 queries) nDCG@10
bge-m3-italian alone 0.796
+ BAAI/bge-reranker-v2-m3 0.821

On MTEB

bge-m3-italian is registered in mteb and its scores on the Italian tasks are in the public results repository β€” so the numbers below are independently reproducible with mteb.get_model("albertobarnabo/bge-m3-italian"). None of these tasks was used for training.

Italian task (main score) BAAI/bge-m3 bge-m3-italian multilingual-e5-large
XPQARetrieval 0.530 0.762 0.473
STS22.v2 0.700 0.799 0.643
SIB200ClusteringS2S 0.256 0.342 0.396
WebFAQRetrieval 0.765 0.778 0.807
JuriFindITRetrieval β€” 0.617 0.594
BelebeleRetrieval β€” 0.932 0.951

Base-model and e5 numbers are the ones published in the same repository. See the MTEB leaderboard filtered on Italian for the full picture.

How it was trained

  • Data: 351,796 Italian query groups (1 positive + 8 hard negatives) built from mMARCO-it, with English cross-encoder teacher margins joined by MS MARCO id. ~9.5% of passages were dropped for mojibake damage in the source translation.
  • Method: FlagEmbedding unified fine-tuning with self-distillation β€” dense, sparse and ColBERT heads trained together. 1 epoch, 5,496 steps, lr 1e-5, 512-token passages, one 48GB GPU, ~6 hours.
  • Long-context pass: a second short run (30k groups, 2048-token synthetic long documents mixed with short ones, lr 5e-6) recovered the long-document ability that 512-token training had cost. This is what the table above measures.
  • Why all three heads: on long Italian documents the sparse head scores 0.78 against dense's 0.61. Language fine-tunes of bge-m3 usually ship dense-only and silently lose it.

Limitations

  • Sparse-only retrieval on short passages is worse than base (0.450 vs 0.474, statistically significant). Hybrid on the same data scores 0.783 β€” use hybrid, or dense, for short passages.
  • Dense-only on long documents is 0.019 below base (not statistically significant after the long-context pass). Hybrid is clearly better than base there (+0.055).
  • Short-passage gains are measured on the same distribution the model was trained on (translated MS MARCO). Expect smaller gains on very different Italian text.
  • Training data derives from MS MARCO, whose original terms are non-commercial research. This is true of most retrieval models trained on MS MARCO; stated here rather than hidden.

Versions

  • v1.1 (current) β€” adds the long-context pass; better than base in hybrid mode on both benchmarks.
  • v1.0 β€” first release, 512-token training only. Still available: revision="v1.0".

Provenance

Weights: MIT, same as the base model. Evaluation is fully reproducible: same script for both models, per-query scores and bootstrap included in the training repo.

@misc{bge-m3-italian-2026,
  author = {Barnabo, Alberto},
  title  = {bge-m3-italian: Italian retrieval fine-tune of bge-m3},
  year   = {2026},
  url    = {https://huggingface.co/albertobarnabo/bge-m3-italian}
}
Downloads last month
413
Safetensors
Model size
0.6B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for albertobarnabo/bge-m3-italian

Base model

BAAI/bge-m3
Finetuned
(573)
this model

Dataset used to train albertobarnabo/bge-m3-italian

Space using albertobarnabo/bge-m3-italian 1

Collection including albertobarnabo/bge-m3-italian