Instructions to use albertobarnabo/bge-m3-italian with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use albertobarnabo/bge-m3-italian with sentence-transformers:
from sentence_transformers import MultiVectorEncoder model = MultiVectorEncoder("albertobarnabo/bge-m3-italian") queries = ["Which planet is known as the Red Planet?"] documents = [ "Venus is often called Earth's twin because of its similar size and proximity.", "Mars, known for its reddish appearance, is often referred to as the Red Planet.", "Jupiter, the largest planet in our solar system, has a prominent red spot.", ] query_embeddings = model.encode_query(queries) document_embeddings = model.encode_document(documents) similarities = model.similarity(query_embeddings, document_embeddings) print(similarities) - Notebooks
- Google Colab
- Kaggle
bge-m3-italian
Italian retrieval, fine-tuned from BAAI/bge-m3. Drop-in replacement: same 1024 dims, same 8192 context, all three heads (dense + sparse + ColBERT) trained and shipped.
Modello di retrieval italiano: sostituisce bge-m3 senza cambiare una riga di codice.
Use it
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("albertobarnabo/bge-m3-italian")
emb = model.encode(["quanto dura la validitΓ della carta d'identitΓ elettronica?"])
For hybrid search (dense + sparse), which is what this model is best at:
from FlagEmbedding import BGEM3FlagModel
model = BGEM3FlagModel("albertobarnabo/bge-m3-italian", use_fp16=True)
out = model.encode(docs, return_dense=True, return_sparse=True)
# score = dense_similarity + 0.3 * lexical_matching_score
ONNX (dense head), for Text Embeddings Inference, fastembed, or the sentence-transformers ONNX backend. onnx/model.onnx is fp32; onnx/model_qint8_avx512_vnni.onnx is dynamically quantized int8 (the same file serves ARM CPUs):
model = SentenceTransformer("albertobarnabo/bge-m3-italian", backend="onnx",
model_kwargs={"file_name": "onnx/model_qint8_avx512_vnni.onnx"})
Both were checked against the PyTorch weights on STS22 (Italian, Spearman): PyTorch 0.7995, ONNX fp32 0.7995 with cosine 1.0000 to the original embeddings, int8 0.8019 on AVX-512 VNNI hardware. Dense embeddings only: the sparse and ColBERT heads stay in sparse_linear.pt / colbert_linear.pt.
fastembed (the library behind Qdrant's client), with the int8 file β CPU-only, 0.57 GB, cosine 0.994 to the PyTorch embeddings:
from fastembed import TextEmbedding
from fastembed.common.model_description import PoolingType, ModelSource
TextEmbedding.add_custom_model(
model="albertobarnabo/bge-m3-italian",
pooling=PoolingType.CLS,
normalization=True,
sources=ModelSource(hf="albertobarnabo/bge-m3-italian"),
dim=1024,
model_file="onnx/model_qint8_avx512_vnni.onnx",
)
model = TextEmbedding(model_name="albertobarnabo/bge-m3-italian")
emb = list(model.embed(["quanto dura la carta d'identitΓ elettronica?"]))
Results
Measured against the base model with identical code; 10,000-sample paired bootstrap.
| Italian retrieval, nDCG@10 | BAAI/bge-m3 | ours v1.0 | ours v1.1 | v1.1 vs base |
|---|---|---|---|---|
| Short passages, dense | 0.751 | 0.777 | 0.777 | +0.027 Β· CI [.021, .032] |
| Short passages, hybrid | 0.752 | 0.787 | 0.783 | +0.031 Β· CI [.026, .037] |
| Long documents, sparse | 0.740 | 0.736 | 0.781 | +0.041 Β· CI [.005, .080] |
| Long documents, hybrid | 0.693 | 0.692 | 0.748 | +0.055 Β· CI [.025, .087] |
| Long documents, dense | 0.627 | 0.595 | 0.609 | β0.019 Β· not significant |
| Short passages, sparse | 0.474 | 0.482 | 0.450 | β0.023 Β· see limitations |
In hybrid mode β the recommended one β this model beats bge-m3 on both short passages and long documents.
Short passages: mMARCO-it dev, 6,304 queries over 206k passages. Long documents: MLDR-it test, 200 queries over 10k documents at 8192 tokens.
In practice, on short passages: the right answer is ranked #1 for 66% of questions (up from 62.5%), and the share of questions where it isn't found in the top 10 at all drops from 11.4% to 8.6%.
Pair it with a reranker for the best results:
| pipeline (2,000 queries) | nDCG@10 |
|---|---|
| bge-m3-italian alone | 0.796 |
| + BAAI/bge-reranker-v2-m3 | 0.821 |
On MTEB
bge-m3-italian is registered in mteb and its scores on the Italian tasks are in the public results repository β so the numbers below are independently reproducible with mteb.get_model("albertobarnabo/bge-m3-italian"). None of these tasks was used for training.
| Italian task (main score) | BAAI/bge-m3 | bge-m3-italian | multilingual-e5-large |
|---|---|---|---|
| XPQARetrieval | 0.530 | 0.762 | 0.473 |
| STS22.v2 | 0.700 | 0.799 | 0.643 |
| SIB200ClusteringS2S | 0.256 | 0.342 | 0.396 |
| WebFAQRetrieval | 0.765 | 0.778 | 0.807 |
| JuriFindITRetrieval | β | 0.617 | 0.594 |
| BelebeleRetrieval | β | 0.932 | 0.951 |
Base-model and e5 numbers are the ones published in the same repository. See the MTEB leaderboard filtered on Italian for the full picture.
How it was trained
- Data: 351,796 Italian query groups (1 positive + 8 hard negatives) built from mMARCO-it, with English cross-encoder teacher margins joined by MS MARCO id. ~9.5% of passages were dropped for mojibake damage in the source translation.
- Method: FlagEmbedding unified fine-tuning with self-distillation β dense, sparse and ColBERT heads trained together. 1 epoch, 5,496 steps, lr 1e-5, 512-token passages, one 48GB GPU, ~6 hours.
- Long-context pass: a second short run (30k groups, 2048-token synthetic long documents mixed with short ones, lr 5e-6) recovered the long-document ability that 512-token training had cost. This is what the table above measures.
- Why all three heads: on long Italian documents the sparse head scores 0.78 against dense's 0.61. Language fine-tunes of bge-m3 usually ship dense-only and silently lose it.
Limitations
- Sparse-only retrieval on short passages is worse than base (0.450 vs 0.474, statistically significant). Hybrid on the same data scores 0.783 β use hybrid, or dense, for short passages.
- Dense-only on long documents is 0.019 below base (not statistically significant after the long-context pass). Hybrid is clearly better than base there (+0.055).
- Short-passage gains are measured on the same distribution the model was trained on (translated MS MARCO). Expect smaller gains on very different Italian text.
- Training data derives from MS MARCO, whose original terms are non-commercial research. This is true of most retrieval models trained on MS MARCO; stated here rather than hidden.
Versions
- v1.1 (current) β adds the long-context pass; better than base in hybrid mode on both benchmarks.
- v1.0 β first release, 512-token training only. Still available:
revision="v1.0".
Provenance
Weights: MIT, same as the base model. Evaluation is fully reproducible: same script for both models, per-query scores and bootstrap included in the training repo.
@misc{bge-m3-italian-2026,
author = {Barnabo, Alberto},
title = {bge-m3-italian: Italian retrieval fine-tune of bge-m3},
year = {2026},
url = {https://huggingface.co/albertobarnabo/bge-m3-italian}
}
- Downloads last month
- 413
Model tree for albertobarnabo/bge-m3-italian
Base model
BAAI/bge-m3