XLM-RoBERTa Toxicity Detection (English)

Overview

This model detects toxic comments in English and classifies the type of toxicity intent.

Model Details

  • Architecture: XLM-RoBERTa-Large (384M parameters)
  • Task: Binary toxicity classification + Multi-label intent classification
  • Intents: severe_toxic, obscene, threat, insult, identity_hate
  • Training Data: Jigsaw Toxic Comment Classification Challenge (161k samples)

Performance

  • Accuracy: 88%
  • F1-Score: 0.88
  • AUC-ROC: 0.9486

Usage

import torch
from transformers import AutoTokenizer
from huggingface_hub import hf_hub_download
from model_utils import MultiTaskXLMR

# Load model
model_path = hf_hub_download("{repo_id_english}", "pytorch_model.bin")
model = MultiTaskXLMR(model_name="xlm-roberta-large", num_intents=5)
model.load_state_dict(torch.load(model_path))
model.eval()

# Inference
tokenizer = AutoTokenizer.from_pretrained("xlm-roberta-large")
text = "You are awesome!"
inputs = tokenizer(text, return_tensors="pt", padding=True, truncation=True)

with torch.no_grad():
    tox_logits, intent_logits = model(inputs['input_ids'], inputs['attention_mask'])
    tox_prob = torch.sigmoid(tox_logits).item()
    print(f"Toxicity: {tox_prob:.4f}")

License

MIT

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support