Instructions to use pshashid/xlmr-toxicity-english with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use pshashid/xlmr-toxicity-english with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("pshashid/xlmr-toxicity-english", device_map="auto") - Notebooks
- Google Colab
- Kaggle
XLM-RoBERTa Toxicity Detection (English)
Overview
This model detects toxic comments in English and classifies the type of toxicity intent.
Model Details
- Architecture: XLM-RoBERTa-Large (384M parameters)
- Task: Binary toxicity classification + Multi-label intent classification
- Intents: severe_toxic, obscene, threat, insult, identity_hate
- Training Data: Jigsaw Toxic Comment Classification Challenge (161k samples)
Performance
- Accuracy: 88%
- F1-Score: 0.88
- AUC-ROC: 0.9486
Usage
import torch
from transformers import AutoTokenizer
from huggingface_hub import hf_hub_download
from model_utils import MultiTaskXLMR
# Load model
model_path = hf_hub_download("{repo_id_english}", "pytorch_model.bin")
model = MultiTaskXLMR(model_name="xlm-roberta-large", num_intents=5)
model.load_state_dict(torch.load(model_path))
model.eval()
# Inference
tokenizer = AutoTokenizer.from_pretrained("xlm-roberta-large")
text = "You are awesome!"
inputs = tokenizer(text, return_tensors="pt", padding=True, truncation=True)
with torch.no_grad():
tox_logits, intent_logits = model(inputs['input_ids'], inputs['attention_mask'])
tox_prob = torch.sigmoid(tox_logits).item()
print(f"Toxicity: {tox_prob:.4f}")
License
MIT
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support