Instructions to use 4nt11/eyenet-incident with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use 4nt11/eyenet-incident with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="4nt11/eyenet-incident")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("4nt11/eyenet-incident") model = AutoModelForSequenceClassification.from_pretrained("4nt11/eyenet-incident", device_map="auto") - Notebooks
- Google Colab
- Kaggle
EYENET Incident Classifier (eyenet-incident)
Multilingual, multi-label classifier that tags short threat-actor messages (forum posts, Telegram/Matrix chatter) with the kind of cyber incident or criminal activity they describe. It is the triage head used by the EYENET observation framework to route raw collector output into incident streams.
- Base model:
jhu-clsp/mmBERT-base(ModernBERT, multilingual, 8192-token context) - Task: multi-label text classification (sigmoid heads, 16 labels)
- Languages: multilingual base; the gold set is weighted toward the languages EYENET collects most, with Spanish as the first calibrated language.
Labels
16 leaf categories (head order = labels.json):
defacement ddos_attack intrusion breach_dump
credentials stealer_logs iab_corporate crimeware_tooling
crime_aas telecom_abuse phishing_delivery fraud_ops
infra_resale recruiting alliance crew_ops
A message can carry several labels at once (a post selling stealer logs that
also recruits is both stealer_logs and recruiting).
Thresholds
This is multi-label: apply a sigmoid, then a per-label decision threshold.
Tuned thresholds ship in calibration.json (keyed by label). Use them instead
of a flat 0.5 - they are what the deployed service uses.
Usage
import json, torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
repo = "4nt11/eyenet-incident"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForSequenceClassification.from_pretrained(repo).eval()
labels = json.load(open("labels.json")) # or hf_hub_download
thr = json.load(open("calibration.json")) # per-label thresholds
text = "Selling fresh RDP access to a EU manufacturing corp, DM for price"
with torch.no_grad():
probs = model(**tok(text, return_tensors="pt", truncation=True)).logits.sigmoid()[0]
hits = [l for l, p in zip(labels, probs.tolist()) if p >= thr.get(l, 0.5)]
print(hits) # e.g. ['iab_corporate', 'access_sale-ish -> iab_corporate']
Training data
Fine-tuned on a human-reviewed gold set of collected threat-actor messages. The shareable version of that corpus is anonymized: opsec and PII identifiers (emails, phones, crypto addresses, target domains, credentials, attachment filenames) are replaced with type-preserving fakes or entity tags before any text leaves the operator. A companion dataset is available under gated access.
Intended use and limitations
- Intended: CTI triage and incident tagging to help a human analyst route and prioritize collector output. An operator-grade assistant, not an oracle.
- Not intended: automated enforcement, attribution of individuals, or any high-stakes decision without human review.
- Limitations: performance is strongest on the languages and board styles
present in the gold set; rare categories (
alliance,crew_ops) have less support. Treat low-confidence multi-label output as a hint, not a verdict.
- Downloads last month
- 4
Model tree for 4nt11/eyenet-incident
Base model
jhu-clsp/mmBERT-base