EYENET Incident Classifier (eyenet-incident)

Multilingual, multi-label classifier that tags short threat-actor messages (forum posts, Telegram/Matrix chatter) with the kind of cyber incident or criminal activity they describe. It is the triage head used by the EYENET observation framework to route raw collector output into incident streams.

  • Base model: jhu-clsp/mmBERT-base (ModernBERT, multilingual, 8192-token context)
  • Task: multi-label text classification (sigmoid heads, 16 labels)
  • Languages: multilingual base; the gold set is weighted toward the languages EYENET collects most, with Spanish as the first calibrated language.

Labels

16 leaf categories (head order = labels.json):

defacement       ddos_attack     intrusion         breach_dump
credentials      stealer_logs    iab_corporate     crimeware_tooling
crime_aas        telecom_abuse   phishing_delivery fraud_ops
infra_resale     recruiting      alliance          crew_ops

A message can carry several labels at once (a post selling stealer logs that also recruits is both stealer_logs and recruiting).

Thresholds

This is multi-label: apply a sigmoid, then a per-label decision threshold. Tuned thresholds ship in calibration.json (keyed by label). Use them instead of a flat 0.5 - they are what the deployed service uses.

Usage

import json, torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification

repo = "4nt11/eyenet-incident"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForSequenceClassification.from_pretrained(repo).eval()
labels = json.load(open("labels.json"))        # or hf_hub_download
thr = json.load(open("calibration.json"))       # per-label thresholds

text = "Selling fresh RDP access to a EU manufacturing corp, DM for price"
with torch.no_grad():
    probs = model(**tok(text, return_tensors="pt", truncation=True)).logits.sigmoid()[0]
hits = [l for l, p in zip(labels, probs.tolist()) if p >= thr.get(l, 0.5)]
print(hits)   # e.g. ['iab_corporate', 'access_sale-ish -> iab_corporate']

Training data

Fine-tuned on a human-reviewed gold set of collected threat-actor messages. The shareable version of that corpus is anonymized: opsec and PII identifiers (emails, phones, crypto addresses, target domains, credentials, attachment filenames) are replaced with type-preserving fakes or entity tags before any text leaves the operator. A companion dataset is available under gated access.

Intended use and limitations

  • Intended: CTI triage and incident tagging to help a human analyst route and prioritize collector output. An operator-grade assistant, not an oracle.
  • Not intended: automated enforcement, attribution of individuals, or any high-stakes decision without human review.
  • Limitations: performance is strongest on the languages and board styles present in the gold set; rare categories (alliance, crew_ops) have less support. Treat low-confidence multi-label output as a hint, not a verdict.
Downloads last month
4
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for 4nt11/eyenet-incident

Finetuned
(175)
this model