BananaMind 2 Remove Nano

BananaMind-2-Remove-Nano

BananaMind-2-Remove-Nano is a compact single-class face detector for blurring or blacking out faces in video. It is YOLO11n fine-tuned from the COCO-pretrained weights on WIDER FACE, built to be fast and to favour recall (finding every face) over leaderboard accuracy.

The model has 2,582,347 parameters (6.4 GFLOPs), a 640x640 input, and ships as PyTorch, ONNX FP32, ONNX FP16 and ONNX INT8 (3.3 MB, for phones and CPUs).

Model Details

Field Value
Parameters 2,582,347 (fused)
Architecture YOLO11n, anchor-free decoupled detection head
Layers (fused) 100
GFLOPs @ 640 6.4
Starting weights COCO-pretrained yolo11n.pt (fine-tuned, not trained from scratch)
Classes 1 (face)
Input size 640x640 (also exported as 384x640 and 640x384 INT8 for 16:9 / 9:16 video)
Weight formats PyTorch .pt, ONNX FP32, ONNX FP16, ONNX INT8 (static, per-channel)
Library Ultralytics 8.4.172

Training Data

Split Images Faces
WIDER FACE train (converted) 12,876 156,994
WIDER FACE val (converted) 3,222 39,112

Converted from CUHK-CSE/wider_face to one-class YOLO format. Boxes flagged invalid were skipped, small faces were kept (blurring needs high recall), images whose only boxes were invalid were dropped (8 images), and images were downscaled to 640 px on the long side.

Training Setup

Field Value
Hardware 1x RTX 5070 Ti (16 GB)
Epochs 40
Image size 640
Batch size 32
Precision AMP
Optimizer AdamW (Ultralytics auto, lr 0.002)
Close mosaic last 5 epochs
Augmentation Ultralytics defaults + random motion blur (p 0.3, 3 to 15 px kernel) + JPEG compression (p 0.3, quality 20 to 90) to imitate video frames
Wall-clock time about 21 minutes
Best epoch 35

Benchmarks

Benchmarks

All models were run through the same evaluator on the same ground truth: the WIDER FACE validation split converted to YOLO format (3,222 images, 39,112 labelled faces, images stored at 640 px or smaller, invalid boxes removed, every face size kept). mAP is COCO-style (101-point AP via ultralytics.utils.metrics.ap_per_class, greedy score-ordered matching, IoU 0.5 and 0.5 to 0.95). "Recall" is the share of labelled faces that receive a box with score >= 0.25 at IoU >= 0.5, which is the setting the blur script uses. Face height is measured in pixels of the stored image. Speed is ONNX Runtime, batch 1, a random 640x640 input: "GPU" is an RTX 5070 Ti (CUDA provider, FP16 files for BananaMind models, FP32 for the others), "CPU" is 4 threads (FP32 files for BananaMind FP32/FP16 rows, the INT8 file for the INT8 row). Timings exclude image decoding and NMS.

* SCRFD-10GF is det_10g.onnx from InsightFace's buffalo_l pack and YuNet is face_detection_yunet_2023mar.onnx from the OpenCV Zoo. Both were run with a score floor of 0.02 and their standard decoding/NMS. They are different architectures trained on different data (not retrained here), so the comparison is indicative, not a controlled study. SCRFD-10GF is the better detector for faces of 20 px and up; YuNet is far smaller and faster but much less accurate.

Model mAP50 mAP50-95 Recall (all) Recall (>=20 px) Recall (<10 px) GPU ms CPU ms File MB
BananaMind-2-Remove-Nano 66.49% 35.54% 61.3% 89.6% 33.1% 1.19 18.9 5.4
BananaMind-2-Remove-Nano INT8 64.95% 34.28% 60.0% 88.5% 31.5% 3.45 12.7 3.3
BananaMind-2-Remove-Small 71.77% 39.26% 67.9% 92.5% 42.1% 1.75 49.5 19.0
SCRFD-10GF* 68.66% 37.18% 66.8% 94.1% 36.3% 2.48 50.5 16.9
YuNet* 47.48% 22.93% 50.0% 86.9% 13.3% 0.77 2.1 0.2

Recall by face size

Face height (px) Faces Remove-Nano Remove-Nano INT8 Remove-Small SCRFD-10GF
under 10 16,305 33.1% 31.5% 42.1% 36.3%
10 to 20 10,521 72.1% 70.7% 79.0% 82.2%
20 and up 12,286 89.6% 88.5% 92.5% 94.1%
40 and up 4,808 94.3% 93.6% 95.9% 96.5%
All faces 39,112 61.3% 60.0% 67.9% 66.8%

Ultralytics val reference

Ultralytics' own validator matches predictions differently and reports slightly higher numbers for the same weights (it is not used in the table above).

Format mAP50 mAP50-95
PyTorch .pt 68.45% 36.24%
ONNX FP32 68.33% 36.21%
ONNX FP16 68.23% 36.06%
ONNX INT8 66.70% 34.77%

Quantization

INT8 costs about 1.6 mAP50 points and 1.4 recall points and makes the file 3.2x smaller than FP32. We also tried lower precision on this model:

Variant mAP50 (Ultralytics val) Note
FP32 68.33% reference
INT8 weights + activations 66.70% shipped as face_yolo11n_int8.onnx (the Detect-head decode is kept in float; plain INT8 returned zero detections)
6-bit weights (simulated, weight-only) 66.51% simulated by rounding conv weights to 6 bits, no 6-bit runtime exists
4-bit weights, 8-bit activations 52.12% too lossy for blurring (face recall 43.0%)

INT8 on a CUDA GPU is not faster (3.45 ms); INT8 is meant for CPUs and phones, where it ran 12.7 ms against 18.9 ms for FP32 in our CPU test.

Privacy notes

Missing a face in a single frame is enough to identify someone, and no detector finds every face: see the recall tables above, especially for tiny (under 20 px), turned-away, covered or motion-blurred faces. Check the output before sharing it.

We also tested our own pixelation (15% padding, blur_video.py logic) against a small learned attacker trained on WIDER FACE crops. At the default 8 blocks the attacker recovered only coarse traits (head pose, hair and skin colour, sometimes glasses or a beard), not identity; with weaker pixelation (16 blocks) plain face recognition on the pixelated image alone already matched 36.7% of faces among 1,500 candidates. The attacker was small and briefly trained, so treat this as a floor on leakage. For strong anonymisation use --mode black.

Limitations

  • Trained only on WIDER FACE. Performance on other kinds of footage (night, heavy compression, unusual cameras, non-human faces) is untested.
  • Tiny and far-away faces are the weak spot: roughly a third to two fifths of faces under 10 px are found. Faces of 20 px and up are found about 90% of the time.
  • Single-image detector. Video behaviour (tracking, hold, interpolation between detections) comes from the script, not the model.
  • Not evaluated against adversarial inputs. Face blurring is not a guarantee of anonymity; voice, clothing, text and surroundings are untouched.
  • Numbers above are self-reported on the WIDER validation split.

Usage

pip install -U ultralytics onnxruntime-gpu lap huggingface_hub
from huggingface_hub import hf_hub_download
from ultralytics import YOLO

model_id = "BananaMind/BananaMind-2-Remove-Nano"
path = hf_hub_download(model_id, "face_yolo11n.pt")

model = YOLO(path)
results = model.predict("photo.jpg", conf=0.25, imgsz=640)
for box in results[0].boxes:
    print(box.xyxy[0].tolist(), float(box.conf[0]))

Blur or black out faces in a video

blur_video.py (included in this repo) detects faces, tracks them with ByteTrack, pixelates or blacks them out, and keeps the audio via ffmpeg.

# pixelate (default), detector on every frame
python blur_video.py in.mp4 out.mp4 --model face_yolo11n_fp16.onnx -n 1 --device 0

# solid black boxes
python blur_video.py in.mp4 out.mp4 --model face_yolo11n_fp16.onnx -n 1 --device 0 --mode black

Defaults: detector every 3 frames (-n), confidence 0.25, 15% box padding, 8 pixel blocks across the shorter side of each face (--blocks, lower is stronger). Use -n 1 for the strictest coverage. With an ONNX model use --device 0; --device cpu makes Ultralytics try to pip-install onnxruntime. The optional --keep mode (leave one person visible) needs face_id.py and InsightFace models that are not included here.

Files

File Notes
face_yolo11n.pt PyTorch weights
face_yolo11n_fp32.onnx, face_yolo11n_fp16.onnx ONNX, fixed 640x640, batch 1
face_yolo11n_int8.onnx ONNX INT8, 640x640
mobile/face_yolo11n_land_int8.onnx, mobile/face_yolo11n_port_int8.onnx INT8, 384x640 landscape and 640x384 portrait: the same weights without the gray padding of a square input (1.6x faster model time on CPU)
blur_video.py detect, track and blur/black-out faces in a video, audio preserved

Intended Use

Fast, local face detection for blurring faces in video and images, research on face-blurring pipelines, and as a small baseline. It is not intended as the sole safeguard where a missed face would cause serious harm.

License

AGPL-3.0, inherited from Ultralytics YOLO11 (the architecture and the COCO-pretrained starting weights). Using these weights or the included script in a product or network service means complying with AGPL-3.0 or obtaining an Ultralytics Enterprise licence.

The training data, WIDER FACE, is listed as CC BY-NC-ND 4.0 / non-commercial research on its Hugging Face card. Do not assume these weights are cleared for commercial use; check the dataset terms before relying on them commercially.

Citation

@inproceedings{yang2016wider,
  author    = {Yang, Shuo and Luo, Ping and Loy, Chen Change and Tang, Xiaoou},
  title     = {WIDER FACE: A Face Detection Benchmark},
  booktitle = {IEEE Conference on Computer Vision and Pattern Recognition (CVPR)},
  year      = {2016}
}

@software{ultralytics_yolo11,
  author  = {Jocher, Glenn and Qiu, Jing},
  title   = {Ultralytics YOLO11},
  year    = {2024},
  url     = {https://github.com/ultralytics/ultralytics},
  license = {AGPL-3.0}
}

🍌

Downloads last month
160
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for BananaMind/BananaMind-2-Remove-Nano

Quantized
(105)
this model

Dataset used to train BananaMind/BananaMind-2-Remove-Nano