YOLO-Master-EsMoE

This repository provides the Axera NPU deployment of YOLO-Master-EsMoE, a COCO object detection model built on a customized Ultralytics pipeline. The model replaces part of the YOLO backbone/head with mixture-of-experts (MoE) modules (ES_MOE). During ONNX export the MoE routing is converted to a dense fallback (all experts participate), producing a static computation graph that can be compiled by Pulsar2 into axmodel for Axera’s NPU-based AX650 Series. Three model sizes are provided: YOLO-Master-EsMoE-S, YOLO-Master-EsMoE-M and YOLO-Master-EsMoE-N, with an input size of 640×640 and 80 COCO classes.

References links:

For those who are interested in model conversion, you can try to export axmodel through

Support Platform

Performance

Model Input Shape Latency (ms) CMM Usage (MB)
YOLO-Master-EsMoE-M.axmodel 1 x 640 x 640 x 3 23.446 70.40
YOLO-Master-EsMoE-S.axmodel 1 x 640 x 640 x 3 9.453 54.81
YOLO-Master-EsMoE-N.axmodel 1 x 640 x 640 x 3 4.406 46.68

Models

Download all files from this repository to the device

root@ax650 ~/root/yolo-master-esmoe # tree -L 3
.
|-- README.md
|-- requirement.txt
|-- models
|   `-- ax650
|       |-- config.json
|       |-- YOLO-Master-EsMoE-M
|       |   |-- 1_host_post_YOLO-Master-EsMoE-M.onnx
|       |   `-- YOLO-Master-EsMoE-M.axmodel
|       |-- YOLO-Master-EsMoE-N
|       |   |-- 1_host_post_YOLO-Master-EsMoE-N.onnx
|       |   `-- YOLO-Master-EsMoE-N.axmodel
|       `-- YOLO-Master-EsMoE-S
|           |-- 1_host_post_YOLO-Master-EsMoE-S.onnx
|           `-- YOLO-Master-EsMoE-S.axmodel
|-- web
|   |-- app.py
|   |-- detector.py
|   `-- README.md
|-- infer
|   |-- infer_axmodel.py
|   `-- infer_onnx.py
`-- tools
    |-- convert_opset.py
    |-- eval_map.py
    |-- export_onnx.py
    `-- split_onnx.py

python env requirement

pip install -r requirement.txt

Inference with AX650 Host

root@ax650 ~/root/yolo-master-esmoe # python web/app.py
* Running on local URL:  http://0.0.0.0:7860
* To create a public link, set `share=True` in `launch()`.

Use the device IP address and port 7860 to access the WebApp, for example http://192.168.1.100:7860. In the page, select the device ax650 and the model YOLO-Master-EsMoE-M, upload an image and click Detect.

Input image

Result:

Inference with YOLO-Master-EsMoE-S and YOLO-Master-EsMoE-N

The demo for YOLO-Master-EsMoE-S and YOLO-Master-EsMoE-N is the same as YOLO-Master-EsMoE-M, just select the corresponding model in the WebApp.

Inference scripts

Besides the Gradio WebApp, two inference scripts are provided.

Run the FP32 ONNX baseline on the development machine:

python infer/infer_onnx.py --model <model.onnx> --image <image.jpg>

Run the axmodel inference on the target board:

python infer/infer_axmodel.py --axmodel models/ax650/YOLO-Master-EsMoE-M/YOLO-Master-EsMoE-M.axmodel --post-onnx models/ax650/YOLO-Master-EsMoE-M/1_host_post_YOLO-Master-EsMoE-M.onnx --image football.jpg
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support