#9111·MONAI

Support for YOLO-style / modern object detection models (2D and 3D)

Author: vikashgCreated Sep 11, 2026Updated Sep 11, 2026

Is your feature request related to a problem? Please describe.

MONAI currently ships a single object-detection architecture — RetinaNet — under monai/apps/detection/ (retinanet_detector.py, retinanet_network.py), with supporting anchor utilities, COCO-style mAP metrics, and box transforms. This is a solid anchor-based, two-stage-style foundation and works in both 2D and 3D.

However, the detector zoo stops there. Users who want fast single-stage or anchor-free detectors most notably the YOLO family — have no native path. The existing guidance (see #903) is essentially "use MONAI transforms inside your own YOLO pipeline," which leaves the detector itself, training loop, box-format handling, and metrics outside MONAI's guarantees. #292 raised YOLO/COCO/Pascal-VOC box-format support in transforms years ago but there is no dedicated tracking issue for native modern detectors, and #8519 lists surgical instrument localization/detection as a target without naming an architecture.

Describe the solution you'd like

Expand monai/apps/detection beyond RetinaNet to include modern detectors, with YOLO as the flagship because of its strong fit for 2D medical / endoscopy / surgical-tool and microscopy use cases (real-time inference, anchor-free variants, mature ecosystem). Concretely:

  1. A YOLO-style detector network (e.g. an anchor-free YOLO head) integrated into the existing DetectorNetwork / detector API so it reuses MONAI's box transforms, anchor/anchor-free utilities, ATSS-style matching, and mAP metrics.
  2. Native handling of the YOLO box format (normalized cx, cy, w, h) in the detection transforms, alongside the existing corner/CCWH conventions (follow-up to #292).
  3. A tutorial mirroring the existing RetinaNet LUNA16 / detection tutorial so the new detector is a drop-in alternative.
  4. (Optional / stretch) Room in the API for other modern detectors such as DETR-family or FCOS, so this is an extensible "detector zoo" rather than a one-off.

Note: interest in YOLO-inspired 3D detection

MONAI's biggest differentiator over general-purpose CV libraries is first-class 3D support — RetinaNet here already runs on volumetric data. If there is community/maintainer interest, a YOLO-inspired 3D detector (a single-stage / anchor-free volumetric detection head operating on 3D feature maps, predicting 6-DoF axis-aligned 3D boxes) would be a genuinely novel and high-value addition: few libraries offer a fast single-stage 3D detector, and use cases like nodule / lesion / landmark detection in CT and MR volumes would benefit directly. I'd be happy to help scope and prototype this if the maintainers see value — please comment if there's appetite for the 3D direction specifically.

Describe alternatives you've considered

  • Continuing to use RetinaNet only (works, but no fast single-stage/anchor-free option, and no YOLO ecosystem interop).
  • Running an external YOLO (ultralytics etc.) alongside MONAI purely for transforms (the current #903 answer) — loses MONAI's metric/box/3D guarantees and reproducibility.

Additional context

  • Existing detection module: monai/apps/detection/ (RetinaNet).
  • Related prior discussions: #292 (box-format transforms), #903 (YOLO usage question), #8519 (MICCAI submissions incl. surgical detection/localization).
  • Willing to contribute an implementation and tutorial, and specifically to prototype the 3D variant if there's interest.