Robotic Bin-Picking Object Detection Using YOLOv11-OBB with SAM2 Auto-Annotation

Authors

  • Very Very Politeknik Negeri Batam
  • Eko Rudiawan Jamzuri Politeknik Negeri Batam

DOI:

https://doi.org/10.30871/jaic.v10i3.12806

Keywords:

Object detection, Oriented bounding box, Robotic bin-picking, Segment Anything Model, YOLOv11

Abstract

Robotic bin-picking requires accurate detection of randomly oriented objects under cluttered conditions. Conventional axis-aligned bounding boxes struggle to distinguish adjacent objects, motivating the use of oriented bounding boxes (OBB). This paper proposes a complete pipeline for bin-picking object detection that combines the Segment Anything Model 2 (SAM2) with YOLOv11-OBB. A three-stage auto-annotation pipeline first applies a YOLOv11s horizontal bounding-box detector to localize each object and assign its class label. SAM2 then performs automatic instance segmentation within each detected bounding-box region without requiring manual point prompts. Last, the resulting masks are converted to OBB annotations via minimum-area rectangle fitting, reducing annotation time by approximately 877× compared with manual labeling. YOLOv11-OBB featuring C2PSA attention, C3k2 convolution blocks, and an anchor-free rotated detection head is trained for 300 epochs on a purpose-built dataset of three cylindrical object classes (white, black, and red) captured in a UR3 collaborative-robot workspace. Experiments demonstrate an overall mAP@0.5 of 0.995 and mAP@0.5:0.95 of 0.949, with an inference time of 138 ms per frame on a consumer CPU. The results indicate that the proposed pipeline is well-suited for oriented object detection in industrial bin-picking applications.

Downloads

Download data is not yet available.

References

[1] M. Ojer, X. Lin, A. Tammaro, and J. R. Sanchez, “PickingDK: A Framework for Industrial Bin-Picking Applications,” Appl. Sci., vol. 12, no. 18, p. 9200, Jan. 2022, doi: 10.3390/app12189200.

[2] K. Kleeberger, R. Bormann, W. Kraus, and M. F. Huber, “A Survey on Learning-Based Robotic Grasping,” Curr. Robot. Rep., vol. 1, no. 4, pp. 239–249, Dec. 2020, doi: 10.1007/s43154-020-00021-6.

[3] S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 39, no. 6, pp. 1137–1149, Jun. 2017, doi: 10.1109/TPAMI.2016.2577031.

[4] W. Liu et al., “SSD: Single Shot MultiBox Detector,” in Computer Vision – ECCV 2016, B. Leibe, J. Matas, N. Sebe, and M. Welling, Eds., Cham: Springer International Publishing, 2016, pp. 21–37. doi: 10.1007/978-3-319-46448-0_2.

[5] P. Dolezel, D. Stursa, and D. Kopecky, “Memory Efficient Deep Learning-Based Grasping Point Detection of Nontrivial Objects for Robotic Bin Picking,” J. Intell. Robot. Syst., vol. 110, no. 3, p. 110, Jul. 2024, doi: 10.1007/s10846-024-02153-9.

[6] J. Ma et al., “Arbitrary-Oriented Scene Text Detection via Rotation Proposals,” IEEE Trans. Multimed., vol. 20, no. 11, pp. 3111–3122, Nov. 2018, doi: 10.1109/TMM.2018.2818020.

[7] X. Yang, J. Yan, Z. Feng, and T. He, “R3Det: Refined Single-Stage Detector with Feature Refinement for Rotating Object,” Proc. AAAI Conf. Artif. Intell., vol. 35, no. 4, pp. 3163–3171, May 2021, doi: 10.1609/aaai.v35i4.16426.

[8] E. Jamzuri, A. Pinandita, R. Analia, and S. Susanto, “Object Detection and Pose Estimation using Rotatable Object Detector DRBox-v2 for Bin-Picking Robot,” presented at the Proceedings of the 5th International Conference on Applied Engineering, ICAE 2022, 5 October 2022, Batam, Indonesia, Jun. 2023. doi: 10.4108/eai.5-10-2022.2326587.

[9] E. R. Jamzuri, R. Analia, and S. Susanto, “Object Detection and Pose Estimation with RGB-D Camera for Supporting Robotic Bin-Picking,” ELKOMIKA J. Tek. Energi Elektr. Tek. Telekomun. Tek. Elektron., vol. 11, no. 1, p. 128, Jan. 2023, doi: 10.26760/elkomika.v11i1.128.

[10] C.-Y. Wang, A. Bochkovskiy, and H.-Y. M. Liao, “YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors,” in 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada: IEEE, Jun. 2023, pp. 7464–7475. doi: 10.1109/CVPR52729.2023.00721.

[11] J. Wei, A. As’arry, K. Anas Md Rezali, M. Zuhri Mohamed Yusoff, H. Ma, and K. Zhang, “A Review of YOLO Algorithm and Its Applications in Autonomous Driving Object Detection,” IEEE Access, vol. 13, pp. 93688–93711, 2025, doi: 10.1109/ACCESS.2025.3573376.

[12] L. He, Y. Zhou, L. Liu, W. Cao, and J. Ma, “Research on object detection and recognition in remote sensing images based on YOLOv11,” Sci. Rep., vol. 15, no. 1, p. 14032, Apr. 2025, doi: 10.1038/s41598-025-96314-x.

[13] M. Geiß, R. Wagner, M. Baresch, J. Steiner, and M. Zwick, “Automatic Bounding Box Annotation with Small Training Datasets for Industrial Manufacturing,” Micromachines, vol. 14, no. 2, p. 442, Feb. 2023, doi: 10.3390/mi14020442.

[14] A. Kirillov et al., “Segment Anything,” in 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France: IEEE, Oct. 2023, pp. 3992–4003. doi: 10.1109/ICCV51070.2023.00371.

[15] M. A. Mazurowski, H. Dong, H. Gu, J. Yang, N. Konz, and Y. Zhang, “Segment anything model for medical image analysis: An experimental study,” Med. Image Anal., vol. 89, p. 102918, Oct. 2023, doi: 10.1016/j.media.2023.102918.

[16] K. Wang et al., “Oriented object detection in optical remote sensing images using deep learning: a survey,” Artif. Intell. Rev., vol. 58, no. 11, p. 350, Aug. 2025, doi: 10.1007/s10462-025-11256-0.

[17] J. Ding, N. Xue, Y. Long, G.-S. Xia, and Q. Lu, “Learning RoI Transformer for Oriented Object Detection in Aerial Images,” in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA: IEEE, Jun. 2019, pp. 2844–2853. doi: 10.1109/CVPR.2019.00296.

[18] X. Xie, G. Cheng, J. Wang, X. Yao, and J. Han, “Oriented R-CNN for Object Detection,” in 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada: IEEE, Oct. 2021, pp. 3500–3509. doi: 10.1109/ICCV48922.2021.00350.

[19] M. Zand, A. Etemad, and M. Greenspan, “Oriented Bounding Boxes for Small and Freely Rotated Objects,” IEEE Trans. Geosci. Remote Sens., vol. 60, pp. 1–15, 2022, doi: 10.1109/TGRS.2021.3076050.

[20] A. Archit et al., “Segment Anything for Microscopy,” Nat. Methods, vol. 22, no. 3, pp. 579–591, Mar. 2025, doi: 10.1038/s41592-024-02580-4.

Downloads

Published

2026-06-24

How to Cite

[1]
V. Very and E. R. Jamzuri, “Robotic Bin-Picking Object Detection Using YOLOv11-OBB with SAM2 Auto-Annotation”, JAIC, vol. 10, no. 3, pp. 3126–3131, Jun. 2026.

Issue

Section

Articles

Similar Articles

<< < 2 3 4 5 6 > >> 

You may also start an advanced similarity search for this article.