Abstract:In the automated harvesting of cluster tomatoes in greenhouse environments, rapid identification of cluster tomatoes and precise localization of picking points are critical. This work addresses the requirements for rapid cluster recognition and accurate picking point localization in automated cluster tomato harvesting within greenhouses by proposing a visual detection and localization method based on an improved YOLOv11-Pose. The method first constructs a dataset containing simultaneous annotations of cluster tomatoes and picking points; then introduces two modules GhostC3ECA and C3k2_CA at different positions within the original YOLOv11-Pose model to achieve network lightweighting and precise picking point localization, respectively. This ultimately enables end-to-end real-time detection of cluster tomatoes and accurate picking point localization. In the network′s backend, we utilize pose-aware non-maximum suppression to optimize detection of densely clustered fruits and key points localization. Experiments validated the approach on a glasshouse-collected cluster tomato dataset, achieving 96.2% recognition accuracy and 86.5% picking point localization accuracy. The model operates at 67.2 fps, meeting real-time processing demands. Verified by real-word greenhouse harvesting scenarios, the proposed cluster tomatoes recognition and picking point localization method based on an improved YOLOv11-Pose provides an efficient and reliable visual detection solution for automated harvesting robots.