thumbnail

Topic

Technologies and technical equipment for agriculture and food industry

Volume

Volume 79 / No. 2 / 2026

Pages : 125-133

Metrics

Volume viewed 0 times

Volume downloaded 0 times

APPLICATION OF ZERO-SHOT LARGE MODELS FOR FRUIT OBJECT DETECTION IN SMART AGRICULTURE

在智慧农业中应用ZERO-SHOT大模型实现水果目标检测

DOI : https://doi.org/10.35633/inmateh-79-11

Authors

(*) Yingdong QIN

South China Agricultural University

Haoyu SONG

South China Agricultural University

Jingyi Li

Guangdong University of Finance & Economics

Keyao WEN

South China Agricultural University

(*) Corresponding authors:

qinyingd@foxmail.com |

Yingdong QIN

Abstract

Efficient and flexible agricultural image annotation is crucial for intelligent crop monitoring, yet conventional detection models are limited by fixed class labels and require extensive manual annotations. This study presents a zero-shot annotation framework that integrates OWLv2, Google’s second-generation open-vocabulary vision model, with large language models (e.g., GPT-3.5, DeepSeek V1) to enable multilingual, natural language-driven fruit recognition in smart agriculture. A user-friendly interface was developed to support individual or batch image annotation with adjustable sensitivity to meet diverse field requirements. Experimental evaluations demonstrated the framework's strong generalizability and semantic understanding capabilities, allowing recognition of unseen fruit categories and attributes such as ripeness or color. The system significantly reduces annotation time and labor costs, while enhancing accessibility through natural language interaction. Compared to traditional models, OWLv2 exhibited superior flexibility and required no task-specific dataset retraining, although its computational demands remain higher than lightweight models like YOLO. These results verify the enormous application potential of OWLv2 and similar zero-shot models in agriculture, providing scalable solutions for automated annotation, real-time monitoring, and large-scale data collection.

Abstract in Chinese

高效且灵活的农业图像标注对于智能作物监测至关重要,然而传统目标检测模型受限于固定的类别标签,并依赖大量人工标注。本研究提出了一种可以语言交互的零样本(zero-shot)标注框架,将Google第二代开放词汇视觉模型OWLv2与大语言模型(GPT-3.5、DeepSeek V1)相结合,以实现面向智慧农业的多语言、自然语言驱动的果实识别。为此,我们开发了一个用户友好的交互界面,支持单张或批量图像标注,并可根据实际田间需求调节检测灵敏度。实验评估结果表明,该框架具备优异的泛化能力与语义理解能力,能够识别未曾在训练集中出现的果品类别及其属性(如成熟度、颜色等)。该系统显著降低了标注所需的时间与人力成本,并通过自然语言交互提升了系统的可及性。相较于传统模型,OWLv2展现出更强的灵活性,且无需针对特定任务重新训练数据集;尽管其计算开销仍高于轻量级模型(如:YOLO)。上述结果验证了OWLv2及同类零样本模型在农业领域具有巨大的应用潜力,为自动化标注、实时监测及大规模数据采集提供了可扩展的技术路径。


Indexed in

Clarivate Analytics.
 Emerging Sources Citation Index
Scopus/Elsevier
Google Scholar
Crossref
Road