Skip to main navigation Skip to search Skip to main content

SAMConvFormer+LLM: Exploring synergistic fusion of Segment Anything Model with joint convolutional transformer and large language model to advance dense agricultural crop analysis

    Research output: Contribution to journalArticlepeer-review

    4 Scopus citations

    Abstract

    Accurate detection of densely fruited crops is essential for yield estimation, ripeness analysis, and actionable recommendations for optimized crop management using large language models. However, real-world agricultural environments present significant challenges, including varying illumination, diverse fruit sizes, and highly complex backgrounds, making robust detection non-trivial. Existing AI models in agriculture, customized on controlled datasets, struggle to generalize to real-field conditions, limiting their practical applicability. To address this, we introduce a real-field dataset capturing the complexities of a dense crop, which serves as a benchmark to expose model limitations and drive the development of more robust methods. Leveraging this, we proposed a novel detection framework (SAMConvFormer) that synergizes the Segment Anything Model (SAM) with a Joint Convolutional Transformer, combining their unique strengths for robust agricultural crop detection. This integration of SAM further significantly improves boundary delineation, particularly for overlapping clusters, blurring, and irregular fruit structures. In addition, we integrate the proposed SAMConvFormer framework with a Large Language Model to develop a comprehensive pipeline for yield estimation, ripeness analysis, and actionable recommendations for optimized crop management. The proposed model achieves state-of-the-art performance with 79.32% sensitivity, 62.64% IoU, and 77.03% Dice coefficient, representing relative improvements of 31%, 18.79% and 11.56% over the baseline model, highlighting its robustness and precision for the detection of densely fruited crops. To advance this research, we will publicly release the proposed framework and related materials, which offer a solid foundation for future studies. GitHub repository at https://github.com/Owais-CodeHub/SAMConvFormer-LLM.

    Original languageBritish English
    Article number111192
    JournalComputers and Electronics in Agriculture
    Volume240
    DOIs
    StatePublished - Jan 2026

    UN SDGs

    This output contributes to the following UN Sustainable Development Goals (SDGs)

    1. SDG 2 - Zero Hunger
      SDG 2 Zero Hunger

    Keywords

    • Densely fruited crop
    • Joint convolutional transformer
    • Large Language Model
    • Precision agriculture
    • Yield estimation

    Fingerprint

    Dive into the research topics of 'SAMConvFormer+LLM: Exploring synergistic fusion of Segment Anything Model with joint convolutional transformer and large language model to advance dense agricultural crop analysis'. Together they form a unique fingerprint.

    Cite this