Skip to main navigation Skip to search Skip to main content

Vision-language models for zero-shot weed detection and visual reasoning in UAV-based precision agriculture

    • United Arab Emirates University

    Research output: Contribution to journalArticlepeer-review

    2 Scopus citations

    Abstract

    Weeds remain a major constraint to row-crop productivity, yet current deep learning approaches for UAV imagery often require extensive annotation, generalize poorly across fields, and provide limited interpretability. We investigate whether modern vision–language models (VLMs) can address these gaps in a zero-shot setting. Using drone images from soybean fields with ground-truth weed boxes, we evaluate six commercial VLMs, ChatGPT-4.1, ChatGPT-4o, Gemini Flash 2.5, Gemini Flash Lite 2.5, LLaMA-4 Scout, and LLaMA-4 Maverick under a unified prompt that elicits (i) weed presence, (ii) spatial localization, (iii) reasoning, (iv) crop growth stage, and (v) crop type. We further introduce Error-Probing Prompting (EPP), a counterfactual follow-up that forces re-analysis under the assumption that weeds are present, and we quantify self-correction with expert-rated interpretability scores (Grounding, Specificity, Plausibility, Non-Hallucination, Actionability). Across models, Gemini Flash 2.5 delivers the most consistent zero-shot performance and highest interpretability, ChatGPT-4.1 provides the strongest reasoning but lower raw detection, ChatGPT-4o offers a balanced profile, and LLaMA-4 variants lag in localization and specificity. Gemini Flash Lite 2.5 is efficient but fails EPP stress tests, revealing brittle reasoning. Visual grounding analysis and a text-to-region overlap metric show that interpretability tracks spatial correctness. Results highlight that explainability and feedback driven adaptability not scale alone best predict reliability for field deployment, and position VLMs as promising, low-annotation tools for precision weed management.

    Original languageBritish English
    Article number1735096
    JournalFrontiers in Plant Science
    Volume16
    DOIs
    StatePublished - 2026

    UN SDGs

    This output contributes to the following UN Sustainable Development Goals (SDGs)

    1. SDG 2 - Zero Hunger
      SDG 2 Zero Hunger
    2. SDG 8 - Decent Work and Economic Growth
      SDG 8 Decent Work and Economic Growth

    Keywords

    • error-probing prompting
    • interpretability
    • multimodal AI
    • precision agriculture
    • UAV imagery
    • vision–language models
    • weed detection
    • zero-shot learning

    Fingerprint

    Dive into the research topics of 'Vision-language models for zero-shot weed detection and visual reasoning in UAV-based precision agriculture'. Together they form a unique fingerprint.

    Cite this