Pith. sign in

Paper Citation Record · LEDGER

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind

As of 19 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 4 inbound Pith citation observations for arXiv:2505.12207.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12207 v3

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:44:43.776033Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T22:31:35.967550Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T22:34:01.406814Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact5
  • verified fuzzy31
  • unresolved30
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f697720b-3c3c-4df7-9cc0-2b23183cb246 · outbound

This paper cites Remote sensing for agricultural applications: A meta-review,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Remote sensing for agricultural applications: A meta-review,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.610919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.538863Z digest=sha256:db2f3144b31bd39e095f0e5002c7cb06c23c86c098df69516acd1ee7e8061cc8

Observation 692f37c2-e072-4b68-a99e-af2c448d4b98 · outbound

This paper cites Zero hunger: future challenges and the way forward towards the achievement of sustainable development goal 2,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Zero hunger: future challenges and the way forward towards the achievement of sustainable development goal 2,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.600406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.543318Z digest=sha256:fc45c4187dd244e3350c50e68da4694f50d0919714732631bbb3cec2ffbbd06a

Observation 125c0312-45dc-4605-9420-5da7e8cc189c · outbound

This paper cites Wheat growth monitoring and yield estimation based on remote sensing data assimilation into the safy crop growth model,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Wheat growth monitoring and yield estimation based on remote sensing data assimilation into the safy crop growth model,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.590362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.547224Z digest=sha256:e0e0073c8908f11d355ba4389c74184c5add003e48257729d72f1f8a91fc8002

Observation ef43612f-ae2d-455c-80b9-868801ec410f · outbound

This paper cites Challenges and opportunities in remote sensing-based crop monitoring: A review,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Challenges and opportunities in remote sensing-based crop monitoring: A review,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.579843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.551042Z digest=sha256:b5a12a5781a693ef98fffab491c94d0e65fa32e580d2565eaa58a904efc87459

Observation 82f1b7cd-e549-4f53-acbb-04b638b80670 · outbound

This paper cites A review of individual tree crown detection and delineation from optical remote sensing images: Current progress and future,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind A review of individual tree crown detection and delineation from optical remote sensing images: Current progress and future,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.569688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.554280Z digest=sha256:f223018f6fc8381e786c892caf3f96050ca5438b1f3318dc23a6f7d1daf10e2b

Observation 258540a6-d650-44f6-bd21-bc559b2efae6 · outbound

This paper cites Progress and prospects of crop diseases and pests monitoring by remote sensing,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Progress and prospects of crop diseases and pests monitoring by remote sensing,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.558105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.557564Z digest=sha256:920498319f4f42f2939276308a2cae54cf9a99553cdc2c61d6a446198c02541f

Observation 845e2d05-b4ed-4664-8551-09a4b1e7c4d2 · outbound

This paper cites Advances in deep learning applications for plant disease and pest detection: A review,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Advances in deep learning applications for plant disease and pest detection: A review,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.547719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.560988Z digest=sha256:6904d4db89332e8e4d68e25757056b011c8c6545c26b9ec4be867b29007bb6aa

Observation ac73e645-33da-40bc-84a7-9ed5c69de8f6 · outbound

This paper cites Evaluation of survey and remote sensing data products used to estimate land use change in the united states: Evolving issues and emerging opportunities,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Evaluation of survey and remote sensing data products used to estimate land use change in the united states: Evolving issues and emerging opportunities,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.537537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.563859Z digest=sha256:1ad054cad3c3c1f9ecc281bedfd41871ff0af091c89e1ba2ef1acf175c00f94a

Observation 72b681ce-6be8-4ce0-8486-400fd17caa30 · outbound

This paper cites FUSU: A Multi-temporal-source Land Use Change Segmentation Dataset for Fine-grained Urban Semantic Understanding.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind FUSU: A Multi-temporal-source Land Use Change Segmentation Dataset for Fine-grained Urban Semantic Understanding

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:44:44.195809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.567092Z digest=sha256:86ffe1d38092b3ae5bcc27016940a01e8cdc53a18a9d378498726aedf9311521

Observation a680ae04-b908-4e2d-a8c6-236359a0473a · outbound

This paper cites Multimodality helps unimodality: Cross-modal few-shot learning with multimodal models,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Multimodality helps unimodality: Cross-modal few-shot learning with multimodal models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.527176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.570686Z digest=sha256:784de86cf313f202f803bb1d7b6e0828740e27dc244556ce6426cdc5d93024f0

Observation ba0563c7-05cc-4e4d-8e45-f4b895bca199 · outbound

This paper cites Hello gpt-4o,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Hello gpt-4o,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.516864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.574271Z digest=sha256:b173ff7d17044c77832f633920f8f395a70b3d08aca3a3beabd0859be661d00f

Observation 5cc661c1-fbec-4fbd-8a8e-6110a841c4cb · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Gemini: A Family of Highly Capable Multimodal Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.577705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.577705Z digest=sha256:54c9a1c5cfb2ffefe223e26aceb2397d086e670f9e0b564341045baa16163698

Observation 95bbac99-1bc3-4569-bf73-1e39a6a70bfa · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind LLaMA: Open and Efficient Foundation Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.581411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.581411Z digest=sha256:c4c2c1b6389b4ddec9713c3cbb8ca689c70a4a8ea20ebfc7dba5ccfad2ac3d9e

Observation a15df018-ae5a-4abb-b778-97e508df4085 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.584689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.584689Z digest=sha256:ead52e6b8fdbae8f81e03d8546a7aa33d162b8c229bcbe043bf2af6eb504ac1f

Observation cec67af7-7727-4142-9644-e13d45039cda · outbound

This paper cites H2rsvlm: Towards helpful and honest remote sensing large vision language model,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind H2rsvlm: Towards helpful and honest remote sensing large vision language model,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.588318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.588318Z digest=sha256:1454ea473be338bf5c0a713f2a3486216d5ef30a5cff6352de7fb8e15f601b1c

Observation de2fc18f-1661-4ada-a68c-8f7a47edd58c · outbound

This paper cites Vision-language models in remote sensing: Current progress and future trends,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Vision-language models in remote sensing: Current progress and future trends,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.592136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.592136Z digest=sha256:f8337816cb4de59d62baa0a03acaaa00adcff66dd0cc347ccca6b66aa83fc708

Observation e77b66b9-093a-414e-bddd-11e6d4fbdffd · outbound

This paper cites Urbench: A comprehensive benchmark for evaluating large multimodal models in multi-view urban scenarios,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Urbench: A comprehensive benchmark for evaluating large multimodal models in multi-view urban scenarios,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.493775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.595638Z digest=sha256:13179c7b740dd038e28689faeb4bfe9133836117533be3571f009ff9010de970

Observation 8fd26173-5223-4deb-a2b0-7700721627cd · outbound

This paper cites XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.599067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.599067Z digest=sha256:8a1f1c238e67757e90afc93be24d92a15370923131436570dbfbb274f80e8084

Observation 51132c66-6bae-46bb-b6d9-f6837c600f03 · outbound

This paper cites Vrsbench: A versatile vision-language benchmark dataset for remote sensing image understanding,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Vrsbench: A versatile vision-language benchmark dataset for remote sensing image understanding,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.482146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.603009Z digest=sha256:0decf1e270b3f8085bd1240e2bad423c29b32157cb57cdfd125cfea7894b3979

Observation 935dba89-3363-4a0f-8481-c9873d081115 · outbound

This paper cites A multimodal benchmark dataset and model for crop disease diagnosis,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind A multimodal benchmark dataset and model for crop disease diagnosis,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.470949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.606428Z digest=sha256:9369ca3b28d83ddf3f09fdc2bd264dc4cb24172d4e0df7f82e49cb72819a18db

Observation 81f7c205-0497-4748-9ec4-3d6cffd92d0a · outbound

This paper cites Visual question answering model for fruit tree disease decision-making based on multimodal deep learning,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Visual question answering model for fruit tree disease decision-making based on multimodal deep learning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.460357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.610828Z digest=sha256:9229272e87cffab4c5c510cb48bef89e8b42cbe1e112170170d72c76870ac2e7

Observation 00bd4ed4-d206-42f4-8312-e759867a65b3 · outbound

This paper cites Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.614271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.614271Z digest=sha256:f3b9afc0295459fb6cd6a7fc603c11620431a2ed63d704d0ec9ea6878d187fde

Observation 981620e5-732a-46bf-897c-d5924e384ec9 · outbound

This paper cites AgMMU: A Comprehensive Agricultural Multimodal Understanding Benchmark.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind AgMMU: A Comprehensive Agricultural Multimodal Understanding Benchmark

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.617655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.617655Z digest=sha256:b4900943fdd650428e45bcfab17bf9b0cc598963bbdcf2d7bddc3ea14499fbe2

Observation cda08dd6-a405-4de8-be9e-add82bd7609c · outbound

This paper cites AgriBench: A Hierarchical Agriculture Benchmark for Multimodal Large Language Models.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind AgriBench: A Hierarchical Agriculture Benchmark for Multimodal Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.621024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.621024Z digest=sha256:c6f1664a5cb15b00e3a7d3daf8b053daf039399ce0d8bc944c67121821ff6edd

Observation 9d1b3ede-d2f8-4672-ae20-fc0d76d86b99 · outbound

This paper cites Rsgpt: A remote sensing vision language model and benchmark,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Rsgpt: A remote sensing vision language model and benchmark,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.448837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.624850Z digest=sha256:76d31ca7781c47c7cb7e3fdaa586d83f5b37bd34c3b48367640e543620a1abc9

Observation d441f3d2-dfc6-4911-a1bc-689d5a486c6b · outbound

This paper cites Earthvqa: towards queryable earth via relational reasoning-based remote sensing visual question answering,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Earthvqa: towards queryable earth via relational reasoning-based remote sensing visual question answering,

Reference 26

Resolution
verified exact
doi, observed 2026-08-15T20:44:43.808993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.628339Z digest=sha256:56170374ccff38bc34e0522eb1a438063abb661c1c84be526f5b0578ebad749b

Observation 9fb18c08-e19b-4ce0-8227-e428771df438 · outbound

This paper cites Lhrs-bot: Empowering remote sensing with vgi- enhanced large multimodal language model,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Lhrs-bot: Empowering remote sensing with vgi- enhanced large multimodal language model,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.437918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.632113Z digest=sha256:18752d5b993160897c142448b64d31f6cce7b04ce2d222086b420a9660c76708

Observation 0dc43abb-5a36-4fee-8687-0f762dd172de · outbound

This paper cites Agrogpt : Efficient agricultural vision-language model with expert tuning,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Agrogpt : Efficient agricultural vision-language model with expert tuning,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.426907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.635625Z digest=sha256:d0f537cb186fbdbb9dc15d0f0b55d1e969f8726a836d592e90fa73b5b28ea389

Observation 85f9deaf-7583-4036-9b6c-fd3f5f4378da · outbound

This paper cites GPT-4 Technical Report.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind GPT-4 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.639089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.639089Z digest=sha256:d62af07b6a3dca45ede029576d512cf596587413b33ddcf862cd2d37f5e7ab94

Observation af598bc1-6abb-416b-ab0b-714e20ae000b · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.642389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.642389Z digest=sha256:4f2f5efc2a5a5721db1ceaac3717e618b82788bf7ff58e33dfe4601caeb16f59

Observation 151b7615-1d8f-4729-85b3-97bca6c70690 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.645653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.645653Z digest=sha256:8b0e7dc91398a025ecc4d031209be44bbc4a1ae8d7f09ad7ee87e9ae9cc9b39a

Observation e2c2b996-2e3e-439d-971d-ea60266a18fc · outbound

This paper cites Improved baselines with visual instruction tuning,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Improved baselines with visual instruction tuning,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.649504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.649504Z digest=sha256:56454d293d6f5541fdf1c9cb35a25c8fdd481c48ab9039c47b24d7ad57554e57

Observation 65cb4ec1-1960-4e1b-a267-031b63af15a9 · outbound

This paper cites Vila: On pre-training for visual language models,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Vila: On pre-training for visual language models,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.653014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.653014Z digest=sha256:788c17c33ecf9bbc47ff2ab0a2083d6b09092b2f3112c3e73f4d097774112f95

Observation b62de6e2-b00a-4a70-a33f-d0509cacfb82 · outbound

This paper cites Efficient prompt tuning of large vision- language model for fine-grained ship classification,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Efficient prompt tuning of large vision- language model for fine-grained ship classification,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.395774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.656459Z digest=sha256:f911f859d201ac169e1f7058cec8b97591a2930010f1c2c567661bde85cbee9d

Observation 2aeed79d-edfe-417f-9f86-7a20cd24b4a1 · outbound

This paper cites Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.385068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.660165Z digest=sha256:31354d01de5d81d31fa1e888a3f5b663189a1dbba8bbd53eb0b5c7f19345cecc

Observation d7f63bc7-3c74-4711-90f2-5424333363cc · outbound

This paper cites Geochat: Grounded large vision-language model for remote sensing,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Geochat: Grounded large vision-language model for remote sensing,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.663669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.663669Z digest=sha256:ffb41cfb6990add72dd0b3c50e725a1f2867e52147767d87cec8852c18309e7b

Observation d90df507-686f-4e43-91f9-4eb9995c5d11 · outbound

This paper cites Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.667182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.667182Z digest=sha256:9943ee82752dfce754b7714050992b3f3cd093d482977526bb659a1f5e3ba783

Observation 6bd7efa1-10d3-4b2a-a952-fb0dc4605ef4 · outbound

This paper cites Integrating deep learning for visual question answering in agricultural disease diagnostics: Case study of wheat rust,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Integrating deep learning for visual question answering in agricultural disease diagnostics: Case study of wheat rust,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.363185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.670821Z digest=sha256:d598280d93af0ffe882ab464eb618bfbdd2d2bada0a131bdbebf91b5f1929796

Observation f4845ac9-d104-45ac-a635-50044e3bb5c3 · outbound

This paper cites Nocaps: Novel object captioning at scale,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Nocaps: Novel object captioning at scale,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.674409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.674409Z digest=sha256:9cf5210d62f3cc7a06ffdce36bd3e9c68fdd38d50fc1387072b91f32a3957c1f

Observation 0dc5f878-943b-4dae-97fb-b276b61b4ebe · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Ok-vqa: A visual question answering benchmark requiring external knowledge,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.678061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.678061Z digest=sha256:5c5605dbec4555873da1d6ba7919a8c4a72e73eeb1a95877e49c09a0f3a564f9

Observation b00d41fc-6530-48da-a56d-dd20f5b10df5 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Making the v in vqa matter: Elevating the role of image understanding in visual question answering,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.339636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.681654Z digest=sha256:74ccdb842d2763f87d200a949c36b4ebf21332d9278fa9b46dfb06d6be169178

Observation b49fbbca-a4f8-4856-8b0d-89237b656308 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Gqa: A new dataset for real-world visual reasoning and compositional question answering,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.328785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.684986Z digest=sha256:fd8b44087bc748fa03ff051814ab798798f6b79d956fba170f3bf73f0e9b3489

Observation 151723f3-09df-42dd-bb9a-9149810740ec · outbound

This paper cites Tap: Text-aware pre-training for text-vqa and text-caption,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Tap: Text-aware pre-training for text-vqa and text-caption,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.317892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.688706Z digest=sha256:a050d1e2980fbd3d3e00cac3fd36e5a949b714ead7c31274ce48ae616af3d736

Observation 0f4016df-612b-45c7-a68f-52145a216021 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.692226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.692226Z digest=sha256:7974c31ebb7d4d97364ecb8d106164ff7fca39dd02b15283bc54257f95fc7426

Observation 15ccb13b-1065-4b47-8a84-b390ece03a39 · outbound

This paper cites Seed-bench: Benchmarking multimodal large language models,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Seed-bench: Benchmarking multimodal large language models,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.301002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.695654Z digest=sha256:d7d52bc7a982972dc9b53119ad66c9701e29fd6c429ffc0abad63fa8132a0f84

Observation 05d5a28d-78ef-46e4-87e7-9f579b2fc602 · outbound

This paper cites MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.699366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.699366Z digest=sha256:db3548cd78ee4f9c02ed90cfbd89f8c7281a23ad96e7978a896ca5e92e4bc489

Observation 11f515d5-fa07-4429-9bc2-c9ee9a943814 · outbound

This paper cites MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.702648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.702648Z digest=sha256:ee140f6dbcd26d7f95968d4f18470cf4340d975f4b6ab688e009ce536a7914f7

Observation 569fbe55-d8de-4efe-a032-48c380ea04cc · outbound

This paper cites OAM-TCD: A globally diverse dataset of high-resolution tree cover maps.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind OAM-TCD: A globally diverse dataset of high-resolution tree cover maps

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:44:44.087748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.705786Z digest=sha256:2a5ef605e0d8584db3b24d6883c61286797d4e4f396b4322d60422b64bb149ff

Observation 64538cee-8b6a-4d3b-b1ce-8d6863d232cc · outbound

This paper cites Growing status observation for oil palm trees using unmanned aerial vehicle (uav) images,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Growing status observation for oil palm trees using unmanned aerial vehicle (uav) images,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.290108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.709332Z digest=sha256:88ae1608527d1902fec2634cd390ce892c709c9005e8928d41bfd9c3846cd6e7

Observation 2d940f55-35de-49d8-8a91-0718809a857c · outbound

This paper cites Agriculture- vision: A large aerial image database for agricultural pattern analysis,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Agriculture- vision: A large aerial image database for agricultural pattern analysis,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.278199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.712944Z digest=sha256:1df0fd8b5ae88f922320d12de3e48c7bb3413dac85204d28adf06059af536bfe

Observation aec19df8-a4c2-4c30-9c23-3404eec3dacb · outbound

This paper cites PhenoBench — A Large Dataset and Benchmarks for Semantic Image Interpretation in the Agricultural Domain,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind PhenoBench — A Large Dataset and Benchmarks for Semantic Image Interpretation in the Agricultural Domain,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.267370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.716921Z digest=sha256:7c812f3c54a59f3c61c1987646e59fb17abf71c1bbbfe58109cabd344564da27

Observation 27d544a8-7844-4f5a-84a0-ae46cebb5ecf · outbound

This paper cites Ip102: A large-scale benchmark dataset for insect pest recognition,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Ip102: A large-scale benchmark dataset for insect pest recognition,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.256424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.720263Z digest=sha256:217a41052dc2446722a0cd8a998e1dd12f2ac73237e9b50a53808dbe9562f2f2

Observation 9363bcde-7669-406c-bbe5-fc68c2159f8f · outbound

This paper cites Deep Fruit Detection in Orchards.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Deep Fruit Detection in Orchards

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:44:44.073003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.723781Z digest=sha256:3d05d340e09b16ca1d06859bf36d1138a945464a836ee55a801d7c80c73f8bef

Observation ea741b03-a55d-4eec-9160-e007f2a070ca · outbound

This paper cites Cropharvest: A global dataset for crop-type classification,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Cropharvest: A global dataset for crop-type classification,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.245693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.727620Z digest=sha256:3a001cd304d5685d1839238eb91d65713c3a135de50f944b23a766778743b15c

Observation 90644816-c143-400a-b292-506a6c4b0232 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind The claude 3 model family: Opus, sonnet, haiku,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.234965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.731366Z digest=sha256:9bec8cb5a62f6960edfc723f593ac8d981abbe86ba535a039d31f50157377cf8

Observation b1d4b9f8-06ba-4a49-b02e-e6e9ece34d7d · outbound

This paper cites TinyLLaVA: A Framework of Small-scale Large Multimodal Models.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind TinyLLaVA: A Framework of Small-scale Large Multimodal Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.735012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.735012Z digest=sha256:fdf38bac4510e8c193693ec4b5b8bb81af9d88a5f331de41749d1072f0dcc127

Observation 550a0bf7-a68f-48e4-b6cc-021b86d96a69 · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.738894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.738894Z digest=sha256:8359d9f166ebe9294490a4af506c4b4651492e1419e19b4167838921befa6f4f

Observation 1d9ef46f-229a-4fa4-9672-61c388c13a4e · outbound

This paper cites InstructBLIP: Towards general-purpose vision-language models with instruction tuning,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind InstructBLIP: Towards general-purpose vision-language models with instruction tuning,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.742793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.742793Z digest=sha256:361729f4c6cf6ba57772a8aba9b27cea24fa254a8286230ef6fcbc8523d95c29

Observation 546fd704-f8b0-41f5-870d-dff1be5acf1c · outbound

This paper cites What matters when building vision- language models?.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind What matters when building vision- language models?

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.747231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.747231Z digest=sha256:cc46eefe8d34b3514808c3c09200daaf380b1576c700a4a7dacc6b9675147c7c

Observation 793974d0-9d2a-4633-984b-027cd81b1ea3 · outbound

This paper cites Geochat: Grounded large vision-language model for remote sensing,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Geochat: Grounded large vision-language model for remote sensing,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.751003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.751003Z digest=sha256:dfbf9f12e3c6c62ae37bd4dff7855e6da4cd00f56ca6052ea0fcca99eae39c79

Observation a6b9c74e-ee63-41aa-ba9f-a870aa24fddf · outbound

This paper cites Geollava-8k: Scaling remote-sensing multimodal large language models to 8k resolution,.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Geollava-8k: Scaling remote-sensing multimodal large language models to 8k resolution,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.754518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.754518Z digest=sha256:b5d0119aea1eb3f96c173cf7a804574cf84a4af945a97bc0d14f907d82ca8db4

Observation 8a12dccf-ae67-4f72-a3ab-62f6d0c5d087 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.758033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.758033Z digest=sha256:ff15d839a5d0ff2ef7761c395d2697b66fbf634ffa48612b97428d7a775e92d6

Observation dbf636e4-2de3-4951-90bb-752049acb7b7 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.761621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.761621Z digest=sha256:d0dc0d68623edab7c2f293758a4f7766bf87fc3dfa7484ee47ab843864cff0d2

Observation 32795418-00df-47b9-9eb9-54cdfafd60a8 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.765494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.765494Z digest=sha256:d6b2868d0e784f529d4038d410d54623c485ed58ff52a79eb946862a152a7027

Observation aebd6c7a-76a4-47ea-80c7-e0cd343c3e11 · outbound

This paper cites an unresolved cited work.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind Unresolved cited work

Reference 65

Resolution
verified exact
raw_fallback, observed 2026-08-15T20:44:43.956799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.769942Z digest=sha256:acbdedb2c9bfa0af6d7c28f188a12a45ad3705691250f041690eff807dec60a5

Observation b9b2a3c0-9be6-46d6-89e5-d92bb9b4556c · outbound

This paper cites com/PRBonn/ phenobench OAM-TCD Dataset.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind com/PRBonn/ phenobench OAM-TCD Dataset

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:44.205485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:43.772861Z digest=sha256:a98019001f65b1a37250d84610282bd54c8cf8071dfb3f189ac9b0d43bebe47d

Observation a97a32d1-1ee7-47ba-bc2a-c0d8f8ed19fb · outbound

This paper cites box-based.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind box-based

Reference 67

Resolution
malformed identifier
no resolver link, observed 2026-08-15T20:44:43.776033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.776033Z digest=sha256:bbe4891d7cf0918085b60cc389eebf52370daf6a2ce6a8b4a809f0d5db6b3d74

Pith citing papers

Observation bb39dfde-fb40-4cd6-a7db-b6b77dc19a05 · inbound

AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture cites this paper.

AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:00:26.927424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T17:59:42.921867Z digest=sha256:4cc5301a993f9819faeb1884d51926bedb8cc41ca0d4f3b2feef9368b2e7a379

Observation 8ab42fb7-eae2-4e35-a89d-faba5bb31b6f · inbound

HM-Bench: A Comprehensive Benchmark for Multimodal Large Language Models in Hyperspectral Remote Sensing cites this paper.

HM-Bench: A Comprehensive Benchmark for Multimodal Large Language Models in Hyperspectral Remote Sensing Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:30:57.695458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T18:06:49.114269Z digest=sha256:5ac1c87527da10b5d470f118996d91d63b9f1f01eec6b9c753847f768f071b1d

Observation eecb98c2-802f-4748-9c77-08aac7d822ef · inbound

AgroTools: A Benchmark for Tool-Augmented Multimodal Agents in Agriculture cites this paper.

AgroTools: A Benchmark for Tool-Augmented Multimodal Agents in Agriculture Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:51:16.383604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T07:46:50.161097Z digest=sha256:dcaf045d8cc526c9d0cdaf4fadc0f103c79912dd79d8b93c683f3aa019807093

Observation 55f06c66-96e7-41a9-8ace-4dffebf7d04b · inbound

VertiCue-Bench: Diagnosing Whether MLLMs Use Height Cues to Resolve 2D Ambiguity in Remote Sensing Natural Scenes cites this paper.

VertiCue-Bench: Diagnosing Whether MLLMs Use Height Cues to Resolve 2D Ambiguity in Remote Sensing Natural Scenes Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:01.409852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T22:31:35.967550Z digest=sha256:0456b15b4d0f84d4bfc137bef3d6ad45b87fbf8fbc1c0b1d0d29bd90c71f1963