Pith. sign in

Paper Citation Record · LEDGER

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors

As of 8 August 2026, this Paper Citation Record lists 100 of 111 outbound references and 0 inbound Pith citation observations for arXiv:2512.15748.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.15748 v2

Coverage vector

measured 100 of 111 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T17:20:57.934252Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 111 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ca819ec2-ca05-49aa-995b-36ada97ca8bd · outbound

This paper cites GPT-4 Technical Report.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:47.815738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:47.815738Z digest=sha256:f2c9f5b03798f4c0523ee516dce09c0cc7455caebb6be33fb118dda851661e68

Observation 0cceec24-9583-43d9-9faf-ef4c50e2808c · outbound

This paper cites Deep learning approaches to the phylogenetic placement of extinct pollen morphotypes.PNAS nexus, 3(1):pgad419,.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Deep learning approaches to the phylogenetic placement of extinct pollen morphotypes.PNAS nexus, 3(1):pgad419,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:47.981109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:47.981109Z digest=sha256:a09f2659db2866bf02468fbd8fe8c1f841beeca2879bd96be8e3a2383de5fb9d

Observation 8cc81f53-a87e-4e57-8bd1-b880e83d3036 · outbound

This paper cites Pollen morphology, deep learning, phylogenetics, and the evolution of environ- mental adaptations in podocarpus.New Phytologist, 2025.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Pollen morphology, deep learning, phylogenetics, and the evolution of environ- mental adaptations in podocarpus.New Phytologist, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:48.160553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:48.160553Z digest=sha256:be4294891a60e78f61c2444e1bd96f61b5041f75690427d22ce9f10e1276d5ec

Observation 02ada333-6979-4eed-bc70-b3eeaded44a3 · outbound

This paper cites Many-shot in-context learning.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Many-shot in-context learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:48.372338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:48.372338Z digest=sha256:a1965d06d39f36fae4ec9faf102445ce4cfeb4e3db0e173ff0de64be548c23c3

Observation c386f622-a4d9-4daf-995c-51d22b97e1ac · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Flamingo: a visual language model for few-shot learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:48.596614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:48.596614Z digest=sha256:b7ef535d99f359c7780f9121e38baf9efb61e6de7cf6ff01c673633e8980840d

Observation c12e10cc-0ed5-4b4a-ae91-e4d7e69f2c6a · outbound

This paper cites Vqa: Visual question answering.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Vqa: Visual question answering

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:48.803358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:48.803358Z digest=sha256:d6d5fe5ebd46ae827b1e06ebafc8100cb829581f88838e5bae98c13ae407d524

Observation 363bb8f4-a260-4013-a072-eec0276d84d8 · outbound

This paper cites Bach, Jesse E.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Bach, Jesse E

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:48.937306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:48.937306Z digest=sha256:9ed17f9a70e290958136960b800d3af6e25c0591c4558b788b530d5ca035f16f

Observation 2ea80824-b1d8-4e38-b37b-e51bdd3a45f2 · outbound

This paper cites Qwen Technical Report.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Qwen Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:49.099404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:49.099404Z digest=sha256:61c0426dc582a4c7b27f6b69b8d380d0dd11cfc7ced6289d87de0ee2ea156eda

Observation 30b64f71-738a-4132-bee3-7d61ab58ebf0 · outbound

This paper cites Qwen2.5-vl technical report, 2025.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Qwen2.5-vl technical report, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:49.288569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:49.288569Z digest=sha256:e52b0d8ccb941f88649ad48aeb8cd1c03cd30f9ea5b824b9a49764d5e79af70b

Observation 063dd15d-e57a-46cc-8888-a1a6b38d0088 · outbound

This paper cites Scaling Biodiversity Monitoring for the Data Age.XRDS: Crossroads, The ACM Magazine for Students, 27(4):14–18, 2021.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Scaling Biodiversity Monitoring for the Data Age.XRDS: Crossroads, The ACM Magazine for Students, 27(4):14–18, 2021

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:49.452368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:49.452368Z digest=sha256:2a0ee115c72c4f37a60ff9a0123affef6d9d0024794609849412d6b24df81d97

Observation 9a00a591-e640-4113-a692-43bd9043338b · outbound

This paper cites The iWildCam 2020 Competition Dataset.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors The iWildCam 2020 Competition Dataset

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:49.597250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:49.597250Z digest=sha256:ec10bfe6ddb241c13ada1bb30fe353ba9e1e7dd6b51ced1e7659aefe21f37934

Observation 025e4585-f6e4-4c8d-bd80-87b9b17680f1 · outbound

This paper cites The iWildCam 2021 Competition Dataset.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors The iWildCam 2021 Competition Dataset

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:49.688619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:49.688619Z digest=sha256:706ac838d38188c3a30dcc36c960b6e699750922157de42de5f4f5099e463137

Observation 1d36f394-b33f-4f4b-b147-4b150541a009 · outbound

This paper cites Deep learning as a tool for ecology and evolution.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Deep learning as a tool for ecology and evolution

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:49.850555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:49.850555Z digest=sha256:2995a334d4b406a502171c7e3f3f0a626d4a0171ad15899be3bbf54e1de259ba

Observation d0442136-8455-493c-b23f-d16add554bf3 · outbound

This paper cites Lan- guage models are few-shot learners.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Lan- guage models are few-shot learners

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:50.021573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:50.021573Z digest=sha256:fc9c1805ab98171280ca4450b6d1bf0dda88d7af88154733b7c72209bd4746d0

Observation 9fbfc295-7d42-4d8b-9a13-6276ce852f40 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Emerging properties in self-supervised vision transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:50.224620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:50.224620Z digest=sha256:8b7c0252d31d5ea677b4d45ff45aa09a69343cd8e7e3fbb9b1e0d79dde133efe

Observation eee0d3a6-fae4-4fa3-b6ee-a02a1d9e5faa · outbound

This paper cites Big self-supervised models are strong semi-supervised learners.Advances in Neural Information Processing Systems (NeurIPS), 2020.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Big self-supervised models are strong semi-supervised learners.Advances in Neural Information Processing Systems (NeurIPS), 2020

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:50.364650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:50.364650Z digest=sha256:a08c732954c0c413469683426f0e3af4a48356414adc70b0ee9ec57d6eaa1017

Observation dd82752a-acc8-444d-9fe2-cbe859cb7ccf · outbound

This paper cites Christian Schmidt, Aditya Jain, Yves Basset, Sara Beery, Maxim Larrivée, and David Rol- nick.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Christian Schmidt, Aditya Jain, Yves Basset, Sara Beery, Maxim Larrivée, and David Rol- nick

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:50.519880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:50.519880Z digest=sha256:d32980f48f4d0478cc4ae0a5a986fb876e0ee0dc2ecbbc722d7d43c0be41de89

Observation d67c3314-a228-4eb5-9e46-6535c52f1e57 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:50.685635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:50.685635Z digest=sha256:35f695a4a6c214a96cee43b5a1a770582110c1fdc571bee5f5f03bb560d56d9b

Observation 9b8cd4f0-47fc-4785-9957-ba9813d0fc6e · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Reproducible scaling laws for contrastive language-image learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:50.852895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:50.852895Z digest=sha256:0011b9f5f549de461969d9a5b2f32c2ec77a968bb24add06c322c82fdfd53d52

Observation 0f1d8706-0ff6-4437-bc38-03fa5a6ab073 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Imagenet: A large-scale hierarchical image database

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:50.937326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:50.937326Z digest=sha256:3088d132a7164c6ec3019a39947c1d204b12611cdd88f7fa8a411cdc50c6d434

Observation 3ae8f157-6c47-41ab-aef3-2707ea304b2c · outbound

This paper cites Discovering localized attributes for fine-grained recog- nition.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Discovering localized attributes for fine-grained recog- nition

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:51.008866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:51.008866Z digest=sha256:f765cf4bce396ef0b372337e6c67050aa7a0aca0a9d15b623e800575db9144ab

Observation d116ea48-52a4-4ba3-9a24-c7b1c6d67ba8 · outbound

This paper cites Ezray, Drew C.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Ezray, Drew C

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:51.126777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:51.126777Z digest=sha256:d3ccb9923409d5f357935cd238651312a03311eddcdd74bc313231fd8669412e

Observation 259325bf-355d-4a3a-8f66-c57b05ec608a · outbound

This paper cites open world.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors open world

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:51.218543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:51.218543Z digest=sha256:303cc8060e733cffd6079959631356675ee297e02438505eb78a1d4b356d859b

Observation 467401c1-866d-456c-aa12-ed0fe1c3b1b7 · outbound

This paper cites Clip-adapter: Better vision-language models with feature adapters.International Journal of Computer Vision, 132(2): 581–595, 2024.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Clip-adapter: Better vision-language models with feature adapters.International Journal of Computer Vision, 132(2): 581–595, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:51.259009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:51.259009Z digest=sha256:08cac88574f4440e36f8c363f39c0b52e0cef2dd8e9b6e7e7992ded4c57db46f

Observation 3f261f1d-74c0-4fc5-aed2-c68c88087b92 · outbound

This paper cites White, James Balhoff, Wasila M Dahdul, Daniel Rubenstein, Hilmar Lapp, Tanya Berger-Wolf, Wei-Lun Chao, and Yu Su.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors White, James Balhoff, Wasila M Dahdul, Daniel Rubenstein, Hilmar Lapp, Tanya Berger-Wolf, Wei-Lun Chao, and Yu Su

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:51.337217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:51.337217Z digest=sha256:57e1aeaeb59e1f01caa878db94da3848c9443b318108a92c4597989aba23c151

Observation 7a3f2a6b-8f9d-4f99-b55b-167d0a3ca7b6 · outbound

This paper cites an unresolved cited work.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:51.488300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:51.488300Z digest=sha256:3543bd4bd09bd4e507368b020dbe6ea9b6c538850f17811ef0738c1594d8f364

Observation 67285c77-58bf-42e2-8f51-fc3fcdacb613 · outbound

This paper cites Deep residual learning for image recognition.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Deep residual learning for image recognition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:51.587844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:51.587844Z digest=sha256:8a3bef7e7f4932418a5c5c1e4260370e6937a1967a5c775bb79a853c84ee127f

Observation e1fdae11-b9df-499c-a7ef-050aad999a91 · outbound

This paper cites Momentum contrast for unsupervised visual repre- sentation learning.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Momentum contrast for unsupervised visual repre- sentation learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:51.737138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:51.737138Z digest=sha256:2ff835122fc816ef4c1f264f56cdb9c7029cea5ada27261cb3b2254176fa42f4

Observation 4bf98107-27c1-4319-99f9-df24137659e3 · outbound

This paper cites Species196: A one-million semi-supervised dataset for fine-grained species recognition.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Species196: A one-million semi-supervised dataset for fine-grained species recognition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:51.877137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:51.877137Z digest=sha256:f6e082ebd5ff42e7d8f0da1a731d1a8b63223d9919f7b5b583eeed451590343e

Observation 020d5389-9d81-4b8b-bdf7-878cfbced8f2 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Distilling the Knowledge in a Neural Network

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:51.987413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:51.987413Z digest=sha256:5c09e36973cbf56cf9cd38944287042501890d3bb72918ea5202ca2a208329e4

Observation 897461c0-a785-410d-8605-65d2e79c455a · outbound

This paper cites Deep learning on butterfly phenotypes tests evolution’s oldest mathematical model.Science advances, 5(8):eaaw4967, 2019.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Deep learning on butterfly phenotypes tests evolution’s oldest mathematical model.Science advances, 5(8):eaaw4967, 2019

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:52.064616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:52.064616Z digest=sha256:ff8aadd07f5e9edf3f1ac23dc4ced8b95d26d3133e2b945454fc342645cfcfd4

Observation a414d1e4-71ac-4481-98e1-1a226ebf001e · outbound

This paper cites Deep learning and computer vision will transform entomology.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Deep learning and computer vision will transform entomology

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:52.147501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:52.147501Z digest=sha256:1375c2925c5588debcb4e8832f1ba8772e302254b7b472f9c1d133b7dd891d4c

Observation 7b2d6fde-a97a-474a-b259-b550e2ad4ea8 · outbound

This paper cites Multimodal learning and reasoning for visual question answering.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Multimodal learning and reasoning for visual question answering

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:52.230950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:52.230950Z digest=sha256:fdd7d2e2ecc7a20ddc4ef980a869e16a7771a16b897d4e6dd5d82be344b898f4

Observation d3dd95a5-12aa-4ada-a4ce-a737a769ab0d · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:52.313777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:52.313777Z digest=sha256:f3fa0af70614488280319851d971efc426e5fbc5edfc0b4edba95c125b805a1f

Observation b1ee9801-cc5d-4f22-89dc-bddb9f53920a · outbound

This paper cites Vi- sual prompt tuning.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Vi- sual prompt tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:52.395188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:52.395188Z digest=sha256:02b4b5a7cfaa5a1e6b307c57e77371a75bd30c7c6851f71d0f244799d1e5ee3b

Observation a01c224e-b043-443d-9d2f-17963259643c · outbound

This paper cites Many-Shot In-Context Learning in Multimodal Foundation Models.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:52.476377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:52.476377Z digest=sha256:6cb0593e60e3504c35b6bac3f18069400c7f5db7fefa5c462ab77232bf612f81

Observation 8d96eb85-8150-483c-b595-27d487048f83 · outbound

This paper cites Maple: Multi-modal prompt learning.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Maple: Multi-modal prompt learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:52.559446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:52.559446Z digest=sha256:b95daadea88263bf8575c6bd28712d46ff3beb95b3333391236b436adcded073

Observation a446798e-7c50-48b5-b0be-f1a1334d152f · outbound

This paper cites Large language models are zero-shot reasoners.Advances in neural information pro- cessing systems (NeurIPS), 2022.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Large language models are zero-shot reasoners.Advances in neural information pro- cessing systems (NeurIPS), 2022

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:52.637961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:52.637961Z digest=sha256:4452027e29f8dca54b507b7dfaa66052d4f5c7169d72d5f818e4edcef319986c

Observation aca7c54a-22b4-4356-9992-69afead9ad97 · outbound

This paper cites Low-rank bilinear pool- ing for fine-grained classification.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Low-rank bilinear pool- ing for fine-grained classification

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:52.724049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:52.724049Z digest=sha256:9dc1d390a7ca5d05f4b2600ecd76b93220570c4636b376727ee1e688632a82a5

Observation 7e573183-e331-4f18-9627-4aa772e0c3d5 · outbound

This paper cites Grounded language-image pre-training.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Grounded language-image pre-training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:52.808500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:52.808500Z digest=sha256:34cc16a01f059b4c1770b25930ff1108fe0b51fa7ca808b7d037477828ebbac2

Observation 58c88dec-83de-4bb1-b497-97c9cfb8ca6a · outbound

This paper cites Bilinear cnn models for fine-grained visual recognition.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Bilinear cnn models for fine-grained visual recognition

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:52.860548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:52.860548Z digest=sha256:b52e73648ca113045fd45fad0204fffbb932d461ac1cb3c30b6533aef75cb8c7

Observation 52922b1b-2bc7-443d-8f78-09216399cd2a · outbound

This paper cites Multimodality helps unimodality: Cross- modal few-shot learning with multimodal models.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Multimodality helps unimodality: Cross- modal few-shot learning with multimodal models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:52.910093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:52.910093Z digest=sha256:d6298d4da08c7cab5dcb966e34e3a69d4d90703244fc916529d5103549e54dbe

Observation ddb5ed0b-43c3-4cef-bfd1-9140ba496a88 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems (NeurIPS), 2023.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Visual instruction tuning.Advances in neural information processing systems (NeurIPS), 2023

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:52.991903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:52.991903Z digest=sha256:135b316c3d0dc951281bb88b1b41a17c619314eda8d8551e7cae7f9498de3ab7

Observation 652f6c9a-56b5-4660-91db-07814f49d229 · outbound

This paper cites Democratizing fine-grained visual recognition with large language models.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Democratizing fine-grained visual recognition with large language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.066383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.066383Z digest=sha256:df15c856fca93d57bfc5fd2ef04b1e1ee2e96d2bbf7ee4f5f8ada4e6130fa144

Observation e68e4722-1374-44f8-963e-cab9829a0a45 · outbound

This paper cites Few-shot recognition via stage-wise retrieval-augmented finetuning.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Few-shot recognition via stage-wise retrieval-augmented finetuning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.115245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.115245Z digest=sha256:86f9a2391b7276122449dc704c20cf1a7e78f6ff0b18b304a64cffb037437fa1

Observation ca1c1d35-3cf7-4eed-bd4e-656b785dc892 · outbound

This paper cites RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.190844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.190844Z digest=sha256:fefc8366c1ecadcbad49fc5dac1475308a1daea56754aae6c1d2c40ffa1a3b9a

Observation 46b1c36b-9f84-4b53-8671-6ac5494d53ef · outbound

This paper cites Rsvp: Reasoning segmentation via visual prompting and multi-modal chain-of-thought.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Rsvp: Reasoning segmentation via visual prompting and multi-modal chain-of-thought

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.230828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.230828Z digest=sha256:a7cbdc155e8b5a06f963ce02023d9936e722b22d73b7532de6d0ed33ab4290c8

Observation e6632fcd-e3b2-49fb-97f4-6ffaaa328211 · outbound

This paper cites Computer vision, machine learning, and the promise of phenomics in ecology and evo- lutionary biology.Frontiers in Ecology and Evolution, 9: 642774, 2021.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Computer vision, machine learning, and the promise of phenomics in ecology and evo- lutionary biology.Frontiers in Ecology and Evolution, 9: 642774, 2021

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.265819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.265819Z digest=sha256:0954d5beb6f714d2e1f8c5d00ab125a726973f23cf0c8540540de5786f977427

Observation 9e1609d5-c897-4beb-ad92-6ebd3999b60f · outbound

This paper cites Lessons learned from a unifying empirical study of parameter-efficient transfer learning (petl) in visual recognition.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Lessons learned from a unifying empirical study of parameter-efficient transfer learning (petl) in visual recognition

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.300062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.300062Z digest=sha256:ad49e1162693550efc46597bb53a95808e6cd060cc53ce5b6259b5ac94493996

Observation 80e96a85-5696-4721-a578-379f078ed6fc · outbound

This paper cites En- hancing clip with gpt-4: Harnessing visual descriptions as prompts.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors En- hancing clip with gpt-4: Harnessing visual descriptions as prompts

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.338422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.338422Z digest=sha256:6308a9b7c78f784842a4613690dcbb050cdba2d30a809ac3e5d7df94f9c89c08

Observation 17d797ad-892d-48c6-9c03-f8511d9a446a · outbound

This paper cites Visual classification via description from large language models.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Visual classification via description from large language models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.379292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.379292Z digest=sha256:b69c9be29cdace8d4a7f046bda796f8ffdb18b765a3e6f664adf1f0de574eb78

Observation 292cad3c-99ac-4139-8646-e5a53418d950 · outbound

This paper cites Re- thinking the role of demonstrations: What makes in-context learning work? InEMNLP, 2022.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Re- thinking the role of demonstrations: What makes in-context learning work? InEMNLP, 2022

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.416678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.416678Z digest=sha256:47924f2e8de6299e0d5724e1e85a49bd31fbb5bf19264bd1fde5b8ed7f42b9ed

Observation 07b303fa-c589-4b18-bff5-51ed6f698aa1 · outbound

This paper cites Compositional chain of thought prompting for large multimodal models.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Compositional chain of thought prompting for large multimodal models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.451792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.451792Z digest=sha256:22c1911ab2589fc304c92fa10c58008115309bf7d24745b9a3a9a414f009bcae

Observation a47d12da-8b70-48f3-a83f-6f6cba1497d5 · outbound

This paper cites Norman, Jason A.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Norman, Jason A

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.486008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.486008Z digest=sha256:60b0a40aa8a699e73a7633e516b6b035590998b0c7681226d58b1f0aeecc91ad

Observation 1aa9f8fb-5af5-45c0-9826-10b2f40e2406 · outbound

This paper cites Inaturalist.Science Scope, 41(7):12–13, 2018.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Inaturalist.Science Scope, 41(7):12–13, 2018

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.510416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.510416Z digest=sha256:59e73a6a6bfb43dfa215a142159cde5a536482227ab1812c6b7deb6bd31bf1f1

Observation 3df99a23-1401-453a-a0a7-a2477cc2c872 · outbound

This paper cites Gpt-5 system card.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Gpt-5 system card

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.545495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.545495Z digest=sha256:b7a503016448555e7582d6d4d78e7b00c9d060d8dc418c9277866d5a9b3831ed

Observation 032ed4b0-959e-4f2b-8180-092b0b1d897d · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors DINOv2: Learning Robust Visual Features without Supervision

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.601756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.601756Z digest=sha256:457f2bad2fd835f19b5e7d55dda93fd3b8af2c71288b2dede5207e33e8b98db7

Observation c3d98390-11bd-4c08-bd51-0c4dcef2eafd · outbound

This paper cites Prompting scientific names for zero-shot species recognition.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Prompting scientific names for zero-shot species recognition

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.638353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.638353Z digest=sha256:0275145a3da82f1bc2b5b81b307d0dd2725f6c8c2fe48397ae09a1edf6c34690

Observation 4ede0685-97ab-40db-9331-57987c6dc25e · outbound

This paper cites The neglected tails of vision-language models.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors The neglected tails of vision-language models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.664309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.664309Z digest=sha256:524538322267b5007d72f4f72770b93d77064a3328169826824f2f6e21fc0fd0

Observation c2834b8f-c987-42f4-be40-f3fa7cf257f9 · outbound

This paper cites Fungitastic: A multi-modal dataset and benchmark for image categorization.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Fungitastic: A multi-modal dataset and benchmark for image categorization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.801502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.801502Z digest=sha256:7ec31eaea41e761466d2fc1be4c215e4ff82a3a15659b8791e4e3bd1a3a19caa

Observation ebfb9336-b53e-4819-a03c-6d57f0006e60 · outbound

This paper cites an unresolved cited work.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.955334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.955334Z digest=sha256:bb17485ead7ff567df4ead2f006bd529ad556a01284af33dff88d69047fad111

Observation 6cc7c3ac-2a96-491e-96ac-8355fb545645 · outbound

This paper cites What does a platypus look like? generating customized prompts 11 for zero-shot image classification.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors What does a platypus look like? generating customized prompts 11 for zero-shot image classification

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:54.081400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:54.081400Z digest=sha256:6825ce536326f4255cfde831e035deac28cefa39bd86cd1d8cee74c255e4b4c1

Observation 6482fc50-3850-4760-943d-e9d36497e825 · outbound

This paper cites Automated identifi- cation of diverse neotropical pollen samples using convolu- tional neural networks.Methods in Ecology and Evolution, 13(9):2049–2064, 2022.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Automated identifi- cation of diverse neotropical pollen samples using convolu- tional neural networks.Methods in Ecology and Evolution, 13(9):2049–2064, 2022

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:54.191504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:54.191504Z digest=sha256:dc6d3a9df1eeec13cc2820421a3f4afa770aee9154f98b0c703935078407b777

Observation 23b6cc76-c598-477a-a133-7dda7da88c8c · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Learn- ing transferable visual models from natural language super- vision

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:54.357240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:54.357240Z digest=sha256:b9e79ae10f5192fa724782cba719f20348be4f2c65de07deaa8756203990a3cb

Observation 4cd5ebc1-8bf7-4367-9fc2-3809b13a232b · outbound

This paper cites an unresolved cited work.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:54.468242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:54.468242Z digest=sha256:821ea3fa4395fe49f1db778bb104b8ce101af6e5607743b11b01353ad00f1fb3

Observation 40fb7c41-b96a-403f-b9ff-229b638f1f9c · outbound

This paper cites Im- proved zero-shot classification by adapting vlms with text descriptions.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Im- proved zero-shot classification by adapting vlms with text descriptions

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:54.628501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:54.628501Z digest=sha256:c285f5f17dd534693a1331233d774d4ff2df9fe25e20809f9194cec6d87600ee

Observation 0e869490-7ea9-4451-b34c-925125b40814 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:54.770955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:54.770955Z digest=sha256:9f79f56eac3520c63b79a21dd27df88c23214805c023cf5597b60ad676c43598

Observation 93a130d3-e95e-41a4-805c-332ba9e1d43b · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:54.841060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:54.841060Z digest=sha256:75ee141d40c0f96f40e377a53f7e65d4b4e079d92a489de72e62b1c6b3750b01

Observation 5640011c-d0d7-4fa3-9300-be652e66b479 · outbound

This paper cites A closer look at the few-shot adaptation of large vision-language models.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors A closer look at the few-shot adaptation of large vision-language models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:54.955759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:54.955759Z digest=sha256:eba3110bf7a6fd45e7b69dc186ff66a24c5aa75ca162237d27876d185119436a

Observation 992727dd-2af2-48c3-9909-b3532eb1a458 · outbound

This paper cites Scaling-up camera traps: monitoring the planet’s biodi- versity with networks of remote sensors.Frontiers in Ecology and the Environment, 15(1):26–34, 2017.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Scaling-up camera traps: monitoring the planet’s biodi- versity with networks of remote sensors.Frontiers in Ecology and the Environment, 15(1):26–34, 2017

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:55.037052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:55.037052Z digest=sha256:da2c8d248a02169cbb2e05efdfb31cfab0593f4de126c0cb838996daea630e6f

Observation c2d8cc07-21b2-4e44-a6f1-e65e05b370da · outbound

This paper cites BioCLIP: A vision foundation model for the tree of life.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors BioCLIP: A vision foundation model for the tree of life

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:55.190028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:55.190028Z digest=sha256:d0ada68913f55b47af032a5b79873125e5913777f2a6d8e41eec2cc66b63f1a9

Observation 8f0ef4bb-d848-45cc-b7ea-5576c8404f04 · outbound

This paper cites The Semi-Supervised iNaturalist-Aves Challenge at FGVC7 Workshop.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors The Semi-Supervised iNaturalist-Aves Challenge at FGVC7 Workshop

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:55.270657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:55.270657Z digest=sha256:d33d63385c8b9d64fb58dc88e37869fcfb38152b6dd7527bbd8151d68e12fc66

Observation 096c0bdc-7cf9-4eb4-b4e5-d925e71f8a94 · outbound

This paper cites A real- istic evaluation of semi-supervised learning for fine-grained classification.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors A real- istic evaluation of semi-supervised learning for fine-grained classification

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:55.356710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:55.356710Z digest=sha256:a77d975701371e1f120735f2f979acf2a899f887b6c24c08602c06dfbdd0fbdd

Observation e0e4b60b-d278-4515-84c5-d03fe65b44c1 · outbound

This paper cites Gemini 2.5: Pushing the frontier with ad- vanced reasoning, multimodality, long context, and next generation agentic capabilities, 2025.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Gemini 2.5: Pushing the frontier with ad- vanced reasoning, multimodality, long context, and next generation agentic capabilities, 2025

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:55.457981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:55.457981Z digest=sha256:ee760082f9fe0f2f6d3740739729f66dccefc6da277f7bc05e756058ca7806b7

Observation 1857c0fd-1ea7-4930-96fb-e21aff3b501d · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Gemini: A Family of Highly Capable Multimodal Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:55.546044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:55.546044Z digest=sha256:e65aa9f3c1e1b24ddd817e068e77ef623d0c693f3d6fbfc8331a72245d1a4641

Observation 4419eb91-a866-412b-8b97-a68289c75fe9 · outbound

This paper cites Glm-4.5v and glm-4.1v-thinking: Towards ver- satile multimodal reasoning with scalable reinforcement learning, 2025.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Glm-4.5v and glm-4.1v-thinking: Towards ver- satile multimodal reasoning with scalable reinforcement learning, 2025

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:55.637210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:55.637210Z digest=sha256:074d8becbecdbb459794482ba65c7475e4e9577218be1353c971f839aced2a39

Observation 257e3c93-b3c0-4497-a4ff-13f90440c814 · outbound

This paper cites Bird Distribution Modelling using Remote Sensing and Citizen Science data.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Bird Distribution Modelling using Remote Sensing and Citizen Science data

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:55.741333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:55.741333Z digest=sha256:884edf2d7cbacc5fd73785e8e4005dbbbb2eb51093750f504ee634b681694f6c

Observation 6e23718d-a86a-47fa-939c-e065e1fedae3 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:55.851986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:55.851986Z digest=sha256:52f667985a7ba08c92fbc604c7526bafd06af58640cc970030ae0b126d0db0c5

Observation 15a1affe-c2e1-468f-aa33-61f9de66ac35 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors LLaMA: Open and Efficient Foundation Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:55.989907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:55.989907Z digest=sha256:36a4515020849a49dd46d5ac771cbea1c40da2d13c3a5b61fd01661688870e44

Observation 1889cb53-ad26-4be2-b948-9b53c541ec54 · outbound

This paper cites Perspectives in machine learning for wildlife conservation.Nature communications, 13(1):792,.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Perspectives in machine learning for wildlife conservation.Nature communications, 13(1):792,

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:56.058377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:56.058377Z digest=sha256:8e20ae49db733f95fa7811ca10a4747db150bcd9d6f4d6734a52a55bc695904e

Observation 08c51221-cede-4954-b399-006042e4eb5d · outbound

This paper cites iNat Challenge 2021 - FGVC8, 2021.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors iNat Challenge 2021 - FGVC8, 2021

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:56.114390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:56.114390Z digest=sha256:693a2eb9495cfb2128bbfb051d1641b53b1449a2b9c615cfea70a7b2a01f511d

Observation aa69fb4a-97d2-4424-bd3c-faccaccb598b · outbound

This paper cites The inaturalist species classification and detection dataset.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors The inaturalist species classification and detection dataset

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:56.214406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:56.214406Z digest=sha256:9245b3e9fb3e40d99ae246a85419eb7988d52b1a61ef92d3c6489da144245580

Observation 7914cdd8-0a0c-4741-84b2-4c340dc22840 · outbound

This paper cites Enabling val- idation for robust few-shot recognition.arXiv preprint arXiv:2506.04713, 2025.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Enabling val- idation for robust few-shot recognition.arXiv preprint arXiv:2506.04713, 2025

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:56.378174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:56.378174Z digest=sha256:e12ad4b0562fcdd50736ed17f03adc7d14866d097d2c6d5ad0e71dd600dce7dd

Observation 041b29a9-730a-44e9-ab79-9a52f2c14551 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:56.556327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:56.556327Z digest=sha256:c2dab3e5bd73da647894da40df2bb1b9790b1856226990ae8f21685630e1cd42

Observation f6ce78e3-5d81-45a2-b963-88a3f6c0d20a · outbound

This paper cites Self-consistency improves chain of thought reason- ing in language models.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Self-consistency improves chain of thought reason- ing in language models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:56.655136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:56.655136Z digest=sha256:7b8619dcdd0a81e3fa2bdcdcc3342fb2470735951cf21cbed95ac0a0a1ab45e1

Observation c47bf2cf-ea3f-4133-a9c1-a069f958b3ee · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Chain-of-thought prompting elicits reasoning in large lan- guage models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:56.759429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:56.759429Z digest=sha256:a000e177a8b20d264cdf6850ec432ba53c9d0a72f9faa04e6cb1fb46396c53ef

Observation 265a5255-3ef1-429a-b037-58327ee3c24f · outbound

This paper cites Large language models are better reasoners with self-verification.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Large language models are better reasoners with self-verification

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:56.856946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:56.856946Z digest=sha256:9ad27affbfab6a7059c1595e9ac8b723484253c90a838a8ee99ddaf5150f6519

Observation 282b710f-f4be-4c31-82be-6f31cd5732b6 · outbound

This paper cites Robust fine-tuning of zero-shot models.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Robust fine-tuning of zero-shot models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:56.969004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:56.969004Z digest=sha256:fecb4181f63cda5e4cc8aee486df52704de436090ffef4ac7731ef34be6952eb

Observation 1f2a6376-5c7b-403a-8929-9a67c87e1be2 · outbound

This paper cites Pro- tect: Prompt tuning for taxonomic open set classification.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Pro- tect: Prompt tuning for taxonomic open set classification

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:57.120474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:57.120474Z digest=sha256:7297ff4bdca762c9f85a34351ddb30dd3a7d6a06022973c112b5890ece2a47d8

Observation bc1791b6-64c2-4fac-9142-e43464e0add9 · outbound

This paper cites Demysti- fying CLIP data.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Demysti- fying CLIP data

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:57.268383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:57.268383Z digest=sha256:ddd4d660985b700830cab7d58b592a15b0d07624183544118df01de28689fd56

Observation 73cc0f6e-18ac-4011-9d2e-9183a1b43033 · outbound

This paper cites Learning concise and descriptive attributes for visual recognition.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Learning concise and descriptive attributes for visual recognition

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:57.355940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:57.355940Z digest=sha256:d17bb6cabaf654478739dab5d3107c07ffda28ad4feacc3def38716528461a18

Observation 0ece8779-db63-4ed4-98c1-a8ef7a6fdcbf · outbound

This paper cites Qwen2.5 Technical Report.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Qwen2.5 Technical Report

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:57.416538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:57.416538Z digest=sha256:8b3c5286269c278d98f61175f2b97e9713d59e86ea835926ec7864a3207af34a

Observation 0123dc96-752f-486c-872d-5a09f020e4d2 · outbound

This paper cites An empirical study of gpt-3 for few-shot knowledge-based vqa.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors An empirical study of gpt-3 for few-shot knowledge-based vqa

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:57.488624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:57.488624Z digest=sha256:09d86fb01985e7d5031c14593758e7a58c042dd8fc3b16329b7837d22844329f

Observation 778cb526-d5d3-404f-89f4-6db2ac3b2869 · outbound

This paper cites Visual- language prompt tuning with knowledge-guided context op- timization.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Visual- language prompt tuning with knowledge-guided context op- timization

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:57.564993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:57.564993Z digest=sha256:119944285e119b43aa2b514149744630ee7f0b0aeaa955c09e05385ee2d5cdeb

Observation a58b1ec4-2516-4354-8b08-cc310e32ee7d · outbound

This paper cites Image captioning with semantic attention.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Image captioning with semantic attention

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:57.636392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:57.636392Z digest=sha256:ee7e16133e48b21d1f86a05bd665f37c4efbf9c5854ac70f40d64870c28cbc32

Observation 58e0763b-799b-4e00-833d-f07ea5e612ac · outbound

This paper cites Task residual for tuning vision-language models.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Task residual for tuning vision-language models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:57.681087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:57.681087Z digest=sha256:55a6fec2c64aeaa4dcf542c1193d2748e916fc7c6c5ddd67361a778460771fc8

Observation 50994f92-3f34-47c9-a1ce-3709b1264338 · outbound

This paper cites Part-based r-cnns for fine-grained category detection.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Part-based r-cnns for fine-grained category detection

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:57.728666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:57.728666Z digest=sha256:9578c669648853670ba78e191f5d8435ce21846739c6aa8edd8bc08927a0e7e0

Observation a1ba8331-1068-4aa8-91cd-20e1cb7247c1 · outbound

This paper cites Revisiting semi-supervised learning in the era of foundation models.Advances in Neural Information Processing Systems (NeurIPS), 2025.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Revisiting semi-supervised learning in the era of foundation models.Advances in Neural Information Processing Systems (NeurIPS), 2025

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:57.770948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:57.770948Z digest=sha256:59d8cf062afd89f49b40a75c8f5242e3d065cfc46c758cea2d7c7980b4827759

Observation 34a446c3-8bff-4cce-9c39-637a136a0c69 · outbound

This paper cites Tip- adapter: Training-free adaption of clip for few-shot clas- sification.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Tip- adapter: Training-free adaption of clip for few-shot clas- sification

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:57.837437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:57.837437Z digest=sha256:9380b307d45ec940e9950c8b108dd5a45c0401dcb4ab4fd39413fe9d6262622a

Observation 5ab9d9f4-4262-4ddf-8d97-80c541c28212 · outbound

This paper cites What makes good examples for visual in-context learning?Advances in Neural Information Processing Systems (NeurIPS), 2023.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors What makes good examples for visual in-context learning?Advances in Neural Information Processing Systems (NeurIPS), 2023

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:57.934252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:57.934252Z digest=sha256:532dcb0060ec5245d9725a1f5b0323b0d54f9aec859176deef0e0ef62c605c0a

Pith citing papers

No inbound Pith citation observations are available.