Pith. sign in

Paper Citation Record · LEDGER

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors

As of 10 August 2026, this Paper Citation Record lists 100 of 111 outbound references and 0 inbound Pith citation observations for arXiv:2512.15748.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.15748 v2

Coverage vector

measured 100 of 111 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T17:20:57.934252Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 111 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ca819ec2-ca05-49aa-995b-36ada97ca8bd · outbound

This paper cites GPT-4 Technical Report.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:47.815738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:47.815738Z digest=sha256:22f59a15dadd7830d55ae942edaa8e7104c198794a8a010db80cfa4f05bc31d8

Observation 0cceec24-9583-43d9-9faf-ef4c50e2808c · outbound

This paper cites Deep learning approaches to the phylogenetic placement of extinct pollen morphotypes.PNAS nexus, 3(1):pgad419,.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Deep learning approaches to the phylogenetic placement of extinct pollen morphotypes.PNAS nexus, 3(1):pgad419,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:47.981109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:47.981109Z digest=sha256:5a49ca5730806589631d7e48a1b5c9627b0684182cdad8534d45153ae5a0f800

Observation 8cc81f53-a87e-4e57-8bd1-b880e83d3036 · outbound

This paper cites Pollen morphology, deep learning, phylogenetics, and the evolution of environ- mental adaptations in podocarpus.New Phytologist, 2025.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Pollen morphology, deep learning, phylogenetics, and the evolution of environ- mental adaptations in podocarpus.New Phytologist, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:48.160553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:48.160553Z digest=sha256:4cff1228528ffb3ef14870245b66b1e40b2ba6ce1c89a8c1bba1d4d03f0e4533

Observation 02ada333-6979-4eed-bc70-b3eeaded44a3 · outbound

This paper cites Many-shot in-context learning.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Many-shot in-context learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:48.372338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:48.372338Z digest=sha256:05ec3c6b644d019970a67ee40e4fa74db5c311ea21c6d2bce789c4bf31344178

Observation c386f622-a4d9-4daf-995c-51d22b97e1ac · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Flamingo: a visual language model for few-shot learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:48.596614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:48.596614Z digest=sha256:4c4393012f61e6b5877bd8b06e67f30752b54a1eb620140a3878d83e711a80d1

Observation c12e10cc-0ed5-4b4a-ae91-e4d7e69f2c6a · outbound

This paper cites Vqa: Visual question answering.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Vqa: Visual question answering

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:48.803358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:48.803358Z digest=sha256:8b3efd686311e6a4b6deea4bf7133bb9c5140a5a02baa27fc59fa3a4ea9591dc

Observation 363bb8f4-a260-4013-a072-eec0276d84d8 · outbound

This paper cites Bach, Jesse E.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Bach, Jesse E

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:48.937306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:48.937306Z digest=sha256:1d6d2a6fe0cccec22929144cd2f49c5627979e7cd04ce30c09f9255fa595d0cf

Observation 2ea80824-b1d8-4e38-b37b-e51bdd3a45f2 · outbound

This paper cites Qwen Technical Report.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Qwen Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:49.099404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:49.099404Z digest=sha256:4953912ad176182fd04fbe9b4a2c67e1f65d5f11de76313e1ff9b368bd0e2017

Observation 30b64f71-738a-4132-bee3-7d61ab58ebf0 · outbound

This paper cites Qwen2.5-vl technical report, 2025.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Qwen2.5-vl technical report, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:49.288569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:49.288569Z digest=sha256:897eac605420a0ffdbd7dd073533d3649c33a8dab817396cd3f374699043e38b

Observation 063dd15d-e57a-46cc-8888-a1a6b38d0088 · outbound

This paper cites Scaling Biodiversity Monitoring for the Data Age.XRDS: Crossroads, The ACM Magazine for Students, 27(4):14–18, 2021.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Scaling Biodiversity Monitoring for the Data Age.XRDS: Crossroads, The ACM Magazine for Students, 27(4):14–18, 2021

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:49.452368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:49.452368Z digest=sha256:8dcc78b177d94d2f184cbffdd1ef95a775d398bab40885cabcf6d76499b2e386

Observation 9a00a591-e640-4113-a692-43bd9043338b · outbound

This paper cites The iWildCam 2020 Competition Dataset.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors The iWildCam 2020 Competition Dataset

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:49.597250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:49.597250Z digest=sha256:22667b98123f307ee36bade8d177ea71a8f75e45e8db9a0a5e8ed7ef07174e2d

Observation 025e4585-f6e4-4c8d-bd80-87b9b17680f1 · outbound

This paper cites The iWildCam 2021 Competition Dataset.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors The iWildCam 2021 Competition Dataset

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:49.688619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:49.688619Z digest=sha256:a6ba338cd13a081703b55604b130b2a5028c2f9b3d10e7824aa54ef56322a0e1

Observation 1d36f394-b33f-4f4b-b147-4b150541a009 · outbound

This paper cites Deep learning as a tool for ecology and evolution.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Deep learning as a tool for ecology and evolution

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:49.850555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:49.850555Z digest=sha256:acaf286d9e95a5f29cf77cb1d10e7c9b53c837eb9809721d21e9620d9ac965be

Observation d0442136-8455-493c-b23f-d16add554bf3 · outbound

This paper cites Lan- guage models are few-shot learners.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Lan- guage models are few-shot learners

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:50.021573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:50.021573Z digest=sha256:6672a5897053f5e010e9eefa5806c5f1d39021d5ad42d056dc1bffb7a7079a85

Observation 9fbfc295-7d42-4d8b-9a13-6276ce852f40 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Emerging properties in self-supervised vision transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:50.224620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:50.224620Z digest=sha256:142fffd87b1726de09b78cc9d59a84cad7406b1dc6f319245f74fac31fc33301

Observation eee0d3a6-fae4-4fa3-b6ee-a02a1d9e5faa · outbound

This paper cites Big self-supervised models are strong semi-supervised learners.Advances in Neural Information Processing Systems (NeurIPS), 2020.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Big self-supervised models are strong semi-supervised learners.Advances in Neural Information Processing Systems (NeurIPS), 2020

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:50.364650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:50.364650Z digest=sha256:af7c210055b4d877348afc61ef0fc84e0dfcbf8959a42c8675df9c764b53b209

Observation dd82752a-acc8-444d-9fe2-cbe859cb7ccf · outbound

This paper cites Christian Schmidt, Aditya Jain, Yves Basset, Sara Beery, Maxim Larrivée, and David Rol- nick.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Christian Schmidt, Aditya Jain, Yves Basset, Sara Beery, Maxim Larrivée, and David Rol- nick

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:50.519880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:50.519880Z digest=sha256:bde7e31e3f10342fb6adc62fbe2ffd8dcacab207eab2ac4510e737f86a3d0568

Observation d67c3314-a228-4eb5-9e46-6535c52f1e57 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:50.685635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:50.685635Z digest=sha256:7b0020489c443c29c319944606feaf1a781d7a21ae27f2aa2a6607735bea1a37

Observation 9b8cd4f0-47fc-4785-9957-ba9813d0fc6e · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Reproducible scaling laws for contrastive language-image learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:50.852895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:50.852895Z digest=sha256:9a2e441eb0b88a44acd41f2566c2bc5c58b02c31d958d119e40a6765243802b1

Observation 0f1d8706-0ff6-4437-bc38-03fa5a6ab073 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Imagenet: A large-scale hierarchical image database

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:50.937326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:50.937326Z digest=sha256:8ca03b77cb4454ba568bf0f99db97f98a42ca8a4a78c1aae52c30c1fdbe344e6

Observation 3ae8f157-6c47-41ab-aef3-2707ea304b2c · outbound

This paper cites Discovering localized attributes for fine-grained recog- nition.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Discovering localized attributes for fine-grained recog- nition

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:51.008866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:51.008866Z digest=sha256:7f44ecf5e3b7b1838ee9f5a232ea051354cca26f55734ab6dba340213eef60e8

Observation d116ea48-52a4-4ba3-9a24-c7b1c6d67ba8 · outbound

This paper cites Ezray, Drew C.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Ezray, Drew C

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:51.126777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:51.126777Z digest=sha256:ac76cddf6079158c3a53da0ca447077a8dbb10677b71887b7c2b6e027e44f793

Observation 259325bf-355d-4a3a-8f66-c57b05ec608a · outbound

This paper cites open world.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors open world

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:51.218543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:51.218543Z digest=sha256:d399704069e0920f3d03241d6ce62e6514171b3064edde1ca12f56aebffc21de

Observation 467401c1-866d-456c-aa12-ed0fe1c3b1b7 · outbound

This paper cites Clip-adapter: Better vision-language models with feature adapters.International Journal of Computer Vision, 132(2): 581–595, 2024.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Clip-adapter: Better vision-language models with feature adapters.International Journal of Computer Vision, 132(2): 581–595, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:51.259009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:51.259009Z digest=sha256:eee416a6e2a2819f30bed6a482fdaa5885d99d3050bb887053b0cd2d57f3a089

Observation 3f261f1d-74c0-4fc5-aed2-c68c88087b92 · outbound

This paper cites White, James Balhoff, Wasila M Dahdul, Daniel Rubenstein, Hilmar Lapp, Tanya Berger-Wolf, Wei-Lun Chao, and Yu Su.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors White, James Balhoff, Wasila M Dahdul, Daniel Rubenstein, Hilmar Lapp, Tanya Berger-Wolf, Wei-Lun Chao, and Yu Su

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:51.337217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:51.337217Z digest=sha256:24f152a8e61d046b3200dd6948b364a2991dcd27a8a7905c4f7cbbdf6c5023bc

Observation 7a3f2a6b-8f9d-4f99-b55b-167d0a3ca7b6 · outbound

This paper cites an unresolved cited work.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:51.488300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:51.488300Z digest=sha256:05a17f3366fb7a12d04bb2a2125f609e79e1d491ef8eb16d71e3790c55e0253c

Observation 67285c77-58bf-42e2-8f51-fc3fcdacb613 · outbound

This paper cites Deep residual learning for image recognition.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Deep residual learning for image recognition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:51.587844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:51.587844Z digest=sha256:3414c5766646220a0b3ad7fb896d95678b52c5788fb6454b28592d8f410e3dc5

Observation e1fdae11-b9df-499c-a7ef-050aad999a91 · outbound

This paper cites Momentum contrast for unsupervised visual repre- sentation learning.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Momentum contrast for unsupervised visual repre- sentation learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:51.737138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:51.737138Z digest=sha256:8c80834a36ea20f8a34dbf303c262011221c793bb342b035237b617b2e755d28

Observation 4bf98107-27c1-4319-99f9-df24137659e3 · outbound

This paper cites Species196: A one-million semi-supervised dataset for fine-grained species recognition.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Species196: A one-million semi-supervised dataset for fine-grained species recognition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:51.877137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:51.877137Z digest=sha256:0198b63004e1eadf7e58256127910f316ee0242dcdd0006de08514ef2f78b87a

Observation 020d5389-9d81-4b8b-bdf7-878cfbced8f2 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Distilling the Knowledge in a Neural Network

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:51.987413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:51.987413Z digest=sha256:e69efe4eb995127e55b75f762851f46ecb9bc555ed68b4b664be47579faceef0

Observation 897461c0-a785-410d-8605-65d2e79c455a · outbound

This paper cites Deep learning on butterfly phenotypes tests evolution’s oldest mathematical model.Science advances, 5(8):eaaw4967, 2019.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Deep learning on butterfly phenotypes tests evolution’s oldest mathematical model.Science advances, 5(8):eaaw4967, 2019

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:52.064616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:52.064616Z digest=sha256:8efc3aa3287e147b7564dc4ad979da14ab3232ab7e9aa63c4c2d92500b471ab8

Observation a414d1e4-71ac-4481-98e1-1a226ebf001e · outbound

This paper cites Deep learning and computer vision will transform entomology.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Deep learning and computer vision will transform entomology

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:52.147501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:52.147501Z digest=sha256:1b09ea72fb112ba6159e4c846c0bfb251f89351728e23024dd0f0749644342b6

Observation 7b2d6fde-a97a-474a-b259-b550e2ad4ea8 · outbound

This paper cites Multimodal learning and reasoning for visual question answering.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Multimodal learning and reasoning for visual question answering

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:52.230950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:52.230950Z digest=sha256:3a0f048b0e46efc2fe9970ffe28951508ef6002b4e64b37e8e6e45abcf6d4a9d

Observation d3dd95a5-12aa-4ada-a4ce-a737a769ab0d · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:52.313777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:52.313777Z digest=sha256:40b69c6467a988baf20aaaa2508d7c9a152aeb031dae116c671e48ebdbbd695b

Observation b1ee9801-cc5d-4f22-89dc-bddb9f53920a · outbound

This paper cites Vi- sual prompt tuning.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Vi- sual prompt tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:52.395188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:52.395188Z digest=sha256:484229a921d9a0f805136e8d395114609a895943d44f75f1967024577eecd615

Observation a01c224e-b043-443d-9d2f-17963259643c · outbound

This paper cites Many-Shot In-Context Learning in Multimodal Foundation Models.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:52.476377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:52.476377Z digest=sha256:ba7401a1fb86ba2ddd25f7636380827af8f9d688b82f4901622a43f341f1ca3e

Observation 8d96eb85-8150-483c-b595-27d487048f83 · outbound

This paper cites Maple: Multi-modal prompt learning.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Maple: Multi-modal prompt learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:52.559446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:52.559446Z digest=sha256:14c12dd615496a3e4c71442f994c2e02ce651bb9c5ec28826d9fd2c8266a3a29

Observation a446798e-7c50-48b5-b0be-f1a1334d152f · outbound

This paper cites Large language models are zero-shot reasoners.Advances in neural information pro- cessing systems (NeurIPS), 2022.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Large language models are zero-shot reasoners.Advances in neural information pro- cessing systems (NeurIPS), 2022

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:52.637961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:52.637961Z digest=sha256:34564742c4b6259acc35ab4a15d4cbe28149adba39aeb1a71e27032038950b4c

Observation aca7c54a-22b4-4356-9992-69afead9ad97 · outbound

This paper cites Low-rank bilinear pool- ing for fine-grained classification.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Low-rank bilinear pool- ing for fine-grained classification

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:52.724049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:52.724049Z digest=sha256:260c452488fcca07ffe9c55baa10832eda1151d39df916a8e282fdaf9ecb9cbf

Observation 7e573183-e331-4f18-9627-4aa772e0c3d5 · outbound

This paper cites Grounded language-image pre-training.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Grounded language-image pre-training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:52.808500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:52.808500Z digest=sha256:f89ffc85965ecfccc6efb3098632cdd0fd01ccfe53e4dc0f0575cc6d266e357a

Observation 58c88dec-83de-4bb1-b497-97c9cfb8ca6a · outbound

This paper cites Bilinear cnn models for fine-grained visual recognition.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Bilinear cnn models for fine-grained visual recognition

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:52.860548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:52.860548Z digest=sha256:d432d258a565f13e881b76cb58f2e11862d3fe2440f6f90b858dec4e0beeb246

Observation 52922b1b-2bc7-443d-8f78-09216399cd2a · outbound

This paper cites Multimodality helps unimodality: Cross- modal few-shot learning with multimodal models.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Multimodality helps unimodality: Cross- modal few-shot learning with multimodal models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:52.910093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:52.910093Z digest=sha256:b522514fa1521dc89f186c22c2959389d9f9627c795799f3c604712380ccc730

Observation ddb5ed0b-43c3-4cef-bfd1-9140ba496a88 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems (NeurIPS), 2023.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Visual instruction tuning.Advances in neural information processing systems (NeurIPS), 2023

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:52.991903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:52.991903Z digest=sha256:0a3c2fea221be9312ddf7c71e826f15f3917bf29a0a884cc7f0bb1a4ff7e31d4

Observation 652f6c9a-56b5-4660-91db-07814f49d229 · outbound

This paper cites Democratizing fine-grained visual recognition with large language models.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Democratizing fine-grained visual recognition with large language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.066383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.066383Z digest=sha256:8ff2862e7844ff8f0b99e3ddfd5b73cfc2bebfb6c6e9b2bbc9f731b27c1b631c

Observation e68e4722-1374-44f8-963e-cab9829a0a45 · outbound

This paper cites Few-shot recognition via stage-wise retrieval-augmented finetuning.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Few-shot recognition via stage-wise retrieval-augmented finetuning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.115245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.115245Z digest=sha256:548d5ef43c40796d55d632f7134cd533fc94d84d30d673aab782489afb8b7844

Observation ca1c1d35-3cf7-4eed-bd4e-656b785dc892 · outbound

This paper cites RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.190844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.190844Z digest=sha256:adceeb34a0d869ebf51b5660f4ed039cc4325571a96398e1cc3e2c4a85e7b351

Observation 46b1c36b-9f84-4b53-8671-6ac5494d53ef · outbound

This paper cites Rsvp: Reasoning segmentation via visual prompting and multi-modal chain-of-thought.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Rsvp: Reasoning segmentation via visual prompting and multi-modal chain-of-thought

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.230828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.230828Z digest=sha256:3b5b177f58ea14700edc75e9cbf31baf6af83c8b7e6a24eeddd203f0eeecf25a

Observation e6632fcd-e3b2-49fb-97f4-6ffaaa328211 · outbound

This paper cites Computer vision, machine learning, and the promise of phenomics in ecology and evo- lutionary biology.Frontiers in Ecology and Evolution, 9: 642774, 2021.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Computer vision, machine learning, and the promise of phenomics in ecology and evo- lutionary biology.Frontiers in Ecology and Evolution, 9: 642774, 2021

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.265819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.265819Z digest=sha256:b300daa816da48ed547edf66083a4c2f88aef3b8a64df0bb32222a575187e489

Observation 9e1609d5-c897-4beb-ad92-6ebd3999b60f · outbound

This paper cites Lessons learned from a unifying empirical study of parameter-efficient transfer learning (petl) in visual recognition.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Lessons learned from a unifying empirical study of parameter-efficient transfer learning (petl) in visual recognition

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.300062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.300062Z digest=sha256:26c667421373abd11241eab7763548e4bbdd6f5319a91da36fe5021415661f73

Observation 80e96a85-5696-4721-a578-379f078ed6fc · outbound

This paper cites En- hancing clip with gpt-4: Harnessing visual descriptions as prompts.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors En- hancing clip with gpt-4: Harnessing visual descriptions as prompts

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.338422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.338422Z digest=sha256:53d160184aea197ffd4e78a0e4d073733feb0bbe09339551bc195926c0ff58a2

Observation 17d797ad-892d-48c6-9c03-f8511d9a446a · outbound

This paper cites Visual classification via description from large language models.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Visual classification via description from large language models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.379292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.379292Z digest=sha256:ea6c9f104a4d34ee28fda00d275a569cebf274036eb8ab11f09e031de5d8be0b

Observation 292cad3c-99ac-4139-8646-e5a53418d950 · outbound

This paper cites Re- thinking the role of demonstrations: What makes in-context learning work? InEMNLP, 2022.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Re- thinking the role of demonstrations: What makes in-context learning work? InEMNLP, 2022

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.416678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.416678Z digest=sha256:7fcdbedbbc55e44b2808121a4d27ad551082ae238aec1f4e210b4b1dcdf9d773

Observation 07b303fa-c589-4b18-bff5-51ed6f698aa1 · outbound

This paper cites Compositional chain of thought prompting for large multimodal models.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Compositional chain of thought prompting for large multimodal models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.451792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.451792Z digest=sha256:c02341ab22cb38f6345238e3508cf073a55a82208806c66a95247ff6985835d5

Observation a47d12da-8b70-48f3-a83f-6f6cba1497d5 · outbound

This paper cites Norman, Jason A.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Norman, Jason A

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.486008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.486008Z digest=sha256:470a9a92b9b6b028318286968c5a7c8202f09cf640f4b0daf4e26f277cd82d4a

Observation 1aa9f8fb-5af5-45c0-9826-10b2f40e2406 · outbound

This paper cites Inaturalist.Science Scope, 41(7):12–13, 2018.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Inaturalist.Science Scope, 41(7):12–13, 2018

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.510416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.510416Z digest=sha256:af3370708cd914c5e89eb094b502d92a505560dffdc9b7548dc1861b3eca9a32

Observation 3df99a23-1401-453a-a0a7-a2477cc2c872 · outbound

This paper cites Gpt-5 system card.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Gpt-5 system card

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.545495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.545495Z digest=sha256:bfe66906b6f0342f036f4ab0ca7eb98c2562bc848a4e1db3ee886c6fd14968d0

Observation 032ed4b0-959e-4f2b-8180-092b0b1d897d · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors DINOv2: Learning Robust Visual Features without Supervision

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.601756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.601756Z digest=sha256:de149ff946eadb0cf758a600c4f82f6b640e678741e8c3b9f5c9be283163ed1e

Observation c3d98390-11bd-4c08-bd51-0c4dcef2eafd · outbound

This paper cites Prompting scientific names for zero-shot species recognition.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Prompting scientific names for zero-shot species recognition

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.638353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.638353Z digest=sha256:996cade1a8041daa9215b9c97b6b4aa35b95fb6f035829554e15bd268b1b4930

Observation 4ede0685-97ab-40db-9331-57987c6dc25e · outbound

This paper cites The neglected tails of vision-language models.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors The neglected tails of vision-language models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.664309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.664309Z digest=sha256:558ca4570a6b8f81dce81ce68a1c17525dbca4bd26efa9a0f0939e75109eb044

Observation c2834b8f-c987-42f4-be40-f3fa7cf257f9 · outbound

This paper cites Fungitastic: A multi-modal dataset and benchmark for image categorization.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Fungitastic: A multi-modal dataset and benchmark for image categorization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.801502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.801502Z digest=sha256:2780f4ec54f8a5479fd2030c40c4dee0b4bf60d1f829a332db6c72ea2274ef98

Observation ebfb9336-b53e-4819-a03c-6d57f0006e60 · outbound

This paper cites an unresolved cited work.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:53.955334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:53.955334Z digest=sha256:8121c80e8e902f78a9f4ff6ccfc30c1c7c7e4b8df740fd9c46015b440eeb11b7

Observation 6cc7c3ac-2a96-491e-96ac-8355fb545645 · outbound

This paper cites What does a platypus look like? generating customized prompts 11 for zero-shot image classification.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors What does a platypus look like? generating customized prompts 11 for zero-shot image classification

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:54.081400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:54.081400Z digest=sha256:bd7f37bd0f1042753771b9ed31a6b15888aaadcc4050777edfa0b1730b662fa6

Observation 6482fc50-3850-4760-943d-e9d36497e825 · outbound

This paper cites Automated identifi- cation of diverse neotropical pollen samples using convolu- tional neural networks.Methods in Ecology and Evolution, 13(9):2049–2064, 2022.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Automated identifi- cation of diverse neotropical pollen samples using convolu- tional neural networks.Methods in Ecology and Evolution, 13(9):2049–2064, 2022

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:54.191504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:54.191504Z digest=sha256:8a58135790c67c4cdffe518e36b0c9ea6b9e2baea7a311ca1b00aeab2e887358

Observation 23b6cc76-c598-477a-a133-7dda7da88c8c · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Learn- ing transferable visual models from natural language super- vision

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:54.357240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:54.357240Z digest=sha256:633bfb9383233c6817e3e12a9d0963f19eed15ea01f66909582fcc25463ea83e

Observation 4cd5ebc1-8bf7-4367-9fc2-3809b13a232b · outbound

This paper cites an unresolved cited work.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:54.468242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:54.468242Z digest=sha256:5b1febdd102c4e4b218ac61fd401138ebd5b44099e90f3a92ae569caebb5ee2d

Observation 40fb7c41-b96a-403f-b9ff-229b638f1f9c · outbound

This paper cites Im- proved zero-shot classification by adapting vlms with text descriptions.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Im- proved zero-shot classification by adapting vlms with text descriptions

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:54.628501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:54.628501Z digest=sha256:5735497767019fcd1ffe09239ad5c88bc6c5c30fa90448414260cc3a3202e8c9

Observation 0e869490-7ea9-4451-b34c-925125b40814 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:54.770955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:54.770955Z digest=sha256:89fa70ad7dfdad3c65972abc13e01cd8413fb742be563f7b386749b3514efe33

Observation 93a130d3-e95e-41a4-805c-332ba9e1d43b · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:54.841060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:54.841060Z digest=sha256:db4a2d0df970d98aa878aca8eb1f614cc438eb989bd3fdd46863f55fca8c6870

Observation 5640011c-d0d7-4fa3-9300-be652e66b479 · outbound

This paper cites A closer look at the few-shot adaptation of large vision-language models.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors A closer look at the few-shot adaptation of large vision-language models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:54.955759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:54.955759Z digest=sha256:175434e0f5a071bd0590dddbb441b5c592594c45fb3e2f65db6ca41843f0a4f7

Observation 992727dd-2af2-48c3-9909-b3532eb1a458 · outbound

This paper cites Scaling-up camera traps: monitoring the planet’s biodi- versity with networks of remote sensors.Frontiers in Ecology and the Environment, 15(1):26–34, 2017.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Scaling-up camera traps: monitoring the planet’s biodi- versity with networks of remote sensors.Frontiers in Ecology and the Environment, 15(1):26–34, 2017

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:55.037052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:55.037052Z digest=sha256:a7218fed21daf5ed09f17bcd4418ef4ce9363b113a7afe7e51241b0ae02b5613

Observation c2d8cc07-21b2-4e44-a6f1-e65e05b370da · outbound

This paper cites BioCLIP: A vision foundation model for the tree of life.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors BioCLIP: A vision foundation model for the tree of life

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:55.190028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:55.190028Z digest=sha256:e87a0aada8a811cdb40662c6a82c9b808959b8873f2e3ee3bf80c99e881b48d9

Observation 8f0ef4bb-d848-45cc-b7ea-5576c8404f04 · outbound

This paper cites The Semi-Supervised iNaturalist-Aves Challenge at FGVC7 Workshop.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors The Semi-Supervised iNaturalist-Aves Challenge at FGVC7 Workshop

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:55.270657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:55.270657Z digest=sha256:030e351579818b6f332dcb8cc0fcb1406b9cfd828da92748cec5ac526126617b

Observation 096c0bdc-7cf9-4eb4-b4e5-d925e71f8a94 · outbound

This paper cites A real- istic evaluation of semi-supervised learning for fine-grained classification.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors A real- istic evaluation of semi-supervised learning for fine-grained classification

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:55.356710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:55.356710Z digest=sha256:9a7d8746cb6bd989a49cbf8bbabce4a80c7517068e9d086989ae2fb54b49ff3c

Observation e0e4b60b-d278-4515-84c5-d03fe65b44c1 · outbound

This paper cites Gemini 2.5: Pushing the frontier with ad- vanced reasoning, multimodality, long context, and next generation agentic capabilities, 2025.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Gemini 2.5: Pushing the frontier with ad- vanced reasoning, multimodality, long context, and next generation agentic capabilities, 2025

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:55.457981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:55.457981Z digest=sha256:8eb4589853487b5575dcfd2acc23d0e95fd20b44a60c22531a6d80e4ce917c98

Observation 1857c0fd-1ea7-4930-96fb-e21aff3b501d · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Gemini: A Family of Highly Capable Multimodal Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:55.546044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:55.546044Z digest=sha256:6f8a911bec703f3f9e6619c6e9d4fd67549e4d4bcccb729f3bc7ecae4d75b55d

Observation 4419eb91-a866-412b-8b97-a68289c75fe9 · outbound

This paper cites Glm-4.5v and glm-4.1v-thinking: Towards ver- satile multimodal reasoning with scalable reinforcement learning, 2025.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Glm-4.5v and glm-4.1v-thinking: Towards ver- satile multimodal reasoning with scalable reinforcement learning, 2025

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:55.637210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:55.637210Z digest=sha256:3db082b29aebda4d890285395093e650a619551b5f73f7bae81ef802154921ab

Observation 257e3c93-b3c0-4497-a4ff-13f90440c814 · outbound

This paper cites Bird Distribution Modelling using Remote Sensing and Citizen Science data.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Bird Distribution Modelling using Remote Sensing and Citizen Science data

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:55.741333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:55.741333Z digest=sha256:8b73828d4d081128c4068d601696ecd0adbe9103e3c25a2afa25e4158ea5c325

Observation 6e23718d-a86a-47fa-939c-e065e1fedae3 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:55.851986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:55.851986Z digest=sha256:c83e9f08da1ea96096f4bc505420b44a5d2e28c88d78bf3d3b9eb866f47cc390

Observation 15a1affe-c2e1-468f-aa33-61f9de66ac35 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors LLaMA: Open and Efficient Foundation Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:55.989907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:55.989907Z digest=sha256:bf7b4f8e825733b72e80b0e0146fe5ec445325e49d91e95ddd75f841fc903175

Observation 1889cb53-ad26-4be2-b948-9b53c541ec54 · outbound

This paper cites Perspectives in machine learning for wildlife conservation.Nature communications, 13(1):792,.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Perspectives in machine learning for wildlife conservation.Nature communications, 13(1):792,

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:56.058377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:56.058377Z digest=sha256:80b16c36513dd07c754441045b4874fd3b46507136bba7fa5667c2a84c6a623f

Observation 08c51221-cede-4954-b399-006042e4eb5d · outbound

This paper cites iNat Challenge 2021 - FGVC8, 2021.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors iNat Challenge 2021 - FGVC8, 2021

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:56.114390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:56.114390Z digest=sha256:cf1934bb0f5567a1ac530c934380d7edcdcc54242c27b14ea802fa583767da1e

Observation aa69fb4a-97d2-4424-bd3c-faccaccb598b · outbound

This paper cites The inaturalist species classification and detection dataset.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors The inaturalist species classification and detection dataset

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:56.214406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:56.214406Z digest=sha256:ec9ddb8de7ef351ea85bcd871a26f6c75aff0e59866566c3feb34cfb478da998

Observation 7914cdd8-0a0c-4741-84b2-4c340dc22840 · outbound

This paper cites Enabling val- idation for robust few-shot recognition.arXiv preprint arXiv:2506.04713, 2025.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Enabling val- idation for robust few-shot recognition.arXiv preprint arXiv:2506.04713, 2025

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:56.378174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:56.378174Z digest=sha256:db73794d2c99b5134e0e141c311a52105eac3c032b4ca56af1974b4f5fbc2785

Observation 041b29a9-730a-44e9-ab79-9a52f2c14551 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:56.556327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:56.556327Z digest=sha256:c7a516a2ff90dd1558389dab56565d5b4bc4dbac58afdde28c067e99a36b2a96

Observation f6ce78e3-5d81-45a2-b963-88a3f6c0d20a · outbound

This paper cites Self-consistency improves chain of thought reason- ing in language models.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Self-consistency improves chain of thought reason- ing in language models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:56.655136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:56.655136Z digest=sha256:19438e6f2ba27dc358f64e4e7ce49650cf60b018d31d98c7a2d315d289ade567

Observation c47bf2cf-ea3f-4133-a9c1-a069f958b3ee · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Chain-of-thought prompting elicits reasoning in large lan- guage models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:56.759429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:56.759429Z digest=sha256:7de10316340ada614908fb867655e5b3e8ef1612208f1b0840adbea7597da756

Observation 265a5255-3ef1-429a-b037-58327ee3c24f · outbound

This paper cites Large language models are better reasoners with self-verification.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Large language models are better reasoners with self-verification

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:56.856946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:56.856946Z digest=sha256:9e6ea9f12fbf3aff1046b97bd76aae4fefff55afe413c4e95a91476382516860

Observation 282b710f-f4be-4c31-82be-6f31cd5732b6 · outbound

This paper cites Robust fine-tuning of zero-shot models.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Robust fine-tuning of zero-shot models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:56.969004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:56.969004Z digest=sha256:d8d085860ab896c21c9bec5e29c087e3d2e7812cceb22d6c2cc08e5ae89bd763

Observation 1f2a6376-5c7b-403a-8929-9a67c87e1be2 · outbound

This paper cites Pro- tect: Prompt tuning for taxonomic open set classification.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Pro- tect: Prompt tuning for taxonomic open set classification

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:57.120474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:57.120474Z digest=sha256:13193e05aba8648e87b1829eeb42ea114ebade61c65cd9dd582b5080f21656da

Observation bc1791b6-64c2-4fac-9142-e43464e0add9 · outbound

This paper cites Demysti- fying CLIP data.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Demysti- fying CLIP data

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:57.268383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:57.268383Z digest=sha256:b1120c0fdf768a4c4f742293470b90307ce35c0e49d45e70e8b1e40216b462e5

Observation 73cc0f6e-18ac-4011-9d2e-9183a1b43033 · outbound

This paper cites Learning concise and descriptive attributes for visual recognition.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Learning concise and descriptive attributes for visual recognition

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:57.355940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:57.355940Z digest=sha256:3cb23fbf4fe415f8d26c586461e4405d794dca6bfad5a650417b663e11272ab8

Observation 0ece8779-db63-4ed4-98c1-a8ef7a6fdcbf · outbound

This paper cites Qwen2.5 Technical Report.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Qwen2.5 Technical Report

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:57.416538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:57.416538Z digest=sha256:665356aab6f9aa2d6259a175462aac1e6d500618e8567ba046b14eadf769afdc

Observation 0123dc96-752f-486c-872d-5a09f020e4d2 · outbound

This paper cites An empirical study of gpt-3 for few-shot knowledge-based vqa.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors An empirical study of gpt-3 for few-shot knowledge-based vqa

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:57.488624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:57.488624Z digest=sha256:b170e71a13e580adcea293fabb974061ebd61c8fb9573e879ee852d476aa8145

Observation 778cb526-d5d3-404f-89f4-6db2ac3b2869 · outbound

This paper cites Visual- language prompt tuning with knowledge-guided context op- timization.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Visual- language prompt tuning with knowledge-guided context op- timization

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:57.564993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:57.564993Z digest=sha256:5ba14cbed12fac4a434c3d92bc2a0608826beb66b91fc3294798c7955be57113

Observation a58b1ec4-2516-4354-8b08-cc310e32ee7d · outbound

This paper cites Image captioning with semantic attention.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Image captioning with semantic attention

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:57.636392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:57.636392Z digest=sha256:06cb82113af4b1a3b74d9da808d90a4bc1ac62e07d692f92e5e50dd25e54e3b5

Observation 58e0763b-799b-4e00-833d-f07ea5e612ac · outbound

This paper cites Task residual for tuning vision-language models.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Task residual for tuning vision-language models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:57.681087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:57.681087Z digest=sha256:79ab0b5c0c63ee99314e7bcc635cad1ebf67dc1c21c6eff55a96587786e83dc2

Observation 50994f92-3f34-47c9-a1ce-3709b1264338 · outbound

This paper cites Part-based r-cnns for fine-grained category detection.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Part-based r-cnns for fine-grained category detection

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:57.728666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:57.728666Z digest=sha256:ca577c7b5a0dbd69d35edeb6afcf513781114900878c064d95db05ed1e668165

Observation a1ba8331-1068-4aa8-91cd-20e1cb7247c1 · outbound

This paper cites Revisiting semi-supervised learning in the era of foundation models.Advances in Neural Information Processing Systems (NeurIPS), 2025.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Revisiting semi-supervised learning in the era of foundation models.Advances in Neural Information Processing Systems (NeurIPS), 2025

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:57.770948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:57.770948Z digest=sha256:2387e088507761f0aa885d54cebe733e42c11e224bd559e8103b959502c0735f

Observation 34a446c3-8bff-4cce-9c39-637a136a0c69 · outbound

This paper cites Tip- adapter: Training-free adaption of clip for few-shot clas- sification.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Tip- adapter: Training-free adaption of clip for few-shot clas- sification

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:57.837437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:57.837437Z digest=sha256:6a00102aa2f96df3e165087f6d419e5d904122d7123821a5d0962f0544bd1446

Observation 5ab9d9f4-4262-4ddf-8d97-80c541c28212 · outbound

This paper cites What makes good examples for visual in-context learning?Advances in Neural Information Processing Systems (NeurIPS), 2023.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors What makes good examples for visual in-context learning?Advances in Neural Information Processing Systems (NeurIPS), 2023

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:57.934252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:57.934252Z digest=sha256:71cc97697da9989ec12b11c5e8781c2c706be10c28b74f4a25a22a1ee22964ad

Pith citing papers

No inbound Pith citation observations are available.