Pith. sign in

Paper Citation Record · LEDGER

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features

As of 19 August 2026, this Paper Citation Record lists 100 of 115 outbound references and 3 inbound Pith citation observations for arXiv:2412.00142.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00142 v3

Coverage vector

measured 100 of 115 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:24:06.917263Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:02:42.424201Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T06:17:26.610695Z

Reference resolution

100 of 115 outbound references displayed

  • verified exact3
  • verified fuzzy13
  • unresolved84
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ac2bcad2-f61d-4522-8da8-0bb0092f96d9 · outbound

This paper cites an unresolved cited work.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.494976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.494976Z digest=sha256:b05afc4e0ae18d8aa5b4f4c975e623214688826cb238538cbb91c9dd00d9981b

Observation 2b0af4f5-1f9b-4e94-ae4e-48600d8aca7f · outbound

This paper cites Vqa: Visual question answering.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Vqa: Visual question answering

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.499604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.499604Z digest=sha256:c02a482d5e09bdba2826fe59bf4cac868ab44c77a6abd595e840fca697713952

Observation c723246f-b926-4169-b19a-6ad76d6bbf5e · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.504091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.504091Z digest=sha256:b2bf2aa88c24d5ada76143ac533ae20157cb6a030ed7f577b088a2471e0f5663

Observation f9217092-aa73-4eb3-8237-d351cacfe238 · outbound

This paper cites Autoencoders, unsupervised learning, and deep architectures.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Autoencoders, unsupervised learning, and deep architectures

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.508859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.508859Z digest=sha256:b09e43bb395837d12a783c3056de34b060f49c64c45418722050ba15330b43c0

Observation a2f0c454-b8df-4a3a-b022-b71a48b53d09 · outbound

This paper cites When XGBoost outperforms GPT-4 on text classification: A case study.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features When XGBoost outperforms GPT-4 on text classification: A case study

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.513025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.513025Z digest=sha256:2578888751a2486a088f5b6efbadf7dc4ce6f33991333cf4808d78be624194c0

Observation 2eb0f0be-de19-49d0-b26a-e5c26d488c06 · outbound

This paper cites Language Models are Few-Shot Learners.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Language Models are Few-Shot Learners

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.517399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.517399Z digest=sha256:dcb9d4f5d6debdd8c0e794c92aa713d473de5e407bf1b807d69725610957d665

Observation ca244d32-4943-4a39-b712-4a115bbb80ac · outbound

This paper cites Fine-Tuned 'Small' LLMs (Still) Significantly Outperform Zero-Shot Generative AI Models in Text Classification.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Fine-Tuned 'Small' LLMs (Still) Significantly Outperform Zero-Shot Generative AI Models in Text Classification

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.522307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.522307Z digest=sha256:dfe681c0eef23a57b1ba57aea35c19aee4cf09ad50774518dbcc3c77a67e17cf

Observation 523fbd7a-de83-4964-8fd9-4630c910dcde · outbound

This paper cites Unified Hallucination Detection for Multimodal Large Language Models.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Unified Hallucination Detection for Multimodal Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.527276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.527276Z digest=sha256:2c87fdbc1c8277579d83b5e94431a86e6e7a81f0bb3ac69f454e160eb70dda49

Observation bb6aebae-89ec-4ba8-ac22-de904549b4c3 · outbound

This paper cites Imagenet: A large-scale hierarchical im- age database.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Imagenet: A large-scale hierarchical im- age database

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.531609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.531609Z digest=sha256:53db50283b985f16f5a5c633975a6245bb286000813bdd440a6a7c7979676a56

Observation fb04a755-f9d4-495d-8378-f998781e6d2e · outbound

This paper cites Llms to the moon? reddit market sentiment analysis with large language mod- els.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Llms to the moon? reddit market sentiment analysis with large language mod- els

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.535646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.535646Z digest=sha256:ed6e80bb501cc4d9d9d71fb50e067e180f5995e6f1d4d8314115ab9032210e97

Observation f3e23a81-ccd2-4f07-9be7-d0e605ac8e6f · outbound

This paper cites Virtex: Learning vi- sual representations from textual annotations.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Virtex: Learning vi- sual representations from textual annotations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.539633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.539633Z digest=sha256:3d8dbd013d635182138b015f72d76808a9f7418b9b62478bf9fbff8a692bdca1

Observation 9f2c5cdc-a32c-49f0-9f38-1a2faea766c4 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.543840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.543840Z digest=sha256:68beea28cbaa6aa7fd4209533f4824df7cad626000c8ec95ab23ed9b7efd090f

Observation eb58b3b7-75ea-4df7-8a7e-2fbf7cb6eb02 · outbound

This paper cites Towards Multimodal In-Context Learning for Vision & Language Models.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Towards Multimodal In-Context Learning for Vision & Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.548314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.548314Z digest=sha256:5053daa999369a82361fae762ccf41e0af996cc630fdd6b28d4164ce55812fd1

Observation df53f3e3-8d29-42fe-9603-2fec24dae47a · outbound

This paper cites Training Vision Transformers for Image Retrieval.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Training Vision Transformers for Image Retrieval

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.552631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.552631Z digest=sha256:6c763f093b77e8a4517eee2b4edc3290537cdab210aa9118df9006a34a2a3108

Observation 5194c9a0-ac5f-40ad-8991-e8ae836711df · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features KTO: Model Alignment as Prospect Theoretic Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.557031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.557031Z digest=sha256:7d6266f944064cc9c70be2bfdc37713ccb14061998baf7106a9be3057d349692

Observation d461ff02-49ca-4f0f-b995-dd15e6850e4e · outbound

This paper cites Behr, and Nancy Kan- wisher.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Behr, and Nancy Kan- wisher

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.561574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.561574Z digest=sha256:9fd6a0710b42ea791b2388fffa66c2040635f72486d354ddb0baadfef52caadf

Observation e878cf0e-87c1-4f84-935a-a03c82514ac9 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.565408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.565408Z digest=sha256:acaedbd23816db56ee6242f1a38bec4915fc67d08e2d1d3015e4d25ec4b9db0a

Observation 8a04b953-ddb6-4016-954f-60eab479891d · outbound

This paper cites SimCSE: Simple Contrastive Learning of Sentence Embeddings.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features SimCSE: Simple Contrastive Learning of Sentence Embeddings

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.569011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.569011Z digest=sha256:a3b56280e55a54c1b446a47b0c29a0a1dcba1b766b0c52386afc298010776f63

Observation ba3caaae-440e-4a34-952d-de863883b8e9 · outbound

This paper cites an unresolved cited work.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.573075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.573075Z digest=sha256:f91b0667cf3926b8a324603a1394ea2f777f14db46ffd1ad33e7fd5f26c6b465

Observation 0d28298b-388f-478a-9ab5-b897417304ed · outbound

This paper cites MERLIN: Multimodal Embedding Refinement via LLM-based Iterative Navigation for Text-Video Retrieval-Rerank Pipeline.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features MERLIN: Multimodal Embedding Refinement via LLM-based Iterative Navigation for Text-Video Retrieval-Rerank Pipeline

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.576490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.576490Z digest=sha256:b54383d60b2701b8685a556ac3cf58bb6bb4d763ddab7277f31e36b978a9cac3

Observation 9451e03c-4b15-4071-9043-e8152a9dd315 · outbound

This paper cites Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.580093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.580093Z digest=sha256:9c8be580d4a86752ae8287c8177ead4c968aa374b5780c0617958053ae0c7716

Observation 4d606d26-7a71-4abf-93d3-d203d269c733 · outbound

This paper cites In-Context Learning Creates Task Vectors.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features In-Context Learning Creates Task Vectors

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.587897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.587897Z digest=sha256:8c613c345c8497f37878ec50716e580b074e4186e8fc177015034d1829e81b15

Observation e74514f7-cbf2-4859-8909-dd1f81a646db · outbound

This paper cites Incorporating structured representations into pretrained vi- sion \& language models using scene graphs.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Incorporating structured representations into pretrained vi- sion \& language models using scene graphs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.591537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.591537Z digest=sha256:fbd07e4bce227d1121842f2edc79fb3531ac11340c54ffe934f6b7c2b6865470

Observation aeda20c5-fa01-45f1-808a-bce0fd2ad5f2 · outbound

This paper cites Hinton G, van der Maaten.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Hinton G, van der Maaten

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.595634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.595634Z digest=sha256:0caa5ba2f096a91f5bef15006b376908f149ccb07cfc9fad277b61775c91ea7c

Observation b3daabcc-1716-4cd4-a045-f9f13176563e · outbound

This paper cites Finding visual task vectors.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Finding visual task vectors

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.599489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.599489Z digest=sha256:39b5c80e53697ec636e8222581b2d354af5f5d565e0d1c392f6af367ca211252

Observation 74179b86-648b-4da6-a344-f76e81323af5 · outbound

This paper cites SugarCrepe: Fixing Hackable Benchmarks for Vision-Language Compositionality.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features SugarCrepe: Fixing Hackable Benchmarks for Vision-Language Compositionality

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.603506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.603506Z digest=sha256:5d03d86c662f21808392c3893a510200256f4e111fec4b6a04b03aafc48e602c

Observation 4b4a7ca8-970e-422b-9686-e7b598e922e9 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features LoRA: Low-Rank Adaptation of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.607573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.607573Z digest=sha256:0891379db786b0d0f969304964359c0d166d0f612687ac2253b4096a56fa028f

Observation 5be0f649-ec2a-418a-abd1-682a70217bc8 · outbound

This paper cites Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.611337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.611337Z digest=sha256:d371bd56e4206d61318b0d96b70b152fdb78f765142066c09704da1570979cd6

Observation f7998ba8-cbd3-4ebe-a15e-14b1a6eadab3 · outbound

This paper cites Llm2clip: Powerful language model unlock richer visual representation.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Llm2clip: Powerful language model unlock richer visual representation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.615370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.615370Z digest=sha256:b3cffa2f46ea00269084829b4214d435cdb22f42f41803cec1d95b8fac039236

Observation 1d62f176-d6c9-4df6-91af-6e84b65d0a47 · outbound

This paper cites Hudson and Christopher D.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Hudson and Christopher D

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.619429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.619429Z digest=sha256:061d717fddf64ad71cd6bcdf09c6404036a2b1b0dbdcc393587406572c507f5a

Observation 5bce9494-9555-495c-83a8-a5e40749a083 · outbound

This paper cites Learning Object Detection from Captions via Textual Scene Attributes.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Learning Object Detection from Captions via Textual Scene Attributes

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:24:07.463692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:24:06.623314Z digest=sha256:05f1322d6b35ab7e424e6d97ef3780e43a85a0200ae72f0afed5a975cf361d9b

Observation 39148b52-46c9-476e-8be6-5ee3e96d056b · outbound

This paper cites Promptbert: Improving bert sen- tence embeddings with prompts.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Promptbert: Improving bert sen- tence embeddings with prompts

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.627346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.627346Z digest=sha256:4775ceb577a505cb0f19eda399db69202625a81524e44dcae797bad80088a754

Observation c6802af6-95e4-4021-927d-ee99eaa91656 · outbound

This paper cites Scaling Sentence Embeddings with Large Language Models.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Scaling Sentence Embeddings with Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.631097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.631097Z digest=sha256:5167c8a949f6c3b28cb581e8b4181c221d079cea7bf3b9b9a5a1b1765a06322d

Observation 237a0d01-dcd2-4b71-b099-f69379f6f25c · outbound

This paper cites E5-V: Universal Embeddings with Multimodal Large Language Models.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features E5-V: Universal Embeddings with Multimodal Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.635140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.635140Z digest=sha256:7ec4e0ab96f3d2dc2460497de915bbda095886d21d129aa39f84f95525ccfa45

Observation d7af35b5-88f0-4098-8bdf-6118b18bba8f · outbound

This paper cites Domain specificity in face perception.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Domain specificity in face perception

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.639187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.639187Z digest=sha256:94f58ac66321457188d9f1087a407e730f8f7ee647ea7b4368209f969a0d595f

Observation 382a63d8-e84e-4863-bc52-dfba30d2f81f · outbound

This paper cites an unresolved cited work.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.642868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.642868Z digest=sha256:b730bc4a2f8987dd4af6b59d122b803c4142b11332357a844b74a42f20caf8ae

Observation 83ef24da-f997-4211-8e96-2834f6894f9c · outbound

This paper cites Auto-Encoding Variational Bayes.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Auto-Encoding Variational Bayes

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.646670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.646670Z digest=sha256:a7887dca94685f13abbc121a7bcaab4e9a08d4734e5d0dc72bfd8f21f63527a0

Observation 910264f5-b8ec-440c-a391-c7395233d91d · outbound

This paper cites Autoregressive image generation us- ing residual quantization.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Autoregressive image generation us- ing residual quantization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.651010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.651010Z digest=sha256:539bfde68a90ae745878ec2313f880a7fbdff2d0870916168be6545e89cc2f80

Observation d4161a6d-6e21-4c15-95eb-f4aebf9e4a94 · outbound

This paper cites Meta-task prompting elicits embeddings from large language models.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Meta-task prompting elicits embeddings from large language models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.654901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.654901Z digest=sha256:edce4ef151ee2d3ac2d9d7a5dc68ae1e4d482c5c640192f1e59b72bbba94b539

Observation 9690ea70-a02c-469a-9577-a87499e1b821 · outbound

This paper cites The power of scale for parameter-efficient prompt tuning.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features The power of scale for parameter-efficient prompt tuning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.659016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.659016Z digest=sha256:aeef4e130ebfda72f2e52ffbfe66ba0c28fbaee532d67710e513cc4b80dca5dd

Observation 9c700684-3f0a-4297-863e-d216f8f1f69c · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.662505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.662505Z digest=sha256:1c38f6f1729f5e6d9c5c54686ad96535b3b5c44674e741cc1e96a46e7ca0dae9

Observation 17c967a0-1d3c-453f-b983-174d6b3af53b · outbound

This paper cites NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.666386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.666386Z digest=sha256:ea1328dd97124353afb0309358ed3e6bbdb51744f54472ff31baea7d0a28ba63

Observation ac48e6c0-5c35-4fbb-8baf-fb982d627ce6 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features LLaVA-OneVision: Easy Visual Task Transfer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.670488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.670488Z digest=sha256:0f2bf68cc302f324656dbab0ffee79053076b4e44e3f0ba3bf4cae5c0ff5cf1d

Observation eed1a4c1-c4c0-487e-837b-c911053a739d · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.674598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.674598Z digest=sha256:79042cc323d1d780e60e81b7e385149ef539cd3be7fea8418248b2a431ef0d8b

Observation a0f68e6c-677c-4fd6-98cc-0496d8fdf931 · outbound

This paper cites BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.678514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.678514Z digest=sha256:17cc15e2a488ef4903c85a80b6a90382a40e9e478be22b4dc64eb355bd89e591

Observation 417cda0d-77b3-4fe0-b33d-6b42dddfdb36 · outbound

This paper cites Conan-embedding: General Text Embedding with More and Better Negative Samples.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Conan-embedding: General Text Embedding with More and Better Negative Samples

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.682186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.682186Z digest=sha256:1cc306ff81833757a1ffe00b082c2326515a66436a0972cad00f5d4ae3b2f71b

Observation 14dceaea-d9d6-4ae5-9a85-08b932e4f859 · outbound

This paper cites Prefix-Tuning: Optimizing Continuous Prompts for Generation.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Prefix-Tuning: Optimizing Continuous Prompts for Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.686196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.686196Z digest=sha256:a88d972c6b6653fe73af8f7ef11b0ee32279b4cf4b01fe666ec9040e6cd66228

Observation 42edebf6-d422-4266-9906-fe60478ae4e0 · outbound

This paper cites Scaling language-image pre- training via masking.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Scaling language-image pre- training via masking

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.690378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.690378Z digest=sha256:c7e179c7e2f4872a41477bf8498c474f418d1e4b1c40cb7406b105a24e8a002c

Observation 970b216e-bf6d-4868-a1c3-aa7efe0d6b62 · outbound

This paper cites Your Mixture-of-Experts LLM Is Secretly an Embedding Model For Free.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Your Mixture-of-Experts LLM Is Secretly an Embedding Model For Free

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.698630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.698630Z digest=sha256:586fa3a44b68b504371f3da07b26b4c00579b47747ede1f2f1ae16bf9fdb9da5

Observation d6aef712-9564-4205-a949-f92b788859cd · outbound

This paper cites GT2Vec: Large Language Models as Multi-Modal Encoders for Text and Graph-Structured Data.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features GT2Vec: Large Language Models as Multi-Modal Encoders for Text and Graph-Structured Data

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:24:07.330290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:24:06.702316Z digest=sha256:60ce6846bce0b3ee4e1de359e386b09a3bc59f30adb0c457fb0786e548086a39

Observation fecdeb87-71e5-4db3-b77f-bc6f4f83444f · outbound

This paper cites Maire, Serge J.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Maire, Serge J

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.707493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.707493Z digest=sha256:0efeb77391b1f31651bf439b26f3bad63aa4dd46270ac4fd5b7c99248d276c05

Observation 32744962-ab80-4aa2-9641-9069a427bd09 · outbound

This paper cites Revisiting the Role of Language Priors in Vision-Language Models.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Revisiting the Role of Language Priors in Vision-Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.711881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.711881Z digest=sha256:f346359510f0a05b954a172a4fbf9fb7537687d0eae6757f00941a597e25a2b8

Observation 7eae4df1-b646-42d4-9a84-4703e2aef339 · outbound

This paper cites Multimodality helps unimodality: Cross- modal few-shot learning with multimodal models.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Multimodality helps unimodality: Cross- modal few-shot learning with multimodal models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.716359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.716359Z digest=sha256:ca4dbd86f76e62ebef92fb870c19061f3eb18bcd388801b88865190d83d0a532

Observation f529054a-834f-4495-80b8-93c892edf478 · outbound

This paper cites Evaluating text-to-visual generation with image- to-text generation.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Evaluating text-to-visual generation with image- to-text generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.720415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.720415Z digest=sha256:c272e31a1c77f144622cdcda56056246ad2f8c329f07ed379ed76e396bd243fe

Observation f3642c95-4b37-43e5-a23d-d6fa293ad50b · outbound

This paper cites Evaluating text-to-visual generation with image- to-text generation.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Evaluating text-to-visual generation with image- to-text generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.724414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.724414Z digest=sha256:a437a639ebfc3dec3cd174d40f3cffa68e87115132a4a606f1c66853fc5c2494

Observation 0ac9665c-8893-4879-a96f-923ed2e325ab · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Improved baselines with visual instruction tuning, 2023

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.728782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.728782Z digest=sha256:e9d86ae4832a21fa90d04479223a1af5b0ce2f3c13deb3613c08e3bc12a09037

Observation 61d8578c-b584-4850-a45e-76d8c05952fc · outbound

This paper cites Visual instruction tuning.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Visual instruction tuning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.733026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.733026Z digest=sha256:e6000cd345ac9d3e2bf7fc733af0393b9c9156974773f5eb4869e37cb9c37368

Observation 81031fed-02db-49cf-93c1-094d24034ad3 · outbound

This paper cites Meaning Representations from Trajectories in Autoregressive Models.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Meaning Representations from Trajectories in Autoregressive Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.737206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.737206Z digest=sha256:d4aea7aa633a996ccedc8bca4161d7de7157ecae94743237aa21c0fe472c3c34

Observation 3808aeeb-1ae9-4b0e-bee6-da5ed897c2b9 · outbound

This paper cites Decision-making with auto-encoding variational bayes.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Decision-making with auto-encoding variational bayes

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.741457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.741457Z digest=sha256:46fd993b4ae9efb61f868f198a5c7a16f89c481160f0a450e56363c8a9c1e026

Observation 860f3a34-4cc4-46b0-a8e6-eba528d8b32f · outbound

This paper cites Decision-making with auto-encoding variational bayes.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Decision-making with auto-encoding variational bayes

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.746257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.746257Z digest=sha256:d80993de550d6c9ae3f523858517ed2508c19b02fb48f83aaa101fcf64e6b3ca

Observation 10691041-26b8-4e48-b5e5-e17e4edce9ba · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.750576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.750576Z digest=sha256:8dff9c4e70c00a535e016621f8946d9e32f08924c55444e42f66db4b9a7d2a3e

Observation 10a89739-248d-433c-a1cc-7c3a993b0531 · outbound

This paper cites an unresolved cited work.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.754674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.754674Z digest=sha256:5feea180439f78fb7d271eae86ce2997e2deb87441821314712db794a8cb60ac

Observation a571fcb7-0b4d-464c-90b0-d04a4c02e882 · outbound

This paper cites Distributed representations of words and phrases and their compositionality.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Distributed representations of words and phrases and their compositionality

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.758709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.758709Z digest=sha256:b92478ba5cafbb5ba575def95e866e6dd876afadbb6e451368de98a079214c5c

Observation a0e5e3f9-82ab-4d97-b0c0-b574445b0d8f · outbound

This paper cites Compositional chain of thought prompting for large multimodal models.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Compositional chain of thought prompting for large multimodal models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.762358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.762358Z digest=sha256:e2af20b513231b3ab425d5640dd132b2d68a710f2f3943fc9f0f05777c2a8793

Observation defbcfcb-9ee4-4916-a941-7e112781b921 · outbound

This paper cites Mteb: Massive text embedding benchmark.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Mteb: Massive text embedding benchmark

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.766043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.766043Z digest=sha256:6a48dd8ddb92bd1b56365ac47c252640108b2326674dacacb52e812c94f262c6

Observation fc664de4-3dde-4be6-9293-770fa48bb75d · outbound

This paper cites Sentence-T5: Scalable Sentence Encoders from Pre-trained Text-to-Text Models.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Sentence-T5: Scalable Sentence Encoders from Pre-trained Text-to-Text Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.770554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.770554Z digest=sha256:b25bbd7bbd3d28eb8fc624748ae4ba75263679f7cb25e28b86541e7951134227

Observation 3091e1f7-985f-49cf-8c5c-7448981a64db · outbound

This paper cites Large Dual Encoders Are Generalizable Retrievers.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Large Dual Encoders Are Generalizable Retrievers

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.775124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.775124Z digest=sha256:746e726a723aa31457a0caeec4f2f1e39bfcd7271a4bc4d70c560997b91063fc

Observation 0670cd0b-a8c1-446e-9b06-3a9ee455f178 · outbound

This paper cites Automated flower classification over a large number of classes.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Automated flower classification over a large number of classes

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.779744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.779744Z digest=sha256:b8e9e3952c0996af2db7e3e9d72a9d82e88acb8e240541a3bf17bf4f254fc953

Observation 0e46d77f-b919-4518-9962-f5637ed0830b · outbound

This paper cites In-context Learning and Induction Heads.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features In-context Learning and Induction Heads

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.784087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.784087Z digest=sha256:9635de58bd81d3cfa8a81cb634157f4b59962d4df6d2f12cb444dc0ea995ebed

Observation 5c1338aa-59ca-4408-93e8-256fda08462f · outbound

This paper cites GPT-4 Technical Report.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features GPT-4 Technical Report

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.787678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.787678Z digest=sha256:f217708f4525ce0e6378904b247a4720e65be10734ab6a8854f137c2c2078199

Observation 010f9609-f6f2-46f5-abea-2eb796dc0a0d · outbound

This paper cites Training language models to follow instructions with human feedback.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Training language models to follow instructions with human feedback

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.791666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.791666Z digest=sha256:fbf3d09d68eb9b30a9e4c19681a9ae154087d4e26c5b2bfdab5aeb4750cee7c1

Observation 8a400b89-6cda-4ef6-ba50-1a486ee49972 · outbound

This paper cites Cats and dogs.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Cats and dogs

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:24:23.006907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:24:06.796238Z digest=sha256:e2d91fd33966ff5fb0af864c0dff724f60f44d363df94bcea586aa79d5b6d1ea

Observation f67ca58f-9e22-46dc-9068-2da43c2402c9 · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Pytorch: An imperative style, high-performance deep learning library

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.800059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.800059Z digest=sha256:fafcaad89f0f6d9843b06ffb0d4d93b3736371f88a5fe2f5c07eb67a640e33e2

Observation 01b999a2-0cd5-4f0e-a976-53c2eaa49278 · outbound

This paper cites Glove: Global vectors for word representation.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Glove: Global vectors for word representation

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:24:22.986807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:24:06.803994Z digest=sha256:bf6ddc3ccf1b861bd517d47fb13410fab86aa4c6461314c9b8b07ac7ebcc2b67

Observation 9e6dcf0b-c3ac-441d-ae24-280d741e57fb · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Learn- ing transferable visual models from natural language super- vision

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.808370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.808370Z digest=sha256:badb7b75d5a5f48c042334e8cf29b3504e80cb97f5cd86bff20a060d65db072c

Observation d8c24296-ec00-47bb-a410-22800db5ee67 · outbound

This paper cites Direct 11 preference optimization: Your language model is secretly a reward model.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Direct 11 preference optimization: Your language model is secretly a reward model

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:24:22.966833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:24:06.813146Z digest=sha256:41e007ad0268bad4665bdc35e23a3517a855b5b6609012cae9e975c63f1296bc

Observation 8cbff334-af6f-467b-8790-ad45513f619c · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.817705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.817705Z digest=sha256:bfa4efa6f06dc10d9191a565ded86583cf67e26ee005b59eccc61855f4f2fa59

Observation 3d3e79ea-988c-4ebf-82db-4d101074fea0 · outbound

This paper cites Sentence-bert: Sen- tence embeddings using siamese bert-networks.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Sentence-bert: Sen- tence embeddings using siamese bert-networks

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:24:22.954210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:24:06.822626Z digest=sha256:4f3d00e54b27d428e7743f5e5f83e74f30f5d3989ee5de3452440c0237709e3b

Observation d5dc1b9c-c76a-4329-a7d7-cba5c40cfc19 · outbound

This paper cites Parallel distributed processing, volume 1: Ex- plorations in the microstructure of cognition: Foundations.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Parallel distributed processing, volume 1: Ex- plorations in the microstructure of cognition: Foundations

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:24:22.942255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:24:06.826447Z digest=sha256:ecfa50070202b25d3cfdde584fdfe0f61c90bf3e1ca393fc53eb215e9d797dcf

Observation a95d2ca8-ef8d-456d-b84a-7e01e8fe4a4e · outbound

This paper cites Facenet: A unified embedding for face recognition and clustering.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Facenet: A unified embedding for face recognition and clustering

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:24:22.928840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:24:06.830396Z digest=sha256:7ffe362e077a5bd64ada65447b6ca9e138d9279d3763b0a6a6b1f1264691b0f7

Observation 48d7cab2-082e-4b9d-95a3-2acea89ca86c · outbound

This paper cites Understanding machine learning: From theory to algorithms.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Understanding machine learning: From theory to algorithms

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.833951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.833951Z digest=sha256:af035a3f8a63166359e6ee9894347e0de68cc9db1824e53e70ede28a9d12c71a

Observation a23325ca-107e-444b-bcef-bb2ae5f3046d · outbound

This paper cites Towards vqa models that can read.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Towards vqa models that can read

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:24:22.908716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:24:06.837472Z digest=sha256:02b788979dd4a3bcd98980523a3e555d1ff2cdbbd86feb36b53c1a51182ebfd3

Observation b7abfb13-0a01-44e0-a800-cbe89f29c988 · outbound

This paper cites Singh, Stephanie C.Y.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Singh, Stephanie C.Y

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:24:22.895527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:24:06.841166Z digest=sha256:88ac9ff56a84236c52eae05dcf16db2c242d9e31e7be5a5a53f289de1090245b

Observation 0e01f113-c27c-4648-8f9c-c3b1c9469ffb · outbound

This paper cites Manning, Andrew Ng, and Christopher Potts.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Manning, Andrew Ng, and Christopher Potts

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:24:22.882259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:24:06.845605Z digest=sha256:47aaf1e8cac2bb6be3943b44137c02dff42acc8c8a4e685e65bd8b00f9c8c1a3

Observation bac12034-a2ca-4b81-bad1-062acff02f6b · outbound

This paper cites Text classification via large language models.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Text classification via large language models

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:24:22.868199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:24:06.849767Z digest=sha256:0da1dbd860527800cdcbe688d7443ee6512887b7d3b3aac826902a19f3364192

Observation d04e8732-fb87-4c32-a3a7-1fe897211233 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.853805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.853805Z digest=sha256:6f10c997a996cb270e0d990cba3d959a68647b595a934f92ed737cc1d79d1762

Observation 72b9cea0-707c-416a-9982-4fed9c1b8160 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Gemini: A Family of Highly Capable Multimodal Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.857635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.857635Z digest=sha256:524444ecc332ce240ca34b24b06f99ac6aa17d8418d8ca48b0d4dbc012c034cc

Observation 1a2f1d75-2de5-4269-b006-4b69ab2ba907 · outbound

This paper cites Winoground: Probing vision and language models for visio-linguistic compositionality.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Winoground: Probing vision and language models for visio-linguistic compositionality

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:24:22.855299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:24:06.861672Z digest=sha256:a002fdadb9a8f2d4ac2c8544622ea26db25fdaf8d0da5fcfa4760a7ec3b76f7f

Observation d86023fd-50ae-405b-9975-d5d6ab24ea65 · outbound

This paper cites Function Vectors in Large Language Models.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Function Vectors in Large Language Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.870004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.870004Z digest=sha256:d6ccae5ae451f733a515db82fce459d57b37af003f8979450ab0c4cd99e32da6

Observation a6f4c5bb-3d3e-4510-b508-11423e45f1ac · outbound

This paper cites Turk and Alex Pentland.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Turk and Alex Pentland

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:24:22.842687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:24:06.873678Z digest=sha256:7ab0a2e96c637a360240e73174a8589ea100ce2fcb14a8fef509c108db773c5b

Observation 442f34b2-5397-4270-beb1-0e3f5c570d8c · outbound

This paper cites Neural Discrete Representation Learning.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Neural Discrete Representation Learning

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.877742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.877742Z digest=sha256:e34e2af249009cf2ae09c626b99fe4edf5cf9c5fb702a5e9c62a68280584b2be

Observation 3ddc2d99-ec51-42ac-87b4-7135453d20db · outbound

This paper cites Belongie.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Belongie

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:24:22.829462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:24:06.882343Z digest=sha256:31ed76acc7bf01669d21473bf7029edd505b2b8f1f463595df85b51083514de7

Observation 61fa933b-adb0-4852-9c4e-78fca8c13fe9 · outbound

This paper cites Improving Text Embeddings with Large Language Models.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Improving Text Embeddings with Large Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.886211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.886211Z digest=sha256:0c5cd2a48068e1e3c86f25567fa685e8f5334948f2fff8e5f0feb9abcb63485e

Observation bc3887f1-26e9-417d-b89f-df56945ae432 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.890364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.890364Z digest=sha256:65c86b0ad47b09cd6319e1426422f4ea9e3337545bc1b6539b56932af62cc52c

Observation 72088dc8-cff8-4351-bf27-2a52e4338ed0 · outbound

This paper cites InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.894730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.894730Z digest=sha256:0abb39d94a80d2a4e3579e689ecc5f52a7f4181697215b8f83baa5fb4854ce5c

Observation 3493a6be-8d70-4e44-b9ac-0abe3377e6cf · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.899322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.899322Z digest=sha256:91ce976e0c16aed9dcf4b4899d279cf7b2e432accb1aff8c6a566f72f5a13f70

Observation 66ed4700-903a-41f0-b7e8-53b3b5f642c3 · outbound

This paper cites Large Language Models Are Zero-Shot Text Classifiers.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Large Language Models Are Zero-Shot Text Classifiers

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.903563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.903563Z digest=sha256:ff9a998ef31723feb12e5ce2b667b9516bba593f25a7ebc0ddf70cd7b9de86bf

Observation 8f5c77e1-fc00-44cd-9ae4-3bf37f86797e · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Finetuned Language Models Are Zero-Shot Learners

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.907862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.907862Z digest=sha256:8e9117178cf07bc205d59c95c8ae6f005afa6985ac6c3e321124cc0b84c1bc42

Observation e0cd3042-7589-4866-b072-5179448ce79b · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.912216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.912216Z digest=sha256:73140ad6f30acd4ca66fb8c6c22f6af7f22ec754e3ef1542b05da4c9467b94ec

Observation 0a9b37a6-4528-40d3-ab9f-fb380a97fed8 · outbound

This paper cites Scaling Laws for Discriminative Classification in Large Language Models.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Scaling Laws for Discriminative Classification in Large Language Models

Reference 103

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:24:07.079988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:24:06.917263Z digest=sha256:ece6084de9291cc6469f347fc145914e52dc4ffdc2fd8af274c17054903f3ca0

Pith citing papers

Observation e87e8d74-2e2e-4461-b28c-1e519c258d34 · inbound

Activation Reward Models for Few-Shot Model Alignment cites this paper.

Activation Reward Models for Few-Shot Model Alignment Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.424201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.424201Z digest=sha256:d917e029879e22ecf95a535cdcacbe8b9320d9106b1ba5d3af96ccd8a87d8b3d

Observation f39de5d2-e178-4293-996e-9311c4bec4e9 · inbound

Filter-And-Refine: A MLLM Based Cascade System for Industrial-Scale Video Content Moderation cites this paper.

Filter-And-Refine: A MLLM Based Cascade System for Industrial-Scale Video Content Moderation Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T14:58:43.281185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:58:43.281185Z digest=sha256:c723ca496baced51dec569662e02067bdcc5f16859fcc4939836e694d646d58d

Observation 10332bb5-406a-41c9-a679-076ba29eadbb · inbound

Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning cites this paper.

Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:17:26.613498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T06:13:33.315525Z digest=sha256:a947bb6202349c5c70e63572701fa664e89c691fc054ef8125c727553656c078