Pith. sign in

Paper Citation Record · LEDGER

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning

As of 20 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 3 inbound Pith citation observations for arXiv:2504.19627.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.19627 v2

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:53:23.309664Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:55:17.207198Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T20:33:17.140335Z

Reference resolution

77 of 77 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7cb46e91-eb48-41c7-92c8-bfa397199e3c · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.013678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.013678Z digest=sha256:31bc4cb89ae772fd98fb34ee81bd059e609d3fbf2d25ab5b423434710a7e8b2e

Observation 5d0a56a5-64e3-4ac9-9108-80443222eaad · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.018475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.018475Z digest=sha256:c3340baf2ecb99bf20d7b7c0d3e17bf7975d4c74dec9283cda28d5887bb67ce2

Observation 0ed35da3-db7b-464c-8295-5a5d6a813147 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.023000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.023000Z digest=sha256:2b9d83ee6fb7aed79cfa3b2453253369fe97ecf9b1118013364306fd6604f004

Observation a46b5166-1d76-433f-81d2-8e62119b7875 · outbound

This paper cites DEEM: Diffusion Models Serve as the Eyes of Large Language Models for Image Perception.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning DEEM: Diffusion Models Serve as the Eyes of Large Language Models for Image Perception

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.026924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.026924Z digest=sha256:f2fb4fb9368032d0c3fb610ebb3ee118b7df38bd074625053861918d926298cf

Observation fbc81211-1750-476d-8e3f-b3e45b621b85 · outbound

This paper cites Few-shot adversarial prompt learning on vision-language models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Few-shot adversarial prompt learning on vision-language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:24.865303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T05:53:23.035776Z digest=sha256:b9c5604df9a738ebdf58b5ae7cb1d5e288fd66ac32c8a200f84b7d052c14520b

Observation 60a87125-fb5b-40e8-8166-923489bf0d4d · outbound

This paper cites A Survey on Vision-Language-Action Models for Embodied AI.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning A Survey on Vision-Language-Action Models for Embodied AI

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.043813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.043813Z digest=sha256:d4ab948d394f8babcf74a4b2e35fc492b39c13001b19672e1ddaa61dc56bd1e7

Observation baa2b6f0-fd0b-41c6-8443-d1caefe4f41c · outbound

This paper cites EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.047792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.047792Z digest=sha256:82f34c9b9f4e73876daa1bb6f7b60985ae8ff051b9a6a83843d26f209863764a

Observation 5de1d7e7-ffa3-488e-a6d6-0523a92297a5 · outbound

This paper cites Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.051893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.051893Z digest=sha256:3579e458b26a868267629b4de6136745c9dbdcfe50c0d3fa1912cb3c1ed74cae

Observation 184d4616-edd7-4750-abc0-a424deda5c62 · outbound

This paper cites Large (vision) language models for autonomous vehicles: Current trends and future directions.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Large (vision) language models for autonomous vehicles: Current trends and future directions

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:24.854388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T05:53:23.055858Z digest=sha256:7c76c24a1675431eef8fc73d41c830641b41747efc6f19a8823a1a497fe64c19

Observation 125a7038-da91-4d41-a57d-fb0b32f86c6c · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.063240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.063240Z digest=sha256:b73ca9dda44e41a932cb6d5fbea593ef5bf1cc924d5a6c9e760a2ac5a2ebd7a3

Observation 1772d2af-35bf-481c-8085-05657b88811c · outbound

This paper cites Visionzip: Longer is better but not necessary in vision language models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Visionzip: Longer is better but not necessary in vision language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.067210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.067210Z digest=sha256:68fc8feaa6c14778130df66ee795930edf86cd857d7d30c6ecab2ea34a215585

Observation 26ca3679-d3c8-4375-a3f7-70c225e2e1b7 · outbound

This paper cites Dynamic programming.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Dynamic programming

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.071287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.071287Z digest=sha256:fabb85e806ce710a71287e85f81b43ec70d43c794ed899e37a1c74015fdde423

Observation c42cde95-d287-4bba-95b6-27f6b109b870 · outbound

This paper cites Learning transferable visual models from natural language supervision.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Learning transferable visual models from natural language supervision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.075009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.075009Z digest=sha256:5f7cdbc7df40dc677f100a5e4cd58b852860fd6bd21df30ac4662ad8493ea9b4

Observation a075c2cb-a895-4cf8-9bc6-9aa595b1b141 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:24.829070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T05:53:23.078596Z digest=sha256:748625e7ad34586140328109fbe5a80a75624d2fcacf000d25520357b0fac2c7

Observation 7d9b35f5-293f-426b-8548-5c0c2b443443 · outbound

This paper cites Visual instruction tuning.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Visual instruction tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.082634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.082634Z digest=sha256:5a7c74995c1c8cd8c432898da58bc41101b9695e4bc06d2e17780f446532ca94

Observation f4d6ef1b-882b-4b0b-87bc-3abfec4fbdcd · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.086330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.086330Z digest=sha256:0756bffd57ff604ec8fe2e9ff512d7b9177ac6bdf00068232286983c929d6d05

Observation 43d2e9e3-b84b-48b8-8a2e-2ced21a70350 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and composi- tional question answering.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Gqa: A new dataset for real-world visual reasoning and composi- tional question answering

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:24.803885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T05:53:23.090053Z digest=sha256:8e28f71583778d2d9db703cfbf5daf1fd2ec4a17bf94e2bdf889ceda7bb7b142

Observation 02ad6099-06b8-4e66-92c6-7d783010e01c · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Vizwiz grand challenge: Answering visual questions from blind people

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:24.791621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T05:53:23.093515Z digest=sha256:ec7f6e76432c8c6340de5fe445dc70f269063d9ac00b8d7a836546ef54493d21

Observation 6693ca22-460b-4a02-81ca-90555dfdac53 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.096850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.096850Z digest=sha256:58828031010c9baaf7853511fbd3fa978360a0c0b1cf37f0a1e743b056dc06bd

Observation 8450252e-628c-4c77-8088-d01c3aa7a2fc · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Evaluating Object Hallucination in Large Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.100232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.100232Z digest=sha256:76a881328219a6e2db8de981471f4f8a734169a912f10d8e47368f7d715921e0

Observation ee254508-bd9d-41b0-91c0-40ef0f8af08a · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.103779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.103779Z digest=sha256:22847b9345574d0e7eb57d5b9f3813064360e6783def63d57f9a16d869355e96

Observation 4b74c2ad-5b6b-43bf-8c90-301577ca761d · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning MMBench: Is Your Multi-modal Model an All-around Player?

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.107383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.107383Z digest=sha256:a8d48ebc57a51395140335f1ba037745d446753b0b74ec2d823fb6165f0fa097

Observation dc96e961-9e3e-4580-8239-753428c6f52c · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.111130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.111130Z digest=sha256:90ba25586179af72ad6aa8086d2ae37aa26008ba0d85a97e3ea3b2fd4d8db03a

Observation 4323062e-8326-4835-8bdc-a0f6384cacdc · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.114832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.114832Z digest=sha256:f492804dbd2e917fa7b64041612c20e3162199fef95fdf6d506e6efb3946c5df

Observation ba6e66c1-2206-4e77-b54b-0f8353a4c00c · outbound

This paper cites Towards vqa models that can read.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Towards vqa models that can read

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.118315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.118315Z digest=sha256:c49e6bf70dc0a3a4a836addd3c90da95460b61ff277270987495e4340aef9690

Observation 5a874e4c-c043-42d0-b833-9300175f110c · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.121996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.121996Z digest=sha256:286bea2a14ba7d6d176a6839eb18137610d0134876f877a56b0648953002d853

Observation d3a8dab6-9012-4923-9ab5-e5fc7f5b2c6b · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Referitgame: Referring to objects in photographs of natural scenes

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.125737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.125737Z digest=sha256:6322f0503e122c60352e4470b35d2d3aba40b2099ac9c9106a8ebad9fedc23a5

Observation ddfe012d-c71c-4bea-88d1-5f697d078454 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.129212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.129212Z digest=sha256:bd86b91cae6644d1f15afd1153e06a8445ef03cb7e89628a78e60eac19915502

Observation a4676939-ae01-466d-b42d-5fe864d31a36 · outbound

This paper cites Open-vocabulary detr with conditional matching.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Open-vocabulary detr with conditional matching

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:24.697737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T05:53:23.132835Z digest=sha256:2e6fbc24319505fa6111ec3d530400c832a432ae413d495e02984220f51c3d38

Observation 5c55bf5d-da0f-46bd-99c0-70085a01d317 · outbound

This paper cites Scene parsing through ade20k dataset.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Scene parsing through ade20k dataset

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.136470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.136470Z digest=sha256:250f384d0704c12ab99579f63425bef37ece83bd3f1c2c15421d76988d2be6e7

Observation 6f8696bc-7c0d-4266-b273-86b86cc46560 · outbound

This paper cites Tgif-qa: Toward spatio-temporal reasoning in visual question answering.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Tgif-qa: Toward spatio-temporal reasoning in visual question answering

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:24.567173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T05:53:23.140146Z digest=sha256:46bb0106a46404db507d5dd14a7673f20aadd88bbcbfa2097c33dae67803f1de

Observation c4945215-6adc-48d7-86dd-8a006d541023 · outbound

This paper cites Video question answering via gradually refined attention over appearance and motion.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Video question answering via gradually refined attention over appearance and motion

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:24.422819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T05:53:23.143671Z digest=sha256:a2c6b3ef3cddcbefaead0bc0ebb3b09d4c8e4c5d4c005c841075b56819d1d800

Observation 8cdabcd1-5c12-471b-a3ea-9b5ddd5cfc3d · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.147269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.147269Z digest=sha256:bbbe114c52659fc664ef33d7c1614f08b7e0a1514be3993f8ee5fa971642cdb8

Observation 01aa8a78-92ef-4483-9c31-afa32bc6a364 · outbound

This paper cites Lmms-eval: Reality check on the evaluation of large multimodal models, 2024.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Lmms-eval: Reality check on the evaluation of large multimodal models, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.150698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.150698Z digest=sha256:604fde6fbe28f1398eed65435df03416877247437f3bbc072011ffea69140766

Observation 044bf9ca-59f7-4e07-be48-3c46853525af · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.153965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.153965Z digest=sha256:448c4bf2f05cab2e1751ef3f4322f312a21ee67b5f53e9940c6f8b4bb2d4fa63

Observation 4950edca-f73c-49f4-8e3e-2c4f31b950b5 · outbound

This paper cites Decoupled Weight Decay Regularization.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Decoupled Weight Decay Regularization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.157826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.157826Z digest=sha256:bdc73730ba1f8cd9b2a9f27b2f465b8130964d4dfa1fd5f4e8cdff104c026e5d

Observation 398a39d2-4839-4375-a88f-c193ae587840 · outbound

This paper cites Least squares quantization in pcm.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Least squares quantization in pcm

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:24.390410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T05:53:23.161487Z digest=sha256:ac2c5c4ee5eb9f61bc2a24d40bea3820af975d63e47cf2260889b9c83afe5fb1

Observation 194cc1a1-ea3f-4f64-815e-13847a58483c · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:24.378541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T05:53:23.165006Z digest=sha256:d6aa0c1c8ae6d6d2e3d2b72109a1a9dd66af166ef61c3e8f7589a359e836ba41

Observation 66aca493-689b-4178-a5f0-76f29e2a6b29 · outbound

This paper cites Introducing idefics: An open reproduction of state-of-the-art visual language model.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Introducing idefics: An open reproduction of state-of-the-art visual language model

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:24.366329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T05:53:23.169026Z digest=sha256:cf6352444e7141700b146cd089dd23639bdd719138db5c305baf3dd00295f358

Observation 3afc06f7-f63c-46b3-98ea-f3831861130c · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.172405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.172405Z digest=sha256:e003259ef1830b82243a179e37114c2787a4c04cb58a346c0e728cd87833dd45

Observation 4f7d9831-2036-445d-acc1-bd80ebf19868 · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.176102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.176102Z digest=sha256:4efbbc18681ac3b308d8490f0cc3f40650c2e7f7f696ea3448e61cd8f4080ba7

Observation 8166c230-2559-4498-8429-5b800357ffea · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.180051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.180051Z digest=sha256:7c8a71567b4642e85e28d74d9f423ff3f9869d3e2b363bcf4a6147bc79ea33e5

Observation 451a1e11-31fe-4c0f-83fc-853e607e4a7b · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.183611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.183611Z digest=sha256:7763de16653b0afc89610a0864dc36708fdcfd63a26153434044e4a08a2f6043

Observation c3a7d57e-628c-45f2-ba7e-df93e34d1504 · outbound

This paper cites Panoptic segmentation.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Panoptic segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:24.276606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T05:53:23.187798Z digest=sha256:d29fa648c527c609f1cc2b8b50662daac176292cafb6bd116fe15319f253100c

Observation a1974b4d-306f-4158-90b9-7e4745b1e338 · outbound

This paper cites Mask r-cnn.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Mask r-cnn

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:24.139140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T05:53:23.191714Z digest=sha256:8a7e4b47e8daeaedffd231ef99a6dc7f85153ad78a114f61c227c4c68ab3593b

Observation 866ae55b-0eaf-4502-89f1-0875fa44cfcb · outbound

This paper cites Side adapter network for open-vocabulary semantic segmentation.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Side adapter network for open-vocabulary semantic segmentation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:24.022353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T05:53:23.195408Z digest=sha256:f3eb567f7cf4b4fd5295425e23cd9a754367eb3a7b8254fbb2b9ae35f302c21d

Observation 59ae4213-c261-4822-8aee-818be2a1d7a7 · outbound

This paper cites F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language Models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.199481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.199481Z digest=sha256:a928c99b4ffa09f800892b550f78cf9a72627e78d3265b55e4545daf0a617775

Observation b57d74d1-f250-4cce-987a-fce997094e50 · outbound

This paper cites Coco-stuff: Thing and stuff classes in context.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Coco-stuff: Thing and stuff classes in context

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:23.994302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T05:53:23.203621Z digest=sha256:b9ad1ebd541ae853b089a69c8c489e3c7e5251ce85f94c1b047b6744b17f45fd

Observation 5629e619-316f-4b4f-8320-a1d254c3f395 · outbound

This paper cites Cat-seg: Cost aggregation for open-vocabulary semantic segmentation.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Cat-seg: Cost aggregation for open-vocabulary semantic segmentation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:23.982575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T05:53:23.207203Z digest=sha256:e7d854857117babc6e0d13abcabaf3a7da0e165e43a65228a622b3505e9a06c0

Observation 65787873-357b-40aa-86f2-0de5a299b0ab · outbound

This paper cites Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.210832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.210832Z digest=sha256:37d3b70900b9cefdc6edf1279400521efd4bb4103690607a09ffa823a21f8734

Observation a9d839f0-6ec9-4250-8664-a2eeb962113c · outbound

This paper cites Token Merging: Your ViT But Faster.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Token Merging: Your ViT But Faster

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.214627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.214627Z digest=sha256:b524b9139d87426d024c8f1cab569132582f96490e894e264d09c7250b544e33

Observation 827f4e3a-4921-4742-850b-2ed579237af7 · outbound

This paper cites QG-VTC: Question-Guided Visual Token Compression in MLLMs for Efficient VQA.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning QG-VTC: Question-Guided Visual Token Compression in MLLMs for Efficient VQA

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-16T05:53:23.431195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T05:53:23.218705Z digest=sha256:aa84bceb17ba6fdcc3755f02b2eeb3a1bbe838da12a75315ca46102cc2a8fed0

Observation a996f7f8-2fca-4bad-abdc-7b3b162d9d90 · outbound

This paper cites Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.222499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.222499Z digest=sha256:80c0a0f52f9e23f954f2e98dd512459c08d9b5c6cffb8a2a6e5a20d3131ea78b

Observation 36dc7495-b72b-4676-9c79-17a83e1ba0e0 · outbound

This paper cites LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.226484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.226484Z digest=sha256:aab2df999198039534a7237a269c5adbec2713a3d905b64ada8fb982019654c1

Observation d2a58e83-7e7c-489e-a5e7-d95904f35d84 · outbound

This paper cites Efficient large multi-modal models via visual context compression.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Efficient large multi-modal models via visual context compression

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:23.970400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T05:53:23.230431Z digest=sha256:47ed2966992f0d3bcedfec3a55a99a89f0303ba923408e0b30da320714ac873b

Observation 4b6e855d-f0e3-4462-937c-2e3a7a51e4be · outbound

This paper cites Matryoshka Query Transformer for Large Vision-Language Models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Matryoshka Query Transformer for Large Vision-Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.233834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.233834Z digest=sha256:99726492e9578d97e841fe55c818bf67d645e8ac865c52747796746d1550cd8e

Observation c971ad23-f87b-4597-ab6f-bc2d6d00ccfe · outbound

This paper cites MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.237184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.237184Z digest=sha256:26f25f13e3031ebe8794f744d99bbd68f71cc4a476dcc079e44ebca288db6fe5

Observation a5618281-9002-4cab-8243-002686e3ef42 · outbound

This paper cites GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.240784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.240784Z digest=sha256:d0d277dc720e40b460cad61a99a5428d4746de05ef3310681f18109a4e6005ce

Observation d554a7df-a338-4878-9613-608e0b47a067 · outbound

This paper cites Vary: Scaling up the vision vocabulary for large vision-language model.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Vary: Scaling up the vision vocabulary for large vision-language model

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:23.955745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T05:53:23.244004Z digest=sha256:dbe2478a37ead2fe569587f760d2c368e018419b21c0228c78f7c657cadcde93

Observation 94c4da0a-b7a8-4151-9222-30312bdfe980 · outbound

This paper cites Distilling large vision-language model with out-of-distribution generalizability.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Distilling large vision-language model with out-of-distribution generalizability

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:23.936213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T05:53:23.247576Z digest=sha256:61b94cba309c331b71d54cea323cdc399566eb2cade8a6c0ec7bdfc83d8b4144

Observation 973a8408-6289-41b6-b827-0f546a63f4da · outbound

This paper cites Rs5m and georsclip: A large scale vision-language dataset and a large vision-language model for remote sensing.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Rs5m and georsclip: A large scale vision-language dataset and a large vision-language model for remote sensing

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:23.924090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T05:53:23.251000Z digest=sha256:89a7a5f315f52da0f4c123cfa74525789d8a05dbde0fde3feb2ad93291da2719

Observation 1b76abfe-60ab-43d3-b62d-8f531ea618e4 · outbound

This paper cites Visual In-Context Learning for Large Vision-Language Models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Visual In-Context Learning for Large Vision-Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.254383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.254383Z digest=sha256:3bb11611b635ca8a024b810b5bfe73d93ea8e42567e097f7bc76bc1c63e02155

Observation 2801aa98-1897-4871-94ba-0997556ae668 · outbound

This paper cites Anomalygpt: Detecting industrial anomalies using large vision-language models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Anomalygpt: Detecting industrial anomalies using large vision-language models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:23.912006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T05:53:23.257827Z digest=sha256:c080e99e05d4a5ed58c1853a3e3a2d7ab9ba1d5918e39f9ff50dd25ede855ff0

Observation 1a0b1ba4-5a23-418e-8f31-8278e8dcf93f · outbound

This paper cites Matryoshka query transformer for large vision-language models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Matryoshka query transformer for large vision-language models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:23.900177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T05:53:23.261043Z digest=sha256:d321e562deabfe92a92b9f6357dfb72d792ea2961618974271b29b22f6b79642

Observation 787510c3-31c8-4478-bb6e-001b07875047 · outbound

This paper cites Pyramidclip: Hierarchical feature alignment for vision-language model pretraining.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Pyramidclip: Hierarchical feature alignment for vision-language model pretraining

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:23.889111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T05:53:23.263859Z digest=sha256:29ba489752c82d6e0c25019609eb465cf8066cc2a2883697c4d4b97cc0566f0b

Observation 3421e68b-421d-4b70-9fbe-95152307f7ad · outbound

This paper cites Matryoshka multimodal models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Matryoshka multimodal models

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:23.877145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T05:53:23.267730Z digest=sha256:f8650b3bd86f5a665735c832c0025a34f1d002bc709c68f4ee89c829e4d81d4b

Observation 7ea3ea87-15dd-47b0-8c9a-030b198b93c0 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.271111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.271111Z digest=sha256:a1a4bd5f5de94181096bbb962ec9e1e5d2aecc3f6f285823e48583f12a509bc5

Observation c0515a80-6d45-40ff-90d7-e42ade2f7460 · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.274389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.274389Z digest=sha256:5c4bf6e9c430ab8d166a9d7e86d7d20ebc4266d4b695116723e0f5e8ee27e817

Observation ab44a47c-17d8-4fce-8e05-3e7632e2a063 · outbound

This paper cites Vision-language models for vision tasks: A survey.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Vision-language models for vision tasks: A survey

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.278220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.278220Z digest=sha256:53b6e3905918d538ef76c413c9d5b5295e01bdeb7736d0dcfc47af40847f13d6

Observation baacb898-2f08-4b8e-ac46-a2cabc026a14 · outbound

This paper cites A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.281964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.281964Z digest=sha256:5946590d6673c1ccf4f630f92df59f42217aefc04a69ff1d34c30eccd686b172

Observation e123e2f9-c737-4629-be61-731edebd03e0 · outbound

This paper cites A Survey of Vision-Language Pre-Trained Models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning A Survey of Vision-Language Pre-Trained Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.286456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.286456Z digest=sha256:5199e8c6d5ae75cc489d8751d1c2d5b708dcbb16382a8a9850690f7054f6cd9e

Observation c852f6d2-61f0-4313-9e83-13b3a24245e0 · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning A Survey on Hallucination in Large Vision-Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.290607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.290607Z digest=sha256:97bb1fa921b506410dd5e69519d11dc3249e04dfbddc17f2aaac621d0e9ea1c1

Observation 51a3d3c5-7bed-4174-98dd-978f5cb96ff2 · outbound

This paper cites Vqa: Visual question answering.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Vqa: Visual question answering

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:23.855623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T05:53:23.294509Z digest=sha256:04d1bcfee418aa0a63063e9184ee267fdbde2af43ece359a05fbc214671c0905

Observation 46e7a9cd-f51f-430d-ac1d-8d54fa3e8eb7 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Vizwiz grand challenge: Answering visual questions from blind people

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.298221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.298221Z digest=sha256:90ba004f3d201488d304e4dfcfad51821252532ae51569401daa83b4c0939eb3

Observation 103d7a38-3bd6-4b96-ba23-4a2a22f5886b · outbound

This paper cites Cider: Consensus-based image description evaluation.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Cider: Consensus-based image description evaluation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.301738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.301738Z digest=sha256:57a49a19555eec5f3ba98a45eb379729b1a4de986f3f68d2271351b0970ad60f

Observation 4aee89f1-4bfa-4a48-890f-db5498a783ab · outbound

This paper cites Fully convolutional networks for semantic segmentation.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Fully convolutional networks for semantic segmentation

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:23.831153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T05:53:23.305013Z digest=sha256:1a454683ad7d94c332bd91dc1d5204f24b8dedeaa6aebf0da9fa974ba59b7b80

Observation 38b7fd2e-bc50-48a5-89fb-25dcfb9270f2 · outbound

This paper cites ℎ#"ℎ$"ℎ%.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning ℎ#"ℎ$"ℎ%

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:23.819377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T05:53:23.309664Z digest=sha256:8b7ec2755d3a9cbf5b8d6851fcf40e3e282657f711df86fad66aeddb3a408647

Pith citing papers

Observation 0a048ce3-2f65-47d8-9e99-1a2d10ec34de · inbound

Towards Modality Generalization: A Benchmark and Prospective Analysis cites this paper.

Towards Modality Generalization: A Benchmark and Prospective Analysis VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T04:55:17.207198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:55:17.207198Z digest=sha256:7fedb53ec7a2b07fa70479acf01b4de357b57f78856ebe7098d64b675e922d31

Observation 439f993c-bf43-46df-ab72-58020bbae29a · inbound

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning cites this paper.

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:33:17.141525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T20:28:30.864143Z digest=sha256:f273c28bca86dd23bbbbc9d35536cc52ccee5832fb20ed3116545c61ab828bb7

Observation 977f4470-facd-449e-89d1-1b2e24913411 · inbound

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding cites this paper.

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:31:01.611015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T14:50:37.022338Z digest=sha256:fdfcc8635871edca2a90b6da45f9b49bda74f5db6e7c4f0f65d8657c3833085e