Pith. sign in

Paper Citation Record · LEDGER

Visual Large Language Models for Generalized and Specialized Applications

As of 11 August 2026, this Paper Citation Record lists 100 of 298 outbound references and 12 inbound Pith citation observations for arXiv:2501.02765.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.02765 v1

Coverage vector

measured 100 of 298 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:08:09.276530Z

measured 112 of 112 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:02:55.719810Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T22:27:12.275903Z

Reference resolution

100 of 298 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 471892d6-bf6e-47be-810c-539c03239be5 · outbound

This paper cites A-fast-rcnn: Hard positive generation via adversary for object detection,.

Visual Large Language Models for Generalized and Specialized Applications A-fast-rcnn: Hard positive generation via adversary for object detection,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.853047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.853047Z digest=sha256:c5ea29664211425119390cd325e8f7425303b6cbcd02cf4c6136e6dc04195911

Observation f9a7a50c-7bc6-434a-8fe5-1e8b45084a0e · outbound

This paper cites Deep residual learning for image recognition,.

Visual Large Language Models for Generalized and Specialized Applications Deep residual learning for image recognition,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.858029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.858029Z digest=sha256:bddb47483c914facc8cd34a0bd7e12dc111601e52f8863211a90efc82ddfc8d5

Observation 645d315f-bd60-4823-a5af-97187a9d16bc · outbound

This paper cites V oxposer: Composable 3d value maps for robotic manipulation with language models,.

Visual Large Language Models for Generalized and Specialized Applications V oxposer: Composable 3d value maps for robotic manipulation with language models,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.863117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.863117Z digest=sha256:786cd86db26248397e3dc337be43c8c192de118512f49669abea09596adf0db6

Observation 37c9aacd-964e-47b5-a557-80b0b4dd4efe · outbound

This paper cites Spatiotemporal multiplier networks for video action recognition,.

Visual Large Language Models for Generalized and Specialized Applications Spatiotemporal multiplier networks for video action recognition,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.867482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.867482Z digest=sha256:c0a3769335a60fa296ff53be58cca565671c4082e473f664f6b88a913e54d6dc

Observation a801af0c-f895-463f-acc3-0444d4bfeaa9 · outbound

This paper cites Temporal action segmentation: An analysis of modern techniques,.

Visual Large Language Models for Generalized and Specialized Applications Temporal action segmentation: An analysis of modern techniques,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.871785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.871785Z digest=sha256:d635ae9e3487f670294879cb309a2ebee172c7f313d5cb6ac6a2ecc24e15ca72

Observation 1ff6cd80-6f21-4682-adb0-9c5e7c560c96 · outbound

This paper cites Vqa: Visual question answering,.

Visual Large Language Models for Generalized and Specialized Applications Vqa: Visual question answering,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.876145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.876145Z digest=sha256:e05b784dd3f913165ca025dce0ad568093e546b0dfc0ff2309211094c070616d

Observation 38750cda-6073-4b12-987d-96820a017c06 · outbound

This paper cites Vision-language models for vision tasks: A survey,.

Visual Large Language Models for Generalized and Specialized Applications Vision-language models for vision tasks: A survey,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.880318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.880318Z digest=sha256:c53e1b1d07fa1044185c45dbf2d6fe68329bf339c32315c9c68c93cd751b2967

Observation 4bc4732e-f008-4be5-8105-351e3acba32f · outbound

This paper cites From captions to visual concepts and back,.

Visual Large Language Models for Generalized and Specialized Applications From captions to visual concepts and back,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.884423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.884423Z digest=sha256:af077f40952d06097c7514c35e55d1a410f723a902bd80403de0d047781bf262

Observation c1454048-42c3-4ede-8f19-5b7e5db08083 · outbound

This paper cites Show and tell: A neural image caption generator,.

Visual Large Language Models for Generalized and Specialized Applications Show and tell: A neural image caption generator,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.888595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.888595Z digest=sha256:e8fa250017263c8f2fd872f44ad86c2a3608fbdd05f41964bd7bd6a8ce7f503a

Observation 11e3e2fe-8184-4092-86f0-27665a8bdd14 · outbound

This paper cites Deep visual-semantic alignments for generating image descriptions,.

Visual Large Language Models for Generalized and Specialized Applications Deep visual-semantic alignments for generating image descriptions,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.892696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.892696Z digest=sha256:c7e2a4199956c46f00ab3993ac306a5e8e005e23b6c1b00b6eb16e1508ace173

Observation 2409a84a-55c2-4431-b8ea-50a397cc87a0 · outbound

This paper cites Guiding the long- short term memory model for image caption generation,.

Visual Large Language Models for Generalized and Specialized Applications Guiding the long- short term memory model for image caption generation,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.896385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.896385Z digest=sha256:45d6ed2d0054f5e293c2d8ae65881ffb52c5751c68336da44363c16a87c3614e

Observation 46560339-5bfa-4400-9572-7afdebb65f17 · outbound

This paper cites Babytalk: Understanding and generating simple image descriptions,.

Visual Large Language Models for Generalized and Specialized Applications Babytalk: Understanding and generating simple image descriptions,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.900544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.900544Z digest=sha256:0e07e5beca0b7d472a20a3421b19114af3b2a3be01d68b9a527dbfab3f7d06d2

Observation 90977128-0520-4f9b-bbf5-8a5237f92251 · outbound

This paper cites Attention is all you need,.

Visual Large Language Models for Generalized and Specialized Applications Attention is all you need,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.905379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.905379Z digest=sha256:2d14467fd7d3b3f21fa564212b35d34edd774f9823471149896a525d7029b770

Observation a029bf49-88bc-42cc-9b1a-440b5fe6000a · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding,.

Visual Large Language Models for Generalized and Specialized Applications Bert: Pre-training of deep bidirectional transformers for language understanding,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.909619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.909619Z digest=sha256:00a431221b29aa4bb485a1be38c34a1b3bf4c3bd26d5fc9dc0e77fbf1aab25b4

Observation 957da850-5e49-4c09-a9b6-0e92c4e56954 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Visual Large Language Models for Generalized and Specialized Applications Learning transferable visual models from natural language supervision,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.916968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.916968Z digest=sha256:5b100b2d1e5b6e4617696c596d730bb7960369725a6c8bdf795397a8901335d3

Observation bed99a16-983d-4140-b533-910a3c601807 · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

Visual Large Language Models for Generalized and Specialized Applications VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.920635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.920635Z digest=sha256:b3f0d9cd2ea09e45f4f7fc1f702eafa6a6f5bc07141e21c979b81f8c5e36d806

Observation 671051be-7985-4b51-b265-2ee5e6d617b9 · outbound

This paper cites Vilbert: Pretraining task- agnostic visiolinguistic representations for vision-and-language tasks,.

Visual Large Language Models for Generalized and Specialized Applications Vilbert: Pretraining task- agnostic visiolinguistic representations for vision-and-language tasks,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.924459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.924459Z digest=sha256:3fa4e55c602060720ca46f808bf945ff2e7cd6fa424a963057be0914e4574d62

Observation 8fe23a5e-67d5-45b8-8975-209776c11845 · outbound

This paper cites Nsp-bert: A prompt-based few-shot learner through an original pre-training task——next sentence prediction,.

Visual Large Language Models for Generalized and Specialized Applications Nsp-bert: A prompt-based few-shot learner through an original pre-training task——next sentence prediction,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.929515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.929515Z digest=sha256:db598f4fd07e4e447ba65a2da1c6feaa8383ce94f72b6a1004cbc6cbc8fafa84

Observation 79685515-2ea2-462b-9b60-db5e52a6c29a · outbound

This paper cites Learning to prompt for vision-language models,.

Visual Large Language Models for Generalized and Specialized Applications Learning to prompt for vision-language models,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.933386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.933386Z digest=sha256:6bebe3ec593aa254ee9f3444dfc5f6bd9c7ad401a55d5a5afd273dfd213dccb0

Observation fe1adfe3-c40d-46ea-bdd4-277122b10a34 · outbound

This paper cites Conditional prompt learning for vision-language models,.

Visual Large Language Models for Generalized and Specialized Applications Conditional prompt learning for vision-language models,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.937810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.937810Z digest=sha256:29ba61d26769695fe8f159129eb16f2c04ab7c5fe3202476313739688caf2858

Observation f44fc438-43c3-45b6-afc4-12574b8090dc · outbound

This paper cites Slip: Self-supervision meets language-image pre-training,.

Visual Large Language Models for Generalized and Specialized Applications Slip: Self-supervision meets language-image pre-training,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.941916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.941916Z digest=sha256:23819cdf15231058d6e80f5eccd8b484e60cf5546d3791fbe97bea3c46169230

Observation 919cb908-8036-474f-b089-97e252b2f95f · outbound

This paper cites Decoupling zero-shot semantic segmentation,.

Visual Large Language Models for Generalized and Specialized Applications Decoupling zero-shot semantic segmentation,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.946084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.946084Z digest=sha256:892c6b3021c369e51a760964800b27a4ea7a1158fb5847f7c73d055bb57d5dac

Observation 8cdc7770-b105-4d6d-a255-3ffbab656aa6 · outbound

This paper cites Open-vocabulary object detection via vision and language knowledge distillation,.

Visual Large Language Models for Generalized and Specialized Applications Open-vocabulary object detection via vision and language knowledge distillation,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.950010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.950010Z digest=sha256:aa333722a6eab658baab4dfe806ae4fca698f8e3c170382eb2a546e1113c43e6

Observation 7a898a00-6d68-43de-bb1d-220564bba826 · outbound

This paper cites Language models are few-shot learners,.

Visual Large Language Models for Generalized and Specialized Applications Language models are few-shot learners,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.954075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.954075Z digest=sha256:9ae012809032e90f64e04fd2af7f0959b75a170c37b79578de0f585bae77462e

Observation 38b183f4-509b-419b-beed-83a614d8bd26 · outbound

This paper cites Training language models to follow instructions with human feedback,.

Visual Large Language Models for Generalized and Specialized Applications Training language models to follow instructions with human feedback,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.958726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.958726Z digest=sha256:56e2150ad7455e65d0d8b1c99ec5394c3ce8c7a2a5f9fe2b83de7d1d8d369061

Observation 57844911-0787-43e1-afdd-909d74cdd694 · outbound

This paper cites Instruction tuning for large language models: A survey,.

Visual Large Language Models for Generalized and Specialized Applications Instruction tuning for large language models: A survey,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.961882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.961882Z digest=sha256:a1543cfb5b134fa29f631fac076125e0513bd75da65ceab60c804fa565c394c3

Observation be38f458-40ab-4a5e-a0d8-a7e028e5fe02 · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

Visual Large Language Models for Generalized and Specialized Applications Flamingo: a visual language model for few-shot learning,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.964884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.964884Z digest=sha256:01bbeb5bdf89763745eca01c2175c464a0e7eeb2dc2131d93e1e7e07df4feed5

Observation 2658480e-728c-493f-ad32-158e8761d2a7 · outbound

This paper cites Vision-and-Language Pretrained Models: A Survey.

Visual Large Language Models for Generalized and Specialized Applications Vision-and-Language Pretrained Models: A Survey

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.967883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.967883Z digest=sha256:57c9a0433b6b61dfb3955dc75d64c232afc0cded7790dd1804cfe7c74be9c989

Observation cfb5b4ed-0690-4b99-9182-029ef12a5fb7 · outbound

This paper cites Vision Language Models in Autonomous Driving: A Survey and Outlook.

Visual Large Language Models for Generalized and Specialized Applications Vision Language Models in Autonomous Driving: A Survey and Outlook

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.972024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.972024Z digest=sha256:12c858b90a129b76c665bb98c56169b8c48da426b7172cda9f62a09f07b77bf8

Observation a8b619b9-9f92-4f8d-84e9-7bbe7dddfde9 · outbound

This paper cites Exploring the frontier of vision-language models: A survey of current methodologies and future directions,.

Visual Large Language Models for Generalized and Specialized Applications Exploring the frontier of vision-language models: A survey of current methodologies and future directions,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.976052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.976052Z digest=sha256:06184e9c2db8dbdb120e653cd48bfdf96929273ed022d067fafbb28f058cbc43

Observation c191914f-cafa-48d2-8239-f5186b9eb42d · outbound

This paper cites MM-LLMs: Recent Advances in MultiModal Large Language Models.

Visual Large Language Models for Generalized and Specialized Applications MM-LLMs: Recent Advances in MultiModal Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.979027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.979027Z digest=sha256:4f833683b663ca48faecb202d5e4ebb8dd6367654d152c166e2777d52cde65ef

Observation ab626ba4-9043-4e99-b595-da6c85bd7601 · outbound

This paper cites Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey.

Visual Large Language Models for Generalized and Specialized Applications Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.983223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.983223Z digest=sha256:bad977ad08c2d31c0e6615199615f86694bfd2429184b6782efc737cc9975293

Observation 05250eb3-ccd6-490c-ae5b-b6566d20ad43 · outbound

This paper cites Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning.

Visual Large Language Models for Generalized and Specialized Applications Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.987964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.987964Z digest=sha256:a61acd709e4e72b083f22e7b58ca76fbb8a5f5c98770be7fe9ccceffb6f50b7d

Observation 29a2ec10-7a9a-47d3-a41a-d496548f5f10 · outbound

This paper cites Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey.

Visual Large Language Models for Generalized and Specialized Applications Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.992543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.992543Z digest=sha256:72763a8de90ca3b6ad999b7a8dda461a1bd6ade92bd721cb5f6cbf30e5ac1377

Observation 58b0312e-73ff-42f8-b310-a2d89b9933b8 · outbound

This paper cites MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs.

Visual Large Language Models for Generalized and Specialized Applications MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.997103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.997103Z digest=sha256:07220b54a84029f9fcad06f7a240aab02209a191719c5dddcb5c1e56904b8b34

Observation 07d3a5f8-a1a3-4bea-b6a5-e97b8e914936 · outbound

This paper cites A survey on multimodal large language models,.

Visual Large Language Models for Generalized and Specialized Applications A survey on multimodal large language models,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.001205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.001205Z digest=sha256:eb8fff600cad650f521765107542652ecf4e20b578401e50eea9b6d78a2f1751

Observation cf2cdd04-5b5d-49d2-ae35-4a9158483203 · outbound

This paper cites Multimodal large language models: A survey,.

Visual Large Language Models for Generalized and Specialized Applications Multimodal large language models: A survey,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.005903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.005903Z digest=sha256:d625b41482c5a7ec8c4695102e439f04a4869799c13236c898ec3793325a8c75

Observation f0225710-3452-405d-ad3d-0768d628a8ed · outbound

This paper cites The (r) evolution of multimodal large language models: A survey,.

Visual Large Language Models for Generalized and Specialized Applications The (r) evolution of multimodal large language models: A survey,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.011318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.011318Z digest=sha256:cafd43844eff60bf7cb9c0b6f42b7aecfa1c3fdbe9eb6137ea6ca90a1141c5a1

Observation d6e45c17-a997-4cda-a95a-fb42c3756e9c · outbound

This paper cites Multimodal foundation models: From specialists to general-purpose assistants,.

Visual Large Language Models for Generalized and Specialized Applications Multimodal foundation models: From specialists to general-purpose assistants,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.015231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.015231Z digest=sha256:8e0de22e2a97e517d6df551be52a32be5a109fcfd3a51bb18b32f2ff0507338c

Observation b052469d-cdb0-4376-a4cd-19b10d8c243f · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

Visual Large Language Models for Generalized and Specialized Applications A Survey on Hallucination in Large Vision-Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.019312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.019312Z digest=sha256:88e57e54ac17473a1dc711c6efd049dafab07df066b81abbfc0830678f6f637d

Observation 826df491-a925-4de3-a07c-76636f75a182 · outbound

This paper cites Language is not all you need: Aligning perception with language models,.

Visual Large Language Models for Generalized and Specialized Applications Language is not all you need: Aligning perception with language models,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.023651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.023651Z digest=sha256:29be131313a0471a495e4c648743eebec3fff1a5ca3c552c031d078c3484ba4f

Observation 29cde9ac-3917-46c9-9cea-2da88292dedd · outbound

This paper cites Visual instruction tuning,.

Visual Large Language Models for Generalized and Specialized Applications Visual instruction tuning,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.027608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.027608Z digest=sha256:738560c7030170c345b46d9a9d0740e560127c55580b7de4ec3fa7251d911ae7

Observation 7b5d3f44-c541-45bb-a854-631c6dc6604c · outbound

This paper cites Improved baselines with visual instruction tuning,.

Visual Large Language Models for Generalized and Specialized Applications Improved baselines with visual instruction tuning,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.031756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.031756Z digest=sha256:fe96ad6f64ff6fa9ae2d4122ec753ce506cfe8ea095465975cd8773e4a10770a

Observation d57865e9-e27f-4dd1-9526-316afe1bdd72 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Visual Large Language Models for Generalized and Specialized Applications MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.036027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.036027Z digest=sha256:7ae229b1a0d8dbdc8ae7b0068103ad2a6c0181ea2ba5d5c92cb055aabf38c2b7

Observation 10838aa6-f0c5-476b-8562-fe08ac8c3e74 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Visual Large Language Models for Generalized and Specialized Applications MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.039703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.039703Z digest=sha256:7948ba182ab3ba2ecbebf237891f3108319adb34111db6073bac2b9a2663cc48

Observation 8631ea6b-81cc-428e-94da-bf4bb37f53c0 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration,.

Visual Large Language Models for Generalized and Specialized Applications mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.049542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.049542Z digest=sha256:3295ffb660dc41fa2240e00c8b6b5c5a60f96ab0feb7b5102764bae01df53ff4

Observation f82b663f-b97d-4d92-b9ca-c7905134c92c · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

Visual Large Language Models for Generalized and Specialized Applications mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.053527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.053527Z digest=sha256:2fd09bdfccae84faeac2fb152427d2df28f01c913dd279e4fac0c0b2156c370a

Observation b61c15f0-da57-44a2-accf-17509c225f4f · outbound

This paper cites MultiModal-GPT: A Vision and Language Model for Dialogue with Humans.

Visual Large Language Models for Generalized and Specialized Applications MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.058272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.058272Z digest=sha256:dccddf45916f19b57f42558a2b919fe4035c7f6842c44b6b5a158e8bbb057800

Observation 78f1acbf-11ee-4a9b-9479-c2658865a39d · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

Visual Large Language Models for Generalized and Specialized Applications Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.061991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.061991Z digest=sha256:82471f11e7183e7d10154d6730020e7aa67d5587e4cdbeb98bed88dbb7527302

Observation 3da2b34a-0b30-4955-825a-754528c89ee4 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Visual Large Language Models for Generalized and Specialized Applications InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.066684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.066684Z digest=sha256:c136ce6f3b7ef3a989b69a1ebdcc72b6483581993f5a6a6446b8b6c5be965c69

Observation 0efb19d4-8cec-4076-8741-facaddc3f7b7 · outbound

This paper cites Cheap and quick: Efficient vision-language instruction tuning for large language models,.

Visual Large Language Models for Generalized and Specialized Applications Cheap and quick: Efficient vision-language instruction tuning for large language models,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.071186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.071186Z digest=sha256:f21f77b5a0dee6e4a8060b01d354e15cefc9504d047bbc1d3af1f5970ff6e121

Observation 077b2f04-1f56-401e-9490-134b47064b2d · outbound

This paper cites Bliva: A simple multimodal llm for better handling of text-rich visual questions,.

Visual Large Language Models for Generalized and Specialized Applications Bliva: A simple multimodal llm for better handling of text-rich visual questions,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.076028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.076028Z digest=sha256:2b5a09624b02f23a13e97a9f0d932d41892ee14dcc83a125cdd38b4db7c436d7

Observation a304fd03-3513-4f8d-b377-20c6df1a35f8 · outbound

This paper cites StableLLaVA: Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data.

Visual Large Language Models for Generalized and Specialized Applications StableLLaVA: Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.080258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.080258Z digest=sha256:52466bd435e8b970d554d6544f5be26e39746135d81aa5337740a94e394f1008

Observation 1e78c950-23cb-4846-b432-b6aa0d9baa46 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Visual Large Language Models for Generalized and Specialized Applications Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.085674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.085674Z digest=sha256:45bd2e5f9ac90e722b36e8e0c68f1525d4b66e19c85f6f5e539fba9bbd23f5ec

Observation e3cbdba7-192c-4ef8-8c3c-30fe5727bf8c · outbound

This paper cites InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition.

Visual Large Language Models for Generalized and Specialized Applications InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.089948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.089948Z digest=sha256:0b8269250629d8d3aa4fed25f2b07b5a8087c742a86329795d2a4cf388225fdf

Observation a1aba056-37e8-4e9c-a31a-c08b751d53c6 · outbound

This paper cites Introducing our multimodal models,.

Visual Large Language Models for Generalized and Specialized Applications Introducing our multimodal models,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.093619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.093619Z digest=sha256:b2c2aa7ebfe10aa23275906a3bd2927934f563d2791704fc45dad8d65da8d69b

Observation b3523e60-4dda-44e2-a029-e6e5a998551a · outbound

This paper cites Monkey: Image resolution and text label are important things JOURNAL OF LATEX CLASS FILES, JANUARY 2025 20 for large multi-modal models,.

Visual Large Language Models for Generalized and Specialized Applications Monkey: Image resolution and text label are important things JOURNAL OF LATEX CLASS FILES, JANUARY 2025 20 for large multi-modal models,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.097440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.097440Z digest=sha256:6b31f6d66d777719b990da39e544b443a824e51af1161db3bff1df6f598ec722

Observation b0df3684-1657-4600-85ec-33de71cef630 · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm,.

Visual Large Language Models for Generalized and Specialized Applications Honeybee: Locality-enhanced projector for multimodal llm,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.101742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.101742Z digest=sha256:4948e694100147dc3879fb81e48adf0e833bbf4902f3360bf96dfcadfc9b007a

Observation 9b41511c-8f44-46d5-80ac-40c03ea1cf5f · outbound

This paper cites Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models.

Visual Large Language Models for Generalized and Specialized Applications Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.106319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.106319Z digest=sha256:ae934e2aaff8330c349f34d07c746e81463996bd58cf1f2e78da727fa6665e00

Observation 7cc276f3-7196-4396-b185-3c04728823ff · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,.

Visual Large Language Models for Generalized and Specialized Applications Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.110962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.110962Z digest=sha256:686e673eb4589f9819b6da3ef54cfcbca264a9478ec56f060eb260d1e877b132

Observation f17b19f4-b99a-4441-9bee-d62bf484118d · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Visual Large Language Models for Generalized and Specialized Applications DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.114597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.114597Z digest=sha256:e32d7ee1bfb64f3db444a7bf6bf335819cb1d873e1d9b5c34de127f8b6c3e5e7

Observation 84c01329-abf5-450f-b226-2e64313fd04c · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

Visual Large Language Models for Generalized and Specialized Applications DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.118747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.118747Z digest=sha256:4cf865100afedaad24e850b740c31884db1ceca5e1579938435518fe6540ea72

Observation 028f0d74-561b-4578-81d8-a40e7a26d598 · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

Visual Large Language Models for Generalized and Specialized Applications MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.123093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.123093Z digest=sha256:27f98db0f849257bcce464e6d3bd76520a23ec4c9c5e16d06d79aea2d323e7d7

Observation c580a315-6f9f-4a4e-82a5-99a0acc31263 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Visual Large Language Models for Generalized and Specialized Applications Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.126439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.126439Z digest=sha256:1454ff04f2e11e0a306f7adfc2cc3e7bbf87699f808537780c33fc16ac2b4c91

Observation 64384dc8-e750-4d2a-891a-fce6c4f356e4 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Visual Large Language Models for Generalized and Specialized Applications Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.130291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.130291Z digest=sha256:6430d1aca4bc29659c61c1af5682d7f891200d99e649891b56dee90a1cf6d9a0

Observation 72b8fb2e-97cb-44d3-935b-e5dbcebb5745 · outbound

This paper cites Lisa: Reasoning segmentation via large language model,.

Visual Large Language Models for Generalized and Specialized Applications Lisa: Reasoning segmentation via large language model,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.134364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.134364Z digest=sha256:4ebc1c1af20ea47568f91aa8cdabd0b47e2795650c78186483cfcaa3de8773b0

Observation 4039403b-4b4c-40e9-9075-3d60c57560ed · outbound

This paper cites Groundhog: Grounding large language models to holistic segmentation,.

Visual Large Language Models for Generalized and Specialized Applications Groundhog: Grounding large language models to holistic segmentation,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.138705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.138705Z digest=sha256:1c091860a9290f2601bc8d63e49d35575c1c6254f2b86dbcd1b6d793db980f9a

Observation 2d2266d1-d6e9-4da1-bc83-b8e2a7bcc5d6 · outbound

This paper cites Jack of all tasks master of many: Designing general-purpose coarse-to-fine vision-language model,.

Visual Large Language Models for Generalized and Specialized Applications Jack of all tasks master of many: Designing general-purpose coarse-to-fine vision-language model,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.147338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.147338Z digest=sha256:a1ff84700bb3ba4d86397453ad09202e9cd4b12fe21684fc2a97f61f4af37f3e

Observation 720f258a-3f55-4b43-ad82-ce55448b6d11 · outbound

This paper cites Contextual Object Detection with Multimodal Large Language Models.

Visual Large Language Models for Generalized and Specialized Applications Contextual Object Detection with Multimodal Large Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.151036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.151036Z digest=sha256:0e326f372028bbad29baef0400f0b01616f75800cfa1b386b52b36986c287fd4

Observation 34bac95e-f369-4a3e-9acc-cf5d70daf4f3 · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model,.

Visual Large Language Models for Generalized and Specialized Applications Pixellm: Pixel reasoning with large multimodal model,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.155891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.155891Z digest=sha256:0d9df43f7528494b9a4960f38cd16643e6eb6a5a71a9be9f084c118241cde4b7

Observation f4ac02f4-25e8-4613-8449-9383856b2d7b · outbound

This paper cites Pixel-aligned language model,.

Visual Large Language Models for Generalized and Specialized Applications Pixel-aligned language model,

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.159006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.159006Z digest=sha256:92da5c559ec7bf4977398f868c956d2055e8e0b4cc202581542ac45d883e839f

Observation 773a9d28-bab2-418b-8967-c88e7e2d0008 · outbound

This paper cites Gsva: Generalized segmentation via multimodal large language models,.

Visual Large Language Models for Generalized and Specialized Applications Gsva: Generalized segmentation via multimodal large language models,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.162393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.162393Z digest=sha256:4b6213d7643c5df5196417a52e394ee64eab6035c2cff560bd0922977556b9ee

Observation 50087dbd-d00d-42b0-a186-a0e84e9fb5af · outbound

This paper cites Llafs: When large language models meet few-shot segmentation,.

Visual Large Language Models for Generalized and Specialized Applications Llafs: When large language models meet few-shot segmentation,

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.166323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.166323Z digest=sha256:2b7c789f2a72ae56695e6e041ae1a55eb5358c28d1e34a1ed3f6b5b24e24fd3d

Observation a35cb8eb-a35e-4c50-a8fe-96b1eea29335 · outbound

This paper cites Glamm: Pixel grounding large multimodal model,.

Visual Large Language Models for Generalized and Specialized Applications Glamm: Pixel grounding large multimodal model,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.170520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.170520Z digest=sha256:f7d680cf7f305bc193de093d8e7a8f82ee14e2eecec159748016ffb1f24b6d5c

Observation c6d090ca-31a8-4356-a9b4-b915f2976c72 · outbound

This paper cites Visionllm: Large language model is also an open-ended decoder for vision-centric tasks,.

Visual Large Language Models for Generalized and Specialized Applications Visionllm: Large language model is also an open-ended decoder for vision-centric tasks,

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.175234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.175234Z digest=sha256:f3bf82c68c92be4710ab4cb8d64dd04ec3519d96ff3a0ca6736eb860d44fe87a

Observation c29b4852-fd46-4d46-a2ca-4cd8a8d820b2 · outbound

This paper cites VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks.

Visual Large Language Models for Generalized and Specialized Applications VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.179783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.179783Z digest=sha256:688d3a7597e073a4c663828b1ac7e629d90f792c3bbb14d2c0297b3fdd69998a

Observation 4cf59903-b8a9-4de8-9906-0e7bad060347 · outbound

This paper cites Osprey: Pixel understanding with visual instruction tuning,.

Visual Large Language Models for Generalized and Specialized Applications Osprey: Pixel understanding with visual instruction tuning,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.183960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.183960Z digest=sha256:47a293416b059a6cbf24b3b2d7f59cab099f272cd385134f2026e70b601da10b

Observation f3a54099-3b22-49d6-87b5-3e72f5081c62 · outbound

This paper cites OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding.

Visual Large Language Models for Generalized and Specialized Applications OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.187364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.187364Z digest=sha256:f149e594ba1ae2a872a18b20db71e5dd6e73141b7671cd83cbecbedc920e96e9

Observation 295759cd-79e0-4dff-929b-98b16726d800 · outbound

This paper cites Llm-seg: Bridging image segmentation and large language model reasoning,.

Visual Large Language Models for Generalized and Specialized Applications Llm-seg: Bridging image segmentation and large language model reasoning,

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.191999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.191999Z digest=sha256:16855df64aa08b14a04a14b7e88728c940f9e7a338de33693910097e59171f52

Observation 12eb69c5-d263-4009-a722-bd60c4d88cd6 · outbound

This paper cites PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model.

Visual Large Language Models for Generalized and Specialized Applications PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.196290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.196290Z digest=sha256:8ff7a4a3e573e9d2b8caab07eb032e7630ec4a6c78ad0117c1661f798d0eb2d4

Observation b7939f39-074d-4d56-9661-49c6465823f6 · outbound

This paper cites High-Quality Entity Segmentation and Grounding.

Visual Large Language Models for Generalized and Specialized Applications High-Quality Entity Segmentation and Grounding

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.200713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.200713Z digest=sha256:91d5d2e0eb6124e5a16ada5bc3d7553750830e3d994152b9942e827eec347e94

Observation 5c8f212d-9576-4172-af34-77b7d3eac30b · outbound

This paper cites DetGPT: Detect What You Need via Reasoning.

Visual Large Language Models for Generalized and Specialized Applications DetGPT: Detect What You Need via Reasoning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.205049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.205049Z digest=sha256:97979a3b28841b9c0ed7975edc5f19a67438ec39e02c8fee4e3c7d01cfae5255

Observation c274bec4-ace3-4ebd-b265-17053a7fb8a6 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Visual Large Language Models for Generalized and Specialized Applications Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.209616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.209616Z digest=sha256:de1e941bd3dede9422bedc21efc08957cd77b459db96e7b96bfd52f8d0479d01

Observation 3757acce-7c71-46a2-94bf-e6f0f6b3881b · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Visual Large Language Models for Generalized and Specialized Applications Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.213246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.213246Z digest=sha256:afedaa8c245e5f2f2e9e4970fffc2674da74069682ebbedd341f747c9c1f5fe4

Observation 7e0c75be-d544-4487-99a1-7095e9845e83 · outbound

This paper cites ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning.

Visual Large Language Models for Generalized and Specialized Applications ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.216959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.216959Z digest=sha256:839bd24b8e742977caf7b9ab384155b2db0fc7ce878cbf588cffd68740c8b631

Observation 821b2ad9-254b-4062-b5fe-f02f30c56f55 · outbound

This paper cites GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest.

Visual Large Language Models for Generalized and Specialized Applications GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.220351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.220351Z digest=sha256:c0f89fdba8262477552c2afe21de6346ca564be7c67a4783ba128beb5b662f76

Observation 9ad3271b-c977-4fb8-8125-94e52b906c45 · outbound

This paper cites BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs.

Visual Large Language Models for Generalized and Specialized Applications BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.224221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.224221Z digest=sha256:b5e8645e3edec44d73911138658b8d2756261d7367e0b24bad2552c4ea8683d0

Observation 97b84671-bcc8-4d56-8f9e-f065a7a5b6b4 · outbound

This paper cites Pink: Unveiling the power of referential comprehension for multi-modal llms,.

Visual Large Language Models for Generalized and Specialized Applications Pink: Unveiling the power of referential comprehension for multi-modal llms,

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.229000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.229000Z digest=sha256:54199ddca5c61efead81cd191dc679d7eefd3160044b1b1c4cb9f169abb22684

Observation 1d7e84db-6650-43bf-b8ab-700823197c4b · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

Visual Large Language Models for Generalized and Specialized Applications Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.233293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.233293Z digest=sha256:6a103fae91dcb953ec8b898dd430038ce9b4d66b33678d7bed279808fd2c0d49

Observation d94408d8-8a68-4b3d-965e-527907aedec8 · outbound

This paper cites Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models.

Visual Large Language Models for Generalized and Specialized Applications Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.237852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.237852Z digest=sha256:c57fc03dcdc31f730a403eed3b63a6692ab862bdf5add167783bb0e9ea06d43e

Observation 9e94c0bd-0618-47da-8a62-9a9240eead46 · outbound

This paper cites InfMLLM: A Unified Framework for Visual-Language Tasks.

Visual Large Language Models for Generalized and Specialized Applications InfMLLM: A Unified Framework for Visual-Language Tasks

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.241365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.241365Z digest=sha256:9bfea4d94c92788aad0d21a5e451b4c216ecb3ffc04aeebdbfa5cfa8e4bffcea

Observation 1fbe80b3-402a-4d1d-ad35-135da6b9e3b6 · outbound

This paper cites Lion: Empowering multimodal large language model with dual-level visual knowledge,.

Visual Large Language Models for Generalized and Specialized Applications Lion: Empowering multimodal large language model with dual-level visual knowledge,

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.245024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.245024Z digest=sha256:32b283d543adb4f134d9cf0aa37bc1fe33dbb772697f290906ad919c7528076b

Observation e90aba5a-d9c9-4b29-85de-0d1999ffbb63 · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

Visual Large Language Models for Generalized and Specialized Applications SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.249066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.249066Z digest=sha256:5e866f41a5218c7dab355b7e399f2b075885a1c662660c57e0c1a3642bc656d7

Observation 36949f8a-0830-4cdb-9115-913f2a23aded · outbound

This paper cites NExT-Chat: An LMM for Chat, Detection and Segmentation.

Visual Large Language Models for Generalized and Specialized Applications NExT-Chat: An LMM for Chat, Detection and Segmentation

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.252943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.252943Z digest=sha256:15882d2afba8af9d9f5749297ba94f361b41f3c3f37aa8dfa7236eb461883d8b

Observation 8f517d6a-4f26-40b5-959d-1ca379553f64 · outbound

This paper cites Griffon: Spelling out All Object Locations at Any Granularity with Large Language Models.

Visual Large Language Models for Generalized and Specialized Applications Griffon: Spelling out All Object Locations at Any Granularity with Large Language Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.257115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.257115Z digest=sha256:31e1c527938071a7d6914091d91daba61b835a07552a7fdb3368ce5e7317b508

Observation dc59dfea-9a5f-4d55-acb7-7be0fffb2415 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Visual Large Language Models for Generalized and Specialized Applications CogVLM: Visual Expert for Pretrained Language Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.260850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.260850Z digest=sha256:943cd725d00a0b84c2db056f728e4961209b717f40b808119daaeeab5f505a63

Observation 1502aaa5-2893-4cbb-8bd8-183d28e27b86 · outbound

This paper cites LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models.

Visual Large Language Models for Generalized and Specialized Applications LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.264452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.264452Z digest=sha256:a972e4c9863357f39ca37306fedaeb3d1914fa838e9a1a4388edac98735cc361

Observation d20aa233-836a-467b-8115-2d971f4d6da0 · outbound

This paper cites Lenna: Language Enhanced Reasoning Detection Assistant.

Visual Large Language Models for Generalized and Specialized Applications Lenna: Language Enhanced Reasoning Detection Assistant

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.268427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.268427Z digest=sha256:4797ca148f2bf4abda9f051f7c3c2aaec03d1318d41c40fba40e145122f2cf69

Observation 0416a874-b557-45af-b665-202975af1755 · outbound

This paper cites ChatterBox: Multi-round Multimodal Referring and Grounding.

Visual Large Language Models for Generalized and Specialized Applications ChatterBox: Multi-round Multimodal Referring and Grounding

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.272318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.272318Z digest=sha256:0be288f2776e0abbf71721b256b28ddf582159015c0b0cea92f61a01c729857f

Observation 24b98bc0-fd9c-445c-8ff8-9b4a8d910f6a · outbound

This paper cites Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models.

Visual Large Language Models for Generalized and Specialized Applications Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.276530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.276530Z digest=sha256:488a79be08df7bca28684eb07e52604d40fa19ded2b4ca48cef5861afb9dcc05

Pith citing papers

Observation 678f2bb4-176f-48b2-b4bf-280e8c7b102c · inbound

Lost in Cultural Translation: Do LLMs Struggle with Math Across Cultural Contexts? cites this paper.

Lost in Cultural Translation: Do LLMs Struggle with Math Across Cultural Contexts? Visual Large Language Models for Generalized and Specialized Applications

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:27:12.278794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T22:27:02.059459Z digest=sha256:afa5d6e9e5b6daaa03b6d0215effd2d530f0306bdd43d732f120704d001d7a11

Observation 58e388ac-6946-4a66-8d9a-accd269b82d1 · inbound

Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects cites this paper.

Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects Visual Large Language Models for Generalized and Specialized Applications

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.719810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.719810Z digest=sha256:3c5c04eb5e95e22ce86c3b50448238a8e1f3090eecb5c650c7dd95a77ab3f71c

Observation c41d5d1e-e5d9-4380-9d1f-1db7de0165a0 · inbound

IndustryEQA: Pushing the Frontiers of Embodied Question Answering in Industrial Scenarios cites this paper.

IndustryEQA: Pushing the Frontiers of Embodied Question Answering in Industrial Scenarios Visual Large Language Models for Generalized and Specialized Applications

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:54:24.604612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:54:24.604612Z digest=sha256:b8bc979e237735ed789c49b9a57611b850c339803ff52dcd3c43c986bb026144

Observation f2bfbacc-3d5b-48d5-a79f-f8303ca9cd2c · inbound

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads cites this paper.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Visual Large Language Models for Generalized and Specialized Applications

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:30.084456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:30.084456Z digest=sha256:a7a934f288d14126c01107cda34c3d7ec0da8c3c8927ad7941b34724a029381b

Observation 78c8ddaa-47de-4843-8a1b-958b2fea1532 · inbound

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models cites this paper.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Visual Large Language Models for Generalized and Specialized Applications

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:13:02.863394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:d721771c3ea7e191debbd0173ea1fe4056cd91313321198d3882150d871f286e

Observation 037ba41e-21a8-45b8-8f95-36081c8fdbb6 · inbound

A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects cites this paper.

A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects Visual Large Language Models for Generalized and Specialized Applications

Reference 215

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:26.815119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:26.815119Z digest=sha256:a824b6359f94195256d172d28816bd4ef72f51116ea4db449431616a98bc02e9

Observation 62210426-7a5e-41af-9e53-ebf45dfe7a3a · inbound

Leveraging Large Language Model for Intelligent Log Processing and Autonomous Debugging in Cloud AI Platforms cites this paper.

Leveraging Large Language Model for Intelligent Log Processing and Autonomous Debugging in Cloud AI Platforms Visual Large Language Models for Generalized and Specialized Applications

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:56.244746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:28:56.244746Z digest=sha256:6269de876dac784225cf17d501e6d63faf6f677e2f978848e1555208ca86d1e9

Observation 8e0db215-d347-40c5-9b13-e62da461b9e4 · inbound

Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens cites this paper.

Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens Visual Large Language Models for Generalized and Specialized Applications

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-19T01:21:58.110331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T01:18:31.661827Z digest=sha256:5496305d041a4384cb249230731c1c3c796fc63056c31b550d035630fdaa99b8

Observation a8ad9122-832a-40e6-a471-fea949cb9352 · inbound

CoEmoGen: Towards Semantically-Coherent and Scalable Emotional Image Content Generation cites this paper.

CoEmoGen: Towards Semantically-Coherent and Scalable Emotional Image Content Generation Visual Large Language Models for Generalized and Specialized Applications

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T04:25:46.368463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:25:46.368463Z digest=sha256:2528d733af51c91c53cc54a75a7491225d5698beeae46ff86ec0a19550efba24

Observation cda5e9a0-1daf-4993-89dd-7463904e7c07 · inbound

IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation cites this paper.

IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation Visual Large Language Models for Generalized and Specialized Applications

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T20:59:17.714781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:59:17.714781Z digest=sha256:20d43ef72f8be8a728f650267fc3f0f4e69370dd8dd2c55d9361771e4a0af84d

Observation d8903798-d4cf-45d4-9ef3-aed4184c790c · inbound

Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making cites this paper.

Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making Visual Large Language Models for Generalized and Specialized Applications

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T06:28:45.289373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:28:45.289373Z digest=sha256:f0b8b05d2257b680a04ff78be4f12c970ed25647e02c4ed7684fe4a7e47b0d61

Observation 187df4d5-837f-4c6a-a968-a90dbe08a29a · inbound

GaLa: Hypergraph-Guided Visual Language Models for Procedural Planning cites this paper.

GaLa: Hypergraph-Guided Visual Language Models for Procedural Planning Visual Large Language Models for Generalized and Specialized Applications

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:26:27.222290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T06:26:06.208036Z digest=sha256:13b09cb9c256d9871854b902d844230a0d55cc1714c079ef98bc1388a3c00bb8