Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:08:09.276530Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 100 of 298 outbound references and 12 inbound Pith citation observations for arXiv:2501.02765.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:08:09.276530Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:02:55.719810Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-22T22:27:12.275903Z
100 of 298 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 471892d6-bf6e-47be-810c-539c03239be5 · outbound
Visual Large Language Models for Generalized and Specialized Applications A-fast-rcnn: Hard positive generation via adversary for object detection,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9a7a50c-7bc6-434a-8fe5-1e8b45084a0e · outbound
Visual Large Language Models for Generalized and Specialized Applications Deep residual learning for image recognition,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 645d315f-bd60-4823-a5af-97187a9d16bc · outbound
Visual Large Language Models for Generalized and Specialized Applications V oxposer: Composable 3d value maps for robotic manipulation with language models,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37c9aacd-964e-47b5-a557-80b0b4dd4efe · outbound
Visual Large Language Models for Generalized and Specialized Applications Spatiotemporal multiplier networks for video action recognition,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a801af0c-f895-463f-acc3-0444d4bfeaa9 · outbound
Visual Large Language Models for Generalized and Specialized Applications Temporal action segmentation: An analysis of modern techniques,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ff6cd80-6f21-4682-adb0-9c5e7c560c96 · outbound
Visual Large Language Models for Generalized and Specialized Applications Vqa: Visual question answering,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38750cda-6073-4b12-987d-96820a017c06 · outbound
Visual Large Language Models for Generalized and Specialized Applications Vision-language models for vision tasks: A survey,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bc4732e-f008-4be5-8105-351e3acba32f · outbound
Visual Large Language Models for Generalized and Specialized Applications From captions to visual concepts and back,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1454048-42c3-4ede-8f19-5b7e5db08083 · outbound
Visual Large Language Models for Generalized and Specialized Applications Show and tell: A neural image caption generator,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11e3e2fe-8184-4092-86f0-27665a8bdd14 · outbound
Visual Large Language Models for Generalized and Specialized Applications Deep visual-semantic alignments for generating image descriptions,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2409a84a-55c2-4431-b8ea-50a397cc87a0 · outbound
Visual Large Language Models for Generalized and Specialized Applications Guiding the long- short term memory model for image caption generation,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46560339-5bfa-4400-9572-7afdebb65f17 · outbound
Visual Large Language Models for Generalized and Specialized Applications Babytalk: Understanding and generating simple image descriptions,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90977128-0520-4f9b-bbf5-8a5237f92251 · outbound
Visual Large Language Models for Generalized and Specialized Applications Attention is all you need,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a029bf49-88bc-42cc-9b1a-440b5fe6000a · outbound
Visual Large Language Models for Generalized and Specialized Applications Bert: Pre-training of deep bidirectional transformers for language understanding,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 957da850-5e49-4c09-a9b6-0e92c4e56954 · outbound
Visual Large Language Models for Generalized and Specialized Applications Learning transferable visual models from natural language supervision,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bed99a16-983d-4140-b533-910a3c601807 · outbound
Visual Large Language Models for Generalized and Specialized Applications VisualBERT: A Simple and Performant Baseline for Vision and Language
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 671051be-7985-4b51-b265-2ee5e6d617b9 · outbound
Visual Large Language Models for Generalized and Specialized Applications Vilbert: Pretraining task- agnostic visiolinguistic representations for vision-and-language tasks,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fe23a5e-67d5-45b8-8975-209776c11845 · outbound
Visual Large Language Models for Generalized and Specialized Applications Nsp-bert: A prompt-based few-shot learner through an original pre-training task——next sentence prediction,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79685515-2ea2-462b-9b60-db5e52a6c29a · outbound
Visual Large Language Models for Generalized and Specialized Applications Learning to prompt for vision-language models,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe1adfe3-c40d-46ea-bdd4-277122b10a34 · outbound
Visual Large Language Models for Generalized and Specialized Applications Conditional prompt learning for vision-language models,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f44fc438-43c3-45b6-afc4-12574b8090dc · outbound
Visual Large Language Models for Generalized and Specialized Applications Slip: Self-supervision meets language-image pre-training,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 919cb908-8036-474f-b089-97e252b2f95f · outbound
Visual Large Language Models for Generalized and Specialized Applications Decoupling zero-shot semantic segmentation,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cdc7770-b105-4d6d-a255-3ffbab656aa6 · outbound
Visual Large Language Models for Generalized and Specialized Applications Open-vocabulary object detection via vision and language knowledge distillation,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a898a00-6d68-43de-bb1d-220564bba826 · outbound
Visual Large Language Models for Generalized and Specialized Applications Language models are few-shot learners,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38b183f4-509b-419b-beed-83a614d8bd26 · outbound
Visual Large Language Models for Generalized and Specialized Applications Training language models to follow instructions with human feedback,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57844911-0787-43e1-afdd-909d74cdd694 · outbound
Visual Large Language Models for Generalized and Specialized Applications Instruction tuning for large language models: A survey,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be38f458-40ab-4a5e-a0d8-a7e028e5fe02 · outbound
Visual Large Language Models for Generalized and Specialized Applications Flamingo: a visual language model for few-shot learning,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2658480e-728c-493f-ad32-158e8761d2a7 · outbound
Visual Large Language Models for Generalized and Specialized Applications Vision-and-Language Pretrained Models: A Survey
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfb5b4ed-0690-4b99-9182-029ef12a5fb7 · outbound
Visual Large Language Models for Generalized and Specialized Applications Vision Language Models in Autonomous Driving: A Survey and Outlook
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8b619b9-9f92-4f8d-84e9-7bbe7dddfde9 · outbound
Visual Large Language Models for Generalized and Specialized Applications Exploring the frontier of vision-language models: A survey of current methodologies and future directions,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c191914f-cafa-48d2-8239-f5186b9eb42d · outbound
Visual Large Language Models for Generalized and Specialized Applications MM-LLMs: Recent Advances in MultiModal Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab626ba4-9043-4e99-b595-da6c85bd7601 · outbound
Visual Large Language Models for Generalized and Specialized Applications Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05250eb3-ccd6-490c-ae5b-b6566d20ad43 · outbound
Visual Large Language Models for Generalized and Specialized Applications Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29a2ec10-7a9a-47d3-a41a-d496548f5f10 · outbound
Visual Large Language Models for Generalized and Specialized Applications Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58b0312e-73ff-42f8-b310-a2d89b9933b8 · outbound
Visual Large Language Models for Generalized and Specialized Applications MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07d3a5f8-a1a3-4bea-b6a5-e97b8e914936 · outbound
Visual Large Language Models for Generalized and Specialized Applications A survey on multimodal large language models,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf2cdd04-5b5d-49d2-ae35-4a9158483203 · outbound
Visual Large Language Models for Generalized and Specialized Applications Multimodal large language models: A survey,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0225710-3452-405d-ad3d-0768d628a8ed · outbound
Visual Large Language Models for Generalized and Specialized Applications The (r) evolution of multimodal large language models: A survey,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6e45c17-a997-4cda-a95a-fb42c3756e9c · outbound
Visual Large Language Models for Generalized and Specialized Applications Multimodal foundation models: From specialists to general-purpose assistants,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b052469d-cdb0-4376-a4cd-19b10d8c243f · outbound
Visual Large Language Models for Generalized and Specialized Applications A Survey on Hallucination in Large Vision-Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 826df491-a925-4de3-a07c-76636f75a182 · outbound
Visual Large Language Models for Generalized and Specialized Applications Language is not all you need: Aligning perception with language models,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29cde9ac-3917-46c9-9cea-2da88292dedd · outbound
Visual Large Language Models for Generalized and Specialized Applications Visual instruction tuning,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b5d3f44-c541-45bb-a854-631c6dc6604c · outbound
Visual Large Language Models for Generalized and Specialized Applications Improved baselines with visual instruction tuning,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d57865e9-e27f-4dd1-9526-316afe1bdd72 · outbound
Visual Large Language Models for Generalized and Specialized Applications MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10838aa6-f0c5-476b-8562-fe08ac8c3e74 · outbound
Visual Large Language Models for Generalized and Specialized Applications MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8631ea6b-81cc-428e-94da-bf4bb37f53c0 · outbound
Visual Large Language Models for Generalized and Specialized Applications mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration,
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f82b663f-b97d-4d92-b9ca-c7905134c92c · outbound
Visual Large Language Models for Generalized and Specialized Applications mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b61c15f0-da57-44a2-accf-17509c225f4f · outbound
Visual Large Language Models for Generalized and Specialized Applications MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78f1acbf-11ee-4a9b-9479-c2658865a39d · outbound
Visual Large Language Models for Generalized and Specialized Applications Otter: A Multi-Modal Model with In-Context Instruction Tuning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3da2b34a-0b30-4955-825a-754528c89ee4 · outbound
Visual Large Language Models for Generalized and Specialized Applications InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0efb19d4-8cec-4076-8741-facaddc3f7b7 · outbound
Visual Large Language Models for Generalized and Specialized Applications Cheap and quick: Efficient vision-language instruction tuning for large language models,
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 077b2f04-1f56-401e-9490-134b47064b2d · outbound
Visual Large Language Models for Generalized and Specialized Applications Bliva: A simple multimodal llm for better handling of text-rich visual questions,
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a304fd03-3513-4f8d-b377-20c6df1a35f8 · outbound
Visual Large Language Models for Generalized and Specialized Applications StableLLaVA: Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e78c950-23cb-4846-b432-b6aa0d9baa46 · outbound
Visual Large Language Models for Generalized and Specialized Applications Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3cbdba7-192c-4ef8-8c3c-30fe5727bf8c · outbound
Visual Large Language Models for Generalized and Specialized Applications InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1aba056-37e8-4e9c-a31a-c08b751d53c6 · outbound
Visual Large Language Models for Generalized and Specialized Applications Introducing our multimodal models,
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3523e60-4dda-44e2-a029-e6e5a998551a · outbound
Visual Large Language Models for Generalized and Specialized Applications Monkey: Image resolution and text label are important things JOURNAL OF LATEX CLASS FILES, JANUARY 2025 20 for large multi-modal models,
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0df3684-1657-4600-85ec-33de71cef630 · outbound
Visual Large Language Models for Generalized and Specialized Applications Honeybee: Locality-enhanced projector for multimodal llm,
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b41511c-8f44-46d5-80ac-40c03ea1cf5f · outbound
Visual Large Language Models for Generalized and Specialized Applications Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cc276f3-7196-4396-b185-3c04728823ff · outbound
Visual Large Language Models for Generalized and Specialized Applications Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f17b19f4-b99a-4441-9bee-d62bf484118d · outbound
Visual Large Language Models for Generalized and Specialized Applications DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84c01329-abf5-450f-b226-2e64313fd04c · outbound
Visual Large Language Models for Generalized and Specialized Applications DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 028f0d74-561b-4578-81d8-a40e7a26d598 · outbound
Visual Large Language Models for Generalized and Specialized Applications MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c580a315-6f9f-4a4e-82a5-99a0acc31263 · outbound
Visual Large Language Models for Generalized and Specialized Applications Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64384dc8-e750-4d2a-891a-fce6c4f356e4 · outbound
Visual Large Language Models for Generalized and Specialized Applications Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72b8fb2e-97cb-44d3-935b-e5dbcebb5745 · outbound
Visual Large Language Models for Generalized and Specialized Applications Lisa: Reasoning segmentation via large language model,
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4039403b-4b4c-40e9-9075-3d60c57560ed · outbound
Visual Large Language Models for Generalized and Specialized Applications Groundhog: Grounding large language models to holistic segmentation,
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d2266d1-d6e9-4da1-bc83-b8e2a7bcc5d6 · outbound
Visual Large Language Models for Generalized and Specialized Applications Jack of all tasks master of many: Designing general-purpose coarse-to-fine vision-language model,
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 720f258a-3f55-4b43-ad82-ce55448b6d11 · outbound
Visual Large Language Models for Generalized and Specialized Applications Contextual Object Detection with Multimodal Large Language Models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34bac95e-f369-4a3e-9acc-cf5d70daf4f3 · outbound
Visual Large Language Models for Generalized and Specialized Applications Pixellm: Pixel reasoning with large multimodal model,
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4ac02f4-25e8-4613-8449-9383856b2d7b · outbound
Visual Large Language Models for Generalized and Specialized Applications Pixel-aligned language model,
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 773a9d28-bab2-418b-8967-c88e7e2d0008 · outbound
Visual Large Language Models for Generalized and Specialized Applications Gsva: Generalized segmentation via multimodal large language models,
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50087dbd-d00d-42b0-a186-a0e84e9fb5af · outbound
Visual Large Language Models for Generalized and Specialized Applications Llafs: When large language models meet few-shot segmentation,
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a35cb8eb-a35e-4c50-a8fe-96b1eea29335 · outbound
Visual Large Language Models for Generalized and Specialized Applications Glamm: Pixel grounding large multimodal model,
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6d090ca-31a8-4356-a9b4-b915f2976c72 · outbound
Visual Large Language Models for Generalized and Specialized Applications Visionllm: Large language model is also an open-ended decoder for vision-centric tasks,
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c29b4852-fd46-4d46-a2ca-4cd8a8d820b2 · outbound
Visual Large Language Models for Generalized and Specialized Applications VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cf59903-b8a9-4de8-9906-0e7bad060347 · outbound
Visual Large Language Models for Generalized and Specialized Applications Osprey: Pixel understanding with visual instruction tuning,
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3a54099-3b22-49d6-87b5-3e72f5081c62 · outbound
Visual Large Language Models for Generalized and Specialized Applications OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 295759cd-79e0-4dff-929b-98b16726d800 · outbound
Visual Large Language Models for Generalized and Specialized Applications Llm-seg: Bridging image segmentation and large language model reasoning,
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12eb69c5-d263-4009-a722-bd60c4d88cd6 · outbound
Visual Large Language Models for Generalized and Specialized Applications PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7939f39-074d-4d56-9661-49c6465823f6 · outbound
Visual Large Language Models for Generalized and Specialized Applications High-Quality Entity Segmentation and Grounding
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c8f212d-9576-4172-af34-77b7d3eac30b · outbound
Visual Large Language Models for Generalized and Specialized Applications DetGPT: Detect What You Need via Reasoning
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c274bec4-ace3-4ebd-b265-17053a7fb8a6 · outbound
Visual Large Language Models for Generalized and Specialized Applications Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3757acce-7c71-46a2-94bf-e6f0f6b3881b · outbound
Visual Large Language Models for Generalized and Specialized Applications Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e0c75be-d544-4487-99a1-7095e9845e83 · outbound
Visual Large Language Models for Generalized and Specialized Applications ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 821b2ad9-254b-4062-b5fe-f02f30c56f55 · outbound
Visual Large Language Models for Generalized and Specialized Applications GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ad3271b-c977-4fb8-8125-94e52b906c45 · outbound
Visual Large Language Models for Generalized and Specialized Applications BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97b84671-bcc8-4d56-8f9e-f065a7a5b6b4 · outbound
Visual Large Language Models for Generalized and Specialized Applications Pink: Unveiling the power of referential comprehension for multi-modal llms,
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d7e84db-6650-43bf-b8ab-700823197c4b · outbound
Visual Large Language Models for Generalized and Specialized Applications Ferret: Refer and Ground Anything Anywhere at Any Granularity
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d94408d8-8a68-4b3d-965e-527907aedec8 · outbound
Visual Large Language Models for Generalized and Specialized Applications Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e94c0bd-0618-47da-8a62-9a9240eead46 · outbound
Visual Large Language Models for Generalized and Specialized Applications InfMLLM: A Unified Framework for Visual-Language Tasks
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fbe80b3-402a-4d1d-ad35-135da6b9e3b6 · outbound
Visual Large Language Models for Generalized and Specialized Applications Lion: Empowering multimodal large language model with dual-level visual knowledge,
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e90aba5a-d9c9-4b29-85de-0d1999ffbb63 · outbound
Visual Large Language Models for Generalized and Specialized Applications SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36949f8a-0830-4cdb-9115-913f2a23aded · outbound
Visual Large Language Models for Generalized and Specialized Applications NExT-Chat: An LMM for Chat, Detection and Segmentation
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f517d6a-4f26-40b5-959d-1ca379553f64 · outbound
Visual Large Language Models for Generalized and Specialized Applications Griffon: Spelling out All Object Locations at Any Granularity with Large Language Models
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc59dfea-9a5f-4d55-acb7-7be0fffb2415 · outbound
Visual Large Language Models for Generalized and Specialized Applications CogVLM: Visual Expert for Pretrained Language Models
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1502aaa5-2893-4cbb-8bd8-183d28e27b86 · outbound
Visual Large Language Models for Generalized and Specialized Applications LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d20aa233-836a-467b-8115-2d971f4d6da0 · outbound
Visual Large Language Models for Generalized and Specialized Applications Lenna: Language Enhanced Reasoning Detection Assistant
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0416a874-b557-45af-b665-202975af1755 · outbound
Visual Large Language Models for Generalized and Specialized Applications ChatterBox: Multi-round Multimodal Referring and Grounding
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24b98bc0-fd9c-445c-8ff8-9b4a8d910f6a · outbound
Visual Large Language Models for Generalized and Specialized Applications Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models
Reference 102
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 678f2bb4-176f-48b2-b4bf-280e8c7b102c · inbound
Lost in Cultural Translation: Do LLMs Struggle with Math Across Cultural Contexts? Visual Large Language Models for Generalized and Specialized Applications
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 58e388ac-6946-4a66-8d9a-accd269b82d1 · inbound
Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects Visual Large Language Models for Generalized and Specialized Applications
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c41d5d1e-e5d9-4380-9d1f-1db7de0165a0 · inbound
IndustryEQA: Pushing the Frontiers of Embodied Question Answering in Industrial Scenarios Visual Large Language Models for Generalized and Specialized Applications
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2bfbacc-3d5b-48d5-a79f-f8303ca9cd2c · inbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Visual Large Language Models for Generalized and Specialized Applications
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78c8ddaa-47de-4843-8a1b-958b2fea1532 · inbound
AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Visual Large Language Models for Generalized and Specialized Applications
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 037ba41e-21a8-45b8-8f95-36081c8fdbb6 · inbound
A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects Visual Large Language Models for Generalized and Specialized Applications
Reference 215
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62210426-7a5e-41af-9e53-ebf45dfe7a3a · inbound
Leveraging Large Language Model for Intelligent Log Processing and Autonomous Debugging in Cloud AI Platforms Visual Large Language Models for Generalized and Specialized Applications
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e0db215-d347-40c5-9b13-e62da461b9e4 · inbound
Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens Visual Large Language Models for Generalized and Specialized Applications
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a8ad9122-832a-40e6-a471-fea949cb9352 · inbound
CoEmoGen: Towards Semantically-Coherent and Scalable Emotional Image Content Generation Visual Large Language Models for Generalized and Specialized Applications
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cda5e9a0-1daf-4993-89dd-7463904e7c07 · inbound
IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation Visual Large Language Models for Generalized and Specialized Applications
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8903798-d4cf-45d4-9ef3-aed4184c790c · inbound
Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making Visual Large Language Models for Generalized and Specialized Applications
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 187df4d5-837f-4c6a-a968-a90dbe08a29a · inbound
GaLa: Hypergraph-Guided Visual Language Models for Procedural Planning Visual Large Language Models for Generalized and Specialized Applications
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.