Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:10:19.297675Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 11 inbound Pith citation observations for arXiv:2501.05901.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:10:19.297675Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:46:40.770467Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T09:45:39.600613Z
92 of 92 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ab615871-4096-4225-9695-e5944592ab1c · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design TallyQA: Answering Complex Counting Questions
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b513fdf7-c303-4b95-a45f-7b65ffcb9363 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb92b421-f6c2-4d8e-8ddc-7ed51b510a0c · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Math-llava: Bootstrapping mathematical reasoning for multimodal large language models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c76ed4c-db84-4feb-96d8-59e7373d7fd5 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Scene text visual question answering
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47a88cda-d9bb-4064-8d9e-08f2f87d76a3 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Language models are few-shot learners
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cb0df00-067f-40a6-a5ae-22a1db441e83 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Coyo-700m: Image-text pair dataset
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4025de86-5d9a-4f13-9045-c08470052d69 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5838efec-a0a4-4e7a-ba76-66ff2d3f4d26 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0718d30d-9be4-45d4-b632-f34df5dbe19a · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07d3c61f-97c8-43a9-b894-6fb073d2041b · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 676dc6fd-a459-41f3-8e90-34323d5397fa · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea6a8da3-dac3-42e3-bd6c-0278617b14c8 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Opencompass: A universal evaluation platform for foundation models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62681333-18af-40fa-993f-ddc558937b75 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Patch n’ pack: Navit, a vision transformer for any aspect ratio and resolution, 2023
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91e0312c-25d2-4e9b-a0f7-87227ae4c170 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80de881f-3042-42e5-b01f-be3b710af1ef · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c899680f-fee7-4c7d-b750-e7abf6e1a6df · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c03a9b22-2881-4bda-b7ae-883bbcd2a9d4 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0c2b559-7bc1-484e-828f-9242f9ac77a8 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1462fed2-711d-4345-b0c5-c3e3e9323149 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Hallusion- bench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef12b89d-8900-4283-9f0e-12db7ea473fd · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Llava-uhd: an lmm perceiving any aspect ratio and high- resolution images
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 556e9947-1be6-4c92-84f4-4eb8b01a1400 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Vizwiz grand challenge: Answering visual questions from blind people
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1e5d2a2-3bff-4857-900b-05d99ab4b647 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbfef737-4be6-48b7-aa95-f7505d2cda4a · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cc1d556-df5e-4abc-b41f-386a1af80e37 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2b36f082-0491-46cc-8bf8-dd42e72110e0 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design A diagram is worth a dozen images
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 97bdcc57-f5c3-4628-896f-39c3de00308c · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design The hateful memes challenge: Detecting hate speech in multimodal memes
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5a668c09-d27c-4093-bd22-12d86d265cb8 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ed5e63c8-1221-44fd-9269-ba4a2db78de3 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Shamma, Michael S
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5e680620-55fa-4e52-a03c-cf52f6ab3c64 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Vqa-rad: Visual question answering dataset for radiology
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a06cd701-f672-4ff1-82ef-f0dc6354bcfc · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Super-clevr: A virtual benchmark to diagnose domain robustness in visual reasoning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 279fc10f-a3e4-4be4-8b7f-5214dc8d0fa3 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Microsoft coco: Common objects in context
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1f34d0e7-f572-4820-a25a-e8b8903f9628 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b87fee95-0856-495a-8e22-3c68b60d4298 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Visual Spatial Reasoning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fdb3330-8930-43c6-945b-ef2c112488e4 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design LLaVA-OneVision: Easy Visual Task Transfer
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43688fbe-a495-494e-b536-0252f7f4a1ea · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Visual Instruction Tuning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74c3b448-56bd-4900-a086-dbf51ec653b1 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c45c1d61-fafc-4bf1-a8d3-e6399630810f · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Visual instruction tuning.Advances in neural information processing systems , 36, 2024
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e4b6a11d-b402-4600-874f-fca0a5ed8fbf · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Optimal Transmitter Design and Pilot Spacing in MIMO Non-Stationary Aging Channels
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8ddac59e-487a-4567-8482-dfe89c57a8dd · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6fcd753-8a4f-4a5f-ad9b-50ec0fa983de · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design POINTS1.5: Building a Vision-Language Model towards Real World Applications
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cea65a07-8ace-4636-8444-dec67b24ff45 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbd25527-43aa-4167-a237-e0ec800df1cd · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Ocrbench: on the hidden mystery of ocr in large multimodal models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 426dea57-00aa-48a3-98fc-486b2eb1b324 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 51954eab-9ef7-4b15-87e6-d813b9c6c99d · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a99e5ca-649b-463b-a91a-dee72ac159ea · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 602a996e-fd75-43f3-ae7e-7f3003a15456 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5afb215e-1594-4a62-a25a-5249612787c4 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Scienceqa: A challenging dataset for multi-modal reasoning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9b315588-3a12-4fba-b859-fd49e30458d6 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faf8283a-7e18-46ca-a27c-1f8a2aeb734b · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df0713eb-5b23-46cd-b0e5-07f2bdd4e536 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c3bb458-c345-4ce0-8ed2-259c8fcd3909 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1eab67f3-c535-4a24-865d-b4530561e025 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Ok-vqa: A benchmark for visual question answering using external knowledge
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 353656bd-fe9b-4c6f-92e1-1b81dbff3304 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Ok-vqa: A benchmark for visual question answering using external knowledge
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation adf12f67-ea75-4759-a1a2-07e94e91a30d · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Marti and H
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 44b7547d-f58e-47a4-a323-65634818ad0d · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design ChartQA: A benchmark for question answering about charts with visual and logical reasoning
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3d2831b3-79aa-4907-bb89-0af2b5b1282f · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Mishra, K
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6346ee5f-5605-4fa1-890f-aaa2178c0369 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Icdar 2019 crohme + tfd: Competition on recognition of handwritten mathematical expressions and typeset formula detection
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 01413134-19ca-472f-b889-74e24ce34bea · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Training language models to follow instructions with human feedback
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb015d12-fe9b-48bf-b82d-eab90ac5e640 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Image Textualization: An Automatic Framework for Creating Accurate and Detailed Image Descriptions
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b479804-6008-4cb5-a524-caa13ed7ffe9 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b9be37e-c212-4e78-a38c-db1e56f6187e · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Improving language understanding by generative pre-training
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54e21e84-9802-4b97-b19f-a4cd5e5b978a · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Unresolved cited work
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8c1421ce-8e09-4402-ad0f-3d0f91432404 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2d48494-a77d-44dd-a784-1cdc906d3bd9 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ec2e600a-074d-4f2f-bf64-3a83eefd9069 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design TextCaps: a Dataset for Image Captioning with Reading Comprehension
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e8077b3-2988-41c0-b4e8-8e1c72a91eb9 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Singh, V
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 09241061-7e7a-45df-aff4-784762ba2d61 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ca5a63c9-7208-4835-95c2-001239932aa2 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Generative multimodal models are in-context learners
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca662c2d-b219-4774-af37-dde5ac2468f8 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Bluelm: An open multilingual 7b language model
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 73800cdf-9925-4ec7-a495-d7b0270492c1 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design G-llava: Solving geometric problem with multi-modal large language model
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 88552b8c-acd1-4b67-955c-0707907b8ffa · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0252062-f4ca-48f6-b1a4-fecbe6d33d66 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57530a53-239c-4ace-8891-15428ad7a436 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97af1f60-fa2e-47f0-9ea1-036141f5622b · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c2defd8-7458-410c-a61a-b57c4d6dd806 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design xgen-mm(blip-3): A family of open large multimodal models
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e3d56420-b439-47f2-a7a7-6d73f7f31fff · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Qwen2.5 Technical Report
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de5346b8-f4ac-4498-8ad4-89cf841e88fb · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9fdfd70-3e8f-4ca7-8e13-458658167b38 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Modeling context in referring expressions
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1806a7ff-388d-4c4b-9707-0687db3b66f6 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Mm-vet: Evaluating large multimodal models for integrated capabilities
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 343f21ad-1e5f-4a60-b8f4-2c44f6368eac · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Syntax-aware network for handwritten mathematical expression recognition
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 267b1168-795c-44e7-954c-d4d5fe374a13 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3d01ea9-d201-4c7d-982f-76d57e889848 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Sigmoid loss for language image pre-training, 2023
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2886988-a028-45f1-b81f-d218cc139a4b · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Raven: A dataset for relational and analogical visual reasoning
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 168265e7-b8c0-42a8-a794-7549f0767f37 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Unimath: A foundational and multimodal mathematical reasoner
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bdcd38c5-f79f-4057-a19b-7d0a72bfd90a · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design InfinityMATH: A Scalable Instruction Tuning Dataset in Programmatic Mathematical Reasoning
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f81bbf74-28f4-490b-832b-344e24487c5f · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e36f441-a28e-497c-9b40-36cd2d2faff2 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73c457fe-6a3c-4da4-a64d-2612acffd3cb · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Improve Vision Language Model Chain-of-thought Reasoning
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0489dc65-fdfd-472b-9299-b388e7b33607 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9de82201-187f-4980-a1ad-680eb986fd42 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Visual7w: Grounded question answering in images
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ce148f0c-2014-4bcd-ac11-1879993774e0 · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design doi: 10.18653/v1/2022.findings-acl.177
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a06ac0d2-14e8-4317-87d0-8bd78ffc564d · outbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Unresolved cited work
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e1190d8e-f3dd-4f60-b6ae-010722314397 · inbound
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Valley2: Exploring Multimodal Models with Scalable Vision-Language Design
Reference 155
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5a16af12-89a8-4a63-b999-feeecd9dd16e · inbound
Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark Valley2: Exploring Multimodal Models with Scalable Vision-Language Design
Reference 136
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ea8ea45-ee51-4ca6-b8c8-5614187626ea · inbound
Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models Valley2: Exploring Multimodal Models with Scalable Vision-Language Design
Reference 116
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf167009-ab5c-4182-90d0-627fb751f613 · inbound
R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO Valley2: Exploring Multimodal Models with Scalable Vision-Language Design
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7ffcbb5-699a-45e6-a7b4-69e78ef5f9e6 · inbound
GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Valley2: Exploring Multimodal Models with Scalable Vision-Language Design
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1165e788-0faf-4ce7-b2d3-b886ed25850a · inbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Valley2: Exploring Multimodal Models with Scalable Vision-Language Design
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baa7f74c-5f0a-411f-8950-4950baa3dbff · inbound
Interpretable Open-Vocabulary Referring Object Detection with Reverse Contrast Attention Valley2: Exploring Multimodal Models with Scalable Vision-Language Design
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8d4a86e-227e-4eb6-b32a-43cb1ac76d0b · inbound
R1-SyntheticVL: Is Synthetic Data from Generative Models Ready for Multimodal Large Language Model? Valley2: Exploring Multimodal Models with Scalable Vision-Language Design
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81c2255b-4e49-474f-8c34-793f6145de02 · inbound
Valley3: Scaling Omni Foundation Models for E-commerce Valley2: Exploring Multimodal Models with Scalable Vision-Language Design
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8097c691-8871-4976-a9d6-5017b52998cc · inbound
Learning to Deny: Action Denial in Multimodal Large Language Models Valley2: Exploring Multimodal Models with Scalable Vision-Language Design
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation de51de91-1470-4da0-9212-b98ff718d6b0 · inbound
VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models Valley2: Exploring Multimodal Models with Scalable Vision-Language Design
Reference 118
Source-reported events for the cited work
Unavailable: canonical work link unavailable.