Pith. sign in

Paper Citation Record · LEDGER

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design

As of 18 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 11 inbound Pith citation observations for arXiv:2501.05901.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.05901 v2

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:10:19.297675Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:46:40.770467Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T09:45:39.600613Z

Reference resolution

92 of 92 outbound references displayed

  • verified exact1
  • verified fuzzy26
  • unresolved64
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ab615871-4096-4225-9695-e5944592ab1c · outbound

This paper cites TallyQA: Answering Complex Counting Questions.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design TallyQA: Answering Complex Counting Questions

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:18.841141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:18.841141Z digest=sha256:f69c30dbf041660a658c30768bb931b8381b040c49504f52afd5c7e207ec583d

Observation b513fdf7-c303-4b95-a45f-7b65ffcb9363 · outbound

This paper cites GPT-4 Technical Report.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:18.847175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:18.847175Z digest=sha256:ad0f103c548ab6d98ee90d2a8efe1626cb1312ed7a08eb531c7f76bd44d554aa

Observation bb92b421-f6c2-4d8e-8ddc-7ed51b510a0c · outbound

This paper cites Math-llava: Bootstrapping mathematical reasoning for multimodal large language models.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Math-llava: Bootstrapping mathematical reasoning for multimodal large language models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:18.852430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:18.852430Z digest=sha256:014427c293d8fd819986087157a8902f8c33877e112ca36b0edbf9b77ddd77fe

Observation 1c76ed4c-db84-4feb-96d8-59e7373d7fd5 · outbound

This paper cites Scene text visual question answering.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Scene text visual question answering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:18.857075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:18.857075Z digest=sha256:0c2570a3ce8b7753f1bb45541b84f5fa22bd24d29f3faf40eb5cef78b6894e3d

Observation 47a88cda-d9bb-4064-8d9e-08f2f87d76a3 · outbound

This paper cites Language models are few-shot learners.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Language models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:18.861346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:18.861346Z digest=sha256:7ec628f44f2a4181aa4b111d8f967a431d4de6ebd142377c282fc712f81607b2

Observation 2cb0df00-067f-40a6-a5ae-22a1db441e83 · outbound

This paper cites Coyo-700m: Image-text pair dataset.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Coyo-700m: Image-text pair dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:18.865508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:18.865508Z digest=sha256:289e94651f4b6161ae8952e8f69556a13bb4f3a76cc6bc6beb81f96e87815a10

Observation 4025de86-5d9a-4f13-9045-c08470052d69 · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:18.869750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:18.869750Z digest=sha256:3fc99304de5188bc3af33c18449514645239f3e54b681d19f80aa49f5e86d1f8

Observation 5838efec-a0a4-4e7a-ba76-66ff2d3f4d26 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:18.874649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:18.874649Z digest=sha256:fc6661bfa6b4c2e651c47af10ce13b60375a0b2c27b13733ff272078c988f627

Observation 0718d30d-9be4-45d4-b632-f34df5dbe19a · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:18.878895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:18.878895Z digest=sha256:55a9304ef43495da2a05b1b6b1d37853c839df314098b790a59774aad23a2757

Observation 07d3c61f-97c8-43a9-b894-6fb073d2041b · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:18.883573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:18.883573Z digest=sha256:7e672a0b5070e86394d703baf1a83168c7541c14ecb160c2aef83f05689104a2

Observation 676dc6fd-a459-41f3-8e90-34323d5397fa · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:18.888636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:18.888636Z digest=sha256:9f477bc2c6948a47e69effc4d8609b7aeb07bd7246744e836db94c653fefad4e

Observation ea6a8da3-dac3-42e3-bd6c-0278617b14c8 · outbound

This paper cites Opencompass: A universal evaluation platform for foundation models.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Opencompass: A universal evaluation platform for foundation models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:18.893598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:18.893598Z digest=sha256:e25829642523500c33841c19adc7cee2f0d72db125f7327f5c5a1c7b9323f58a

Observation 62681333-18af-40fa-993f-ddc558937b75 · outbound

This paper cites Patch n’ pack: Navit, a vision transformer for any aspect ratio and resolution, 2023.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Patch n’ pack: Navit, a vision transformer for any aspect ratio and resolution, 2023

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:18.897609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:18.897609Z digest=sha256:13f658c7c7a0cc1da683c6103146f557f47b2a3a78e7470370ff93f611b31051

Observation 91e0312c-25d2-4e9b-a0f7-87227ae4c170 · outbound

This paper cites Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:18.901451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:18.901451Z digest=sha256:43df9913241b013865cfdb407d763d3a74754339e2a899b3de668030ec72b514

Observation 80de881f-3042-42e5-b01f-be3b710af1ef · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:18.905756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:18.905756Z digest=sha256:dfce75077a47ed78ba6375eef1bea914d68105f4de6a5342cb19fd25d155ef7f

Observation c899680f-fee7-4c7d-b750-e7abf6e1a6df · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:18.910490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:18.910490Z digest=sha256:5209bee72b19c665d196269b7eec1a7bf5bf4808110ac74814ffeee5b5366dc4

Observation c03a9b22-2881-4bda-b7ae-883bbcd2a9d4 · outbound

This paper cites Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:18.919339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:18.919339Z digest=sha256:fece5ba2c72701a8f3993f8a7d6ff0743b65e58dbb803dc7ed8f0abc0b501476

Observation d0c2b559-7bc1-484e-828f-9242f9ac77a8 · outbound

This paper cites Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:18.923630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:18.923630Z digest=sha256:5db0682139002047c70286192262624c44e2e50225cce4de89f87fd34b46a100

Observation 1462fed2-711d-4345-b0c5-c3e3e9323149 · outbound

This paper cites Hallusion- bench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Hallusion- bench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:18.927686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:18.927686Z digest=sha256:d96a65aecafbebd311cf20938da75897b8eff2739dd3871d6cf8513529213109

Observation ef12b89d-8900-4283-9f0e-12db7ea473fd · outbound

This paper cites Llava-uhd: an lmm perceiving any aspect ratio and high- resolution images.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Llava-uhd: an lmm perceiving any aspect ratio and high- resolution images

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:18.931724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:18.931724Z digest=sha256:67ffb65c4640d19665318a7aa9ec129c1e2cbe0e81e070bc8ba4ea37379addee

Observation 556e9947-1be6-4c92-84f4-4eb8b01a1400 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Vizwiz grand challenge: Answering visual questions from blind people

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:18.936887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:18.936887Z digest=sha256:c942c1a11da8d0063955d57ba37ebd3440ecb962acba90ee72bb55607f4e4cf8

Observation c1e5d2a2-3bff-4857-900b-05d99ab4b647 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:18.941965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:18.941965Z digest=sha256:abd93668caaefc184e1018822f95ef517467c398a08e4fa4a20f27494ac5cf77

Observation bbfef737-4be6-48b7-aa95-f7505d2cda4a · outbound

This paper cites an unresolved cited work.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:18.946148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:18.946148Z digest=sha256:ab211f45ebcd1e8ec3d07824b987613ec38927c7413be3591e13167a7613305c

Observation 4cc1d556-df5e-4abc-b41f-386a1af80e37 · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Clevr: A diagnostic dataset for compositional language and elementary visual reasoning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:20.525148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:18.950270Z digest=sha256:1ad1e74f9cec7c260c3ea0cf88be6c717c298ed9d429a4bb6a6219c94fcbc6dc

Observation 2b36f082-0491-46cc-8bf8-dd42e72110e0 · outbound

This paper cites A diagram is worth a dozen images.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design A diagram is worth a dozen images

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:20.512612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:18.954354Z digest=sha256:bf416de01db5351076ea0e9396bbf6b330f70c1efd597c7e5945c576c6a0e99e

Observation 97bdcc57-f5c3-4628-896f-39c3de00308c · outbound

This paper cites The hateful memes challenge: Detecting hate speech in multimodal memes.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design The hateful memes challenge: Detecting hate speech in multimodal memes

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:20.500901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:18.958752Z digest=sha256:6379d3b02c9d452a6893a36ad13e2efc3d06404485a196ccbad4d0ffe80c1069

Observation 5a668c09-d27c-4093-bd22-12d86d265cb8 · outbound

This paper cites an unresolved cited work.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:10:20.489501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:18.962818Z digest=sha256:a7edf56d1d470ae5dc389423da300bfc00c0c4fc71b88e300a21f7f38e4af8b0

Observation ed5e63c8-1221-44fd-9269-ba4a2db78de3 · outbound

This paper cites Shamma, Michael S.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Shamma, Michael S

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:20.478120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:18.967044Z digest=sha256:79f1a1d771aec2d4894bb6b277baa7d66df7fcff2606825c95735c9c8db3e91a

Observation 5e680620-55fa-4e52-a03c-cf52f6ab3c64 · outbound

This paper cites Vqa-rad: Visual question answering dataset for radiology.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Vqa-rad: Visual question answering dataset for radiology

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:20.464301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:18.972717Z digest=sha256:92838d3dc64662bb5683ab8841e40d44a04a583b27b367fabb01c8fed59f5455

Observation a06cd701-f672-4ff1-82ef-f0dc6354bcfc · outbound

This paper cites Super-clevr: A virtual benchmark to diagnose domain robustness in visual reasoning.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Super-clevr: A virtual benchmark to diagnose domain robustness in visual reasoning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:20.451654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:18.982673Z digest=sha256:320746ea3f0a916198f2ef95c81fc3612b59f8627c3efcdbeda1b068bba03f11

Observation 279fc10f-a3e4-4be4-8b7f-5214dc8d0fa3 · outbound

This paper cites Microsoft coco: Common objects in context.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Microsoft coco: Common objects in context

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:20.438500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:18.987614Z digest=sha256:3faa8ec08f35262565c96d8d2b433402d47c08f6272d39a0594bc99907a007dd

Observation 1f34d0e7-f572-4820-a25a-e8b8903f9628 · outbound

This paper cites CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:18.992036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:18.992036Z digest=sha256:5ac1e8bbf43dc3387f539dca6543557d3803eb0312fd419fe373230525e61daf

Observation b87fee95-0856-495a-8e22-3c68b60d4298 · outbound

This paper cites Visual Spatial Reasoning.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Visual Spatial Reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:18.996958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:18.996958Z digest=sha256:a76b4a925fd9614172be891df3cfff5041f06241113d14d21ce65da9bf01672a

Observation 7fdb3330-8930-43c6-945b-ef2c112488e4 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design LLaVA-OneVision: Easy Visual Task Transfer

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.001752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.001752Z digest=sha256:b5e753a52d98a08e71fddc6855f7c5ce579745a1778ae40c880b14e6ea1b7eff

Observation 43688fbe-a495-494e-b536-0252f7f4a1ea · outbound

This paper cites Visual Instruction Tuning.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Visual Instruction Tuning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.010663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.010663Z digest=sha256:fd11320a4e2ab82c00d457b92609196d6988148683025624dffce1242f4ba772

Observation 74c3b448-56bd-4900-a086-dbf51ec653b1 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.015439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.015439Z digest=sha256:64c13f44d20eef5610a5f2195dde18b7fcf7bfa61958ca930fea97b9a14f3049

Observation c45c1d61-fafc-4bf1-a8d3-e6399630810f · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems , 36, 2024.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Visual instruction tuning.Advances in neural information processing systems , 36, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:20.419011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:19.019537Z digest=sha256:c390ccb1f36b6bd2ec0eeaa7cdafa1c58f326965044adf2ba2f281f20e231642

Observation e4b6a11d-b402-4600-874f-fca0a5ed8fbf · outbound

This paper cites Optimal Transmitter Design and Pilot Spacing in MIMO Non-Stationary Aging Channels.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Optimal Transmitter Design and Pilot Spacing in MIMO Non-Stationary Aging Channels

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T21:10:19.700408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:19.023914Z digest=sha256:4a3b409d97244aa3a18024aa459220c8faadca4b70cc710f0e9202cb651612ab

Observation 8ddac59e-487a-4567-8482-dfe89c57a8dd · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.028341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.028341Z digest=sha256:3be4dc403d323e547b58314c8254d02e9bae4c372f3ac6a37bdf78176c1ef884

Observation d6fcd753-8a4f-4a5f-ad9b-50ec0fa983de · outbound

This paper cites POINTS1.5: Building a Vision-Language Model towards Real World Applications.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design POINTS1.5: Building a Vision-Language Model towards Real World Applications

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.033317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.033317Z digest=sha256:ac0f00adf101400123cade21106bd3ecc343f1ead32febb067ed94e7e76602d3

Observation cea65a07-8ace-4636-8444-dec67b24ff45 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.038904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.038904Z digest=sha256:99a87e71e80bab37a5e929024a082e61482ac60bb0ee898b002e81fe8f62f7b1

Observation fbd25527-43aa-4167-a237-e0ec800df1cd · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Ocrbench: on the hidden mystery of ocr in large multimodal models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.045966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.045966Z digest=sha256:c624f82468e0deaedcb335380a08c0f394a9a4a9dc87509ceadf601d624edea3

Observation 426dea57-00aa-48a3-98fc-486b2eb1b324 · outbound

This paper cites an unresolved cited work.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:10:20.398456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:19.051761Z digest=sha256:f4f454d125b7a016822d10e567fbaef511d4723e0b14f6a58046ec6dc9a84be3

Observation 51954eab-9ef7-4b15-87e6-d813b9c6c99d · outbound

This paper cites Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.060895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.060895Z digest=sha256:aae6952bc52331576d95a91e448282c9bd4aae70e683f524ee4db9d4c2c3a95b

Observation 1a99e5ca-649b-463b-a91a-dee72ac159ea · outbound

This paper cites An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.065456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.065456Z digest=sha256:a73d71aa5582aa323edf417025fc04e2a8f890a926a5109058e1e29ac02d69fa

Observation 602a996e-fd75-43f3-ae7e-7f3003a15456 · outbound

This paper cites Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.071219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.071219Z digest=sha256:b6f14f11b5eeb8d050864bba400b011aa375fba4277b52e636b4d3b29e68374d

Observation 5afb215e-1594-4a62-a25a-5249612787c4 · outbound

This paper cites Scienceqa: A challenging dataset for multi-modal reasoning.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Scienceqa: A challenging dataset for multi-modal reasoning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:20.387378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:19.076081Z digest=sha256:a49a64385b9f3ef4a74c6d64a41a006635f5760dbeead8c28fcb438e36d358ed

Observation 9b315588-3a12-4fba-b859-fd49e30458d6 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.080916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.080916Z digest=sha256:4afbbaf9ceb9d5197704c8a691c2f089bf3ffbaf6054037d344642f540a64c17

Observation faf8283a-7e18-46ca-a27c-1f8a2aeb734b · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.086305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.086305Z digest=sha256:eb7a0a7fca496db26152b40253d7a631ccb3373b520d2e515fced95a665cb725

Observation df0713eb-5b23-46cd-b0e5-07f2bdd4e536 · outbound

This paper cites BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.092613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.092613Z digest=sha256:1dd5242e8dce4efca0fb66c22a2de29f5ca76d70e76ad9b48afc0af896992f5c

Observation 2c3bb458-c345-4ce0-8ed2-259c8fcd3909 · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.097996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.097996Z digest=sha256:370a828c623b09c0778cb9e91b0cf5e735ffb09ffd249b07c471e68584cee279

Observation 1eab67f3-c535-4a24-865d-b4530561e025 · outbound

This paper cites Ok-vqa: A benchmark for visual question answering using external knowledge.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Ok-vqa: A benchmark for visual question answering using external knowledge

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:20.358350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:19.102417Z digest=sha256:8701b2d755fe043c6342ef64ba56e08cc2dd0b54c046bd5fe8550327739f774e

Observation 353656bd-fe9b-4c6f-92e1-1b81dbff3304 · outbound

This paper cites Ok-vqa: A benchmark for visual question answering using external knowledge.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Ok-vqa: A benchmark for visual question answering using external knowledge

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:20.340532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:19.106298Z digest=sha256:5932e65add1e11f75b52e4aa4949d67a7323f0878cbef69d113ed80b6e935a8d

Observation adf12f67-ea75-4759-a1a2-07e94e91a30d · outbound

This paper cites Marti and H.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Marti and H

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:20.324986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:19.110900Z digest=sha256:fa2811feb2520c8cee7ffeb46e3620b9668ffaf37230e9eb83889de95a24c6b8

Observation 44b7547d-f58e-47a4-a323-65634818ad0d · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design ChartQA: A benchmark for question answering about charts with visual and logical reasoning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:20.311658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:19.115965Z digest=sha256:1b92e8cee3117698cacdeaaca9f2f9ac26f8b20f4d8909c5d800b36221d59d7f

Observation 3d2831b3-79aa-4907-bb89-0af2b5b1282f · outbound

This paper cites Mishra, K.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Mishra, K

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:20.298661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:19.124809Z digest=sha256:c682168b5caff95e1f2ec48fc3b1fb197c7cf30aa69f0ee0e8ebfa2b9de79343

Observation 6346ee5f-5605-4fa1-890f-aaa2178c0369 · outbound

This paper cites Icdar 2019 crohme + tfd: Competition on recognition of handwritten mathematical expressions and typeset formula detection.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Icdar 2019 crohme + tfd: Competition on recognition of handwritten mathematical expressions and typeset formula detection

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:20.285723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:19.128725Z digest=sha256:1a2b8746bdc1e8efbcf1e2675a07dee049b5b59f5f04ae25d6001941f056e2a4

Observation 01413134-19ca-472f-b889-74e24ce34bea · outbound

This paper cites Training language models to follow instructions with human feedback.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Training language models to follow instructions with human feedback

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.132849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.132849Z digest=sha256:a7a21445f2fb4ac7af28d2bcb0c8d2a26744370df0a57d131d6e3f198ec849aa

Observation eb015d12-fe9b-48bf-b82d-eab90ac5e640 · outbound

This paper cites Image Textualization: An Automatic Framework for Creating Accurate and Detailed Image Descriptions.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Image Textualization: An Automatic Framework for Creating Accurate and Detailed Image Descriptions

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.136628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.136628Z digest=sha256:88c8d19eb798f6909e821c286d60b658d1a6712b8df5b1a09a870df0a215bd20

Observation 8b479804-6008-4cb5-a524-caa13ed7ffe9 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.141439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.141439Z digest=sha256:fc8244f1cae5ed11a963f6b534c2a014e19087eca1a5c95e0d3bf02d68e5d657

Observation 3b9be37e-c212-4e78-a38c-db1e56f6187e · outbound

This paper cites Improving language understanding by generative pre-training.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Improving language understanding by generative pre-training

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.146845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.146845Z digest=sha256:2b17509c002d01061cfb4fa9ee2127ee503f42393cf682600ea503e87bed09dc

Observation 54e21e84-9802-4b97-b19f-a4cd5e5b978a · outbound

This paper cites an unresolved cited work.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:10:20.240517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:19.151179Z digest=sha256:3780f9fa40f6b64413ca790c90f6ae09d5268516f49018628e11584c3aea2c80

Observation 8c1421ce-8e09-4402-ad0f-3d0f91432404 · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.156640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.156640Z digest=sha256:282395eb15d15bef959d0c9be5d9b3c07fa30c7d4ca664eb99bc4efeebe27bf2

Observation f2d48494-a77d-44dd-a784-1cdc906d3bd9 · outbound

This paper cites Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:20.228777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:19.162085Z digest=sha256:e9c0c77505fe31d6f26224b01be52b4b66e58b922aefa6c1bd2a0e12e1724c43

Observation ec2e600a-074d-4f2f-bf64-3a83eefd9069 · outbound

This paper cites TextCaps: a Dataset for Image Captioning with Reading Comprehension.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design TextCaps: a Dataset for Image Captioning with Reading Comprehension

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.166540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.166540Z digest=sha256:07c664855bc8fdfabeb94f71d8d353cdea42460109fb2b78ebb7292c373146aa

Observation 0e8077b3-2988-41c0-b4e8-8e1c72a91eb9 · outbound

This paper cites Singh, V.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Singh, V

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:20.216740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:19.170656Z digest=sha256:fdd67274ae64af52640d26beec24389a84cfa09f24a496cfe78207bce1c236d2

Observation 09241061-7e7a-45df-aff4-784762ba2d61 · outbound

This paper cites Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:20.204075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:19.175230Z digest=sha256:a8e3c58d1fba8dd31aab254c3c5367fa49553a122cec1bdc41b9907ceab7e115

Observation ca5a63c9-7208-4835-95c2-001239932aa2 · outbound

This paper cites Generative multimodal models are in-context learners.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Generative multimodal models are in-context learners

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.180364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.180364Z digest=sha256:cb54cedeea03db446d6689d301bfd3894f0e3d4031c8d0eb564841843b399d38

Observation ca662c2d-b219-4774-af37-dde5ac2468f8 · outbound

This paper cites Bluelm: An open multilingual 7b language model.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Bluelm: An open multilingual 7b language model

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:20.182778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:19.188306Z digest=sha256:2df7cfd5a00443a6d7738ab3a6707794a31c266c73f3bf414b58b4343d930a9a

Observation 73800cdf-9925-4ec7-a495-d7b0270492c1 · outbound

This paper cites G-llava: Solving geometric problem with multi-modal large language model.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design G-llava: Solving geometric problem with multi-modal large language model

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:20.170288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:19.193069Z digest=sha256:fe720b89109a25ea64bac23b39d3a18ff724093fa57c65a3809b5a2c877c4e74

Observation 88552b8c-acd1-4b67-955c-0707907b8ffa · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.197628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.197628Z digest=sha256:5af13f9a1d6505597a6e98e83da41ddc384dbef5d012e26cb98d96ac33f3e2e4

Observation a0252062-f4ca-48f6-b1a4-fecbe6d33d66 · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.202065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.202065Z digest=sha256:df85f9425353d4aadb32c19bbb6a6a62e77b1ff860d611ada13afe52538ce1f4

Observation 57530a53-239c-4ace-8891-15428ad7a436 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.208549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.208549Z digest=sha256:33882e5f53526fe6980de511e81d89235662dc5248221009bd02dc4d7c0e86a8

Observation 97af1f60-fa2e-47f0-9ea1-036141f5622b · outbound

This paper cites Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.215334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.215334Z digest=sha256:191794c6002faae56380dc8ac17072fc8e88f5b546fc160a78ca5a976dfee08d

Observation 7c2defd8-7458-410c-a61a-b57c4d6dd806 · outbound

This paper cites xgen-mm(blip-3): A family of open large multimodal models.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design xgen-mm(blip-3): A family of open large multimodal models

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:20.155096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:19.220050Z digest=sha256:63dba3f99af8914cdb5b8a839def2f916e5a0507ebd80e9ea6395709a7df121e

Observation e3d56420-b439-47f2-a7a7-6d73f7f31fff · outbound

This paper cites Qwen2.5 Technical Report.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Qwen2.5 Technical Report

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.224500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.224500Z digest=sha256:b28891923bf5764bfa716d575d957ddf901b206eb2be66e80b7169386a603822

Observation de5346b8-f4ac-4498-8ad4-89cf841e88fb · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.228951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.228951Z digest=sha256:0424cc54b094be8a94cf2cbfa2bcca7cd935a7b75f5f8a9a015c21babdc48f77

Observation c9fdfd70-3e8f-4ca7-8e13-458658167b38 · outbound

This paper cites Modeling context in referring expressions.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Modeling context in referring expressions

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:20.142272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:19.233265Z digest=sha256:fd85bc1932bc146bdac96ab27b7d2d86182f192934b4d71d1e50153f611061fd

Observation 1806a7ff-388d-4c4b-9707-0687db3b66f6 · outbound

This paper cites Mm-vet: Evaluating large multimodal models for integrated capabilities.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Mm-vet: Evaluating large multimodal models for integrated capabilities

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.237553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.237553Z digest=sha256:db3dd8507789872add43cea83d314acacf11fbc1bede4221c7bfcd606e85599a

Observation 343f21ad-1e5f-4a60-b8f4-2c44f6368eac · outbound

This paper cites Syntax-aware network for handwritten mathematical expression recognition.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Syntax-aware network for handwritten mathematical expression recognition

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:20.121985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:19.242436Z digest=sha256:f52b3d1a2c40acd682b7c53366997fb55d593b482083d57436b60fa94944fdbc

Observation 267b1168-795c-44e7-954c-d4d5fe374a13 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.247526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.247526Z digest=sha256:f155147a2df0f621377e3ec0de1e4e56f2e7efdec1752393712b053f74047a36

Observation a3d01ea9-d201-4c7d-982f-76d57e889848 · outbound

This paper cites Sigmoid loss for language image pre-training, 2023.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Sigmoid loss for language image pre-training, 2023

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.252682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.252682Z digest=sha256:59d2cab9e796fcd2e9740bddcc08ed7ecd3ee1afbeee136c272ee2d74d7fea11

Observation c2886988-a028-45f1-b81f-d218cc139a4b · outbound

This paper cites Raven: A dataset for relational and analogical visual reasoning.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Raven: A dataset for relational and analogical visual reasoning

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:20.088977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:19.256873Z digest=sha256:48e5d5c6dece9cc5af2e5815d180e642a5804b0347f2568a7dd3a2b2c1a900a9

Observation 168265e7-b8c0-42a8-a794-7549f0767f37 · outbound

This paper cites Unimath: A foundational and multimodal mathematical reasoner.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Unimath: A foundational and multimodal mathematical reasoner

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:20.076146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:19.261072Z digest=sha256:5d6c6d2b91aa5aaa208ca65912977b19116ddffe9c6a67f3d32707160dbf2d21

Observation bdcd38c5-f79f-4057-a19b-7d0a72bfd90a · outbound

This paper cites InfinityMATH: A Scalable Instruction Tuning Dataset in Programmatic Mathematical Reasoning.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design InfinityMATH: A Scalable Instruction Tuning Dataset in Programmatic Mathematical Reasoning

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:10:19.427271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:19.273158Z digest=sha256:c49e9002913ddd67e09674c68686bf3c583c142abb33e2b0bbf12d848aec71ee

Observation f81bbf74-28f4-490b-832b-344e24487c5f · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.279181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.279181Z digest=sha256:a7ad0f695cf553c102ea56e88b579083f14d5c084b48f46e9ec3056a68eaf65f

Observation 6e36f441-a28e-497c-9b40-36cd2d2faff2 · outbound

This paper cites MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.283464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.283464Z digest=sha256:195fb5ce60e78ef5c87310dcf26d95e7390a600fd6333163514f514bf177c5a5

Observation 73c457fe-6a3c-4da4-a64d-2612acffd3cb · outbound

This paper cites Improve Vision Language Model Chain-of-thought Reasoning.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Improve Vision Language Model Chain-of-thought Reasoning

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.288710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.288710Z digest=sha256:ddf3c1afd898b2a1b8ab9249aa8edf40c6cfb65d4853671a6067515adef77be0

Observation 0489dc65-fdfd-472b-9299-b388e7b33607 · outbound

This paper cites Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.293396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.293396Z digest=sha256:587a948b958d5d6535e67fadafa3dbbf99136f1963c21028f0fde969188735db

Observation 9de82201-187f-4980-a1ad-680eb986fd42 · outbound

This paper cites Visual7w: Grounded question answering in images.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Visual7w: Grounded question answering in images

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:20.049394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:19.297675Z digest=sha256:32bfc5b4f76f736abb742c8a5888773c383778e2083d48856801e3a682904d35

Observation ce148f0c-2014-4bcd-ac11-1879993774e0 · outbound

This paper cites doi: 10.18653/v1/2022.findings-acl.177.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design doi: 10.18653/v1/2022.findings-acl.177

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.120428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.120428Z digest=sha256:d47526d6242c6b6c109dabf082be88569f25793c19829a0abd1da443f7b1a805

Observation a06ac0d2-14e8-4317-87d0-8bd78ffc564d · outbound

This paper cites an unresolved cited work.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:10:20.062689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T21:10:19.267832Z digest=sha256:1a015f10c235c0cdc1811c33e3d7fc74d090c73341ba56a03d0ff45ed6b6737b

Pith citing papers

Observation e1190d8e-f3dd-4f60-b6ae-010722314397 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Valley2: Exploring Multimodal Models with Scalable Vision-Language Design

Reference 155

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:20:00.119265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:b76e9946c1c360e53619cbcb7f8f30f0a6ed6d5742f7d01b3b6472fc5faca664

Observation 5a16af12-89a8-4a63-b999-feeecd9dd16e · inbound

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark cites this paper.

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark Valley2: Exploring Multimodal Models with Scalable Vision-Language Design

Reference 136

Resolution
unresolved
no resolver link, observed 2026-08-16T11:46:40.770467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:46:40.770467Z digest=sha256:ac5f284026e2d99d6688c324cd97109d7bd07d665e59ec051564d76aef51719a

Observation 8ea8ea45-ee51-4ca6-b8c8-5614187626ea · inbound

Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models cites this paper.

Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models Valley2: Exploring Multimodal Models with Scalable Vision-Language Design

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-16T05:12:18.597250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:12:18.597250Z digest=sha256:620b4a1cd6ba05f2f767df10c53f9238dcb6cd5918ded378d6e13a04dca4525f

Observation bf167009-ab5c-4182-90d0-627fb751f613 · inbound

R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO cites this paper.

R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO Valley2: Exploring Multimodal Models with Scalable Vision-Language Design

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:01:33.718044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:01:33.718044Z digest=sha256:ce5436162fd768aeef0c2f32bd33b0a0dd3027cbe3205e9ee90bf66998a7513f

Observation f7ffcbb5-699a-45e6-a7b4-69e78ef5f9e6 · inbound

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking cites this paper.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Valley2: Exploring Multimodal Models with Scalable Vision-Language Design

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.123082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.123082Z digest=sha256:450823bc64d14e1121a4b170e47904302660f0ede035b6f865b6c2c0787abc82

Observation 1165e788-0faf-4ce7-b2d3-b886ed25850a · inbound

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? cites this paper.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Valley2: Exploring Multimodal Models with Scalable Vision-Language Design

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:34.362773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:34.362773Z digest=sha256:92e824a5050f54aeec1787c55715928bf49e09725c9e4a12108db84b0bfdb922

Observation baa7f74c-5f0a-411f-8950-4950baa3dbff · inbound

Interpretable Open-Vocabulary Referring Object Detection with Reverse Contrast Attention cites this paper.

Interpretable Open-Vocabulary Referring Object Detection with Reverse Contrast Attention Valley2: Exploring Multimodal Models with Scalable Vision-Language Design

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:07.467028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:07.467028Z digest=sha256:39ec06a8f73ede8cbc7f170a303580f71351f85a102e317450152f365ff8d139

Observation a8d4a86e-227e-4eb6-b32a-43cb1ac76d0b · inbound

R1-SyntheticVL: Is Synthetic Data from Generative Models Ready for Multimodal Large Language Model? cites this paper.

R1-SyntheticVL: Is Synthetic Data from Generative Models Ready for Multimodal Large Language Model? Valley2: Exploring Multimodal Models with Scalable Vision-Language Design

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T05:06:46.616733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:06:46.616733Z digest=sha256:1e6f135b3bec5cada86428645eaf65b70189b81225b94bfb1f071c1e303e03e7

Observation 81c2255b-4e49-474f-8c34-793f6145de02 · inbound

Valley3: Scaling Omni Foundation Models for E-commerce cites this paper.

Valley3: Scaling Omni Foundation Models for E-commerce Valley2: Exploring Multimodal Models with Scalable Vision-Language Design

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:51:06.067123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-09T14:53:55.160230Z digest=sha256:ce075c74b2bd7bbf549da5b955b267601edb0269e54fce878cb6ca5f5548eff6

Observation 8097c691-8871-4976-a9d6-5017b52998cc · inbound

Learning to Deny: Action Denial in Multimodal Large Language Models cites this paper.

Learning to Deny: Action Denial in Multimodal Large Language Models Valley2: Exploring Multimodal Models with Scalable Vision-Language Design

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:45:39.602043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-01T06:21:09.996386Z digest=sha256:8b148526e70407f4a3c3f3d5f71a69054b81065ae48d54aa62052fa8c7736e1b

Observation de51de91-1470-4da0-9212-b98ff718d6b0 · inbound

VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models cites this paper.

VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models Valley2: Exploring Multimodal Models with Scalable Vision-Language Design

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:49.025460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:35:49.025460Z digest=sha256:1cb5bc671ae03b6e784f0738ce83b15e4e619b143a94f88f2aaad1c5bfd0da15