Pith. sign in

Paper Citation Record · LEDGER

POINTS1.5: Building a Vision-Language Model towards Real World Applications

As of 14 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 6 inbound Pith citation observations for arXiv:2412.08443.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.08443 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:53:36.381720Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:10:19.033317Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T10:41:04.336073Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 98f326bd-d824-46b8-8c6f-03f5a4358a38 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.576467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.576467Z digest=sha256:41cd1e27b3eec38fce6f7bd2c205cb863028a7b3cde0afd1f20e3fc4ac306482

Observation fffee9bc-21d1-4932-8b30-6199ea5151be · outbound

This paper cites InternLM2 Technical Report.

POINTS1.5: Building a Vision-Language Model towards Real World Applications InternLM2 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.582276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.582276Z digest=sha256:b646595896252140faf3dd3fa1b815feb8dee9a4f99883e63069d9c481ad8319

Observation 4bbf6cd1-7f28-4e8e-8fd2-d03e9d4c631d · outbound

This paper cites MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks.

POINTS1.5: Building a Vision-Language Model towards Real World Applications MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.587371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.587371Z digest=sha256:94fbd85e6a5caf68ce5d8a993bf02289cd064197fc729f438af7bc7c784fd1c2

Observation 8a7755b8-4a17-44f0-ba00-f66d6d19457e · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.591330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.591330Z digest=sha256:6d6b7375c9de17479b41ba834d163614910eb018af643858c31590f72506fc7c

Observation 2b13e1ea-8a7d-4329-aff7-e811edf72d4b · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

POINTS1.5: Building a Vision-Language Model towards Real World Applications How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.595179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.595179Z digest=sha256:9f96acf07e423686304e155d68cc89ae6dd422aa2391b225f7c795b98553516e

Observation 56929c1f-128c-4875-8c03-7304f608f8fd · outbound

This paper cites Opencompass: A universal evaluation platform for foundation models.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Opencompass: A universal evaluation platform for foundation models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.598976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.598976Z digest=sha256:d87cf0e7f4aa89171765224f82d563f8b08a8719081664e39b6f42b1d3875d40

Observation d3b610b1-d779-4dda-a0fe-b16a35b5f244 · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

POINTS1.5: Building a Vision-Language Model towards Real World Applications NVLM: Open Frontier-Class Multimodal LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.602293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.602293Z digest=sha256:61faf9db99be0aa93db54836bc4bf4a7f84b1a08901fb897daf50002d57578b5

Observation 63705407-a9b3-45c3-871e-4abcdab7546a · outbound

This paper cites Flash A ttention-2: Faster attention with better parallelism and work partitioning.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Flash A ttention-2: Faster attention with better parallelism and work partitioning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.606187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.606187Z digest=sha256:c69e1719b4b2e798edc8da11718c2fb5657e0c41c3942315fcdbd498406ea5df

Observation fa7d6077-5c4c-40ae-aaf9-5149cf7741da · outbound

This paper cites Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.609367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.609367Z digest=sha256:5e0aa4f7b75c7d66bb8299e7a8c43bd72845b54e26ddb466b7ac2eaadedf6b10

Observation 6f7540e0-44f8-4006-9bdc-593e85a1777a · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

POINTS1.5: Building a Vision-Language Model towards Real World Applications InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.612707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.612707Z digest=sha256:791e729e8dc95a0fcf766482387926baf8dbd1a1e18ceda3a2d4377a2386f50c

Observation 5a6064f6-6ffd-448e-bde2-da149ef06fc6 · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

POINTS1.5: Building a Vision-Language Model towards Real World Applications InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.618990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.618990Z digest=sha256:71e8549bf8e0610f474a01aa91d4dedf43630b55c800caeafb45a7eae1580e69

Observation 8aa0ed58-c860-4f81-890e-b2152312c996 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

POINTS1.5: Building a Vision-Language Model towards Real World Applications An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.623331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.623331Z digest=sha256:261d93e0b65ff39af65a9a63bfa7214cf4b7254ded1b1baabb6342152d9404e2

Observation a8af7fbc-f283-488f-937e-0cb0f76f67e7 · outbound

This paper cites VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models.

POINTS1.5: Building a Vision-Language Model towards Real World Applications VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.627431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.627431Z digest=sha256:fea501330d1b97ada45a5ce632f625e1e69334ba6cdf00f2aebf9d79638140ff

Observation a49dc5aa-2bcc-4e25-a66d-81af93fc73fd · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.659004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.659004Z digest=sha256:89dd9ec63cf59d031cba5ee4e740e558e1f50bfa0699d53f1ab170c9563d422d

Observation 7b1f3a07-e8c8-4486-841f-f83e4bda3359 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

POINTS1.5: Building a Vision-Language Model towards Real World Applications Gaussian Error Linear Units (GELUs)

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.707889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.707889Z digest=sha256:04c4035351098d37242590194fdc4fdd2e17975c1556f47910325a7e20b46ca6

Observation e42ad58f-133b-4397-9030-1a3c2d1ad19c · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.756491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.756491Z digest=sha256:05760c8d3b3e0efd843d0f9b9bddac732c9569322ce20d7aa8517c8ffa289a54

Observation 1dcf57e1-9e31-480f-8549-2dd20fc0b965 · outbound

This paper cites A Diagram Is Worth A Dozen Images.

POINTS1.5: Building a Vision-Language Model towards Real World Applications A Diagram Is Worth A Dozen Images

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.785439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.785439Z digest=sha256:d1302a35f710fd2c6a582243e068ef0fea802924a343c5c2b3c0c44b64c60e31

Observation b7faf391-27ab-4a72-ae2d-11ac41cfa454 · outbound

This paper cites Visual information extraction in the wild: practical dataset and end-to-end solution.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Visual information extraction in the wild: practical dataset and end-to-end solution

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:53:37.041554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T17:53:35.824583Z digest=sha256:5b5fc3e26a0d876157012c56c16afa230c35879ff9b69ba35b6e3462837dd911

Observation 9e8b26e7-b403-4696-a5c3-cb5433f133bf · outbound

This paper cites What matters when building vision-language models?.

POINTS1.5: Building a Vision-Language Model towards Real World Applications What matters when building vision-language models?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.917989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.917989Z digest=sha256:c6effce374413e967427216941dc6ba8f2f3093140d6962f6bb48e51cfebc8c0

Observation 78734d24-d5fc-4569-a79a-fd35d6b0bb42 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

POINTS1.5: Building a Vision-Language Model towards Real World Applications LLaVA-OneVision: Easy Visual Task Transfer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.932406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.932406Z digest=sha256:43221abefd61aa35abf8443e17ef064fe949a34849ade3388e5fb7eebd7f4612

Observation 36e10ffa-6e9b-4907-935e-5c3a439df210 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

POINTS1.5: Building a Vision-Language Model towards Real World Applications SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.936447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.936447Z digest=sha256:b4ce4fd2cff400860edd6767807b1fdfe4463cedeedbd8e52fb8a08f4c129300

Observation 0b6dc376-ad75-4176-82ef-fe4be5652c06 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.940354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.940354Z digest=sha256:581b34696049497f40842acb38da87389cee28424a9c4e487480ea450ccfc024

Observation 66370c34-82cf-4eb5-820f-232b96f1f9b0 · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

POINTS1.5: Building a Vision-Language Model towards Real World Applications HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.944722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.944722Z digest=sha256:320f0aa02b9882210f059665fd24aec82958a5bb50e717cf0d67fb12ddc72331

Observation 071b681f-98c4-493c-8cdc-a50cb5a3afe0 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023 b.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Improved baselines with visual instruction tuning, 2023 b

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.948644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.948644Z digest=sha256:710d6fd6f08ecda1cc422076813461b5c6032333dfc22791c343e475f9379d1f

Observation 77f76f56-d232-4a00-9c98-02436a04d336 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024 a.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Llava-next: Improved reasoning, ocr, and world knowledge, January 2024 a

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.952889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.952889Z digest=sha256:ad289507fbd1864290d95017a198467aaf54d76578ef7c4ff87192f86118ccde

Observation 972eac2e-bf3e-49cd-834e-783b58a4a2b6 · outbound

This paper cites Visual instruction tuning.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Visual instruction tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.957433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.957433Z digest=sha256:a93dbdc0fd98543a2b45dec6df808f428ebeff1ae14bdc9bd0a41603d867a863

Observation f89b701e-93f5-4293-bae0-c0cfbdd33724 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

POINTS1.5: Building a Vision-Language Model towards Real World Applications MMBench: Is Your Multi-modal Model an All-around Player?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.962357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.962357Z digest=sha256:71b0832d649a0747b808e25babf8ca67dc15ca300b22c3327e2085a89db35341

Observation a384600c-cc4b-499b-84dc-3b5cf5d34d46 · outbound

This paper cites Rethinking Overlooked Aspects in Vision-Language Models.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Rethinking Overlooked Aspects in Vision-Language Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-11T17:53:36.597445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T17:53:35.966122Z digest=sha256:bc1323ddd72d5812526458108cce5b3af42246d0959cc02ac5b7bb27875eb37f

Observation 7951d8c8-f64f-4ccf-8f2b-5775eec7cce1 · outbound

This paper cites POINTS: Improving Your Vision-language Model with Affordable Strategies.

POINTS1.5: Building a Vision-Language Model towards Real World Applications POINTS: Improving Your Vision-language Model with Affordable Strategies

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.969925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.969925Z digest=sha256:c5d242a8f94d5058e3d272d0457b37be3ba76a972b7786cdde165acb9b6618ea

Observation 50e3ba2d-8416-4470-b73b-fd21746959c4 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

POINTS1.5: Building a Vision-Language Model towards Real World Applications OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.973695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.973695Z digest=sha256:d9ac7ce87230498671ff92b909bb2fa25e0832e7381ca75c458b5033cf1ed2d3

Observation 34f1136b-ad4d-4994-853d-951ba864696f · outbound

This paper cites A convnet for the 2020s.

POINTS1.5: Building a Vision-Language Model towards Real World Applications A convnet for the 2020s

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.977503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.977503Z digest=sha256:3d336cda199db24b7f911e5e4bf5b72720725ac45e4f866bf8aef7a1dfe01473

Observation 89baec70-f09e-44b5-93b8-14e1d436eaf7 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

POINTS1.5: Building a Vision-Language Model towards Real World Applications DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.981062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.981062Z digest=sha256:33905d4f8df2bb8816056daa8300c87c09eac192fe5ee2563a19cec013c17417

Observation b327698e-fac0-4c8c-9161-75f40f9c3ca5 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.984662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.984662Z digest=sha256:5bee491f4256794b52da80bac04967096ade1d5120337acc1a1699c3527be087

Observation 2cd4dffa-a152-49c7-bf65-25aef7937ab0 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

POINTS1.5: Building a Vision-Language Model towards Real World Applications MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.988316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.988316Z digest=sha256:5e9497ea035dd08edabcd40befdd36e21a9119d227584942ea43cd4d46fa33f8

Observation e226fb37-2f60-4d15-a93c-75819721536c · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.992520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.992520Z digest=sha256:7efb9bbd0ef4b736622f3592b672176a3e9dbb40fe7d9b4fc5147aebdf5813df

Observation cbfbed3c-8f3a-4141-b704-b929aba53ed1 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:53:36.970221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T17:53:36.008926Z digest=sha256:5f0492c7a90a2dc8d44a9bd1fff56e6e66356a0f2a99c59b2852fb27d6a0a41a

Observation 5ad0b438-50a9-4792-af7c-3f1d36dec613 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

POINTS1.5: Building a Vision-Language Model towards Real World Applications ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:36.092779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:36.092779Z digest=sha256:49e7115481b1de943f3a38703cb6395f2ba988ff05b13944460b8e126bb55aa9

Observation 8e8460d5-5ad4-421a-ac79-21eb1a002e51 · outbound

This paper cites Gpt-4 technical report.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Gpt-4 technical report

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:53:36.916367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T17:53:36.140019Z digest=sha256:7a2dd28e339d41215058143ce5a08f773975d46493b231de31bdd911fb39926c

Observation fe849fcd-952f-4170-ab54-5f5683f44c81 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:36.181095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:36.181095Z digest=sha256:2898922c31bdf8de96b54f75068a0d31d86c5cc1cd18b5ffef5b7157c7f35dbe

Observation f64f42bd-cfde-4509-8d8b-190a3f9fcf55 · outbound

This paper cites Neural Machine Translation of Rare Words with Subword Units.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Neural Machine Translation of Rare Words with Subword Units

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:36.228107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:36.228107Z digest=sha256:abb374f883fdec28ebd7f05e715e03aace91f8c6346882e24257e9fd4255aeb6

Observation 31dbbec2-e0c7-409c-964f-eb7c2e654e8b · outbound

This paper cites Fast WordPiece Tokenization.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Fast WordPiece Tokenization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:36.267935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:36.267935Z digest=sha256:243de3142ce853c42f84bfa39c1f352095062125a9b142a9e2bdd694b3bec648

Observation 6f9a65c3-c770-4e2f-81ef-a00b446ac23d · outbound

This paper cites Attention is all you need.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Attention is all you need

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:36.313185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:36.313185Z digest=sha256:90b78440015928a42ce4a72cfc66a0e188fcc0633be3d663a410ccc2bbd86876

Observation c14f66df-556f-4525-b716-5a3e9cc4b655 · outbound

This paper cites To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning.

POINTS1.5: Building a Vision-Language Model towards Real World Applications To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:36.335982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:36.335982Z digest=sha256:91e03919282919a8c673fd2ed3de8599a538477d2b6d9743745aebe64479a85b

Observation 6318f71a-845c-4003-a7c7-672af8b109eb · outbound

This paper cites Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:36.339878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:36.339878Z digest=sha256:b2ae30cd871f50c56d4480ea27fe8c84cdb3d7d13a2d4326a60e8712f5829089

Observation 5b8fc23f-4cbb-4437-b08e-bff5cf8c03ac · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:36.343951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:36.343951Z digest=sha256:d33f20efc87f844e36901eda73e86ed5a821111c90964d44f700d50da1920d21

Observation 50bf95f2-032e-4064-a534-321dc9d7cfd6 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Emu3: Next-Token Prediction is All You Need

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:36.347655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:36.347655Z digest=sha256:c801d90cdf9a28e85ed6b717862c3274e2d54708422085eade4f4fbaecbd8bae

Observation f8b086a8-c3ea-4660-859b-9ba04b5f2c90 · outbound

This paper cites Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:36.351685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:36.351685Z digest=sha256:8e7712bddaf29f288d1dcb983aa88e6a34e3073dcd8d8447a752e9d0945a2736

Observation 5ad075e0-148a-4e0e-8a89-62ddac381dfa · outbound

This paper cites Qwen2 Technical Report.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Qwen2 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:36.355481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:36.355481Z digest=sha256:339402e4a81f168c0a358aea850e8284cb140c77a84dcb0deff921166d98adab

Observation d87c14d9-e5a4-4bf8-8b4d-0a8404ebbba0 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

POINTS1.5: Building a Vision-Language Model towards Real World Applications MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:36.359251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:36.359251Z digest=sha256:ae923044e5f4f7fd7ef94b32d2d6294a89e37be9d910e46405de03369a48636b

Observation af0f15c6-49b2-4c09-9a44-03c60e20f07b · outbound

This paper cites A Survey on Multimodal Large Language Models.

POINTS1.5: Building a Vision-Language Model towards Real World Applications A Survey on Multimodal Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:36.363033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:36.363033Z digest=sha256:5e1ce2c256aa85dbfbecdd53dd85a4c88de17d623ac5c344fb7a57a8d5fc0b64

Observation 3ea851cd-f6f8-4c30-92d2-e3df83e2fb48 · outbound

This paper cites Capsfusion: Rethinking image-text data at scale.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Capsfusion: Rethinking image-text data at scale

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:53:36.849572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T17:53:36.367287Z digest=sha256:2c013260803d7bcd6ba13ed2eed5ff28fa1bd47e1008f2a4ab8c9e74cb3eac7f

Observation 97b5a4a0-ffdf-47c8-8113-9a7a11592997 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

POINTS1.5: Building a Vision-Language Model towards Real World Applications MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:36.370595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:36.370595Z digest=sha256:0d0139e761ac2500161135ee63246dcc9cbb03ad720ecca0311d8c3682d3d1e0

Observation 48a51f39-2a3d-4042-8e4f-a36e24b4bfc3 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:36.374453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:36.374453Z digest=sha256:428e40261a0eb51dcfe1d9a2be8ccdd4aa4398279872b6659dbbbffe88389152

Observation 1c05cdad-49c4-4650-bc8f-06c059e545cc · outbound

This paper cites MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning.

POINTS1.5: Building a Vision-Language Model towards Real World Applications MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:36.377753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:36.377753Z digest=sha256:8e8f0e180b85e4f9dbda007f2ec400fa77067005b51f9a46ce96d50486dca246

Observation ff9fabc9-1e7e-4575-8c0a-7e6aa616fac1 · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169--186.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169--186

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:53:36.789081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T17:53:36.381720Z digest=sha256:db1d4d855408ce14ffe49df8a4e303314006390c2c374896d65e605557d4fe8f

Pith citing papers

Observation d6fcd753-8a4f-4a5f-ad9b-50ec0fa983de · inbound

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design cites this paper.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design POINTS1.5: Building a Vision-Language Model towards Real World Applications

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.033317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.033317Z digest=sha256:6d5dfe9164de66e7fbcb972954bb54d68a478267e843ad25127e2409aa0c0f9f

Observation 8995f9b6-2354-4758-aae2-58978eb21bde · inbound

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model cites this paper.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model POINTS1.5: Building a Vision-Language Model towards Real World Applications

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.522131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.522131Z digest=sha256:cbaf2d3d121eadb91d739a43b5660014eea69c993037483eedb5473469da7249

Observation 2ebea4e1-dbe9-44cd-aa48-b66c76f329cc · inbound

Ocean-OCR: Towards General OCR Application via a Vision-Language Model cites this paper.

Ocean-OCR: Towards General OCR Application via a Vision-Language Model POINTS1.5: Building a Vision-Language Model towards Real World Applications

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T14:14:54.729437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:14:54.729437Z digest=sha256:1f81c01aa4290d060cdc485778a302a2aaea6413eedc5ae491e7f82ef6334c12

Observation f1f44d6b-23d1-46ab-b8a5-1dc68c20ceb4 · inbound

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs cites this paper.

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs POINTS1.5: Building a Vision-Language Model towards Real World Applications

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:41:04.339239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T15:23:08.671342Z digest=sha256:c3c96aec3abda967fdc4a56dd56c0ef4298db1c2fcbbf046e567c890fde5e215

Observation 978fed47-5263-4088-8de9-d7612181ca27 · inbound

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management cites this paper.

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management POINTS1.5: Building a Vision-Language Model towards Real World Applications

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:10:26.364451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T13:09:24.304696Z digest=sha256:352083258372a9a66f0fffa0c59d123715f6df8910ad052ce8bf06d1cf55927a

Observation 8bf490b1-1558-4081-a933-75d550175b34 · inbound

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management cites this paper.

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management POINTS1.5: Building a Vision-Language Model towards Real World Applications

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T16:18:22.058395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:18:22.058395Z digest=sha256:26450a3d23e85b188aaa398427a3e91188921a8ce6162dd8dedce89bc1a0001c