Pith. sign in

Paper Citation Record · LEDGER

CoMemo: LVLMs Need Image Context with Image Memory

As of 7 August 2026, this Paper Citation Record lists 100 of 114 outbound references and 2 inbound Pith citation observations for arXiv:2506.06279.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06279 v1

Coverage vector

measured 100 of 114 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:02:30.468386Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-11T01:49:15.136031Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T16:01:23.018232Z

Reference resolution

100 of 114 outbound references displayed

  • verified exact0
  • verified fuzzy26
  • unresolved74
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 79d773ee-f286-40d3-bd8a-4f084176938d · outbound

This paper cites write newline.

CoMemo: LVLMs Need Image Context with Image Memory write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.127567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.127567Z digest=sha256:310a79523515c3bbec3547219334bb471086c15001400a9ef4b068ea596683ee

Observation 1b37105c-5bfe-48d5-a435-308641a2b8b8 · outbound

This paper cites Nocaps: Novel object captioning at scale.

CoMemo: LVLMs Need Image Context with Image Memory Nocaps: Novel object captioning at scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.132542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.132542Z digest=sha256:f7ebd2e82a0f02192328792c367e82991981c226c7281ad99e8240f75bb70e5b

Observation 7646e2aa-2a92-4878-a2de-6d3de73342b7 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

CoMemo: LVLMs Need Image Context with Image Memory Flamingo: a visual language model for few-shot learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.136019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.136019Z digest=sha256:82fb577a93e36de01751aea084f672c788d83fbef37151f8351c2cd8bd119b1e

Observation f750bb2f-0dcb-4ec5-bf3c-732c7d48c15b · outbound

This paper cites MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms.

CoMemo: LVLMs Need Image Context with Image Memory MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.139466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.139466Z digest=sha256:b8636b195c0d3d74f447f0e56e3527585ed78c4670e47505d2f44634df9fca34

Observation def89704-aae5-45d8-a5ac-d1f50160c601 · outbound

This paper cites A., Datla, V.

CoMemo: LVLMs Need Image Context with Image Memory A., Datla, V

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.143129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.143129Z digest=sha256:b3c44aff3ecd9bf1fa228a3dd8c2444681106fc0960343b1734154c50fa267af

Observation 2e30b908-a5a0-41d9-9b70-8ab804a9e781 · outbound

This paper cites F., Tito, R., Mafla, A., Gomez, L., Rusinol, M., Valveny, E., Jawahar, C., and Karatzas, D.

CoMemo: LVLMs Need Image Context with Image Memory F., Tito, R., Mafla, A., Gomez, L., Rusinol, M., Valveny, E., Jawahar, C., and Karatzas, D

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.146609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.146609Z digest=sha256:4aee27e60127114db8246abd50b8f0cc7eea028a1cfb0a04a4c06488c50f649e

Observation b8d49dac-16a7-4aa6-825c-c39d28fffe1e · outbound

This paper cites Coyo-700m: Image-text pair dataset.

CoMemo: LVLMs Need Image Context with Image Memory Coyo-700m: Image-text pair dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.150059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.150059Z digest=sha256:bb7783db348ec64d47d1c829db16291d6f0784dafb3bfa2d72a6a1a8f4224766

Observation 8a9ee46c-ef58-4c1f-ae7b-45ef3d6f2469 · outbound

This paper cites and Xiao, J.

CoMemo: LVLMs Need Image Context with Image Memory and Xiao, J

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.154200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.154200Z digest=sha256:097755429bc30ea7bdcc70c295cae040da4153773bf041ca128f5fc37b2b93fe

Observation 05435e20-be77-4581-9970-3db59625d8a3 · outbound

This paper cites Textocr-gpt4v.

CoMemo: LVLMs Need Image Context with Image Memory Textocr-gpt4v

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.157676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.157676Z digest=sha256:962f8dce98936463e4d22df6461fe4aee73e16ac70e56f290ba92f1b7b36347c

Observation ea0cc5cf-0707-45eb-84f8-9d044095792f · outbound

This paper cites MapQA: A Dataset for Question Answering on Choropleth Maps.

CoMemo: LVLMs Need Image Context with Image Memory MapQA: A Dataset for Question Answering on Choropleth Maps

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.160890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.160890Z digest=sha256:7faff7f25bb4af58f1993de65b18a2733090b6a2141cbed2099e8c774ee0f3c3

Observation e07abd60-3b54-47fa-9573-77bf4aa61664 · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

CoMemo: LVLMs Need Image Context with Image Memory ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.164986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.164986Z digest=sha256:29a6e66d0590739dcf97231d4cfea538a9e6ab13ee408c991f0c37ec02d7f8b3

Observation 59f67add-d782-4305-b878-8925ec32c11e · outbound

This paper cites UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression.

CoMemo: LVLMs Need Image Context with Image Memory UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.168722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.168722Z digest=sha256:c9fd9ccede504f8072ff6c820bee34c5e07279db2279c34403f59bd619a02b83

Observation ae2a3243-366e-4ca9-bc2f-638aae7bf3b1 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

CoMemo: LVLMs Need Image Context with Image Memory Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.172306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.172306Z digest=sha256:0831774ba8f8a778f5ca409a59bff827460344172e6954699be0579595091818

Observation aa14a1d1-29be-4aeb-89cf-76468b19416b · outbound

This paper cites EVLM: An Efficient Vision-Language Model for Visual Understanding.

CoMemo: LVLMs Need Image Context with Image Memory EVLM: An Efficient Vision-Language Model for Visual Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.175851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.175851Z digest=sha256:527fe1788ee4f58eee832ec8cd7d32fae96095aded747c03484752d475f02c96

Observation bbc65fc0-80bd-44e0-a9d6-e376db1066f4 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

CoMemo: LVLMs Need Image Context with Image Memory Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.179318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.179318Z digest=sha256:fe7922d4f25c2064a343f69d308a8776bfe70d37a593cbba7ca7a41bc063692d

Observation de2765e1-0281-4a57-8c39-37fe9e9ebe9e · outbound

This paper cites Complicated Table Structure Recognition.

CoMemo: LVLMs Need Image Context with Image Memory Complicated Table Structure Recognition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.182794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.182794Z digest=sha256:dd8595388c6f2cfcddabb8f73a37b0d34a70fd7662eac278f48080d3eb1e5e82

Observation 8b42769e-724d-4fc0-9d0c-1b09306a9c0f · outbound

This paper cites K., Liu, Y., Sun, Y., Ng, C.

CoMemo: LVLMs Need Image Context with Image Memory K., Liu, Y., Sun, Y., Ng, C

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.187512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.187512Z digest=sha256:4f1b7414c1cc714c97211985f41f1f71b80168d7c2a8401d81adbbebcb119d69

Observation cd94c8a4-0aba-436e-a7a4-f13cea61dff6 · outbound

This paper cites Simple and Effective Multi-Paragraph Reading Comprehension.

CoMemo: LVLMs Need Image Context with Image Memory Simple and Effective Multi-Paragraph Reading Comprehension

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.190767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.190767Z digest=sha256:7086aade2552f3bc5ce189bbcee0d3b192708e8721cc06973b49818c1e373010

Observation 33df9bd5-09bf-4ca6-8a80-9c1818466996 · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

CoMemo: LVLMs Need Image Context with Image Memory NVLM: Open Frontier-Class Multimodal LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.194240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.194240Z digest=sha256:2fbc4698f903225006d7fdea0660ee59e036669ff60dbb89e57aac1bb5b682a1

Observation ff4a96ac-f9d2-4b61-bde0-f07a2c0002e3 · outbound

This paper cites Deep visual template-free form parsing.

CoMemo: LVLMs Need Image Context with Image Memory Deep visual template-free form parsing

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.197809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.197809Z digest=sha256:c35471ed33be07ec4809b7405fd30515b408a88d71552cbe748a02c0705d30a7

Observation 1af4d204-7229-42ac-9f8c-b5f8dca7c0f4 · outbound

This paper cites The Llama 3 Herd of Models.

CoMemo: LVLMs Need Image Context with Image Memory The Llama 3 Herd of Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.201168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.201168Z digest=sha256:17311d7945d754581cdd6c91db2a428599dd10995eb6bcff9045093d69d5f6de

Observation 6ae6ab93-5751-4fb2-98a7-9ddc968e5796 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

CoMemo: LVLMs Need Image Context with Image Memory MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.204597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.204597Z digest=sha256:91d946b8486522ba3cd7849ea49fc68e8c227ef26b8ff268f8f7f523a6eb0c75

Observation 763727de-789a-4026-a5f0-6c930777ce2c · outbound

This paper cites A., Ma, W.-C., and Krishna, R.

CoMemo: LVLMs Need Image Context with Image Memory A., Ma, W.-C., and Krishna, R

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.208209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.208209Z digest=sha256:14ca09b92fb8396d86dbf86fddb495ee064859e49067509f5e570c16c959344e

Observation 4544efaf-0813-4a67-9ba4-cbd5ad03d612 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

CoMemo: LVLMs Need Image Context with Image Memory Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.211435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.211435Z digest=sha256:c1334ef1d10c7ac5036b3ac895b9c537966102507583fec11375318958153c17

Observation 12fc4599-f778-4b46-a339-87cc2d6b3369 · outbound

This paper cites Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark.

CoMemo: LVLMs Need Image Context with Image Memory Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.214909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.214909Z digest=sha256:6b51c1cfe7790f30611517e8f723360e41ae2be5d81a2eaa3ba0c66290652c86

Observation 52e019b4-ae1e-4d76-a27d-ed70e62f19f6 · outbound

This paper cites Eaten: Entity-aware attention for single shot visual text extraction.

CoMemo: LVLMs Need Image Context with Image Memory Eaten: Entity-aware attention for single shot visual text extraction

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.218140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.218140Z digest=sha256:7378fa57a6abac9704cdb1b77f541134b5eb7cb3153ba9b5c0b65bfb1e15aa82

Observation 2430febf-ee02-49cd-b57f-6ac84cc8b47f · outbound

This paper cites Icpr2018 contest on robust reading for multi-type web images.

CoMemo: LVLMs Need Image Context with Image Memory Icpr2018 contest on robust reading for multi-type web images

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.221368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.221368Z digest=sha256:ce04bc7e450a695c48518f17d1cba78caa0a00a1858ace0f1a9a5195bf04a95c

Observation 209f2471-82e4-450e-8fe7-45bb600d254e · outbound

This paper cites PathVQA: 30000+ Questions for Medical Visual Question Answering.

CoMemo: LVLMs Need Image Context with Image Memory PathVQA: 30000+ Questions for Medical Visual Question Answering

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.224674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.224674Z digest=sha256:cc1c9301316874ccb69258f038dd6e919bb5245f0b6e6f9e0f0cb8d9cbd31b82

Observation 669db1a4-b54e-49e4-8bfe-4d8b997731fb · outbound

This paper cites an unresolved cited work.

CoMemo: LVLMs Need Image Context with Image Memory Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.228260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.228260Z digest=sha256:f17f8d1b285bb6f951a7c4483b86f1ec326b41cccb930efdf6b0c221237ca4f6

Observation 6ca98ae8-e705-436f-9692-368ec67430e2 · outbound

This paper cites Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment.

CoMemo: LVLMs Need Image Context with Image Memory Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.231512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.231512Z digest=sha256:03ecc86eedfaae2911f916af60930e487215bea18e7152bb9f5f8658bf1fa16b

Observation 13ce348c-6f4e-4ff0-a75c-36eed2ec4bee · outbound

This paper cites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding.

CoMemo: LVLMs Need Image Context with Image Memory mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.234729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.234729Z digest=sha256:393ec831980c8d035150ab66a35a1d317ee0f0ef3dbf2c18c0e14735e5151a2f

Observation 8c75405c-2f20-4413-a079-1700b6a00b72 · outbound

This paper cites Medical-diff-vqa: a large-scale medical dataset for difference visual question answering on chest x-ray images, 2023.

CoMemo: LVLMs Need Image Context with Image Memory Medical-diff-vqa: a large-scale medical dataset for difference visual question answering on chest x-ray images, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.238109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.238109Z digest=sha256:16e229ffea58e1c553cdd31333a0f42d30691f7dfcf54282e394b1cc11212ee2

Observation beed86e7-8ef0-4c38-b874-721f76f683c1 · outbound

This paper cites Movienet: A holistic dataset for movie understanding.

CoMemo: LVLMs Need Image Context with Image Memory Movienet: A holistic dataset for movie understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.241444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.241444Z digest=sha256:8194d6ec12c08b26610e9563f41edeaeeb66f03e6358193653f5a4c80c6b7c64

Observation a3283491-6cde-4272-a8e9-6a8b47581690 · outbound

This paper cites Icdar2019 competition on scanned receipt ocr and information extraction.

CoMemo: LVLMs Need Image Context with Image Memory Icdar2019 competition on scanned receipt ocr and information extraction

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.244739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.244739Z digest=sha256:4c81c8d5abab6ed9ec3360ce63c0610450457c168465f2305356fa2ab9f54105

Observation 4dfa8682-7aa0-4224-8ea3-1847b10f9a60 · outbound

This paper cites an unresolved cited work.

CoMemo: LVLMs Need Image Context with Image Memory Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.247942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.247942Z digest=sha256:4f7847bd3536cca3543dc29e3cfc884493857ff6f444a4b38d220085c51b093e

Observation 174f77cd-8052-47e7-9770-9eba5b8372c6 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

CoMemo: LVLMs Need Image Context with Image Memory MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.251165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.251165Z digest=sha256:dee813f7522a408d31041ef6834ac43e5c84babfda04cbd9af7d0b4fdcb058b4

Observation 80d25d45-dc03-4fc4-a455-9dc0d7bf3c4c · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning.

CoMemo: LVLMs Need Image Context with Image Memory Clevr: A diagnostic dataset for compositional language and elementary visual reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.254523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.254523Z digest=sha256:e1bafade5aead716f9298979b2c56d5184908a3e7fa611342a1060fa21600cdb

Observation b5cdb447-51fa-47f8-a580-c592747f2c7e · outbound

This paper cites Dvqa: Understanding data visualizations via question answering.

CoMemo: LVLMs Need Image Context with Image Memory Dvqa: Understanding data visualizations via question answering

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.258111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.258111Z digest=sha256:b0d7d67c81d1de2ac313ef65a00f22769861d7e6abc59cb0c49ec22501c95464

Observation d8be9ef5-2188-4f56-a41a-7957667c4179 · outbound

This paper cites FigureQA: An Annotated Figure Dataset for Visual Reasoning.

CoMemo: LVLMs Need Image Context with Image Memory FigureQA: An Annotated Figure Dataset for Visual Reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.261341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.261341Z digest=sha256:71c593ec5b3e40b4b7792acdbe7d75748e34a3426bb99ee5d2f7f09491b7a5a7

Observation c1474023-82b6-4b32-be90-265dbf797445 · outbound

This paper cites Chart-to-Text: A Large-Scale Benchmark for Chart Summarization.

CoMemo: LVLMs Need Image Context with Image Memory Chart-to-Text: A Large-Scale Benchmark for Chart Summarization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.264586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.264586Z digest=sha256:81f6759c15540505d178c8fa0c9117102b4f49ce6597b283c26fba8a42ed441c

Observation 8277c4b0-01fe-4f01-b02a-6ea0cf332505 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

CoMemo: LVLMs Need Image Context with Image Memory Referitgame: Referring to objects in photographs of natural scenes

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.268275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.268275Z digest=sha256:a4583ea1263f4cf22b4a7e590d6a09047831cf181c5d0386d8f8f31cac90e2ec

Observation 8f63e000-09da-4d75-827b-ff1de325a7e9 · outbound

This paper cites A diagram is worth a dozen images.

CoMemo: LVLMs Need Image Context with Image Memory A diagram is worth a dozen images

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.271379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.271379Z digest=sha256:63d876ade3e63d3f5593aacf8f46a261be6aaec04bbe8f1251262904b81a719f

Observation f1cc699c-fcae-4550-87ec-cf8c16069995 · outbound

This paper cites Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension.

CoMemo: LVLMs Need Image Context with Image Memory Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.274553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.274553Z digest=sha256:e7f2b6cf07fc209e09dddcf29db9e576a32a59a983bd5e607cee888f4ee9128f

Observation 57ad395e-beeb-4c37-a87b-ed1205ca2522 · outbound

This paper cites Visual information extraction in the wild: practical dataset and end-to-end solution.

CoMemo: LVLMs Need Image Context with Image Memory Visual information extraction in the wild: practical dataset and end-to-end solution

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.278069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.278069Z digest=sha256:f6d2d935d18b0845eb0f4ba9d4bb36be49747529b32512c32ff83a81989f65db

Observation 810c6589-e670-4112-a806-0b314790c112 · outbound

This paper cites Laion-gpt4v dataset.

CoMemo: LVLMs Need Image Context with Image Memory Laion-gpt4v dataset

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.281255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.281255Z digest=sha256:f3a3aaf11465d0e9ad63262bf9fcd0d52fdfc0554b2e457791e2697473ab518e

Observation dfc59059-e7aa-4d95-a722-6d38b63e3d24 · outbound

This paper cites J., Gayen, S., Ben Abacha, A., and Demner-Fushman, D.

CoMemo: LVLMs Need Image Context with Image Memory J., Gayen, S., Ben Abacha, A., and Demner-Fushman, D

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.284375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.284375Z digest=sha256:d68fd4bf1863040699f7cf5742fd47799f62b3e1f3eaac33e119ef4ba210d867

Observation 9a24351c-1ad4-46be-9297-e4b072587cfb · outbound

This paper cites What matters when building vision-language models?.

CoMemo: LVLMs Need Image Context with Image Memory What matters when building vision-language models?

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.287592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.287592Z digest=sha256:bdffea31c7fea1579a830e8579735d345689e6e0707bc3ea6ef509cf1725e5e0

Observation cc7246f4-6e8f-4389-a348-e5da0dbb0458 · outbound

This paper cites G., and Lov \'o n Melgarejo, J.

CoMemo: LVLMs Need Image Context with Image Memory G., and Lov \'o n Melgarejo, J

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.290994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.290994Z digest=sha256:aed8dda4bcc6c27a833c773b03038b4702b0b9bae03d077533bb2adfd6bab7b6

Observation 36f591ec-7866-4929-8473-e3c6ff0d02ff · outbound

This paper cites Chemvlm: Exploring the power of multimodal large language models in chemistry area.

CoMemo: LVLMs Need Image Context with Image Memory Chemvlm: Exploring the power of multimodal large language models in chemistry area

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.294273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.294273Z digest=sha256:d401eb2332918ce870248963ebe16cb68c53b333e850b2dbfcda4c1c9ef6f28f

Observation 9322a4e8-564e-40c4-a543-064dc566e5e2 · outbound

This paper cites an unresolved cited work.

CoMemo: LVLMs Need Image Context with Image Memory Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.297412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.297412Z digest=sha256:6b509e6cfd0653802aecc649d5755a2880ef4c468472ce1b86fe8a9041aa8c22

Observation dd1e125a-4572-4e4e-b84a-48a5cf860a21 · outbound

This paper cites Vila: On pre-training for visual language models.

CoMemo: LVLMs Need Image Context with Image Memory Vila: On pre-training for visual language models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.300577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.300577Z digest=sha256:8593198de0540c86dddacaf58a9247e4efbccdf8349d83693e3795492c454137

Observation 0efdde03-aabb-4c5e-9981-273822210940 · outbound

This paper cites an unresolved cited work.

CoMemo: LVLMs Need Image Context with Image Memory Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.303815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.303815Z digest=sha256:49337e5e7c52366c7e427d19e3e894617d051b3f4921dd8705f3ba38411d89c3

Observation 01704a5f-c125-49bc-881a-eb17e9f8b86c · outbound

This paper cites an unresolved cited work.

CoMemo: LVLMs Need Image Context with Image Memory Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.307014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.307014Z digest=sha256:9cc2bbcfbe69e591e27ef773370f80e0fc51d3a4ec70d997aee0be9ff156905a

Observation 99dea663-00ac-46ea-babb-7b35d882486d · outbound

This paper cites Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering.

CoMemo: LVLMs Need Image Context with Image Memory Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.518169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.310104Z digest=sha256:3147555ac4e8090388f80f78d6ce446c65a8afc69c4c41953ad2e49aed4d95a6

Observation b189077f-4b5e-4840-ad8e-bea1bbe31d7d · outbound

This paper cites Casia online and offline chinese handwriting databases.

CoMemo: LVLMs Need Image Context with Image Memory Casia online and offline chinese handwriting databases

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.506657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.313225Z digest=sha256:306fa7a2afaa9b210327d285bfdd720bca5b9a5cc100490c1dcc09454ff81c8f

Observation 946e01d2-f093-45ca-a824-c0bf5ccc3cc5 · outbound

This paper cites Visual spatial reasoning.

CoMemo: LVLMs Need Image Context with Image Memory Visual spatial reasoning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.495266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.316435Z digest=sha256:3a22adac1169c03d99bed8a0e56ed0bd698655fce1cef6dc77a625590ccaef58

Observation b07f0ca7-744d-4623-b1e8-e0f98c805451 · outbound

This paper cites Mitigating hallucination in large multi-modal models via robust instruction tuning.

CoMemo: LVLMs Need Image Context with Image Memory Mitigating hallucination in large multi-modal models via robust instruction tuning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.484152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.319716Z digest=sha256:7b5b74a84b5865421480af0a293844862dd1bbd3087673bdfedd7381d1de1080

Observation 17c71ad7-9f4f-4159-9ce9-1a6764bd694f · outbound

This paper cites MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning.

CoMemo: LVLMs Need Image Context with Image Memory MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.322925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.322925Z digest=sha256:216423cf12c5deaa9b379d996a4883fca53b76bbaaa79c23936aa883d891aa17

Observation f00e6b2e-9063-4c1a-8cd1-2a7c6a92d1a5 · outbound

This paper cites an unresolved cited work.

CoMemo: LVLMs Need Image Context with Image Memory Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.326278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.326278Z digest=sha256:644845afcfe27f27bc90090030b385c74cb9e92d9295e955e67f7ca0552510d6

Observation 604930f5-0b88-4c59-95d0-1ee410d9de32 · outbound

This paper cites an unresolved cited work.

CoMemo: LVLMs Need Image Context with Image Memory Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.329367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.329367Z digest=sha256:9f7499140959b2f7677d30651ea31c7c72b4f3e50560423dc39f7a7831a76313

Observation 8546b904-8a57-41ab-a652-42fed20649b7 · outbound

This paper cites F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P.

CoMemo: LVLMs Need Image Context with Image Memory F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.332532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.332532Z digest=sha256:d14e00325b659e365cccae9ea416255f82761f85dd079c6ec96380f041a3846f

Observation 90241c51-b471-4207-a6b0-416a39b7b9e8 · outbound

This paper cites Paying more attention to image: A training-free method for alleviating hallucination in lvlms.

CoMemo: LVLMs Need Image Context with Image Memory Paying more attention to image: A training-free method for alleviating hallucination in lvlms

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.449278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.335865Z digest=sha256:a22a9003a03b4b509d7466cc9439dfefa4e6ea9b2691928e9ea2239a5eebcf03

Observation 733505ed-4a4d-490f-9430-7783d7a73374 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pp.\ 216--233.

CoMemo: LVLMs Need Image Context with Image Memory Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pp.\ 216--233

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.339582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.339582Z digest=sha256:b4b797ad3d8e3fb285a0fc4c2cc6d9c170922199f95c3cbfcf8e42be5590760d

Observation 5695cc77-06fc-4aca-a66e-e9a988521821 · outbound

This paper cites MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs.

CoMemo: LVLMs Need Image Context with Image Memory MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.346839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.346839Z digest=sha256:008a44cb6a989097e83fea0ed13fece3484e0b5a289540dae345a3d4abaa45e0

Observation e3b3ff6b-8025-4c8a-98d7-14cdf5e688b2 · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

CoMemo: LVLMs Need Image Context with Image Memory Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.350057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.350057Z digest=sha256:898d0828872a9cf51e86b8a4b14d9ff1807b5ed738197534ce27ac12a3b547de

Observation 6e92e147-407f-4e5b-8b8f-1e49ef1e8d61 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

CoMemo: LVLMs Need Image Context with Image Memory Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.429290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.353502Z digest=sha256:4ad7a50591afc6c07528605d6050757e62a158c5e2f1569adaff123a4f107cc8

Observation 4f30c604-bac2-41b0-9774-f0e2832b62cd · outbound

This paper cites Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning.

CoMemo: LVLMs Need Image Context with Image Memory Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.357233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.357233Z digest=sha256:ac4559e68dc9fc80391273f741e3bfec0f9888eee5672abef654053f3720f7be

Observation 156f6993-0a26-418b-bb6f-867c05b90e02 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

CoMemo: LVLMs Need Image Context with Image Memory MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.361051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.361051Z digest=sha256:415905b95ae97c8e38637d5ac3491b8f67259b6685c65e379b8ddb6859eb2869

Observation f661c154-e39b-42b0-980a-4f0457f3041a · outbound

This paper cites Deepart: Learning joint representations of visual arts.

CoMemo: LVLMs Need Image Context with Image Memory Deepart: Learning joint representations of visual arts

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.418177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.364523Z digest=sha256:24078ad1e9c2a91b24adf6a8418495b632cf7229583907682d06ae7b711493c4

Observation c77c4483-ffa4-4447-b75b-c3042ff0f6a1 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

CoMemo: LVLMs Need Image Context with Image Memory Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.406702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.367657Z digest=sha256:afb0a03231bcf5e9a5cf5f2a14a37bd114a513defe23286684bdc4ba9957c0d5

Observation dfa4350b-4f3c-46ca-9875-93bcd26166eb · outbound

This paper cites and Bunke, H.

CoMemo: LVLMs Need Image Context with Image Memory and Bunke, H

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.394019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.370785Z digest=sha256:b922c433676f5017287df4631e681364de84d439a41c0683d75b2f5311c648b1

Observation 18c6213e-1545-47cb-916a-ab4fe68274be · outbound

This paper cites L., Tan, J.

CoMemo: LVLMs Need Image Context with Image Memory L., Tan, J

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.382660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.374097Z digest=sha256:49052a6e4b6fc075de7fe2a47e2c7087eb17b30edd695e489b1ce45572295e1c

Observation 137e982f-9a0d-49a5-8a9e-134d0c8a0659 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

CoMemo: LVLMs Need Image Context with Image Memory ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.377251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.377251Z digest=sha256:5b23e6536c35f8c047d61c93539db834156ab1558cce8eca1706b525fa71143f

Observation be2f4648-53c9-4e91-8ab0-fc059295805c · outbound

This paper cites UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning.

CoMemo: LVLMs Need Image Context with Image Memory UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.380654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.380654Z digest=sha256:2dbb5324e4fb99efc8134706bece4cebfbcda34350300af25c75bdd44c96725e

Observation e6bcc7c4-897b-429b-ad41-109aa8beff90 · outbound

This paper cites Infographicvqa.

CoMemo: LVLMs Need Image Context with Image Memory Infographicvqa

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.371284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.384119Z digest=sha256:0da79ee40d1b6c5fc0aac5ec4acd33c8d9915d3e2494b5f48063800a0f38b822

Observation 9ff00e35-d786-456c-8814-5f0078660c84 · outbound

This paper cites M., and Kumar, P.

CoMemo: LVLMs Need Image Context with Image Memory M., and Kumar, P

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.360029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.387472Z digest=sha256:bb98f556c68ce0deddd772c0d8a5eaabd7e638f2c3e67f6ec1c9554b3a69a203

Observation 3ecfc09d-1ebb-4353-a622-1da18e7a13fe · outbound

This paper cites Opengvlab/internvl-chat-v1-2-sft-data, Jan 2024.

CoMemo: LVLMs Need Image Context with Image Memory Opengvlab/internvl-chat-v1-2-sft-data, Jan 2024

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.348665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.391046Z digest=sha256:336d56119f0668cfd76bbfcd23fae9029f85a7e5f193e37028c4fc1a8e8aed07

Observation 63479621-6e1a-4838-8694-93568d5fa6e7 · outbound

This paper cites Training language models to follow instructions with human feedback.

CoMemo: LVLMs Need Image Context with Image Memory Training language models to follow instructions with human feedback

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.394227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.394227Z digest=sha256:d7f3ada720c16a0660429c0d5615c2c45c89d096fa34bb43f9d7cb6d82bc5226

Observation 0c1e3acc-6270-4b13-aa75-e1b031b50306 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

CoMemo: LVLMs Need Image Context with Image Memory Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.397502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.397502Z digest=sha256:e00cf1a8d867591028f0d480bba883a16ee1831faae1532728055eac11f06351

Observation 01a410e3-8cf4-4fae-b973-5d36686e14b5 · outbound

This paper cites A., Wang, L., Cervantes, C.

CoMemo: LVLMs Need Image Context with Image Memory A., Wang, L., Cervantes, C

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.329398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.400986Z digest=sha256:e67fa93062032f090ff8bc52eb61a10a082b1897c92930a83fdf14a67ee548e4

Observation 1cc2ab6b-5754-4d27-8241-182a2418c82a · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

CoMemo: LVLMs Need Image Context with Image Memory Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.318852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.404242Z digest=sha256:222be8e063c9e7dd171f490f8db88555569f1a2146f506161ed22bed10f906ce

Observation bb4b8522-8214-418c-a2ff-1c533289bb58 · outbound

This paper cites Laion coco: 600m synthetic captions from laion2b-en.

CoMemo: LVLMs Need Image Context with Image Memory Laion coco: 600m synthetic captions from laion2b-en

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.307704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.407507Z digest=sha256:02e68b89b806f1d923dd188987ea9afab74915d695aac57dc27cf205cc2b7fc4

Observation 1b84a50c-ef90-4ba9-bdef-b3783f87e3ad · outbound

This paper cites Solving geometry problems: Combining text and diagram interpretation.

CoMemo: LVLMs Need Image Context with Image Memory Solving geometry problems: Combining text and diagram interpretation

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.296729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.410789Z digest=sha256:198d322a99e7b35f91091bf18713e30fe80fc4d7672cea462f87457e4ea753d1

Observation fbe9a7b8-59c9-41b3-99d9-a4dfb01ad904 · outbound

This paper cites an unresolved cited work.

CoMemo: LVLMs Need Image Context with Image Memory Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:02:31.285582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.414053Z digest=sha256:22f1df7b91ac51cf0cb68d522ff7c293303fe9fba92e0990e582efdf35bdb0aa

Observation 1a3ad93e-ed8e-4a09-9d05-0bf9b7515aec · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

CoMemo: LVLMs Need Image Context with Image Memory Objects365: A large-scale, high-quality dataset for object detection

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.274988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.417252Z digest=sha256:964f867360668384088712ffba8f2bcc660b6a03b4cd2efaa7c333fb27afce45

Observation 0a2e401d-22de-4408-b203-8ce576a0d7e5 · outbound

This paper cites Towards vqa models that can read.

CoMemo: LVLMs Need Image Context with Image Memory Towards vqa models that can read

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.264334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.420406Z digest=sha256:c018fd78564b5f55f937c63ef648360d6e3c8c2e2f45e6bce34a90989cf8e30e

Observation be159f74-34f0-4efd-9ca0-9bfb09e11a71 · outbound

This paper cites Towards vqa models that can read.

CoMemo: LVLMs Need Image Context with Image Memory Towards vqa models that can read

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.252411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.423527Z digest=sha256:b58669dd67703c9a8de169169b4dd07bb8413eed35bf0cbdcb87249d6bc4f1b2

Observation 3179f98c-0cf1-44ad-ab44-ee10b2d0dbce · outbound

This paper cites Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text.

CoMemo: LVLMs Need Image Context with Image Memory Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.241606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.426830Z digest=sha256:90b3da5f10392823dc4b68a77151c7b1f47ec6e17c491b6c4030023a5a66e702

Observation d554ce3f-0191-4126-aca7-caf3d993cddc · outbound

This paper cites MileBench: Benchmarking MLLMs in Long Context.

CoMemo: LVLMs Need Image Context with Image Memory MileBench: Benchmarking MLLMs in Long Context

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.430076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.430076Z digest=sha256:5530b5963144bef6904bd8909b0bb65ddbda6e4f0663b017f0464f9947209e14

Observation 15374b95-0c4e-410f-af9a-b0853b967627 · outbound

This paper cites Transformer roadmap: 2.

CoMemo: LVLMs Need Image Context with Image Memory Transformer roadmap: 2

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.229796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.433564Z digest=sha256:213b2d011dceb9acfb27acc1c74b37994ad2941859d03c60728aabe3d24496dd

Observation 9d98cf6c-e5e3-4c5c-980b-b5d1a4852d4f · outbound

This paper cites C., Han, J., Ding, E., Liu, J., Karatzas, D., et al.

CoMemo: LVLMs Need Image Context with Image Memory C., Han, J., Ding, E., Liu, J., Karatzas, D., et al

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.218537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.436705Z digest=sha256:eaf010b5c6088ad98af87938d9539865d9f8701de8cfea215b956be1d46bc24f

Observation 18ec7216-01ed-42d5-85ed-341a4c6cd833 · outbound

This paper cites Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy, 2024.

CoMemo: LVLMs Need Image Context with Image Memory Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy, 2024

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.207757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.439836Z digest=sha256:e081977a5b54bc2be4dad4b382a0c0cf00a162441c3dc370f1fae955a3de7c0d

Observation 295ed525-12db-4af6-bcb7-6dfc706f61ef · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

CoMemo: LVLMs Need Image Context with Image Memory Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.444417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.444417Z digest=sha256:b69fefa99016a2e42741a8007ee044807dd8239f6b32986c1fdcdc0f9c5003c0

Observation a1f5b464-6f14-461f-aedd-9d6cbb42dc60 · outbound

This paper cites COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images.

CoMemo: LVLMs Need Image Context with Image Memory COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.447765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.447765Z digest=sha256:b336ea267c47e6bd1f9a3fe15e2a56ab06f3f8099ac32ae3e3057ce7b7c5dcbb

Observation 9f93a81d-03ac-4199-8d57-73741ee9121c · outbound

This paper cites V3det: Vast vocabulary visual detection dataset.

CoMemo: LVLMs Need Image Context with Image Memory V3det: Vast vocabulary visual detection dataset

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.196589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.451246Z digest=sha256:99ccff8c93303f68ad9b9495838c70b4b5bdb4437d5b33ac60d963e077270bb7

Observation ae545fb5-9c31-4421-9b57-625e6aeba795 · outbound

This paper cites Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset.

CoMemo: LVLMs Need Image Context with Image Memory Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.454664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.454664Z digest=sha256:966198d466cb5bbef09ade2c3023158cfc26a8e191fad8b5c52133d028fc26c6

Observation 6c69544c-f78e-4907-bb82-9e20ca304583 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

CoMemo: LVLMs Need Image Context with Image Memory Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.458205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.458205Z digest=sha256:204198576d9d00026ddb711665905d20383ae7a66f3b6038195483718570e3da

Observation 74ab545b-117c-4047-824a-57ee643cddd5 · outbound

This paper cites The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World.

CoMemo: LVLMs Need Image Context with Image Memory The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.461439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.461439Z digest=sha256:82281658ad96a140db37b366ff2f2b5bff8cd24e54cdfb566a88dc8900348562

Observation 4f7dfd5e-5b7a-4f5b-90cb-ac18b352031d · outbound

This paper cites Needle In A Multimodal Haystack.

CoMemo: LVLMs Need Image Context with Image Memory Needle In A Multimodal Haystack

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.464875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.464875Z digest=sha256:2ce9deed26ddca623ff72a60711d4339f3c658af0d89477ccfec27b597032e0d

Observation 60177022-5d53-47f2-baf6-075177bef644 · outbound

This paper cites C., Luo, C., Jin, L., Chan, C.

CoMemo: LVLMs Need Image Context with Image Memory C., Luo, C., Jin, L., Chan, C

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.184921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.468386Z digest=sha256:ea0249430057d14bbebbd64ca73e7274d0a09d2dc8510a0374ada56e1aa841ef

Pith citing papers

Observation 141c552d-38db-4546-9c1b-cf0bfe83f149 · inbound

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs cites this paper.

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs CoMemo: LVLMs Need Image Context with Image Memory

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:01:23.021630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T18:53:06.494640Z digest=sha256:0328592ada865cc70c0f690b08192d1f5bdb3b2eb706084cc3a88ba8527b8551

Observation a6c78e64-9adc-4225-94c0-35ecef142c6c · inbound

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs cites this paper.

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs CoMemo: LVLMs Need Image Context with Image Memory

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:50:51.525657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:49:15.136031Z digest=sha256:14b62f6dcd0e66555999ca271ca7f4901a66bbeb09739a25ef5cb849d9b27ab7