Pith. sign in

Paper Citation Record · LEDGER

CoMemo: LVLMs Need Image Context with Image Memory

As of 17 August 2026, this Paper Citation Record lists 100 of 114 outbound references and 2 inbound Pith citation observations for arXiv:2506.06279.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06279 v1

Coverage vector

measured 100 of 114 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:02:30.468386Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-11T01:49:15.136031Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T16:01:23.018232Z

Reference resolution

100 of 114 outbound references displayed

  • verified exact0
  • verified fuzzy26
  • unresolved74
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 79d773ee-f286-40d3-bd8a-4f084176938d · outbound

This paper cites write newline.

CoMemo: LVLMs Need Image Context with Image Memory write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.127567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.127567Z digest=sha256:77c8210531be2d58a9155fd40a34a653e38dd91d7e7c0f458437cef3e8ff795b

Observation 1b37105c-5bfe-48d5-a435-308641a2b8b8 · outbound

This paper cites Nocaps: Novel object captioning at scale.

CoMemo: LVLMs Need Image Context with Image Memory Nocaps: Novel object captioning at scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.132542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.132542Z digest=sha256:c8a8817a0cd65bb4296a0a7c30d37c76d26efa3a027879f9bea2fa26075f4357

Observation 7646e2aa-2a92-4878-a2de-6d3de73342b7 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

CoMemo: LVLMs Need Image Context with Image Memory Flamingo: a visual language model for few-shot learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.136019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.136019Z digest=sha256:9e40c71e4b110fd01a5ca902e69962bcca1aefaa812f9833cc36e5a2e4451f40

Observation f750bb2f-0dcb-4ec5-bf3c-732c7d48c15b · outbound

This paper cites MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms.

CoMemo: LVLMs Need Image Context with Image Memory MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.139466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.139466Z digest=sha256:99f0c0472426bd214e672181f0882f2ad2ba21078f08b576182f77f8fa61b66b

Observation def89704-aae5-45d8-a5ac-d1f50160c601 · outbound

This paper cites A., Datla, V.

CoMemo: LVLMs Need Image Context with Image Memory A., Datla, V

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.143129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.143129Z digest=sha256:029dab92977d2f20fd35f596b291b7adafb0ae8b9e2a17e236a2049f80397a05

Observation 2e30b908-a5a0-41d9-9b70-8ab804a9e781 · outbound

This paper cites F., Tito, R., Mafla, A., Gomez, L., Rusinol, M., Valveny, E., Jawahar, C., and Karatzas, D.

CoMemo: LVLMs Need Image Context with Image Memory F., Tito, R., Mafla, A., Gomez, L., Rusinol, M., Valveny, E., Jawahar, C., and Karatzas, D

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.146609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.146609Z digest=sha256:68c82475426f64755a6e618fae3f023fa0006ed0fdef75a86dec888fc53967b4

Observation b8d49dac-16a7-4aa6-825c-c39d28fffe1e · outbound

This paper cites Coyo-700m: Image-text pair dataset.

CoMemo: LVLMs Need Image Context with Image Memory Coyo-700m: Image-text pair dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.150059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.150059Z digest=sha256:e20cd1f33d3b03ddd046eaf58ed0e7b79d34bac8bcaa376f7e717d1ea45b4d4e

Observation 8a9ee46c-ef58-4c1f-ae7b-45ef3d6f2469 · outbound

This paper cites and Xiao, J.

CoMemo: LVLMs Need Image Context with Image Memory and Xiao, J

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.154200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.154200Z digest=sha256:2364e672df8b6fd3c9060467eb2f756b10cf68d1a8641ac863cbd3434b944751

Observation 05435e20-be77-4581-9970-3db59625d8a3 · outbound

This paper cites Textocr-gpt4v.

CoMemo: LVLMs Need Image Context with Image Memory Textocr-gpt4v

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.157676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.157676Z digest=sha256:decd2cda0f93ff254d3279fc36066b3fecc76ea8862fbebfea36703e81f1914d

Observation ea0cc5cf-0707-45eb-84f8-9d044095792f · outbound

This paper cites MapQA: A Dataset for Question Answering on Choropleth Maps.

CoMemo: LVLMs Need Image Context with Image Memory MapQA: A Dataset for Question Answering on Choropleth Maps

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.160890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.160890Z digest=sha256:de44aa345388c756c306c70818040950cb77b18d783a71b37587ded6faf30cac

Observation e07abd60-3b54-47fa-9573-77bf4aa61664 · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

CoMemo: LVLMs Need Image Context with Image Memory ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.164986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.164986Z digest=sha256:a3d11899d916b2acb8647339513450ce0a49647b4aeae9e5b3b956d4e1650e24

Observation 59f67add-d782-4305-b878-8925ec32c11e · outbound

This paper cites UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression.

CoMemo: LVLMs Need Image Context with Image Memory UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.168722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.168722Z digest=sha256:b1a0357ec7f771a0b9a59a315d39e7c6ee2c1fee07fd8a63985cb19d55f2eed3

Observation ae2a3243-366e-4ca9-bc2f-638aae7bf3b1 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

CoMemo: LVLMs Need Image Context with Image Memory Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.172306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.172306Z digest=sha256:ff6c2ae7db403ee613fc3af607b52dadbef651b85f4316a19aae780daf23547b

Observation aa14a1d1-29be-4aeb-89cf-76468b19416b · outbound

This paper cites EVLM: An Efficient Vision-Language Model for Visual Understanding.

CoMemo: LVLMs Need Image Context with Image Memory EVLM: An Efficient Vision-Language Model for Visual Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.175851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.175851Z digest=sha256:ef58aeefcf80bfeaff62736dcb61890853aece12749fd83d74c13ecac0da8e71

Observation bbc65fc0-80bd-44e0-a9d6-e376db1066f4 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

CoMemo: LVLMs Need Image Context with Image Memory Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.179318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.179318Z digest=sha256:79d668b0bcdf2caa17887a8ffb6d70523eca2580945b48ef511fc3072531edf3

Observation de2765e1-0281-4a57-8c39-37fe9e9ebe9e · outbound

This paper cites Complicated Table Structure Recognition.

CoMemo: LVLMs Need Image Context with Image Memory Complicated Table Structure Recognition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.182794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.182794Z digest=sha256:0345c15d009d174fac58a2ef5931f90eb86d7100aae007984175dcc8c2a08fe3

Observation 8b42769e-724d-4fc0-9d0c-1b09306a9c0f · outbound

This paper cites K., Liu, Y., Sun, Y., Ng, C.

CoMemo: LVLMs Need Image Context with Image Memory K., Liu, Y., Sun, Y., Ng, C

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.187512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.187512Z digest=sha256:519a1ea342109da14e34ae06f9ddd26ab05c67e718121990a600c20c82309b2a

Observation cd94c8a4-0aba-436e-a7a4-f13cea61dff6 · outbound

This paper cites Simple and Effective Multi-Paragraph Reading Comprehension.

CoMemo: LVLMs Need Image Context with Image Memory Simple and Effective Multi-Paragraph Reading Comprehension

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.190767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.190767Z digest=sha256:add336e1f42564c17e4de0231dd5fad191569caefaeba1da954bf980618b5e26

Observation 33df9bd5-09bf-4ca6-8a80-9c1818466996 · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

CoMemo: LVLMs Need Image Context with Image Memory NVLM: Open Frontier-Class Multimodal LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.194240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.194240Z digest=sha256:5e79e31cf76a69afad2a9c5e4c2b5f469e71eb7d8ca738908b518aa9ff1a4fee

Observation ff4a96ac-f9d2-4b61-bde0-f07a2c0002e3 · outbound

This paper cites Deep visual template-free form parsing.

CoMemo: LVLMs Need Image Context with Image Memory Deep visual template-free form parsing

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.197809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.197809Z digest=sha256:dfc5d3d60ea669a547540a2d018122127540ed1b234968c00f45f861df3650e0

Observation 1af4d204-7229-42ac-9f8c-b5f8dca7c0f4 · outbound

This paper cites The Llama 3 Herd of Models.

CoMemo: LVLMs Need Image Context with Image Memory The Llama 3 Herd of Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.201168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.201168Z digest=sha256:183554b47849dbb8a99fe92695d806313966907e614fc66e5cf1f2a1f52b42c1

Observation 6ae6ab93-5751-4fb2-98a7-9ddc968e5796 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

CoMemo: LVLMs Need Image Context with Image Memory MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.204597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.204597Z digest=sha256:f596d550f5e3a281f37b926a025129975cfaf364f00c7ed2a46eb03396d11c3d

Observation 763727de-789a-4026-a5f0-6c930777ce2c · outbound

This paper cites A., Ma, W.-C., and Krishna, R.

CoMemo: LVLMs Need Image Context with Image Memory A., Ma, W.-C., and Krishna, R

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.208209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.208209Z digest=sha256:3f6c1374be2e48b07999deacf829bf12de18774bc18024df66b04f352f432ea7

Observation 4544efaf-0813-4a67-9ba4-cbd5ad03d612 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

CoMemo: LVLMs Need Image Context with Image Memory Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.211435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.211435Z digest=sha256:76b948857b749fb2381053f2b082f39625a0f32b782d978a72faa0c437ac3e03

Observation 12fc4599-f778-4b46-a339-87cc2d6b3369 · outbound

This paper cites Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark.

CoMemo: LVLMs Need Image Context with Image Memory Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.214909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.214909Z digest=sha256:1d3d867c2f01589300fe19e036bbfb76e2b8e4b1b88e4fc6fac70791d456ea04

Observation 52e019b4-ae1e-4d76-a27d-ed70e62f19f6 · outbound

This paper cites Eaten: Entity-aware attention for single shot visual text extraction.

CoMemo: LVLMs Need Image Context with Image Memory Eaten: Entity-aware attention for single shot visual text extraction

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.218140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.218140Z digest=sha256:40fb82f4b526a843e5c70028879a861b618a41883b693b52051463b3d186071a

Observation 2430febf-ee02-49cd-b57f-6ac84cc8b47f · outbound

This paper cites Icpr2018 contest on robust reading for multi-type web images.

CoMemo: LVLMs Need Image Context with Image Memory Icpr2018 contest on robust reading for multi-type web images

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.221368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.221368Z digest=sha256:6daa920ddc9b123cd7156d47d1a3b4175829ccb8fe3ebf0146d800a8a4a81833

Observation 209f2471-82e4-450e-8fe7-45bb600d254e · outbound

This paper cites PathVQA: 30000+ Questions for Medical Visual Question Answering.

CoMemo: LVLMs Need Image Context with Image Memory PathVQA: 30000+ Questions for Medical Visual Question Answering

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.224674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.224674Z digest=sha256:84f9e1d2fdef6f328d7dbb766ae0f30c9e94a11341fd0d85f7a0258ac8115f12

Observation 669db1a4-b54e-49e4-8bfe-4d8b997731fb · outbound

This paper cites an unresolved cited work.

CoMemo: LVLMs Need Image Context with Image Memory Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.228260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.228260Z digest=sha256:f9bd909d7634829f6bb46bcc57fbbee3b070f85ec879bac5ebbcbad637e9c3f3

Observation 6ca98ae8-e705-436f-9692-368ec67430e2 · outbound

This paper cites Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment.

CoMemo: LVLMs Need Image Context with Image Memory Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.231512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.231512Z digest=sha256:96f4833aa1d75bc8f5438895d1f3edc1a3a03318984e62759c447ceba2949823

Observation 13ce348c-6f4e-4ff0-a75c-36eed2ec4bee · outbound

This paper cites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding.

CoMemo: LVLMs Need Image Context with Image Memory mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.234729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.234729Z digest=sha256:824af33aa609c53fe68ec96824df3b8c25fe658c943f6e6ebb686a8366588870

Observation 8c75405c-2f20-4413-a079-1700b6a00b72 · outbound

This paper cites Medical-diff-vqa: a large-scale medical dataset for difference visual question answering on chest x-ray images, 2023.

CoMemo: LVLMs Need Image Context with Image Memory Medical-diff-vqa: a large-scale medical dataset for difference visual question answering on chest x-ray images, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.238109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.238109Z digest=sha256:e05b1b464d6f3b4a90b88bd7dc7ca37da503861f8774d49803a88ca86912be5d

Observation beed86e7-8ef0-4c38-b874-721f76f683c1 · outbound

This paper cites Movienet: A holistic dataset for movie understanding.

CoMemo: LVLMs Need Image Context with Image Memory Movienet: A holistic dataset for movie understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.241444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.241444Z digest=sha256:089b58662a86a2531c5a425b706ed8e6e78775d4ebd7a3e57fadb85e2baebfef

Observation a3283491-6cde-4272-a8e9-6a8b47581690 · outbound

This paper cites Icdar2019 competition on scanned receipt ocr and information extraction.

CoMemo: LVLMs Need Image Context with Image Memory Icdar2019 competition on scanned receipt ocr and information extraction

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.244739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.244739Z digest=sha256:21e2605917c2f2e537b8da53febf4b8467b528cdccdea1521ca1ad494c106ee1

Observation 4dfa8682-7aa0-4224-8ea3-1847b10f9a60 · outbound

This paper cites an unresolved cited work.

CoMemo: LVLMs Need Image Context with Image Memory Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.247942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.247942Z digest=sha256:871ecd4788485de82e8972c9dd3e759a0416426873108b3479d66662a750edf5

Observation 174f77cd-8052-47e7-9770-9eba5b8372c6 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

CoMemo: LVLMs Need Image Context with Image Memory MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.251165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.251165Z digest=sha256:5872471490cd369f99009509953c0d5f8ccf818173cafcffa5977d0cd9e36b6c

Observation 80d25d45-dc03-4fc4-a455-9dc0d7bf3c4c · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning.

CoMemo: LVLMs Need Image Context with Image Memory Clevr: A diagnostic dataset for compositional language and elementary visual reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.254523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.254523Z digest=sha256:5ae4c962dd6c418da77af4808afe98d70b5642e86669a9b35bb14a8e70c2e498

Observation b5cdb447-51fa-47f8-a580-c592747f2c7e · outbound

This paper cites Dvqa: Understanding data visualizations via question answering.

CoMemo: LVLMs Need Image Context with Image Memory Dvqa: Understanding data visualizations via question answering

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.258111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.258111Z digest=sha256:eb08c7eb97f4bfc49bd8315d4eead4f51eb7ff7c9b32a36d97cdbe17b396bd92

Observation d8be9ef5-2188-4f56-a41a-7957667c4179 · outbound

This paper cites FigureQA: An Annotated Figure Dataset for Visual Reasoning.

CoMemo: LVLMs Need Image Context with Image Memory FigureQA: An Annotated Figure Dataset for Visual Reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.261341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.261341Z digest=sha256:2559e519a8728d07de4c9cf6334433e14cbec269b1e22a4afc9bed21bf05a580

Observation c1474023-82b6-4b32-be90-265dbf797445 · outbound

This paper cites Chart-to-Text: A Large-Scale Benchmark for Chart Summarization.

CoMemo: LVLMs Need Image Context with Image Memory Chart-to-Text: A Large-Scale Benchmark for Chart Summarization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.264586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.264586Z digest=sha256:d6fd36b35ccbe5edd1be90b071ce0a948697459c09544daced4b2791b2f4a36c

Observation 8277c4b0-01fe-4f01-b02a-6ea0cf332505 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

CoMemo: LVLMs Need Image Context with Image Memory Referitgame: Referring to objects in photographs of natural scenes

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.268275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.268275Z digest=sha256:580dc064dfdd77824b2be42deb43688baba8c1c6004ac7a4992d5daf853bd0e5

Observation 8f63e000-09da-4d75-827b-ff1de325a7e9 · outbound

This paper cites A diagram is worth a dozen images.

CoMemo: LVLMs Need Image Context with Image Memory A diagram is worth a dozen images

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.271379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.271379Z digest=sha256:efc635e8007b6467b26ba5ec4fd2222309eb8d8f609523099f624cde296a9112

Observation f1cc699c-fcae-4550-87ec-cf8c16069995 · outbound

This paper cites Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension.

CoMemo: LVLMs Need Image Context with Image Memory Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.274553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.274553Z digest=sha256:a964b3dd6d8c85fcc8609ab6e30243e378d115de8ad36455cf94eb8a7b334f3f

Observation 57ad395e-beeb-4c37-a87b-ed1205ca2522 · outbound

This paper cites Visual information extraction in the wild: practical dataset and end-to-end solution.

CoMemo: LVLMs Need Image Context with Image Memory Visual information extraction in the wild: practical dataset and end-to-end solution

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.278069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.278069Z digest=sha256:1f9caff13c7be57f140684db4689dacc6fd50bdbbd0da3217f404dccebebd42d

Observation 810c6589-e670-4112-a806-0b314790c112 · outbound

This paper cites Laion-gpt4v dataset.

CoMemo: LVLMs Need Image Context with Image Memory Laion-gpt4v dataset

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.281255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.281255Z digest=sha256:79d33a117e62f1124738785c7b1fd6caa9530c2ccd74c9722f7b552d89979bbd

Observation dfc59059-e7aa-4d95-a722-6d38b63e3d24 · outbound

This paper cites J., Gayen, S., Ben Abacha, A., and Demner-Fushman, D.

CoMemo: LVLMs Need Image Context with Image Memory J., Gayen, S., Ben Abacha, A., and Demner-Fushman, D

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.284375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.284375Z digest=sha256:9e8485d114adbea9d59dfc9a9f1f6f9e2dffe235a705637adccb13987944ec0d

Observation 9a24351c-1ad4-46be-9297-e4b072587cfb · outbound

This paper cites What matters when building vision-language models?.

CoMemo: LVLMs Need Image Context with Image Memory What matters when building vision-language models?

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.287592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.287592Z digest=sha256:ad8f978aff5725e95ac1f66057bfb86ec9808a2ebf2c995bfcd5f962fa63c89f

Observation cc7246f4-6e8f-4389-a348-e5da0dbb0458 · outbound

This paper cites G., and Lov \'o n Melgarejo, J.

CoMemo: LVLMs Need Image Context with Image Memory G., and Lov \'o n Melgarejo, J

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.290994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.290994Z digest=sha256:d3b7a7a62a420871c34f9f91d4efef5ab5fa2824174113a01ae8c4b954d6c4e4

Observation 36f591ec-7866-4929-8473-e3c6ff0d02ff · outbound

This paper cites Chemvlm: Exploring the power of multimodal large language models in chemistry area.

CoMemo: LVLMs Need Image Context with Image Memory Chemvlm: Exploring the power of multimodal large language models in chemistry area

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.294273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.294273Z digest=sha256:f2a4a5dc327242285c76a00da791b6b325fc9f95d023d222e0af31668cda3630

Observation 9322a4e8-564e-40c4-a543-064dc566e5e2 · outbound

This paper cites an unresolved cited work.

CoMemo: LVLMs Need Image Context with Image Memory Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.297412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.297412Z digest=sha256:552e790726190977121cba398c9b540674f526943ad7481dde3659091518ebb1

Observation dd1e125a-4572-4e4e-b84a-48a5cf860a21 · outbound

This paper cites Vila: On pre-training for visual language models.

CoMemo: LVLMs Need Image Context with Image Memory Vila: On pre-training for visual language models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.300577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.300577Z digest=sha256:c4811372f59babcc7d53300418123024b8489f1285fd307ad4b24ba0ecaf69f1

Observation 0efdde03-aabb-4c5e-9981-273822210940 · outbound

This paper cites an unresolved cited work.

CoMemo: LVLMs Need Image Context with Image Memory Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.303815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.303815Z digest=sha256:e49781ee9223042c09534d830ac97daf3af0b1d9cc8a6f7f3716b2aad75352d0

Observation 01704a5f-c125-49bc-881a-eb17e9f8b86c · outbound

This paper cites an unresolved cited work.

CoMemo: LVLMs Need Image Context with Image Memory Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.307014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.307014Z digest=sha256:01bf4d891a5280d66d2e7d1d92bc01c7d640b9e294ff5f4038d869ba21ad162c

Observation 99dea663-00ac-46ea-babb-7b35d882486d · outbound

This paper cites Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering.

CoMemo: LVLMs Need Image Context with Image Memory Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.518169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.310104Z digest=sha256:d326a51fe11ed406061fbd3e0a4f3fecb16cec094f818a68d6b94e6394647585

Observation b189077f-4b5e-4840-ad8e-bea1bbe31d7d · outbound

This paper cites Casia online and offline chinese handwriting databases.

CoMemo: LVLMs Need Image Context with Image Memory Casia online and offline chinese handwriting databases

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.506657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.313225Z digest=sha256:60ae41dd12f02dffdf038ff22b54faf4970da8d083712bc55523cc11451a83cb

Observation 946e01d2-f093-45ca-a824-c0bf5ccc3cc5 · outbound

This paper cites Visual spatial reasoning.

CoMemo: LVLMs Need Image Context with Image Memory Visual spatial reasoning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.495266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.316435Z digest=sha256:ac522fddbf82bf361f7cc1fd0d95691b1c4d525f2a6d2803c820139014cb0f57

Observation b07f0ca7-744d-4623-b1e8-e0f98c805451 · outbound

This paper cites Mitigating hallucination in large multi-modal models via robust instruction tuning.

CoMemo: LVLMs Need Image Context with Image Memory Mitigating hallucination in large multi-modal models via robust instruction tuning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.484152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.319716Z digest=sha256:3224b5c8635d18dcc4dd65d99b4703604bea7db2b932bdf5512c78b6444c2504

Observation 17c71ad7-9f4f-4159-9ce9-1a6764bd694f · outbound

This paper cites MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning.

CoMemo: LVLMs Need Image Context with Image Memory MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.322925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.322925Z digest=sha256:f0e711e1907ec42242a574fa15b820a4efa7801e016ab294484bdcca3988f4f8

Observation f00e6b2e-9063-4c1a-8cd1-2a7c6a92d1a5 · outbound

This paper cites an unresolved cited work.

CoMemo: LVLMs Need Image Context with Image Memory Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.326278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.326278Z digest=sha256:04103b4378db8449c8bf3ed9e7730716a81fdbc3c2c6b42a89022d6ef77c4b81

Observation 604930f5-0b88-4c59-95d0-1ee410d9de32 · outbound

This paper cites an unresolved cited work.

CoMemo: LVLMs Need Image Context with Image Memory Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.329367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.329367Z digest=sha256:602035404e73484e3f0751c6d0508a83e2fe0556985746046cd317e6f2f9584b

Observation 8546b904-8a57-41ab-a652-42fed20649b7 · outbound

This paper cites F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P.

CoMemo: LVLMs Need Image Context with Image Memory F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.332532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.332532Z digest=sha256:4cf47a7e43210048fb3043a6b44d4cc4e773c9a0eb33b648ca1832af5dbb3757

Observation 90241c51-b471-4207-a6b0-416a39b7b9e8 · outbound

This paper cites Paying more attention to image: A training-free method for alleviating hallucination in lvlms.

CoMemo: LVLMs Need Image Context with Image Memory Paying more attention to image: A training-free method for alleviating hallucination in lvlms

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.449278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.335865Z digest=sha256:0e392f7b878295624fdcebd7904000003494fcbb6599732c348e0956266f5339

Observation 733505ed-4a4d-490f-9430-7783d7a73374 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pp.\ 216--233.

CoMemo: LVLMs Need Image Context with Image Memory Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pp.\ 216--233

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.339582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.339582Z digest=sha256:e70dfb4d91e4260c241ba2089fed07e16f8b7530b60fe7b06e5aa10ee1e965f3

Observation 5695cc77-06fc-4aca-a66e-e9a988521821 · outbound

This paper cites MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs.

CoMemo: LVLMs Need Image Context with Image Memory MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.346839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.346839Z digest=sha256:f04df96db217d4853ce0a7dc4fb8e08063efaf242fd0521c9a801edce2c4295c

Observation e3b3ff6b-8025-4c8a-98d7-14cdf5e688b2 · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

CoMemo: LVLMs Need Image Context with Image Memory Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.350057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.350057Z digest=sha256:052481dde495d5a63aba7939c5e4f0ed14545a75593791a053cc3a605e7c07ad

Observation 6e92e147-407f-4e5b-8b8f-1e49ef1e8d61 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

CoMemo: LVLMs Need Image Context with Image Memory Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.429290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.353502Z digest=sha256:58e6c80007999f988bbce590b4f52eb030ca89dcc671ded9fda526d0960c9555

Observation 4f30c604-bac2-41b0-9774-f0e2832b62cd · outbound

This paper cites Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning.

CoMemo: LVLMs Need Image Context with Image Memory Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.357233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.357233Z digest=sha256:f9ca6df706a65ed8bf4035b3476220f2f5d69aa0f4d304eedf5baa41b3c1a51f

Observation 156f6993-0a26-418b-bb6f-867c05b90e02 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

CoMemo: LVLMs Need Image Context with Image Memory MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.361051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.361051Z digest=sha256:2e30f5ed19162fce7fa03ca6b787bf1e562ed5058bf96b7dc21b6bd4f09e8fff

Observation f661c154-e39b-42b0-980a-4f0457f3041a · outbound

This paper cites Deepart: Learning joint representations of visual arts.

CoMemo: LVLMs Need Image Context with Image Memory Deepart: Learning joint representations of visual arts

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.418177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.364523Z digest=sha256:f426a0a2fd5c873e009e4838f625f9298ae5e5a35ab5ec35a911a67756f3b7db

Observation c77c4483-ffa4-4447-b75b-c3042ff0f6a1 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

CoMemo: LVLMs Need Image Context with Image Memory Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.406702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.367657Z digest=sha256:616c54f72f84323e0bf6c9d0e79bfda94dbbf79ba1efd9482675a7ed7358cf8b

Observation dfa4350b-4f3c-46ca-9875-93bcd26166eb · outbound

This paper cites and Bunke, H.

CoMemo: LVLMs Need Image Context with Image Memory and Bunke, H

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.394019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.370785Z digest=sha256:7d74266f337794f0ae27442c93f09495913d3615ea3b974ffd8b52aa255a49b9

Observation 18c6213e-1545-47cb-916a-ab4fe68274be · outbound

This paper cites L., Tan, J.

CoMemo: LVLMs Need Image Context with Image Memory L., Tan, J

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.382660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.374097Z digest=sha256:363f294e6651eb7d1ccab6834e23d4544967efe08ff98710b7b2e77577d8ba31

Observation 137e982f-9a0d-49a5-8a9e-134d0c8a0659 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

CoMemo: LVLMs Need Image Context with Image Memory ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.377251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.377251Z digest=sha256:b4a8bc0c71e282d4e91c1a9242bc40cad1c2a95bf4620cec86cfb761d3488f08

Observation be2f4648-53c9-4e91-8ab0-fc059295805c · outbound

This paper cites UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning.

CoMemo: LVLMs Need Image Context with Image Memory UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.380654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.380654Z digest=sha256:ea71a19324fd7d5ade7a2549b44f47023ecafea92d88c952d1e0cb7be68e98cc

Observation e6bcc7c4-897b-429b-ad41-109aa8beff90 · outbound

This paper cites Infographicvqa.

CoMemo: LVLMs Need Image Context with Image Memory Infographicvqa

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.371284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.384119Z digest=sha256:bd843efcce84b30e5310de9d31cd658f87d511e289f76a3c4f0b86e7c9a10a81

Observation 9ff00e35-d786-456c-8814-5f0078660c84 · outbound

This paper cites M., and Kumar, P.

CoMemo: LVLMs Need Image Context with Image Memory M., and Kumar, P

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.360029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.387472Z digest=sha256:6d3339ec5eb697a20f21a344007d01d18c5287a8245aa382b13889bf2fcf697e

Observation 3ecfc09d-1ebb-4353-a622-1da18e7a13fe · outbound

This paper cites Opengvlab/internvl-chat-v1-2-sft-data, Jan 2024.

CoMemo: LVLMs Need Image Context with Image Memory Opengvlab/internvl-chat-v1-2-sft-data, Jan 2024

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.348665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.391046Z digest=sha256:572e508e5743a189f32aec78db9eb2212d4262ebea58e841c52be3d446c30a76

Observation 63479621-6e1a-4838-8694-93568d5fa6e7 · outbound

This paper cites Training language models to follow instructions with human feedback.

CoMemo: LVLMs Need Image Context with Image Memory Training language models to follow instructions with human feedback

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.394227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.394227Z digest=sha256:e3d011bb5a1b267a372a84fb918a7e81d8747668bdf8e46de4433c71e99b19e8

Observation 0c1e3acc-6270-4b13-aa75-e1b031b50306 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

CoMemo: LVLMs Need Image Context with Image Memory Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.397502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.397502Z digest=sha256:f710b42a54e5aee2fb8a55330c0cc8ba7bb4d70ca006d5ec0ea439f2544a9499

Observation 01a410e3-8cf4-4fae-b973-5d36686e14b5 · outbound

This paper cites A., Wang, L., Cervantes, C.

CoMemo: LVLMs Need Image Context with Image Memory A., Wang, L., Cervantes, C

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.329398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.400986Z digest=sha256:ae9c128951582a8944f5183829e4545ac7412f8e504d72f02f98a623b5198b2a

Observation 1cc2ab6b-5754-4d27-8241-182a2418c82a · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

CoMemo: LVLMs Need Image Context with Image Memory Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.318852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.404242Z digest=sha256:d014c80108b9685bd2286cc51e76ac8e456ddf62c77a2205a66793cdc633fa88

Observation bb4b8522-8214-418c-a2ff-1c533289bb58 · outbound

This paper cites Laion coco: 600m synthetic captions from laion2b-en.

CoMemo: LVLMs Need Image Context with Image Memory Laion coco: 600m synthetic captions from laion2b-en

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.307704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.407507Z digest=sha256:a18e3ef972a4906d5dbb0875af91a291de9a52aa1d3f57e6fe9e4bf44ab94e72

Observation 1b84a50c-ef90-4ba9-bdef-b3783f87e3ad · outbound

This paper cites Solving geometry problems: Combining text and diagram interpretation.

CoMemo: LVLMs Need Image Context with Image Memory Solving geometry problems: Combining text and diagram interpretation

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.296729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.410789Z digest=sha256:5f3970f468fd262820ecf0349403ea2bd9554f6002dc199b3888465cd4bfe7c4

Observation fbe9a7b8-59c9-41b3-99d9-a4dfb01ad904 · outbound

This paper cites an unresolved cited work.

CoMemo: LVLMs Need Image Context with Image Memory Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:02:31.285582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.414053Z digest=sha256:7a8f3418c91ee3245c001bdae148ccf9041b941e0574126906cdf737e7751fd6

Observation 1a3ad93e-ed8e-4a09-9d05-0bf9b7515aec · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

CoMemo: LVLMs Need Image Context with Image Memory Objects365: A large-scale, high-quality dataset for object detection

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.274988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.417252Z digest=sha256:273c302ff1c7c58a2ca628a7c67a7d792c225756dee8484cc80d6ae0f985abaa

Observation 0a2e401d-22de-4408-b203-8ce576a0d7e5 · outbound

This paper cites Towards vqa models that can read.

CoMemo: LVLMs Need Image Context with Image Memory Towards vqa models that can read

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.264334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.420406Z digest=sha256:4150a780187c5b3fb0ca5d50c514f1bf791c0a898c130049aca3294172dbbd3b

Observation be159f74-34f0-4efd-9ca0-9bfb09e11a71 · outbound

This paper cites Towards vqa models that can read.

CoMemo: LVLMs Need Image Context with Image Memory Towards vqa models that can read

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.252411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.423527Z digest=sha256:35c41a1dd738eb7684223f4a2e6259ea535c3f71d2b98bcdd15493deb6ed9de1

Observation 3179f98c-0cf1-44ad-ab44-ee10b2d0dbce · outbound

This paper cites Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text.

CoMemo: LVLMs Need Image Context with Image Memory Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.241606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.426830Z digest=sha256:094038d2931a4955e6f9da4c0cec03a7ba11b569e342175d2d04263d9771b9d8

Observation d554ce3f-0191-4126-aca7-caf3d993cddc · outbound

This paper cites MileBench: Benchmarking MLLMs in Long Context.

CoMemo: LVLMs Need Image Context with Image Memory MileBench: Benchmarking MLLMs in Long Context

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.430076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.430076Z digest=sha256:7634ee40b68a77571a6953e04733e982ad0d57001cd79e1843167c521d9a38e8

Observation 15374b95-0c4e-410f-af9a-b0853b967627 · outbound

This paper cites Transformer roadmap: 2.

CoMemo: LVLMs Need Image Context with Image Memory Transformer roadmap: 2

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.229796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.433564Z digest=sha256:5a7542f27239e4a5b6a6aac3379fcf682ab313a4abaac43688f43ba8b65509ad

Observation 9d98cf6c-e5e3-4c5c-980b-b5d1a4852d4f · outbound

This paper cites C., Han, J., Ding, E., Liu, J., Karatzas, D., et al.

CoMemo: LVLMs Need Image Context with Image Memory C., Han, J., Ding, E., Liu, J., Karatzas, D., et al

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.218537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.436705Z digest=sha256:b169fed760185824a32c0d69b709ff8f1892e8c7b78e2ba6f498e3dd7a9023a9

Observation 18ec7216-01ed-42d5-85ed-341a4c6cd833 · outbound

This paper cites Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy, 2024.

CoMemo: LVLMs Need Image Context with Image Memory Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy, 2024

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.207757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.439836Z digest=sha256:67cd405e5105b9cbac5a98cbd7c5e0539169b6740dfa83df428bc34b5c10aad1

Observation 295ed525-12db-4af6-bcb7-6dfc706f61ef · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

CoMemo: LVLMs Need Image Context with Image Memory Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.444417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.444417Z digest=sha256:3e272162860b813cdff2484ca75a2b5216d55d6b2d984aba0bbcc66dbbe8ff92

Observation a1f5b464-6f14-461f-aedd-9d6cbb42dc60 · outbound

This paper cites COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images.

CoMemo: LVLMs Need Image Context with Image Memory COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.447765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.447765Z digest=sha256:2055c35ce3b73f628a87bc55d94c025a1b3ddda3a5b54e2bd77bdeab59d50aa9

Observation 9f93a81d-03ac-4199-8d57-73741ee9121c · outbound

This paper cites V3det: Vast vocabulary visual detection dataset.

CoMemo: LVLMs Need Image Context with Image Memory V3det: Vast vocabulary visual detection dataset

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.196589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.451246Z digest=sha256:16886ba399ac8ea4fa6d29a2aae7da51eafae7f7653c1dff753cae1072175607

Observation ae545fb5-9c31-4421-9b57-625e6aeba795 · outbound

This paper cites Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset.

CoMemo: LVLMs Need Image Context with Image Memory Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.454664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.454664Z digest=sha256:4451a0fa71351d8db79b6adc6a8f133c3c27dfb845fb8b7bc231ef6b75368233

Observation 6c69544c-f78e-4907-bb82-9e20ca304583 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

CoMemo: LVLMs Need Image Context with Image Memory Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.458205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.458205Z digest=sha256:dbf47438da5dd833c73b4b4d4b2b26746b0fbb2a8ed845944ee68eb881f93de3

Observation 74ab545b-117c-4047-824a-57ee643cddd5 · outbound

This paper cites The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World.

CoMemo: LVLMs Need Image Context with Image Memory The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.461439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.461439Z digest=sha256:bc81a0a98e1dd808a6241edb386a8394c7d3e6c8d965608a92a47f4f85894dba

Observation 4f7dfd5e-5b7a-4f5b-90cb-ac18b352031d · outbound

This paper cites Needle In A Multimodal Haystack.

CoMemo: LVLMs Need Image Context with Image Memory Needle In A Multimodal Haystack

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.464875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.464875Z digest=sha256:88a50f3e98d028e47fdd782f316850c24a3afa0799cfecd1f40a397d92fefcf3

Observation 60177022-5d53-47f2-baf6-075177bef644 · outbound

This paper cites C., Luo, C., Jin, L., Chan, C.

CoMemo: LVLMs Need Image Context with Image Memory C., Luo, C., Jin, L., Chan, C

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.184921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.468386Z digest=sha256:1aa192a36eb091495ab54de86b7de4da8c72de79365776fa3f2e493ac81973fc

Pith citing papers

Observation 141c552d-38db-4546-9c1b-cf0bfe83f149 · inbound

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs cites this paper.

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs CoMemo: LVLMs Need Image Context with Image Memory

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:01:23.021630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-09T18:53:06.494640Z digest=sha256:2611aa6f4412957669004e7b9ef193ccdbada118800e1c6513de753f8efda45f

Observation a6c78e64-9adc-4225-94c0-35ecef142c6c · inbound

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs cites this paper.

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs CoMemo: LVLMs Need Image Context with Image Memory

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:50:51.525657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T01:49:15.136031Z digest=sha256:d40a52a48de26c11cff3e7d8113d4e90d08a0d9cd30722878b61b67a8ab31d70