Pith. sign in

Paper Citation Record · LEDGER

Region-Level Context-Aware Multimodal Understanding

As of 13 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 1 inbound Pith citation observation for arXiv:2508.12263.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.12263 v2

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T19:39:48.245781Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:35:37.095627Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T10:11:09.424972Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact2
  • verified fuzzy27
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c6aea20c-74fd-41b4-953c-32687809851b · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Region-Level Context-Aware Multimodal Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.022168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.022168Z digest=sha256:68bd43df8dbbb53b5cd7fc26f255226117fb1e1ba200d99e63686a457b8f63f3

Observation 7df1ea6a-cf53-4abe-8f61-2d24ae8de02a · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Region-Level Context-Aware Multimodal Understanding Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.032858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.032858Z digest=sha256:ca85af08704a79673afa3869a8c192a1f3d60b10aa1169fda9168a1a9f86d98c

Observation 8424158f-e0c5-4b3d-9f5f-c5aab1d835ac · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Region-Level Context-Aware Multimodal Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.041259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.041259Z digest=sha256:93ff978d42e6a4574385d17aa673d4589fde0149bcfd81fbce2a3766871258af

Observation 6a73e7ef-8b2c-40c4-b159-6154601ab78b · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Region-Level Context-Aware Multimodal Understanding InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.049375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.049375Z digest=sha256:b0480ffb7e8189b9a43a61fd3642afcb3474ae03624bda5fafd54d785dca15f5

Observation ec805674-2e99-49b6-8ff8-56e581a3b0f5 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Region-Level Context-Aware Multimodal Understanding DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.055078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.055078Z digest=sha256:ff6dc1a3e86f1de0f0284edb6551b30e38f52cd6d06ad58822592125f598daa4

Observation f5b6ccda-328f-47e0-895c-d34039522811 · outbound

This paper cites MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and Text.

Region-Level Context-Aware Multimodal Understanding MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and Text

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.059765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.059765Z digest=sha256:a457d22f046e1acbf6d09b3f4a2755a4faaa2e0570e8041cb5f9ee0935c50516

Observation dc0acddf-0452-4542-8916-34dbb5c29744 · outbound

This paper cites MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning.

Region-Level Context-Aware Multimodal Understanding MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.064324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.064324Z digest=sha256:865c40f096d37f9f3fd59b552a6289cc6f9efd7867527b44bb1cd9acdbab83ec

Observation 16e4f698-fc0d-4b37-80d7-dc44182c46ea · outbound

This paper cites CaMML: Context-Aware Multimodal Learner for Large Models.

Region-Level Context-Aware Multimodal Understanding CaMML: Context-Aware Multimodal Learner for Large Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-05T19:39:48.648811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.069121Z digest=sha256:a6f57de4971da1676ae2ce30896833af0b2fb02df2e8c3dbf41b3495fb2fe973

Observation 8a816166-f40f-4779-b0a8-40d5702a3811 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,.

Region-Level Context-Aware Multimodal Understanding Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:49.075111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.078598Z digest=sha256:62a7e79e968d49dd490f637d6d92731136acb3f32f7f578047604cdf5fc64f84

Observation 064f0dad-04c1-40ce-a4be-87c156d73b20 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes,.

Region-Level Context-Aware Multimodal Understanding Referitgame: Referring to objects in photographs of natural scenes,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:49.062001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.082164Z digest=sha256:a42e13c56a2c2cc823b861222b5616227156b13ff53d75f5f54c5706d243f40c

Observation 7da6b354-52cc-4587-a071-ef4c2aebda7c · outbound

This paper cites Panoptic scene graph generation,.

Region-Level Context-Aware Multimodal Understanding Panoptic scene graph generation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:49.049397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.087505Z digest=sha256:961e4f4ef0e0e53ac1e852bf70f6c611fbe4c59a3e60898551fdf096be760f85

Observation 0f216edd-f8c6-4466-954b-3e5636654f35 · outbound

This paper cites Glamm: Pixel grounding large multimodal model,.

Region-Level Context-Aware Multimodal Understanding Glamm: Pixel grounding large multimodal model,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:49.037852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.092365Z digest=sha256:5cec7a355128cf2744195fcb07a6e210088dd6695b9a7f0a14e5a2c0d33ee7d1

Observation 2157a260-b549-480a-8763-0fa70a8c963e · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

Region-Level Context-Aware Multimodal Understanding Bleu: a method for automatic evaluation of machine translation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:49.025900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.097391Z digest=sha256:24fd428a84babcce800bb1690d1651a189eaa1aeaef6284f6b0095c632f25111

Observation 56ae1682-2407-4411-8588-67915790807f · outbound

This paper cites Rouge: A package for automatic evaluation of summaries,.

Region-Level Context-Aware Multimodal Understanding Rouge: A package for automatic evaluation of summaries,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.100736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.100736Z digest=sha256:bd7cc4f8ee378a1f391d9177172bbf455e56cfc9e2761ce7ed73fc1b1f08f84f

Observation d28d5f7c-fb8d-454a-a840-b822c9a5e598 · outbound

This paper cites Cider: Consensus- based image description evaluation,.

Region-Level Context-Aware Multimodal Understanding Cider: Consensus- based image description evaluation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:49.005462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.104232Z digest=sha256:8ab7b3565efe830ea8987e79047517e0aa9f3694d5c90246afef6ad09d331e23

Observation a41a9f92-bcd7-46f8-a4dd-f57476b6cbe9 · outbound

This paper cites Spice: Semantic propositional image caption evaluation,.

Region-Level Context-Aware Multimodal Understanding Spice: Semantic propositional image caption evaluation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.984661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.108957Z digest=sha256:f8c5246572ee4a91b54ce8d38aac89067d9e2fa8b60700c72486e5173517a221

Observation ba4e72e0-11cf-4a11-a933-129254c97da0 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,.

Region-Level Context-Aware Multimodal Understanding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.971525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.113607Z digest=sha256:ac35e3f336dcb72921be9ddfe93853761e579959a9c91a5639dfe34f16c7d096

Observation a70b2165-d355-46ca-91c3-e18b54b1079e · outbound

This paper cites Visual Instruction Tuning.

Region-Level Context-Aware Multimodal Understanding Visual Instruction Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.121173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.121173Z digest=sha256:79a69a0c930ea7cd3b82be105438943c968504bb6f73ad187f70a13a726c72fe

Observation 069b9a4b-b512-452b-ab20-7cb3fc479039 · outbound

This paper cites Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models.

Region-Level Context-Aware Multimodal Understanding Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.124792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.124792Z digest=sha256:6ac22901443fc3aee283c4b52433620d41410f7fd0f91f0dd98e87a0f380be32

Observation 4e50f82b-d859-4e4a-b256-471307846976 · outbound

This paper cites Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models.

Region-Level Context-Aware Multimodal Understanding Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.130365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.130365Z digest=sha256:f04e2cd56f3432678b51b1ea16ba4ff84bc79f47035340919a85d44fb2cd8f1a

Observation 5a10a371-7cdf-43c3-bbfa-86c72a2a00cc · outbound

This paper cites Available: https://api.semanticscholar.org/CorpusID: 256390509.

Region-Level Context-Aware Multimodal Understanding Available: https://api.semanticscholar.org/CorpusID: 256390509

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.958837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.117178Z digest=sha256:fe2f9304a1f33f103a4f54c4b418dbc832226eda21626eb88b3cee29467ef894

Observation f4621501-30fc-44ef-897a-3d9c6edede2f · outbound

This paper cites Mmict: Boosting multi-modal fine-tuning with in-context examples,.

Region-Level Context-Aware Multimodal Understanding Mmict: Boosting multi-modal fine-tuning with in-context examples,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.948011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.137165Z digest=sha256:3aa37d66a31de92a2657c9a7c545fadadda5ba7b7e1c118201758059b3f2e6c5

Observation 5ed4d786-02e6-4e98-bbb7-f525b886ef24 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering,.

Region-Level Context-Aware Multimodal Understanding Learn to explain: Multimodal reasoning via thought chains for science question answering,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.936899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.140247Z digest=sha256:612741d8faa9a89672d7c1abe7e90772f07b60d7d92045e850349da9d8591ad0

Observation ac638776-fc03-48f3-85c0-9b918924fde1 · outbound

This paper cites Cantor: Inspiring multimodal chain- of-thought of mllm,.

Region-Level Context-Aware Multimodal Understanding Cantor: Inspiring multimodal chain- of-thought of mllm,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.916037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.144298Z digest=sha256:000b5bb371a46e74aa9a5e72e437de1b0d4797b3264b9e3ef2dc931a0494d609

Observation 5a2cff7f-f92a-4cfd-9d31-a79215644a5a · outbound

This paper cites INF-LLaVA: Dual-perspective Perception for High-Resolution Multimodal Large Language Model.

Region-Level Context-Aware Multimodal Understanding INF-LLaVA: Dual-perspective Perception for High-Resolution Multimodal Large Language Model

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-05T19:39:48.570485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.133771Z digest=sha256:78cfe3fc9ef71c4eaaa6fe924322d1d36349f7182b4ea7f43f007fe4a187d973

Observation 2fcc1d4d-6c3a-42db-a9a9-0668919903c7 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video understanding,.

Region-Level Context-Aware Multimodal Understanding Video-llama: An instruction-tuned audio-visual language model for video understanding,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.874397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.159498Z digest=sha256:a3fa0356cee9d9d0a6cdb7cf386b8bd61f29c41fda3d60da691ff68ab63a3c60

Observation cf21f313-22a6-403a-b487-4054478e175c · outbound

This paper cites Video-rag: Visually-aligned retrieval-augmented long video comprehension,.

Region-Level Context-Aware Multimodal Understanding Video-rag: Visually-aligned retrieval-augmented long video comprehension,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.163553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.163553Z digest=sha256:a92a9e845daed3877282f3d9f4fc626a8dd4d89c37eb14b250b33f342e7979be

Observation ec322ad7-3d90-4a70-ac13-c2e78f070e97 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Region-Level Context-Aware Multimodal Understanding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.167701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.167701Z digest=sha256:6fc9714237964e689234ed3d9c9603075626d6dce4ea971a6c512dd3edbe28df

Observation c4866031-3d7d-46ce-968e-a6d1cb721dd1 · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection,.

Region-Level Context-Aware Multimodal Understanding Video-llava: Learning united visual representation by alignment before projection,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.902281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.149659Z digest=sha256:d0dcf438701adf3dbc50f41ba6c46daecf2da8785c77a7ae2a1106d196ac9fc4

Observation c0506359-bb7d-41f3-a44d-c2e962694888 · outbound

This paper cites Available: https://api.semanticscholar.org/CorpusID: 265281544.

Region-Level Context-Aware Multimodal Understanding Available: https://api.semanticscholar.org/CorpusID: 265281544

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.888958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.153196Z digest=sha256:ed1ec070037477ca4689b7195c7ea0e62438db28e8a338ee0e350988840e79b0

Observation 66411ede-fb32-46f8-9fc1-6074fe3a8c78 · outbound

This paper cites Manipllm: Embodied multimodal large language model for object-centric robotic manipulation,.

Region-Level Context-Aware Multimodal Understanding Manipllm: Embodied multimodal large language model for object-centric robotic manipulation,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.834882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.181046Z digest=sha256:50b9dfd56e105c22a8c97415d97b880c287d046d52c20d21211df29c2470e5af

Observation 7d7eda24-ed12-4dfe-90c5-206349a46147 · outbound

This paper cites Rap: Retrieval-augmented personalization for multimodal large language models,.

Region-Level Context-Aware Multimodal Understanding Rap: Retrieval-augmented personalization for multimodal large language models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.804984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.188905Z digest=sha256:f761a3918add92d842f3b9df6a765fe35aa1e342dc4baec3898e9173eca7631a

Observation 863d828b-bcf9-4ba0-af2f-ac13b716c004 · outbound

This paper cites Yo’llava: Your personalized language and vision assistant,.

Region-Level Context-Aware Multimodal Understanding Yo’llava: Your personalized language and vision assistant,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.790306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.191574Z digest=sha256:bb59852cde1d7c2a4b9a45a698a1fdcf519d926c5bfb39cbec6b18095c4913c8

Observation db9dc312-26fa-4102-9fae-f82b7ab730a3 · outbound

This paper cites Jm3d & jm3d- llm: Elevating 3d representation with joint multi-modal cues,.

Region-Level Context-Aware Multimodal Understanding Jm3d & jm3d- llm: Elevating 3d representation with joint multi-modal cues,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.861748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.171506Z digest=sha256:ba99975245ad8ed8d01e4c3a2ebfdc11c84f86276908788f9483d988c2a16467

Observation 2631a3bc-7b7e-4b32-bb7d-bb594ec41218 · outbound

This paper cites Palm-e: An embodied multimodal language model,.

Region-Level Context-Aware Multimodal Understanding Palm-e: An embodied multimodal language model,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.850483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.176293Z digest=sha256:9070edca67d914813bc30d05e02f5fca893f1d2e1e5ec993f9786b60d536f7db

Observation 36873bb4-dd1a-45ee-940a-2d65d5c98aa0 · outbound

This paper cites Llm2clip: Powerful language model unlock richer visual representation,.

Region-Level Context-Aware Multimodal Understanding Llm2clip: Powerful language model unlock richer visual representation,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.207295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.207295Z digest=sha256:3f3fc5bd21b44731b08bdbc4594b1873fb7baf4dc2a4f1ae0a93a9d29bc7c38d

Observation 0ae8dd8e-75a3-4d8d-9bc0-668bf4a658cd · outbound

This paper cites Available: https://api.semanticscholar.org/CorpusID: 266573457.

Region-Level Context-Aware Multimodal Understanding Available: https://api.semanticscholar.org/CorpusID: 266573457

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.819456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.185416Z digest=sha256:26a0b32aa33e1dc2090e617a1365a993c98a7aaba282fad66fa5b1a2f576de98

Observation 5f8819b5-6c47-44eb-a65c-fbf2719e09ef · outbound

This paper cites GPT-4o System Card.

Region-Level Context-Aware Multimodal Understanding GPT-4o System Card

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.214532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.214532Z digest=sha256:3cc0443b9bc37d76becc000c000c5e653ac0c0c14e86f21438f8cb480d169d7f

Observation ea84cffa-cf95-4d85-b959-fcb037b10bb2 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Region-Level Context-Aware Multimodal Understanding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.225203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.225203Z digest=sha256:bdae87856bfd9baa6eaf525f7c4a60395d52f9e9f7d235396b4cc4897eaec480

Observation 06819efa-51cd-4240-b180-13c4f8121ed3 · outbound

This paper cites Myvlm: Personalizing vlms for user-specific queries,.

Region-Level Context-Aware Multimodal Understanding Myvlm: Personalizing vlms for user-specific queries,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.778005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.198688Z digest=sha256:201febe318f86a227e330bd3a600fc1ec56bdacb2d9691f6cdae4a3ed6929b04

Observation 5e8108fc-7083-4939-85cd-9dcfe042a13d · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Region-Level Context-Aware Multimodal Understanding CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.202916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.202916Z digest=sha256:efcf1bc0504c630c884839792667be093535ce776ea8364f52778f29fb6821c6

Observation 80b1a083-761c-416a-9ece-e9b080dc6912 · outbound

This paper cites DeepSeek-V3 Technical Report.

Region-Level Context-Aware Multimodal Understanding DeepSeek-V3 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.237682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.237682Z digest=sha256:bf94dc5382f9f0d6adf6a0f729cab0e9b15c3fccaff8a6ced693028dbdb3d6c3

Observation aacd94b3-7518-4400-b327-feb67497b463 · outbound

This paper cites Gemini 2.0: A new era of multimodal models,.

Region-Level Context-Aware Multimodal Understanding Gemini 2.0: A new era of multimodal models,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.766616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.210718Z digest=sha256:916acc7b3f7197c76db8b977ec411e912f9115439c904799149449aed0d70258

Observation 7a31e000-7b9a-412b-9ce5-ab7d02769624 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Region-Level Context-Aware Multimodal Understanding Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.245781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.245781Z digest=sha256:c76854ce09d09a559fccbf2f03a1d4e0d82e0962ac6c03767bf1452257ba4b42

Observation 7caa56f4-8a90-467c-939b-7101d79bffb0 · outbound

This paper cites Intern vl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,.

Region-Level Context-Aware Multimodal Understanding Intern vl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:48.756620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.229795Z digest=sha256:c18c5c4b7444c409ec492db7db30106dd9e7456b1bff96bce76881daf5849d58

Observation 5f8d553e-db90-487d-b911-114ff3e7f983 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Region-Level Context-Aware Multimodal Understanding MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.233486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.233486Z digest=sha256:a089862a69ae899e62e37bacf969a0f35b806d16ba0c575f82e017ed1108bfe8

Observation eb47bbe7-f3a4-47da-b8e7-ab3b418008ce · outbound

This paper cites Enabling Large Language Models to Generate Text with Citations.

Region-Level Context-Aware Multimodal Understanding Enabling Large Language Models to Generate Text with Citations

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.241070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.241070Z digest=sha256:e2cd770a258b5f7574dfff1f066b3ac208215adb17331d6264597b4a77bad0e0

Observation be0e5648-70cf-4467-88ea-e2d6934d2848 · outbound

This paper cites Available: https://api.semanticscholar.org/CorpusID: 248476411.

Region-Level Context-Aware Multimodal Understanding Available: https://api.semanticscholar.org/CorpusID: 248476411

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:49.242414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.037500Z digest=sha256:d55d1e09aa608492b5fd75bfeaa95100fda0fe4b65429d499c5fb67f57e521c1

Observation af0d61fe-6251-47e2-b377-ccfecb726f3d · outbound

This paper cites Available: https://api.semanticscholar.org/CorpusID: 258615266.

Region-Level Context-Aware Multimodal Understanding Available: https://api.semanticscholar.org/CorpusID: 258615266

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:49.168952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.052310Z digest=sha256:ba832e0705ea78c611291bbf3fe140c1bbec1a9d86dd3233a97fb5daf8d88f0f

Observation eba9c55e-cb75-44ff-adae-69bb5301e5f3 · outbound

This paper cites Available: https://api.semanticscholar.org/CorpusID: 266844925.

Region-Level Context-Aware Multimodal Understanding Available: https://api.semanticscholar.org/CorpusID: 266844925

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:39:49.110845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T19:39:48.074850Z digest=sha256:5961add5b4c715c90c7562e431e3db251196e1cd216a1b75351752d46e382b86

Pith citing papers

Observation 9865e114-cb2d-446a-9799-6fb45b263c7f · inbound

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation cites this paper.

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation Region-Level Context-Aware Multimodal Understanding

Reference 183

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:09.472995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T15:35:37.095627Z digest=sha256:cfe6317efdeb729a8d838d0d59d503a7f0f0112641398119e5a65d316ce579fe