Pith. sign in

Paper Citation Record · LEDGER

D-Attn: Decomposed Attention for Large Vision-and-Language Models

As of 11 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 2 inbound Pith citation observations for arXiv:2502.01906.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01906 v2

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T14:05:14.472756Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T18:36:56.248838Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T14:55:48.338931Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3b031cd4-e97c-437e-8af5-2f2722b14244 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.290681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.290681Z digest=sha256:70449fe015b061359b79ad7e9c3638404486b57adf457b60b7a70feb1ef30f0a

Observation adceb6f2-2f75-4290-bab9-57dc011072c9 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.295768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.295768Z digest=sha256:439d8be4d172042d522b7ade97a54f2a7f85e78383e45bf5f669d8c9cf88701a

Observation 80e6294c-3685-4764-ba54-1a13bfcc929e · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

D-Attn: Decomposed Attention for Large Vision-and-Language Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.300189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.300189Z digest=sha256:e62352ee9b68f26a0b537c88acd5b0c549a1301361cc4126685aebe51d0f81e7

Observation 88d3230d-c0d5-4073-8b1f-2d8c5e271075 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.304523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.304523Z digest=sha256:a5dabe60730a662a6ba4903b06d5787a77161ab0916c71eb05e993450712d00f

Observation e7e7d843-7f38-41b1-83b1-2eb4595de604 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

D-Attn: Decomposed Attention for Large Vision-and-Language Models ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.308714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.308714Z digest=sha256:c647b0052a9a37ec6c299b8d6577a659e31938e870ab722621c58b303c5dc398

Observation fa9ddf79-68dd-418a-90bc-41a8155328d9 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.312984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.312984Z digest=sha256:6e63acde5851c7dd532a275906e5a182a7d3a26a90f9b3591edb2696e970160b

Observation d3ff51b0-5a65-494f-881f-e2d71c255733 · outbound

This paper cites InstructBLIP: Towards general-purpose vision-language models with instruction tuning.

D-Attn: Decomposed Attention for Large Vision-and-Language Models InstructBLIP: Towards general-purpose vision-language models with instruction tuning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:05:15.211198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T14:05:14.317620Z digest=sha256:2df64caab5f9c37753bf8771313f22d2623be5dfaaab9cdbc341fc9ca93b004c

Observation 192e3792-14de-446e-9c37-4b189678d361 · outbound

This paper cites Flashattention: Fast and memory-efficient exact at- tention with io-awareness.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Flashattention: Fast and memory-efficient exact at- tention with io-awareness

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.321948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.321948Z digest=sha256:9e34a8073f507825aa8ab2968f352fd566b4dcf4ed84c4faf97a565ea1e086ea

Observation c9cd7b00-ac84-40cc-88ef-97ac92ff8c1d · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.326040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.326040Z digest=sha256:f3deb3bb1659f7eeabed82c95882309e34e012c05c0eb0582dfb4cbbef8d0738

Observation 17e4642b-7869-4e3d-8473-003ef59eb356 · outbound

This paper cites The Llama 3 Herd of Models.

D-Attn: Decomposed Attention for Large Vision-and-Language Models The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.330494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.330494Z digest=sha256:53edb6fc820d97db5aea27b0a69e6d7a70afc45f7bacf5070ee1bae8607e24a7

Observation c8464cba-c180-44c0-9735-26aad1a99455 · outbound

This paper cites Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:05:15.191778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T14:05:14.334854Z digest=sha256:c93a4cabbad30513e313da8ce14bf2cfd80aa9cbb1660553be904379b42694f7

Observation ea370fc4-6ddb-4a39-9fae-a7d71f1b706e · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:05:15.180349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T14:05:14.338911Z digest=sha256:d65ce093ec1e786183aa51492217923336a8dca0f61093d69422702f6125d2ee

Observation f001ea89-c505-4f05-8e3e-aa88ff44479f · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:05:15.168505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T14:05:14.343679Z digest=sha256:c2d69d6612d19df8ff1ded66524b19a841f0dfab5206a28f549af3ebd57ba316

Observation 09fd5c60-7e62-45a1-a745-39ff9e9c0b50 · outbound

This paper cites Mistral 7B.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Mistral 7B

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.347857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.347857Z digest=sha256:d2be4c724773bf7c897c59b065a7ccfa98f3a0f74e2716c54168975187e7700e

Observation a6589d66-e0a2-47d8-8bc5-26891dcdeee7 · outbound

This paper cites Obelics: An open web-scale filtered dataset of interleaved image-text documents.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Obelics: An open web-scale filtered dataset of interleaved image-text documents

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:05:15.157328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T14:05:14.352337Z digest=sha256:e4bc6875a9e04461e87445019a4b4e12e1ae0f1a2ffa5ad223d548132e208aa4

Observation e9b0923e-e273-4b18-816d-c3bf85ffb1f5 · outbound

This paper cites What matters when building vision-language models?.

D-Attn: Decomposed Attention for Large Vision-and-Language Models What matters when building vision-language models?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.356190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.356190Z digest=sha256:e6d4f63cba83a091abc2ddcbb71a2ca4f5e04a87c38977335e13a6474e2bd341

Observation 173997d1-1f41-4af4-a811-6f1d730820b3 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

D-Attn: Decomposed Attention for Large Vision-and-Language Models SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.359838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.359838Z digest=sha256:10e258927fe4e5a001a72d071d1725302ac640997b10976a19ac56af1334b1ad

Observation 9e54a25f-f822-42c6-aa14-b949d6a8f2cb · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

D-Attn: Decomposed Attention for Large Vision-and-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.363667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.363667Z digest=sha256:705177a56dbd13295c54480cf3b9bb0c169741c7b4c1acdf1780cd758ef9c1f1

Observation 08265120-3b95-49cb-970a-de14a62c02b2 · outbound

This paper cites CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts.

D-Attn: Decomposed Attention for Large Vision-and-Language Models CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.367567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.367567Z digest=sha256:892e3e0775848f4331ee8507d9fbc32e512966375b04333cb941d37021c97efe

Observation 223e6291-37bf-4480-abcf-566f856630a3 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.371060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.371060Z digest=sha256:3d2c617464290f9a1cae7309a6cf5cbc51a61c787d3faab45cb933c1b844d6b5

Observation 417357ba-842f-48af-a660-a44b75d72d6d · outbound

This paper cites Vila: On pre-training for vi- sual language models.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Vila: On pre-training for vi- sual language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:05:15.145461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T14:05:14.375645Z digest=sha256:67f971236ce67eeceed7be4637d9451a62c625425e25fb0503fd6e5965aad18b

Observation d4add187-c6aa-4103-beff-e29fcc2406c8 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:05:15.134058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T14:05:14.379082Z digest=sha256:41d33a96c9c1529cf2f1b651054497a212fe3fdc2e51e4abdc2c50f5c3b5ec06

Observation 433b69d9-cffa-421c-946e-a6b537f2c9ed · outbound

This paper cites Visual instruction tuning.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Visual instruction tuning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:05:15.122663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T14:05:14.382727Z digest=sha256:bc14254f7247e0fbfe0debc32de90ce5ddca6a0998a5bc8f2dffe35c5cf65b88

Observation 159aa463-374d-4d50-bf25-dc282f902de0 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

D-Attn: Decomposed Attention for Large Vision-and-Language Models MMBench: Is Your Multi-modal Model an All-around Player?

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.386562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.386562Z digest=sha256:004dbee1dab7700598bed9b677c7b6bb203837fb1a91ada5fc65092a85497b00

Observation 454579ff-d616-4d8e-a4ae-2d84fc731b64 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.390358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.390358Z digest=sha256:63c4d1798e2770a84e71f869b745ab1cccece2cf0c14a8bcc9cdc80406eb61b6

Observation 2087160c-c19c-4a68-b326-69fbe4f4b50f · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Generation and comprehension of unambiguous object descriptions

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:05:15.103440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T14:05:14.394583Z digest=sha256:f569bf207d115bbc2fce3cd242e3255749409b14535b35e5d435fa25449c00e9

Observation 0a838e0a-a8ac-4e61-a576-cb261aadd4cf · outbound

This paper cites Im2text: Describing images using 1 million captioned pho- tographs.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Im2text: Describing images using 1 million captioned pho- tographs

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:05:15.091372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T14:05:14.398367Z digest=sha256:96eaf6e9b8775ab0f1b0ec7aed2f95efab38d578693fca8a097f2b19fe53df61

Observation 298d547f-edea-4791-9c8a-c3972b2f4acd · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Learning transferable visual models from natural language supervi- sion

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.401848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.401848Z digest=sha256:c3f0e16aa04b616c7aecaa3c33f9036292a14c19457d6372b14120f740401591

Observation 4f977e5a-79b0-4a48-ba80-7f0ccac233ef · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Direct preference optimization: Your language model is secretly a reward model

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:05:15.072188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T14:05:14.405503Z digest=sha256:1c2ff79fe48bd5e37cb4516dbab375991055df820475f670ddcfe733af9a9f63

Observation 02a2530b-e778-49ae-a070-33c3f99b6b96 · outbound

This paper cites Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parame- ters.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parame- ters

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:05:15.060526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T14:05:14.408785Z digest=sha256:6223820b09501fd3218c261ab67e9b799edfc5f978654fb5efb07668b24a3843

Observation ee7fec9c-38a3-423e-bd96-273086dfc69c · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.412579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.412579Z digest=sha256:a64f9c5351c939a0bdedafffb278954ab124a82cd3e33a27b93c31d9575adf29

Observation de7d92c9-5105-4721-ad2a-f8374590bd50 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:05:15.042113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T14:05:14.416558Z digest=sha256:f0d52177dbce8926c7361273899a607b38bcce1101da4f90afe67ec290c4f4ae

Observation 9960457e-7d6a-4517-b50a-bbd2e35a012f · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.420713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.420713Z digest=sha256:9b52454ec8e1809554040f9b2ae6d0e36a1f4c2777e970857db2bde5cae03625

Observation b44540fb-c5fc-4b17-b7b3-8a751062a856 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.424620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.424620Z digest=sha256:1f15db58f67e02b34342b9dd8eac4efdc18ce5c63d61db9fc38c396da52e2a84

Observation ba9cbdde-fc6f-4357-a5c5-a4c2309de7d8 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Gemma 2: Improving Open Language Models at a Practical Size

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.429018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.429018Z digest=sha256:48a00cc6aa2b5c49d8a7d04f9b095bb7393ee6bea770b583860f1361ff8f70ff

Observation 89f04384-5289-4c8e-bff2-5b4a4ec89be1 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.433108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.433108Z digest=sha256:6ccd3e70365ea43a025c475b44d80a552d91101e2081a37b5368cc67e4722843

Observation 03869c0e-baec-44e3-8b76-a4125807fb12 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:05:15.030824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T14:05:14.437779Z digest=sha256:e4957e98b847e30802eea749cea828422bf55d6fda7d2ab102fc77d5cecaf8c0

Observation ca2618b5-70a5-477b-9beb-1d596975b5e7 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

D-Attn: Decomposed Attention for Large Vision-and-Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.441568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.441568Z digest=sha256:7d2b01884a74b8f3b51054cfc188a356411b043d8ddcfdd2806f3ed9199e2a74

Observation 3841db45-5b6e-4a84-b4cb-5a59cc3f7c75 · outbound

This paper cites Attention is all you need.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Attention is all you need

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.445639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.445639Z digest=sha256:50cf81511141dc4a27181a432cc4027b6739400fd4d3bb66588da9d17f4f2572

Observation 9376278d-7200-45de-a595-6f5db885139e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.450118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.450118Z digest=sha256:8c0e945aa38eafa0ced935a2a04963de6e3c819de1a2564a6e6082abd2424c1c

Observation f1abfd77-5bdb-42aa-813c-93b7497d5419 · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.

D-Attn: Decomposed Attention for Large Vision-and-Language Models xgen-mm (blip-3): A family of open large multimodal models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.454269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.454269Z digest=sha256:807b4fafbb87d59950b724cc95082538492f70916607b316d6885b29939323b4

Observation 3fd189c9-7df5-44e4-9779-fc85c463905f · outbound

This paper cites Sigmoid loss for language image pre-training.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Sigmoid loss for language image pre-training

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.458648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.458648Z digest=sha256:42e15d3e4e2ac738d2a846ee636f2cdc19f7c5aea36d41724e7a8833db3095aa

Observation 5354be34-96a8-4711-963a-93cfce087c8d · outbound

This paper cites Root mean square layer nor- malization.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Root mean square layer nor- malization

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:05:15.004175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T14:05:14.463392Z digest=sha256:030cc9ae13d822ac0fe48f50c13cb0ba442f12c69a3e38922b273d144ac05798

Observation a6dbfcef-af61-40a4-b18e-79bfe137ddfb · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

D-Attn: Decomposed Attention for Large Vision-and-Language Models InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.467898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.467898Z digest=sha256:6a689156a1dd6f15a808197efbd67460cc8b6483c4ee0b6b22bed04532f8a9a2

Observation 9a114ea8-42d5-4c90-a50d-934f99259a63 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

D-Attn: Decomposed Attention for Large Vision-and-Language Models Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:05:14.992477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T14:05:14.472756Z digest=sha256:89e2fd403312c13ab740c8ee0934f57a4c57ef051f9c57a57d7d10c461ae86cf

Pith citing papers

Observation 5edf4a9f-7223-4950-9f63-3642631244a3 · inbound

RAVE: Re-Allocating Visual Attention in Large Multimodal Models cites this paper.

RAVE: Re-Allocating Visual Attention in Large Multimodal Models D-Attn: Decomposed Attention for Large Vision-and-Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:13:13.456501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T11:10:58.225078Z digest=sha256:1800fc549af1a094b1ab84d4af0b6ddefba3957bac3dc61be7a1a3caf181e5f9

Observation 6161ec78-0fa6-4416-991c-7ef33d749c76 · inbound

RAVE: Re-Allocating Visual Attention in Large Multimodal Models cites this paper.

RAVE: Re-Allocating Visual Attention in Large Multimodal Models D-Attn: Decomposed Attention for Large Vision-and-Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:55:48.340671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T18:36:56.248838Z digest=sha256:e5510df7e5c87df261e13df5517d5a3b1f5e92df6df148c2eb7e344a456e72c9