Pith. sign in

Paper Citation Record · LEDGER

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models

As of 10 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2502.02406.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02406 v3

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T12:23:40.585910Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 23cbddc7-6441-4224-8f25-fc14a084dbc9 · outbound

This paper cites write newline.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.403541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.403541Z digest=sha256:18231c459d540b47efdde9d00eeeb33b32b6e2c0d8cb6f22a50d3e91c8f9ff2f

Observation e68b4704-9172-4ff5-a8f0-b34c1ac43dd8 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.294271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.409366Z digest=sha256:de349e6c57ded95113855895eadbf01c2c609ed84b00015649a1eca0e2d1cc3b

Observation b4c60257-454e-4756-866d-3a71467c3e66 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.414321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.414321Z digest=sha256:9cfe36c18a9ea0f7f909460baba7cacaaf8be222f9c5726de23339eb421e4bfc

Observation aebbd7d9-714b-4fd8-9d13-9663f9ba0898 · outbound

This paper cites Longformer: The Long-Document Transformer.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Longformer: The Long-Document Transformer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.420888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.420888Z digest=sha256:d8d53eb72c58f1c7855ed67dedcc38add2b09d485972950ca6355a3a8a5cb8f7

Observation 5318c903-e261-442d-ae39-3add2963c798 · outbound

This paper cites an unresolved cited work.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-09T12:23:41.283500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.425814Z digest=sha256:a6ca47db9b6150be01c78f2e0f4840e67c05eaf8af567c71f3c8e5f9d6805b0c

Observation 441df307-bd25-4c85-acb1-14b589bc797c · outbound

This paper cites Striped Attention: Faster Ring Attention for Causal Transformers.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Striped Attention: Faster Ring Attention for Causal Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.430129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.430129Z digest=sha256:4f303fe14a53c937519d22c859662b15815296f66785af7fdfc945efde1559cb

Observation 2e4acc3f-d938-4512-af51-bd0dea97e9e4 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.434769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.434769Z digest=sha256:784c26273bf3f2f2371deb6dbb4e9bce711ac6687e3b9ea6f5aa2004d79f71c2

Observation 9df26055-1cd8-4052-94ac-2ce8b2a4c3f7 · outbound

This paper cites Adapting language models to compress contexts.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Adapting language models to compress contexts

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.266075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.439586Z digest=sha256:0de3dded7149defe768f22644ff44abf9516dc17d6513e771fee1024c0c9493c

Observation 7c55aa77-e64d-4ed6-afcb-cd2196b2a9dd · outbound

This paper cites M., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models M., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.443931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.443931Z digest=sha256:8bba12ad7e3e6e7d04b85f3f3810c13ea7bdeca2961cadb913bd85821173821f

Observation be44c55a-50ac-41c7-b535-6ee346f5d340 · outbound

This paper cites Nvlm: Open frontier-class multimodal llms.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Nvlm: Open frontier-class multimodal llms

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.248285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.448162Z digest=sha256:73c02bf3212d4e31dd16f4ebfadfce72ed1e535ca1b2d3ea0d858bc2a0f0b06a

Observation 55554418-4064-4083-ad50-9ffaf6b3814e · outbound

This paper cites Flash A ttention-2: Faster attention with better parallelism and work partitioning.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Flash A ttention-2: Faster attention with better parallelism and work partitioning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.452814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.452814Z digest=sha256:d633c526e795ce2e64bc187b0473898d88dcf107a82a64702e0b947c7d32c0d9

Observation cb45b3b9-ddf3-4fa4-8e70-4bd743cb4629 · outbound

This paper cites Y., Ermon, S., Rudra, A., and R \'e , C.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Y., Ermon, S., Rudra, A., and R \'e , C

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.456933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.456933Z digest=sha256:4f5afacc9ff57eecb69d578f9a77ea34f7188d3e465190df4001dc6cd3a57904

Observation 94261df9-98b0-48d8-bc7a-1e5e3ce7da21 · outbound

This paper cites S., Monga, R., Chen, K., Devin, M., Le, Q.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models S., Monga, R., Chen, K., Devin, M., Le, Q

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.224522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.460640Z digest=sha256:09e6633c2936fe34b5d3414e4556d0c19c323d767f08f8e2845d2ef1ff03d978

Observation aedfadff-b42b-4e90-9a3d-fa558079e64d · outbound

This paper cites LongNet: Scaling Transformers to 1,000,000,000 Tokens.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models LongNet: Scaling Transformers to 1,000,000,000 Tokens

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.464868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.464868Z digest=sha256:e0f30135bc931b1da658a45deebdc54fa36e32a51ab1feb651b30b6928ef4cde

Observation d60cc123-5f52-4e57-800f-b4b14048891c · outbound

This paper cites The design and operation of CloudLab.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models The design and operation of CloudLab

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.213838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.469166Z digest=sha256:beeda30135930a60ea8c9a2c655876eda4fcba8948eb686cae0b58aaed262f16

Observation 3224933a-38ff-426d-b9eb-2025c44d3273 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.472983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.472983Z digest=sha256:3e0007955553c9c5eeaf15ba93334e77c42267507ac9858b6d572e0b0c20bfb8

Observation c7b5788d-0ecb-4cae-808c-301145995424 · outbound

This paper cites The Llama 3 Herd of Models.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models The Llama 3 Herd of Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.477312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.477312Z digest=sha256:7c418e28d04a19fd1cbf9860ba9845cb07f5f831bbeab85757ca89b6b7208a48

Observation eeeeecb0-6ee5-478b-bb91-1eac35ee2559 · outbound

This paper cites Llava-uhd: An lmm perceiving any aspect ratio and high-resolution images.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Llava-uhd: An lmm perceiving any aspect ratio and high-resolution images

Reference 18

Resolution
verified exact
doi, observed 2026-08-09T12:23:40.634688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.481945Z digest=sha256:75961ae7db9432e9b2646fc0f27af4120757fee149655058612f6f79b7de4cc8

Observation 1d2b86f3-3ce3-49ee-bac1-a38a8925cbb3 · outbound

This paper cites K., Jia, M., Cao, X., Shah, A., Shrivastava, A., and Lim, S.-N.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models K., Jia, M., Cao, X., Shah, A., Shrivastava, A., and Lim, S.-N

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.203438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.485838Z digest=sha256:30c382c7038f4946ebb4a26fb01c867c6c08c960253c3b418497d68888ab25c3

Observation e1ddd6bc-bfa5-4fc8-9098-56725910e111 · outbound

This paper cites Video ReCap: Recursive Captioning of Hour-Long Videos.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Video ReCap: Recursive Captioning of Hour-Long Videos

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.491200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.491200Z digest=sha256:98d1f1f34cb146ea4b1fc523377e0bc6fae19f678382ee2e4da7cc8d9bab8694

Observation 63d306a6-146e-47f2-b971-f1a38d53bb57 · outbound

This paper cites A., Tanaka, M., Zhang, C., Zhang, M., Aminadabi, R.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models A., Tanaka, M., Zhang, C., Zhang, M., Aminadabi, R

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.495082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.495082Z digest=sha256:74a13e5e85c8e24853eed4684a685180f532526668606b6f9606b3d41a0ec35b

Observation 9ee31d37-2d85-4086-9cd2-7a58eb10fbae · outbound

This paper cites Reformer: The efficient transformer.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Reformer: The efficient transformer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.498688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.498688Z digest=sha256:815aa2e56e837a5a7910a26e3456233b5823ae352cd3001acc504ae659ec2bd3

Observation dd60d2e3-059a-452a-a70d-ae25d0daed60 · outbound

This paper cites A., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models A., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.186213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.502235Z digest=sha256:da8f388df1d88bbd0c933067860dbbba4ef4d6c468468d96ae8b2600a3cb1470

Observation 7f6189eb-792e-44ad-a149-da35816a6ff4 · outbound

This paper cites A., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models A., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.175383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.505912Z digest=sha256:679293cd434e01f2762ebcfece2a55ca5872d162dd8cb01abd7a11bf36f45450

Observation 83f57cc7-d32f-477a-9dda-d00ed98de030 · outbound

This paper cites M., Kiela, D., Cord, M., and Sanh, V.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models M., Kiela, D., Cord, M., and Sanh, V

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.164568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.509643Z digest=sha256:2b4c2ad429bab23675ef6d998fd545f7f45df7b13e4c0bf9b4bbe8d617d8946b

Observation 2460e5a7-2fbb-46b3-a919-30a1280278e6 · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.513418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.513418Z digest=sha256:1d2f9335ba852af01d4db7d9db098da93a428e0fd630f79458dfe22991788976

Observation d60026c8-84b5-428f-9b58-d1c9ec52e47d · outbound

This paper cites P., Ma, X., Stoica, I., Gonzalez, J.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models P., Ma, X., Stoica, I., Gonzalez, J

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.153717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.517456Z digest=sha256:cd9efb56e0ef294c084908b0f0855442a8222692ca3ef5490d0a8563e57118f5

Observation be1f82a0-b4b0-4d61-bfd9-1e27e0cb9e6d · outbound

This paper cites Blip-2: bootstrapping language-image pre-training with frozen image encoders and large language models.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Blip-2: bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.142932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.521288Z digest=sha256:cd68ff0da95040e62b9af938389c8ad263dd935102e768366cb288d101991b4c

Observation fd1e80ac-e63a-4310-a9a9-39bf3e74e556 · outbound

This paper cites an unresolved cited work.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-09T12:23:41.132089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.525030Z digest=sha256:c01ea191aca348ff9600c60bf633027e514b6518a41de19f47030f9e5b1c9c9f

Observation d7881f19-2f4d-4c98-a88c-a836db4bedd2 · outbound

This paper cites Ringattention with blockwise transformers for near-infinite context.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Ringattention with blockwise transformers for near-infinite context

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.121090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.528809Z digest=sha256:ea7a5cf01f22b1081850c6142b26ed33f1562fae14597233902653bee5a52e83

Observation e6297378-2b6d-428a-b06c-488d05e77eee · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models NVILA: Efficient Frontier Visual Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.532423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.532423Z digest=sha256:e4db9bf56cb8e3f354c05e02c6c481c3f1a1c7ccf7e3abd455413885f291b1ae

Observation a8989cd5-06b8-40c3-b946-9947e56522ca · outbound

This paper cites Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.536435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.536435Z digest=sha256:6b67763882c51731376c188e6d7d87001407b02ef79be2c53c5dfa3d8c95363f

Observation b5554898-8713-4103-9a92-9742951063dc · outbound

This paper cites R., Ganger, G.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models R., Ganger, G

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.540566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.540566Z digest=sha256:a3c49fd404af976570cf06e50cbe704e60a2b8c519d6aeb28071a7cc0b9e3560

Observation 74ddb37f-4cc6-4e0e-9864-635d472ea920 · outbound

This paper cites Momentor: advancing video large language model with fine-grained temporal reasoning.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Momentor: advancing video large language model with fine-grained temporal reasoning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.110012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.544509Z digest=sha256:f7cf48e271b81bb4807379e70b05ed42181b6e9b0f890903b5104d6f371dc4f7

Observation db78e0b2-276a-4709-99f1-bd534e6608c5 · outbound

This paper cites ModServe : Scalable and resource-efficient large multimodal model serving, 2025.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models ModServe : Scalable and resource-efficient large multimodal model serving, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.548348Z digest=sha256:cac55d092ed2add654b9f0bc47390cfd6a707e50da6864121c530084a793f0b6

Observation 06db48e3-66b4-4590-8a0d-9129585b9bbc · outbound

This paper cites Zero: memory optimizations toward training trillion parameter models.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Zero: memory optimizations toward training trillion parameter models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.098690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.552540Z digest=sha256:32cc90788af37910e0e1a1798e105a7439004a4401340eb3bc8c357baa19f885

Observation 379d8cdf-8b39-4175-8596-df4581f53c59 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.556139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.556139Z digest=sha256:a5466f87ff9de0b1e3557d60ec76bfda337a880eb4426d5b80569119641cbb0a

Observation ba04c592-e6b8-4068-9537-dc810042bf83 · outbound

This paper cites Repository-level prompt generation for large language models of code.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Repository-level prompt generation for large language models of code

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.087584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.560141Z digest=sha256:dd0568e21011d0186cbc1cae9c1922bcfe1b4f80b99fc14faaa61a2d72de5058

Observation 6895e989-0ca5-4de7-a58c-b953c8279411 · outbound

This paper cites PEARL : Prompting large language models to plan and execute actions over long documents.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models PEARL : Prompting large language models to plan and execute actions over long documents

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.075061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.563745Z digest=sha256:7e25eb025cbed64ae9e42e5490bfef0754181083490b36fdf1d58ab495cb0e05

Observation a06ddf6a-335b-4569-b1f9-33d6e3a8227d · outbound

This paper cites T., and Cox, D.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models T., and Cox, D

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.567345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.567345Z digest=sha256:4f936a5bf0624d438a5ddc190db932ab0ff55d1c01d2ae89d7316663fa92c8a6

Observation 835f1694-a240-4085-bdad-6310f2cb8dbf · outbound

This paper cites N., Kaiser, L., and Polosukhin, I.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models N., Kaiser, L., and Polosukhin, I

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.570913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.570913Z digest=sha256:84d2f17afd122d6e3c7eae8eb716ba03e38490bbc1e8653082d3b5cf38b8aeef

Observation c1dc54ac-ab78-4c1f-acee-58b39e7581c2 · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.574541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.574541Z digest=sha256:74f063b882848979bebf2332b954d7134f0035c063cfe8ecd38ebb056bcc30c6

Observation 5c79bfde-886f-459b-bfcf-e69046163793 · outbound

This paper cites Big bird: transformers for longer sequences.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Big bird: transformers for longer sequences

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.057941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.578333Z digest=sha256:ac96950950953f924ae856764ca731c34915a53ee12c8c8b28e12f680fe7a6c0

Observation 5bd7e58e-7aff-40c9-ac3c-fb150b07ea85 · outbound

This paper cites R epo C oder: Repository-level code completion through iterative retrieval and generation.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models R epo C oder: Repository-level code completion through iterative retrieval and generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.582047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.582047Z digest=sha256:7d57184c42d691975f60fb063649ad1a8f515657a501b9558dd98bda21e8fccf

Observation da7a0ffc-6d1b-4c9c-9f65-a080c283e80c · outbound

This paper cites Mini GPT -4: Enhancing vision-language understanding with advanced large language models.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Mini GPT -4: Enhancing vision-language understanding with advanced large language models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.585910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.585910Z digest=sha256:72e20d007bc07523bc756ce6e7a7d92e117669bd2a9bf19a8865b92596782f93

Pith citing papers

No inbound Pith citation observations are available.