Pith. sign in

Paper Citation Record · LEDGER

Efficient Multi-modal Long Context Learning for Training-free Adaptation

As of 19 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 1 inbound Pith citation observation for arXiv:2505.19812.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19812 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:11:35.513418Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T07:23:19.428348Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T06:36:43.892355Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6c929e86-9daa-4a0a-ae4e-f983f0075e83 · outbound

This paper cites Many-Shot In-Context Learning.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Many-Shot In-Context Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:32.304232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:32.304232Z digest=sha256:d6f8b18d65b994a7e35ab558ad5c307420f57952b8ef59071271509281e294ee

Observation 7593ca6e-b579-4dcc-81f8-f9092f82acbf · outbound

This paper cites Flamingo : A visual language model for few-shot learning.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Flamingo : A visual language model for few-shot learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:11:37.660562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:11:32.344644Z digest=sha256:c827b5a9d4953057ffa3d12399f3c534c5029a31752918d1d655c38239a329a4

Observation ef4c7895-a837-44e8-866d-182a5d776de0 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:32.427273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:32.427273Z digest=sha256:4e93438f8480d70b126610d4bb4ab97ba483a2bbe29eeb95ef5f7bbe65523beb

Observation 74436313-c209-4d5f-80ea-8c01ac2c68fc · outbound

This paper cites In-Context Learning with Long-Context Models: An In-Depth Exploration.

Efficient Multi-modal Long Context Learning for Training-free Adaptation In-Context Learning with Long-Context Models: An In-Depth Exploration

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:32.503441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:32.503441Z digest=sha256:453a349e5ddcc63f7d010beee6915efb85f679ee0d68cdba012abcad248595d0

Observation 34c7919d-a2ca-4cf7-af04-7fcd4afbe359 · outbound

This paper cites an unresolved cited work.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:11:37.567201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:11:32.585907Z digest=sha256:c20291cd9e990cb6df8d4bd3035da3ab9e84dc8db520317be4b79d8c5468a5d6

Observation 67e192e5-0c0f-4ed9-8f34-2a263735c92e · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

Efficient Multi-modal Long Context Learning for Training-free Adaptation PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:32.665501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:32.665501Z digest=sha256:719e8ea3b780427774bf048b7d88672767bbce7a7318a97f51aa79f114049127

Observation 2167d069-857a-41eb-9072-65fe975ff2e5 · outbound

This paper cites Efficient Large Multi-modal Models via Visual Context Compression.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Efficient Large Multi-modal Models via Visual Context Compression

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:32.736571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:32.736571Z digest=sha256:1b813a75f2a1f215ed9f4c274bb84a46a8cf9c053eee0faf2e473e7b703eddca

Observation 7bc0f75e-6125-493f-8c57-10d178b0b7b9 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:32.854096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:32.854096Z digest=sha256:4c8efbffc3fd075c4e1738b2880d179a444592c3696df440e11da797d3b9e2be

Observation 860b85bd-4838-41f7-99ef-2752919474db · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

Efficient Multi-modal Long Context Learning for Training-free Adaptation An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:11:37.423877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:11:32.943225Z digest=sha256:5a0955210788f37273d736835c301dbc6cd65ba89c51e8ec563ba6d66c6ffdc0

Observation 0b8b3a9a-4426-43dc-8af1-703943d46d37 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Efficient Multi-modal Long Context Learning for Training-free Adaptation How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:33.035519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:33.035519Z digest=sha256:509b1cf7bda7c23fe08559a83a73b8750aed35f0222f0366bb0bcd964a2c4fd3

Observation 643dfc05-6154-42ff-b97c-a625472b495f · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

Efficient Multi-modal Long Context Learning for Training-free Adaptation SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:33.134085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:33.134085Z digest=sha256:ba3ee07051744bc9465ab68914233e309efcee69e382fec3bd6e6474cc79fec7

Observation 15528579-1929-409d-acd8-fd534cf5f0cf · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Imagenet: A large-scale hierarchical image database

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:33.219466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:33.219466Z digest=sha256:9f3acdafe04d213e4b34a91588cf5730482bf4ea53d231515767d93164b76288

Observation e1975f8b-7e2c-40f8-b5cf-d91c1e6c1c2e · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

Efficient Multi-modal Long Context Learning for Training-free Adaptation InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:33.331132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:33.331132Z digest=sha256:5dc8a5fc16912f6c578a062feefd6bfbd10974ae921d5ac509a70a52121fbd4e

Observation 67dc8539-a0cb-4d7a-90c9-6d742bf8a407 · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:33.409925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:33.409925Z digest=sha256:7c9c407b6f261f57745dd955fad4ed2a7ba9a2dc49b3b85ef8151f0ad5537ae6

Observation c1299413-da39-438a-bf0e-0c18d0dc8edd · outbound

This paper cites E2vpt: An effective and efficient approach for visual prompt tuning.

Efficient Multi-modal Long Context Learning for Training-free Adaptation E2vpt: An effective and efficient approach for visual prompt tuning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:11:37.241367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:11:33.483567Z digest=sha256:935b80f273d876b0e95d4c521b6eba9010a2102136aee23e98757908b09de92c

Observation 0ec7e0b0-a40f-4510-b9a3-2b647d61eced · outbound

This paper cites Facing the Elephant in the Room: Visual Prompt Tuning or Full Finetuning?.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Facing the Elephant in the Room: Visual Prompt Tuning or Full Finetuning?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:33.554951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:33.554951Z digest=sha256:7fe82da66db38835e1cdb7ebee971c97ccbb0617963edaf9842afeed4f542ad7

Observation 3c152c7e-5fd1-4c63-ba77-0bcba38ae54c · outbound

This paper cites How Well Does GPT-4V(ision) Adapt to Distribution Shifts? A Preliminary Investigation.

Efficient Multi-modal Long Context Learning for Training-free Adaptation How Well Does GPT-4V(ision) Adapt to Distribution Shifts? A Preliminary Investigation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:33.627203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:33.627203Z digest=sha256:760433063b612a450ce7b6552814a1808e6cc0352056c62eb0b30c252dfcb62d

Observation d2d172a6-b150-4e43-a188-e8811223a222 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Training Compute-Optimal Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:33.707321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:33.707321Z digest=sha256:3b8369fcbadbe5b39b2d7c2960341523ad73e906f7532fb8f4d954febc83da4d

Observation a97cbb4a-0e38-412c-8503-cfe33ccabf46 · outbound

This paper cites J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W.

Efficient Multi-modal Long Context Learning for Training-free Adaptation J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:33.763272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:33.763272Z digest=sha256:082a8a614b896c5956086f2fa8cc562c6c2d8af61f5658b285aace34ec329e0f

Observation 519e0c92-c527-4f64-b436-a095b9845eb7 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

Efficient Multi-modal Long Context Learning for Training-free Adaptation MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:33.828506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:33.828506Z digest=sha256:6f3c80df90547a3813268eb97f8ba912aeee34628e8ac81af15a2231a57dab84

Observation 6c479d0e-d8d5-402a-9232-e27a2d20978e · outbound

This paper cites Multimodal task vectors enable many-shot multimodal in-context learning.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Multimodal task vectors enable many-shot multimodal in-context learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:11:37.171478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:11:33.910662Z digest=sha256:3eedb022938f7a2633f6941298117cd4466ffa79d67c663539f65a3da97b862a

Observation 2418698b-fb25-4137-a8ad-8379cf3a1a2a · outbound

This paper cites K., Patra, B., et al.

Efficient Multi-modal Long Context Learning for Training-free Adaptation K., Patra, B., et al

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:11:37.072222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:11:33.954634Z digest=sha256:8ffd52dcb36f207c71aedade45ad16c35682051bf0018982cc9a58e34d05ba92

Observation babc7053-0aa1-404c-9f40-dea4a049286d · outbound

This paper cites Visual prompt tuning.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Visual prompt tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:34.015730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:34.015730Z digest=sha256:1f71ca95cbf35de2e935df516ad26d80d8c8dd609c3b69ecd534e1f8395bc0bb

Observation bb0a9261-7215-4c54-943f-05f6d289123c · outbound

This paper cites A., Wang, J.

Efficient Multi-modal Long Context Learning for Training-free Adaptation A., Wang, J

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:11:36.940675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:11:34.109443Z digest=sha256:17d60810adcfb872cbd87711c27233b5d19a62033e2a267ef96ff1111956d411

Observation 23bc6067-5560-47dc-8812-daeea5d49daa · outbound

This paper cites In-Context Learning with Many Demonstration Examples.

Efficient Multi-modal Long Context Learning for Training-free Adaptation In-Context Learning with Many Demonstration Examples

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:34.160389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:34.160389Z digest=sha256:443aba6bc823984152bb1682e73b555602c79ffc98ac1f5ef62a335ac5f910c5

Observation a00b5135-1371-4ce3-bbb3-dd0acbe987e5 · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

Efficient Multi-modal Long Context Learning for Training-free Adaptation SnapKV: LLM Knows What You are Looking for Before Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:34.233628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:34.233628Z digest=sha256:4e7ac13b2f72e964e1e42b434d84d3fa84adf935169f11eb7868709547b1d0db

Observation f9bb4932-1b06-4bd8-b864-9d21e6b712bf · outbound

This paper cites an unresolved cited work.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:11:36.840451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:11:34.287038Z digest=sha256:04793ceda9e96b986139819368f17a3486010fd41af520cb7aa681ad32b401d0

Observation 1d887552-a737-4a45-baac-c887b14432db · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Efficient Multi-modal Long Context Learning for Training-free Adaptation DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:34.352277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:34.352277Z digest=sha256:78bafe86440303b3d1c4d713b87dc38c28c912868f54f7165ed8f37978181475

Observation a8e3e0eb-4a9f-4b9a-84f1-c5ae5072c6a1 · outbound

This paper cites K., and Buehler, M.

Efficient Multi-modal Long Context Learning for Training-free Adaptation K., and Buehler, M

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:11:36.761632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:11:34.398078Z digest=sha256:5811eff321d1592dd5c7ff5c30941116268f57c7abb9123ede479ed6c5c25406

Observation 0318de08-c0f8-4412-9937-4342799fd6f7 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:34.467761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:34.467761Z digest=sha256:39aa4af650f9e50e6979bc9536cec5125aaa5e8a9ea3111b5cb4330f1943177b

Observation 6a77dcc2-de59-46ea-b7dd-68f6fba2315f · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:34.518914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:34.518914Z digest=sha256:17d07523af4dfd40123416cd1ffac53cab85908e06cc2959880ad564936a7624

Observation 97260de0-4913-4b9e-9525-16a8bd43c526 · outbound

This paper cites S., Sayeed, K.

Efficient Multi-modal Long Context Learning for Training-free Adaptation S., Sayeed, K

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:11:36.610526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:11:34.552988Z digest=sha256:c7da677c6f604a75651c9e7481c61a1b75833b2cec20492c9ed00eb5e01d97b4

Observation e63ac90b-c753-42c1-9c9e-ef41c2101dea · outbound

This paper cites Locllm: Exploiting generalizable human keypoint localization via large language model.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Locllm: Exploiting generalizable human keypoint localization via large language model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:11:36.533295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:11:34.581357Z digest=sha256:eacbacb19baa8e069b588e273301381bcf8d065709a1c1f4b830a19115bfa5c8

Observation e3704243-3f69-43bc-9b09-978e6aca2fee · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:34.642345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:34.642345Z digest=sha256:5a2b58fe453993484789bd742e3da03af73a801ca55b528e3069c26a0de7a2b4

Observation 11826797-261a-40f2-b98f-66a8f21e0a04 · outbound

This paper cites M2pt: Multimodal prompt tuning for zero-shot instruction learning.

Efficient Multi-modal Long Context Learning for Training-free Adaptation M2pt: Multimodal prompt tuning for zero-shot instruction learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:11:36.408293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:11:34.698886Z digest=sha256:9c20a286efe426bf9cff5f72db2e3400834d1988178ad71b97ac7f73d6488b26

Observation 6e0f3a88-6f01-436d-83d3-2e7ed9ae359b · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Efficient Multi-modal Long Context Learning for Training-free Adaptation CogVLM: Visual Expert for Pretrained Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:34.762926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:34.762926Z digest=sha256:732cf6d1014be997fba306949f590403917fb761f19dbf12c7a372f9182deccd

Observation 6820257a-da9a-4614-8419-f4f9f994887d · outbound

This paper cites Efficient streaming language models with attention sinks.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Efficient streaming language models with attention sinks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:34.819044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:34.819044Z digest=sha256:9d7ddd38e114d1ccdd5498feb4c10ac6ac376c10306ec1e533e7e78612a3b764

Observation ebd177b5-26e7-4f03-8e4f-73a59f334cf0 · outbound

This paper cites Pink: Unveiling the power of referential comprehension for multi-modal llms.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Pink: Unveiling the power of referential comprehension for multi-modal llms

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:11:36.347206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:11:34.872853Z digest=sha256:0c0f577b261cefeca96457378eeb6083e51ce025246f4768dc187fad03da82af

Observation 362d4ae8-9baa-4683-894e-d66714b5ba9a · outbound

This paper cites Pyramidinfer: Pyramid kv cache compression for high-throughput llm inference.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Pyramidinfer: Pyramid kv cache compression for high-throughput llm inference

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:11:36.194550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:11:34.919982Z digest=sha256:75e6c4ef33d59c127aeec67d5b809fe66787b2a1a3eb546cc756f1aa188c61db

Observation 79afe392-d574-4bc9-af5a-75d37f59269f · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Yi: Open Foundation Models by 01.AI

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:34.997434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:34.997434Z digest=sha256:82b7bec3dd4e4c3821c1fcd919196ca586996dc7cbd0f642a78d72668d9c956f

Observation 2f7a8c52-13f7-49be-926a-c10ac801f6d5 · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

Efficient Multi-modal Long Context Learning for Training-free Adaptation InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:35.075180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:35.075180Z digest=sha256:8377d6a96e5942e6d7cd9a2553f010454c012903955e3bc72839d5adb8d9a3cc

Observation 5e26cdee-dfc5-4d36-bf20-3a7f973e4bd1 · outbound

This paper cites On the Out-Of-Distribution Generalization of Multimodal Large Language Models.

Efficient Multi-modal Long Context Learning for Training-free Adaptation On the Out-Of-Distribution Generalization of Multimodal Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:35.156013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:35.156013Z digest=sha256:edcdabd26ad8f9cb81c96151827e69517577a64289280c5c6c7379efe9960627

Observation 6f215e4e-184c-4263-8b20-3192cfbb75cc · outbound

This paper cites MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?.

Efficient Multi-modal Long Context Learning for Training-free Adaptation MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:35.208090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:35.208090Z digest=sha256:741e7509b79d70bd2b8a9c27d9305f0690a085ce2bbb050313b535176d6fabb5

Observation 2b06defb-91f6-4e98-b706-31b9ce580a2d · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.

Efficient Multi-modal Long Context Learning for Training-free Adaptation H2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:35.267520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:35.267520Z digest=sha256:6240156bfe026cb75b82879b0e92ee837ce3989e7a4f25bb5d98d0a898767a4f

Observation 8ff75235-870d-43e0-aa6b-fe336c89a79f · outbound

This paper cites Towards automatic learning of procedures from web instructional videos.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Towards automatic learning of procedures from web instructional videos

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:11:36.089663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:11:35.326708Z digest=sha256:c983aee41bed1737e419f79b83e8694050f9f38e982cfa1f5c25c59d5fc96456

Observation f215767b-ef12-4d92-b20f-2bd45afa47d4 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Efficient Multi-modal Long Context Learning for Training-free Adaptation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:35.385052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:35.385052Z digest=sha256:ea03c07dab22d3c846b422caa6afb6681d11cbdfff5d79f08b85e2b8caa04b07

Observation 023b0629-f762-439d-8158-f004e14b5aad · outbound

This paper cites Mini GPT -4: Enhancing vision-language understanding with advanced large language models.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Mini GPT -4: Enhancing vision-language understanding with advanced large language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:11:35.849194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:11:35.454571Z digest=sha256:cb7f1f22395e8ec042c622c2a1e5f339692769538aed1ad87aff6b228d163d00

Observation 82575eb0-9963-44bb-be2d-ef1d287bce29 · outbound

This paper cites write newline.

Efficient Multi-modal Long Context Learning for Training-free Adaptation write newline

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:35.513418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:35.513418Z digest=sha256:3261d9098e2c7391d3e31d1cc6eef9b8e3fbc74ee7064664fefeab84243bf38d

Pith citing papers

Observation cd89840e-8c0f-42d1-ae9c-b9b74dc55f22 · inbound

Hyper-ICL: Attention Calibration with Hyperbolic Anchor Distillation for Multimodal In-Context Learning cites this paper.

Hyper-ICL: Attention Calibration with Hyperbolic Anchor Distillation for Multimodal In-Context Learning Efficient Multi-modal Long Context Learning for Training-free Adaptation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:36:43.893801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T07:23:19.428348Z digest=sha256:8f881c118f5dc60a8b86ecc4aa11aad830bf2800dfb56b4857fda7026dcb6d13