Pith. sign in

Paper Citation Record · LEDGER

Efficient Multi-modal Long Context Learning for Training-free Adaptation

As of 10 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 1 inbound Pith citation observation for arXiv:2505.19812.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19812 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:11:35.513418Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T07:23:19.428348Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T06:36:43.892355Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6c929e86-9daa-4a0a-ae4e-f983f0075e83 · outbound

This paper cites Many-Shot In-Context Learning.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Many-Shot In-Context Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:32.304232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:32.304232Z digest=sha256:08327594d742779ae4176e498d4f9b287038ccd22a9319b2fd247e0cbd5bddff

Observation 7593ca6e-b579-4dcc-81f8-f9092f82acbf · outbound

This paper cites Flamingo : A visual language model for few-shot learning.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Flamingo : A visual language model for few-shot learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:11:37.660562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:11:32.344644Z digest=sha256:0968a39ac0cc707707c6afc394b90be466769e04bd2a9c6a3016c315801d0ffc

Observation ef4c7895-a837-44e8-866d-182a5d776de0 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:32.427273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:32.427273Z digest=sha256:776ad631172c88f54a7a2343d293d17a9973fa85ccd0a98d985455380149faf4

Observation 74436313-c209-4d5f-80ea-8c01ac2c68fc · outbound

This paper cites In-Context Learning with Long-Context Models: An In-Depth Exploration.

Efficient Multi-modal Long Context Learning for Training-free Adaptation In-Context Learning with Long-Context Models: An In-Depth Exploration

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:32.503441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:32.503441Z digest=sha256:d34c3b6cbf669d63a549ffafcbf201cad26110cdaadc062dd5717f165fa3ae24

Observation 34c7919d-a2ca-4cf7-af04-7fcd4afbe359 · outbound

This paper cites an unresolved cited work.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:11:37.567201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:11:32.585907Z digest=sha256:03c1d2f5a3d697826108bdc22938292d92acecc412d5102a0f39474a691976b0

Observation 67e192e5-0c0f-4ed9-8f34-2a263735c92e · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

Efficient Multi-modal Long Context Learning for Training-free Adaptation PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:32.665501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:32.665501Z digest=sha256:cf28479ebbdaba7ecc163278adf8135adc94fc029146b696445151962ed90cd8

Observation 2167d069-857a-41eb-9072-65fe975ff2e5 · outbound

This paper cites Efficient Large Multi-modal Models via Visual Context Compression.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Efficient Large Multi-modal Models via Visual Context Compression

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:32.736571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:32.736571Z digest=sha256:c2bc2cb863aec29c7c79f78dc347d8e92fc5a4fb78b69932e7f5b676abf69c02

Observation 7bc0f75e-6125-493f-8c57-10d178b0b7b9 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:32.854096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:32.854096Z digest=sha256:6e6f0909a7088bffd5dcc2f76fa13dad438396eb4703f1c7152aa1bfd0429493

Observation 860b85bd-4838-41f7-99ef-2752919474db · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

Efficient Multi-modal Long Context Learning for Training-free Adaptation An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:11:37.423877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:11:32.943225Z digest=sha256:6d4faf5546eba1931e902536e0c3db92c0e0c10ea6419df6d8151b93a058b556

Observation 0b8b3a9a-4426-43dc-8af1-703943d46d37 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Efficient Multi-modal Long Context Learning for Training-free Adaptation How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:33.035519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:33.035519Z digest=sha256:a8fa7aae453233479f5eb088c1553eaf09728b6d7423a271f19284f9c92f032d

Observation 643dfc05-6154-42ff-b97c-a625472b495f · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

Efficient Multi-modal Long Context Learning for Training-free Adaptation SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:33.134085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:33.134085Z digest=sha256:2d0176615f7f475c6c4f48ea3a421e6ca243906fab51060d84d3d539257f2f82

Observation 15528579-1929-409d-acd8-fd534cf5f0cf · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Imagenet: A large-scale hierarchical image database

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:33.219466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:33.219466Z digest=sha256:fc045288bcfe1312781553564e5cb01738b7ca2800fdf8a0c8bad1311f8df30b

Observation e1975f8b-7e2c-40f8-b5cf-d91c1e6c1c2e · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

Efficient Multi-modal Long Context Learning for Training-free Adaptation InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:33.331132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:33.331132Z digest=sha256:6f522de9931c2a1629c0cce2951da8dc71309a8df6b8d5ebbd9cbdf63845ecdd

Observation 67dc8539-a0cb-4d7a-90c9-6d742bf8a407 · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:33.409925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:33.409925Z digest=sha256:b545702e0ae181c013c8b42f4267328c2715d457c19b7a5f051438f16782d883

Observation c1299413-da39-438a-bf0e-0c18d0dc8edd · outbound

This paper cites E2vpt: An effective and efficient approach for visual prompt tuning.

Efficient Multi-modal Long Context Learning for Training-free Adaptation E2vpt: An effective and efficient approach for visual prompt tuning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:11:37.241367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:11:33.483567Z digest=sha256:0c020038cda6f4344b2c12dd2fc1bd45141750b70b83b3caf5d70546089b0bb0

Observation 0ec7e0b0-a40f-4510-b9a3-2b647d61eced · outbound

This paper cites Facing the Elephant in the Room: Visual Prompt Tuning or Full Finetuning?.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Facing the Elephant in the Room: Visual Prompt Tuning or Full Finetuning?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:33.554951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:33.554951Z digest=sha256:efb4ee7b3ac6a06d3b4690d834dbbdd6206e9a57732838a2b7669cb6edf8d567

Observation 3c152c7e-5fd1-4c63-ba77-0bcba38ae54c · outbound

This paper cites How Well Does GPT-4V(ision) Adapt to Distribution Shifts? A Preliminary Investigation.

Efficient Multi-modal Long Context Learning for Training-free Adaptation How Well Does GPT-4V(ision) Adapt to Distribution Shifts? A Preliminary Investigation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:33.627203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:33.627203Z digest=sha256:8d35c6e52639cb145ab3faaec257385d1036a673c4a753026c3704208d0d0d70

Observation d2d172a6-b150-4e43-a188-e8811223a222 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Training Compute-Optimal Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:33.707321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:33.707321Z digest=sha256:489285998f0c27830abb3f6900ff30d705a0928d68bc1ebd10fce6669a8efddb

Observation a97cbb4a-0e38-412c-8503-cfe33ccabf46 · outbound

This paper cites J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W.

Efficient Multi-modal Long Context Learning for Training-free Adaptation J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:33.763272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:33.763272Z digest=sha256:704ff02c6240a382bd8e39509c2602f9e72a2e3e07aa4b644b93aed92942b981

Observation 519e0c92-c527-4f64-b436-a095b9845eb7 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

Efficient Multi-modal Long Context Learning for Training-free Adaptation MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:33.828506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:33.828506Z digest=sha256:5303b07a9cc988769cc36b3c61ebda8f16a3a323ad0f77d8e95823f644c1b742

Observation 6c479d0e-d8d5-402a-9232-e27a2d20978e · outbound

This paper cites Multimodal task vectors enable many-shot multimodal in-context learning.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Multimodal task vectors enable many-shot multimodal in-context learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:11:37.171478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:11:33.910662Z digest=sha256:5d272407f27ed542e0fe45cca4e2b51c4e64d4813c5e61322ae53e9143c8c06d

Observation 2418698b-fb25-4137-a8ad-8379cf3a1a2a · outbound

This paper cites K., Patra, B., et al.

Efficient Multi-modal Long Context Learning for Training-free Adaptation K., Patra, B., et al

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:11:37.072222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:11:33.954634Z digest=sha256:fb0ece22e492a3de679446ebec4260d86cbc391baea5f6c371cebad248ed9b78

Observation babc7053-0aa1-404c-9f40-dea4a049286d · outbound

This paper cites Visual prompt tuning.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Visual prompt tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:34.015730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:34.015730Z digest=sha256:02c1e595996306cc046e43928026eb1c8c5c67fad93f7343dbfa50fb7d9f32f7

Observation bb0a9261-7215-4c54-943f-05f6d289123c · outbound

This paper cites A., Wang, J.

Efficient Multi-modal Long Context Learning for Training-free Adaptation A., Wang, J

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:11:36.940675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:11:34.109443Z digest=sha256:65bb951f3c9902552f95dc5518ab705ca7407e53a6ac96ed1c3273252d2fb094

Observation 23bc6067-5560-47dc-8812-daeea5d49daa · outbound

This paper cites In-Context Learning with Many Demonstration Examples.

Efficient Multi-modal Long Context Learning for Training-free Adaptation In-Context Learning with Many Demonstration Examples

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:34.160389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:34.160389Z digest=sha256:556cf2fa6c5f0935c074a43b953ae1bfdc0176d484907e33b1b2f98bf507ca96

Observation a00b5135-1371-4ce3-bbb3-dd0acbe987e5 · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

Efficient Multi-modal Long Context Learning for Training-free Adaptation SnapKV: LLM Knows What You are Looking for Before Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:34.233628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:34.233628Z digest=sha256:00664b569b6177d5ee63aac5e95390ce4c442de7389d303920dc60f117abbdf7

Observation f9bb4932-1b06-4bd8-b864-9d21e6b712bf · outbound

This paper cites an unresolved cited work.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:11:36.840451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:11:34.287038Z digest=sha256:f5b9ac9bbc262b90157d3d0488f75edd7741f2edbb4477ff6e2cd736a3b3bce6

Observation 1d887552-a737-4a45-baac-c887b14432db · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Efficient Multi-modal Long Context Learning for Training-free Adaptation DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:34.352277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:34.352277Z digest=sha256:f0f6eedd02c13527344fbfacf97e189c3dc1a6e47700f5a3cb01a7a0dc7bde10

Observation a8e3e0eb-4a9f-4b9a-84f1-c5ae5072c6a1 · outbound

This paper cites K., and Buehler, M.

Efficient Multi-modal Long Context Learning for Training-free Adaptation K., and Buehler, M

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:11:36.761632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:11:34.398078Z digest=sha256:b90fe91549e6477fe8d06bfa4d354e76e10f06ddfb8f50e634618708175bebbb

Observation 0318de08-c0f8-4412-9937-4342799fd6f7 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:34.467761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:34.467761Z digest=sha256:ed10fa4a90c79482cb1b23028ce1bb6f2024cfc5596129283a66aea16a7bc84b

Observation 6a77dcc2-de59-46ea-b7dd-68f6fba2315f · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:34.518914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:34.518914Z digest=sha256:8012b3bf46d371bb7b779a7f3acec701eb5bd1778fe73210bfc95cec73fcc081

Observation 97260de0-4913-4b9e-9525-16a8bd43c526 · outbound

This paper cites S., Sayeed, K.

Efficient Multi-modal Long Context Learning for Training-free Adaptation S., Sayeed, K

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:11:36.610526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:11:34.552988Z digest=sha256:0ceaeeb90bd75b9bc91fe7566aa2523ce0869463ae900b09116782b685b0a24b

Observation e63ac90b-c753-42c1-9c9e-ef41c2101dea · outbound

This paper cites Locllm: Exploiting generalizable human keypoint localization via large language model.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Locllm: Exploiting generalizable human keypoint localization via large language model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:11:36.533295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:11:34.581357Z digest=sha256:0f7140845dc35c3bde13c9b6dc7b341647c7f50999acbd9ad6353bdc3ece5ac7

Observation e3704243-3f69-43bc-9b09-978e6aca2fee · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:34.642345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:34.642345Z digest=sha256:283cf37f59cf891901f900b1907c3b72fb962800cb5bb790011e08a094eb0a2b

Observation 11826797-261a-40f2-b98f-66a8f21e0a04 · outbound

This paper cites M2pt: Multimodal prompt tuning for zero-shot instruction learning.

Efficient Multi-modal Long Context Learning for Training-free Adaptation M2pt: Multimodal prompt tuning for zero-shot instruction learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:11:36.408293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:11:34.698886Z digest=sha256:f6bb45f120521ff5121e9129fda649011b69935d36cf637ee4e24ab1b2228f8e

Observation 6e0f3a88-6f01-436d-83d3-2e7ed9ae359b · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Efficient Multi-modal Long Context Learning for Training-free Adaptation CogVLM: Visual Expert for Pretrained Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:34.762926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:34.762926Z digest=sha256:da0b65769119baaafae88e525c20d2413f8f639f728cadcd497188cdc82cf462

Observation 6820257a-da9a-4614-8419-f4f9f994887d · outbound

This paper cites Efficient streaming language models with attention sinks.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Efficient streaming language models with attention sinks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:34.819044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:34.819044Z digest=sha256:1d576e9f2ff1426bddcd48333a9adf6af708a529e79cd670f02bb2d8b19522e9

Observation ebd177b5-26e7-4f03-8e4f-73a59f334cf0 · outbound

This paper cites Pink: Unveiling the power of referential comprehension for multi-modal llms.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Pink: Unveiling the power of referential comprehension for multi-modal llms

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:11:36.347206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:11:34.872853Z digest=sha256:967c30f7648b9801e9440d22ab43299670413a376a1a4970bc360a0be0f3de44

Observation 362d4ae8-9baa-4683-894e-d66714b5ba9a · outbound

This paper cites Pyramidinfer: Pyramid kv cache compression for high-throughput llm inference.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Pyramidinfer: Pyramid kv cache compression for high-throughput llm inference

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:11:36.194550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:11:34.919982Z digest=sha256:f11278c1122e7d3d0a2c834e6e5cd6875ee75f4020bc1c050fc44ff629612430

Observation 79afe392-d574-4bc9-af5a-75d37f59269f · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Yi: Open Foundation Models by 01.AI

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:34.997434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:34.997434Z digest=sha256:7780d5653a436faa405a149e1e80e6ef96a0f0eda1afd7518ddae7728e15098d

Observation 2f7a8c52-13f7-49be-926a-c10ac801f6d5 · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

Efficient Multi-modal Long Context Learning for Training-free Adaptation InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:35.075180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:35.075180Z digest=sha256:829462d1f9c0ef9354ff0a712328f5eff20bbffe5dc5b12b1d06066bf292c812

Observation 5e26cdee-dfc5-4d36-bf20-3a7f973e4bd1 · outbound

This paper cites On the Out-Of-Distribution Generalization of Multimodal Large Language Models.

Efficient Multi-modal Long Context Learning for Training-free Adaptation On the Out-Of-Distribution Generalization of Multimodal Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:35.156013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:35.156013Z digest=sha256:98e4d6c540db15d95c4be65c894575d279650f27d42c49850fc3419d67e7217c

Observation 6f215e4e-184c-4263-8b20-3192cfbb75cc · outbound

This paper cites MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?.

Efficient Multi-modal Long Context Learning for Training-free Adaptation MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:35.208090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:35.208090Z digest=sha256:785897f21a82c35c663eab79914078dcfec5e43d09823763ece75f76d55cf3cf

Observation 2b06defb-91f6-4e98-b706-31b9ce580a2d · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.

Efficient Multi-modal Long Context Learning for Training-free Adaptation H2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:35.267520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:35.267520Z digest=sha256:3b19cf0622eb668d35d0e792eda3d83f5bdfcf5958cabebba7b50d772c860232

Observation 8ff75235-870d-43e0-aa6b-fe336c89a79f · outbound

This paper cites Towards automatic learning of procedures from web instructional videos.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Towards automatic learning of procedures from web instructional videos

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:11:36.089663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:11:35.326708Z digest=sha256:05f6b58d69bab64c32db8e10a9cee56394456ef1e162a1681c64e646ef5f675b

Observation f215767b-ef12-4d92-b20f-2bd45afa47d4 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Efficient Multi-modal Long Context Learning for Training-free Adaptation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:35.385052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:35.385052Z digest=sha256:53fbdf9f299f3523d18f75761b324408c1da3869cb0bcb31940622c6aee298e6

Observation 023b0629-f762-439d-8158-f004e14b5aad · outbound

This paper cites Mini GPT -4: Enhancing vision-language understanding with advanced large language models.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Mini GPT -4: Enhancing vision-language understanding with advanced large language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:11:35.849194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:11:35.454571Z digest=sha256:3593498fb7fb52f461dd3c98e5dfba11ad9d521de4eeb1a37b74ce0ed9517364

Observation 82575eb0-9963-44bb-be2d-ef1d287bce29 · outbound

This paper cites write newline.

Efficient Multi-modal Long Context Learning for Training-free Adaptation write newline

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:35.513418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:35.513418Z digest=sha256:89174f0be314e75caf3c6ea7f01d7ecfc17a1026670765de1ea6491526b7fb1c

Pith citing papers

Observation cd89840e-8c0f-42d1-ae9c-b9b74dc55f22 · inbound

Hyper-ICL: Attention Calibration with Hyperbolic Anchor Distillation for Multimodal In-Context Learning cites this paper.

Hyper-ICL: Attention Calibration with Hyperbolic Anchor Distillation for Multimodal In-Context Learning Efficient Multi-modal Long Context Learning for Training-free Adaptation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:36:43.893801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T07:23:19.428348Z digest=sha256:688b8896f6219cc382441b94f97bccf602d32a1da322446f16c62fa955e605a1