Pith. sign in

Paper Citation Record · LEDGER

Clapper: Compact Learning and Video Representation in VLMs

As of 14 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2505.15529.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15529 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:20:49.012698Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8f111aa1-f365-4e08-a008-587d8d0b50db · outbound

This paper cites Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Kar \' e n Simonyan.

Clapper: Compact Learning and Video Representation in VLMs Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Kar \' e n Simonyan

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:20:51.732946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T15:20:44.435871Z digest=sha256:979732308da5f0823f93e0222d08de189cfaeb2cb6a0c93a3f3ac5ce39ea369b

Observation 4757746a-bce9-4cbc-b8e7-4def18f07d21 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:20:51.581798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T15:20:44.511993Z digest=sha256:0a5df7da60a17f64293479a272542ed5f81ef2a43ed72ff8d7664bc1377df5c0

Observation 7e5aaeaf-d987-421e-b41d-72332f633ec6 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:20:51.407479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T15:20:44.586369Z digest=sha256:7ca0f2a2f5b4aa4c82468368e29655998666a021040240778c69cefb6bdde455

Observation aeb95aec-7c5b-48a8-9143-5c9c25861d5b · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Clapper: Compact Learning and Video Representation in VLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:44.667515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:44.667515Z digest=sha256:815545e0e6a28aaef1760ae1ca091ef3c8a9570cd5d2effe4ea382733eb12f98

Observation 89c9b623-186a-4330-be78-cabcf333b42b · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Clapper: Compact Learning and Video Representation in VLMs How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:44.764951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:44.764951Z digest=sha256:af6fc801751d97dcb3b25fc28903a9232bc60ab2f57bea2833f725cc9cd8b37d

Observation 042da8ab-4e83-4132-bc45-8a49da6f3094 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Clapper: Compact Learning and Video Representation in VLMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:44.871031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:44.871031Z digest=sha256:a197c1fa750e58e66be7f81904cd883c0d65d2db42f191f3c30236395076f227

Observation 9df7712b-7efb-4210-95f3-c02909292127 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:20:51.190608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T15:20:44.997198Z digest=sha256:496dfe3db628d0a7c6027e6cb67d4510a0866248f92dac248db30a6732507790

Observation 6537dfb7-4d62-42b2-8b9e-152561038d2e · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Clapper: Compact Learning and Video Representation in VLMs Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:45.092112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:45.092112Z digest=sha256:c90d611a1264b46cd162dafefd69c1eb4475eb51a4a7c687010edbc54c9a7591

Observation 32001848-f00c-4b73-b278-eb58c1f57253 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Clapper: Compact Learning and Video Representation in VLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:45.191691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:45.191691Z digest=sha256:1d1a382feb4c3e7aac7b8ddcacf921c20517cfee4805ecbcaafce1041d07b475

Observation fc5946e2-dd72-48a1-a1bc-98da3fea1e84 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:45.283481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:45.283481Z digest=sha256:cd28bb776f38efc4a91c60d433e7042354cdaffc33f59d98b7eaeb7e856555f5

Observation 568b2f86-13e7-4fb3-bedd-e721c409f31d · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:20:51.011985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T15:20:45.401124Z digest=sha256:575c31158aabdf57a07d452e17478757572618edb31beaff8167a3d1f03fb5c7

Observation 6dcc8ae6-4df9-4b87-a376-5dbf66e45257 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:20:50.816531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T15:20:45.516975Z digest=sha256:4c2a5885d29680649934d3d100065c48b4d748d9dae494de28e6bc1adb530add

Observation 0833c6ca-2656-44aa-93b4-173aa01292d6 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Clapper: Compact Learning and Video Representation in VLMs Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:45.618643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:45.618643Z digest=sha256:31cd9a236d8084c6ba11b276c8f5cfd9a053ffab96b4923e7f23573f0dc6b10b

Observation 81ccf352-fc30-40ad-a176-ed4d766704f4 · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

Clapper: Compact Learning and Video Representation in VLMs Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:45.733144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:45.733144Z digest=sha256:1c0a362eab258cedcd8b94dc7cbf640bfde770468afd3748781bfb09d8fb34f0

Observation 69f0435a-ef32-44d7-b5ee-a8290edf9e0f · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

Clapper: Compact Learning and Video Representation in VLMs TempCompass: Do Video LLMs Really Understand Videos?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:45.814866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:45.814866Z digest=sha256:4ea8d00ef6e3bf492bd4d70064947bd876f61a016e9fd8c36763a7ff51af2b56

Observation 68cd9ea3-10f6-419d-b479-3cc699219b56 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:45.926076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:45.926076Z digest=sha256:b5ebc9e75de5650a7fe98feab974f4f12777cb7f0398e3d2a3f08a5f97140fbd

Observation 5e32be1b-96f0-47c5-a829-3e799dc63adc · outbound

This paper cites OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation.

Clapper: Compact Learning and Video Representation in VLMs OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:46.025304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:46.025304Z digest=sha256:ff7ef6538fde93fc7d194a79809e93f6c8a4df5efbd349ee90784b137fbb2b9d

Observation 2c42a8a7-e8b0-4b89-a580-674d9dd00723 · outbound

This paper cites GPT-4 Technical Report.

Clapper: Compact Learning and Video Representation in VLMs GPT-4 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:46.130255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:46.130255Z digest=sha256:19347aadf8879efea59a21072a42322a572a108009a09b8b364ee9558b047df2

Observation 29148629-9ec8-41a5-8c18-a42e3e65cfc3 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:46.240845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:46.240845Z digest=sha256:b054c68fe6b43bc4d1c7fd9c4963f1196de68204346106c1b1cfdbd8d81daad8

Observation ccdf8af0-55ec-4984-87a4-2d1ccafd23d7 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:46.356809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:46.356809Z digest=sha256:0d2ee27c241ad78cb45da3f71082d39b5fa21f490bbf88f115e090ae30da249b

Observation 11bfe300-e9c2-47e6-91a2-4e1d3cdb879a · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:46.438533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:46.438533Z digest=sha256:0f1b258942c721cfc5e1dc1b1cfe476bfc8f125e65198f5b2fef385a7875b4ba

Observation 53be6b26-fba9-45f5-9341-c07df4c0678b · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Clapper: Compact Learning and Video Representation in VLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:46.538312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:46.538312Z digest=sha256:dba640cfac8e299fa0e07f37e5c76a26c7c4221a8f96bb38159c150f9ed0672e

Observation df10b4bf-19b9-4cc1-9cda-ea7ba31abd8d · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Clapper: Compact Learning and Video Representation in VLMs LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:46.656613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:46.656613Z digest=sha256:1a445caa672a91bf8f7aff8b5526bc8d80185789f220639ba0c158bb89f64394

Observation 8e97a8a8-ca31-4697-95fb-48e4305a115e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Clapper: Compact Learning and Video Representation in VLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:46.774925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:46.774925Z digest=sha256:08ce47e2fb24d0c6273f6b98c101963df89a550067056ab7e61bbec3e54b98d6

Observation 4e8687f3-97cb-4618-809d-b22cb0a964e5 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:46.882268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:46.882268Z digest=sha256:0afbf443adb476cad62d05b777143c915c9587fabe4451dc5918f89e8fae4a7e

Observation 52373ded-3d67-44cb-b8ae-820b898f4e2f · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:20:50.489650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T15:20:46.992993Z digest=sha256:cfa3d252544d675fd1dd0587ca06193f2a031a01261e86151cf01cc8a6300a54

Observation e80577a3-53a0-4e6a-b701-074b0ce0ca31 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:20:50.314362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T15:20:47.097954Z digest=sha256:77e6bb5beb6d050da960f075a49c716c1790099b75397d28cc2e53422259bbd4

Observation 6d1c5864-6482-4522-bde2-8406b87655e1 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:20:50.061845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T15:20:47.209607Z digest=sha256:1144e44d7662776bd8eae2ad5576767c26b637299f33b5229b1ddb65c61746bc

Observation 65d26525-3b22-4b76-8c52-335ada6560d3 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:47.360957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:47.360957Z digest=sha256:2f2b05cf8570724532064ebe0cdfca5b83f8d40f3eae80f234fadc046cf97294

Observation 4e19bcb9-4faa-4af7-8a25-9f01f45e34d9 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Clapper: Compact Learning and Video Representation in VLMs PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:47.442019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:47.442019Z digest=sha256:df9e048fe02dd23012af35166058ae43da3756d747e625c5912faca1a6d558f3

Observation 0a170583-a703-4bcf-9e43-4138824d1160 · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

Clapper: Compact Learning and Video Representation in VLMs SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:47.549466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:47.549466Z digest=sha256:bf1c4635461f0313ed6fbb122219fb79d23bc4b74e22c0292d8439d166e1e4bd

Observation 504d6f4d-54f0-4a77-8703-264b2636ec34 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Clapper: Compact Learning and Video Representation in VLMs LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:47.582468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:47.582468Z digest=sha256:ed090d97e982ea5c6a8b93e97e2c0e6ecc15cd76a7e5f74f06e9e9d5f6c54cb8

Observation 3b7b5785-9727-433c-8acf-043e29bc57cb · outbound

This paper cites Qwen2 Technical Report.

Clapper: Compact Learning and Video Representation in VLMs Qwen2 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:47.666280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:47.666280Z digest=sha256:a0691c87777632adef0f67e72c2fae5595a9b625b99a07ae8768558128deab41

Observation dcf3998a-0299-44c7-99b1-60537273aa35 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Clapper: Compact Learning and Video Representation in VLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:47.766618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:47.766618Z digest=sha256:a1ea1ad41fc3ba8ca7b15aa90d8f7239a20463be7931a6168bd28d0894d79cc7

Observation 6123c3b3-0e54-484e-b7dd-3352b9d48e95 · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

Clapper: Compact Learning and Video Representation in VLMs mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:47.848486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:47.848486Z digest=sha256:c17a0750732b4190e83a001567052abf24ad14e3645d1e707f9b980dc3a98232

Observation b7036e20-0a50-4793-a074-fb882ae4e945 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:47.964092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:47.964092Z digest=sha256:fa5a9ccbb91d2d129dbc019695f47f08db671ca48e8f9877afad6129c681dc12

Observation b080d67c-5c39-4dbb-8d51-116ecac20bc5 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:48.040187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:48.040187Z digest=sha256:b5a871c8e9cd43fa80e421606c6005a7bc9f9ff659f0d919b713235fe405ac45

Observation e0f45e9c-7100-445f-a5c3-ea4118e6ebbc · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:48.150431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:48.150431Z digest=sha256:a9a2080ad38c84c2fdb354a2a569b9fdd4460b440628aed12ebcdccb003295e3

Observation cdfe6a61-b8c6-4b66-8526-ef453138e929 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

Clapper: Compact Learning and Video Representation in VLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:48.289056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:48.289056Z digest=sha256:961dd4ef69611b522076481c4a0dff0de5b7032ee9d302f3a09be2bc46640efe

Observation 0fec3900-d476-4a76-b440-735d2d295676 · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

Clapper: Compact Learning and Video Representation in VLMs InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:48.439015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:48.439015Z digest=sha256:af6c0af8e6d01ded4ffc19fb975f07e3a96b344f0b9ddeb3ab6fc5815b432482

Observation 896f522a-1fd3-4c1f-9435-e1462eb735e9 · outbound

This paper cites Long Context Transfer from Language to Vision.

Clapper: Compact Learning and Video Representation in VLMs Long Context Transfer from Language to Vision

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:48.560323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:48.560323Z digest=sha256:3c8a83f21900033eaa15f1c6ebbeb2986027c8bd211f1ce870c2706dcd192b8f

Observation a2c8dda5-b1a7-4752-a01d-d624987f1540 · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

Clapper: Compact Learning and Video Representation in VLMs Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:48.669699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:48.669699Z digest=sha256:168b35efee8b9a54d7e5bb54c33369e969324c02ebfde1cb40dfd10c4c3bfc15

Observation 61da8d38-d993-463d-b042-23d06d7753db · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:20:49.835526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T15:20:48.791387Z digest=sha256:26b929ea0181ca536b53752164e45852673a449d19b6d7b714b1ea0906e56630

Observation 8eacf7cc-13ab-4b30-9181-67af9725b7e1 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Clapper: Compact Learning and Video Representation in VLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:48.883282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:48.883282Z digest=sha256:580aa21f21acc04fd24b7ca6792bcb9a9d50b0bcb8cfe0daa64e803b3042965b

Observation 20c7a929-9e13-4928-860f-e2097167ac49 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Clapper: Compact Learning and Video Representation in VLMs MLVU: Benchmarking Multi-task Long Video Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:49.012698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:49.012698Z digest=sha256:ca464d257fb2dd6029fcc150826b62e977d897ac9e22bcb1321a3e52dd039286

Pith citing papers

No inbound Pith citation observations are available.