Pith. sign in

Paper Citation Record · LEDGER

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs

As of 14 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 2 inbound Pith citation observations for arXiv:2508.21044.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.21044 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:41:15.586140Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T17:05:52.493660Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T17:08:00.954550Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f29c4805-84c0-4a34-92c7-6508dc49d3ec · outbound

This paper cites PruneVid: Visual Token Pruning for Efficient Video Large Language Models.

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T14:41:15.535604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:41:15.535604Z digest=sha256:b3a1a75f32b793e586d8b54c4a9dbf12850322a9ce9d8072f8a8f9f86e4ba847

Observation ce4e96b2-8cef-4572-8cf7-50b4b1ef4c25 · outbound

This paper cites In European Conference on Computer Vision, 323–340.

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs In European Conference on Computer Vision, 323–340

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T14:41:15.541099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:41:15.541099Z digest=sha256:025399f22b2dbb9c9ddad6a2a1df270eeb36bd4e4c46dd0d2cb9106ab12e4f89

Observation cbfd5e3d-863a-4930-afec-362e3a42d90c · outbound

This paper cites arXiv preprint arXiv:2403.15388.

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs arXiv preprint arXiv:2403.15388

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T14:41:15.547335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:41:15.547335Z digest=sha256:4049cf45ffda152cd03a804f35299be9585b18e0683dbc4321f3b6a3deeac04b

Observation 1fb9ee88-8002-412b-98db-60ff0724ad84 · outbound

This paper cites arXiv preprint arXiv:2505.21334.

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs arXiv preprint arXiv:2505.21334

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T14:41:15.552875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:41:15.552875Z digest=sha256:9ffc3614a9b55d25cb66560a3e210e06c0477a05074278d7ef3f4bda29051167

Observation e844714e-1986-46ad-abc5-1084ac886157 · outbound

This paper cites arXiv preprint arXiv:2503.11187.

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs arXiv preprint arXiv:2503.11187

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T14:41:15.559338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:41:15.559338Z digest=sha256:fc6515cf40bd4e5f6db7739dd8915b7c66ccccbe1607ff1d8df1c50b01431767

Observation 6e288fbc-89bf-4447-b45b-1edfde484ffc · outbound

This paper cites LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs.

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T14:41:15.565494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:41:15.565494Z digest=sha256:c9f5fae28206649e1d148ab55b41a2139fed2a9c22c9e3ecfa87fcb83933c0f4

Observation ea89fa3a-6c78-4761-b489-6d1bd969cae8 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T14:41:15.570938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:41:15.570938Z digest=sha256:a3042f96eaff3ae69d2d70b46ce573f7527c6e5bc5d7444673bda4538f85de2e

Observation e26d286f-8408-4b14-bce5-80b80213c53a · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T14:41:15.575985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:41:15.575985Z digest=sha256:58d43ce02adc649e446cae0edbbb3c46bb94e8390f1615b25939a40a6adbe4c9

Observation f5871a72-838b-4606-a26c-86655c41fd12 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T14:41:15.580800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:41:15.580800Z digest=sha256:044122bd3909ccc8956c2969f1b33e00606fb154e1faf8fe60394b685339d3f6

Observation 3852654f-126a-4d63-bc82-d74c946be87f · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs MLVU: Benchmarking Multi-task Long Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T14:41:15.586140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:41:15.586140Z digest=sha256:eda2e644bafe28a50591939365a52811442c8314e9e52d7706aa36ce496b6203

Observation 5901147d-987a-471d-a87d-180b646201cc · outbound

This paper cites arXiv preprint arXiv:2411.17686.

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs arXiv preprint arXiv:2411.17686

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T14:41:15.530096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:41:15.530096Z digest=sha256:3a6975fa2ae213ea8ae4316481f0d8b04d3cd39fd2c0f53bb7b292df24dd3e23

Observation 09ded1e6-3c85-45f7-881a-d26649f03049 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T14:41:15.525180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:41:15.525180Z digest=sha256:98f4828b09b250b29104dd678e9019353b73c104cc511547da9b39898fd95edf

Pith citing papers

Observation 29f1ba24-1c86-4ce9-8209-8e21e967bea8 · inbound

DINO-VO: Learning Where to Focus for Enhanced State Estimation cites this paper.

DINO-VO: Learning Where to Focus for Enhanced State Estimation MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:08:00.956426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T17:05:52.493660Z digest=sha256:f3d0adf538c456fe4bb78ddd5ea26f123289d844fee975cb2aa35c4ae6c72eaa

Observation d356d0ca-6110-45a7-ab72-1ed6e6e22b76 · inbound

Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs cites this paper.

Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:49:48.956060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T00:43:44.921189Z digest=sha256:8d5365660c88c7eda86488afd29ed2c4f3319fd5ebb62b3e21538276bbf453f2