Pith. sign in

Paper Citation Record · LEDGER

PruneVid: Visual Token Pruning for Efficient Video Large Language Models

As of 14 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 11 inbound Pith citation observations for arXiv:2412.16117.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.16117 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T10:49:52.282672Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:11:49.166898Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T19:16:31.884188Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a8ed0265-8631-425c-94c2-ad34909d2cde · outbound

This paper cites write newline.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:52.079871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:49:52.079871Z digest=sha256:d210bbc515efd2a149423df8be5f7c2dd69413c94bc54503029d2132d24c11e2

Observation 30473d58-f116-4053-bd1a-05e46582f1a3 · outbound

This paper cites Gpt-4 technical report.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Gpt-4 technical report

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.926719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.086279Z digest=sha256:2d72c066d8beb155b99bd7160cc0c284fb42ba3483e40f8abab08f56dd5a3ad6

Observation 24cf1b02-3b0c-4215-83c3-b0644fce1bb7 · outbound

This paper cites Token merging: Your vit but faster.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Token merging: Your vit but faster

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.910981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.091355Z digest=sha256:1b05457d1a5fe60a7d549d54fa96cde1dbe437f0b8a9f84f919285364d40c605

Observation d9f85c15-d746-46a3-9c7a-7fa9c099bb4a · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.894952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.096116Z digest=sha256:1cf4f3ff64a23db00813e0500590ca6d5363735cc15adbc2c0cf1544baa1428c

Observation 9b5385b5-af5a-4a49-a855-7686442ec1f5 · outbound

This paper cites Study on density peaks clustering based on k-nearest neighbors and principal component analysis.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Study on density peaks clustering based on k-nearest neighbors and principal component analysis

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.878958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.101360Z digest=sha256:796f107f43d1ee075e1de9da8a2dc123c9a24bf06e46271c59721ced34f79259

Observation 88ec6223-67ae-4794-8672-c06896802dce · outbound

This paper cites Slowfast networks for video recognition.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Slowfast networks for video recognition

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:52.106189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:49:52.106189Z digest=sha256:87f4fca0149e1520026d9014f7cede587b96b9385c9329fc9bc93c1657725805

Observation 80bb6965-da4c-4a9e-bcc2-b4610da4b881 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.852860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.110962Z digest=sha256:4c181d0a849b3ca2547d6760c4c3c2ecf8bff475af447a0fe0dafcc74856d13e

Observation d3366fe1-9738-43b6-bced-577444a029bd · outbound

This paper cites Visual hallucinations of multi-modal large language models.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Visual hallucinations of multi-modal large language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.836547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.116404Z digest=sha256:7f6d537532c490b3c70728d8ec012b5e71c478fa5015a5eebed1b457a4440112

Observation 09420e7a-1b5e-45ac-89f0-207dcf335507 · outbound

This paper cites Chat-univi: Unified visual representation empowers large language models with image and video understanding.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Chat-univi: Unified visual representation empowers large language models with image and video understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:52.121305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:49:52.121305Z digest=sha256:8234e81770af32cdfce4f4938f28a8f88c515164a5db43003764f1fa35f974ec

Observation 362e2a01-8708-4b6e-8e27-43cf5acbc5a9 · outbound

This paper cites An image grid can be worth a video: Zero-shot video question answering using a vlm.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models An image grid can be worth a video: Zero-shot video question answering using a vlm

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.809348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.126167Z digest=sha256:f45200141ad657138397348babbfacb021662311602fae7bd3f130bbaff3b757

Observation f7a02555-cc91-4288-892d-9518d6300505 · outbound

This paper cites Spvit: Enabling faster vision transformers via latency-aware soft token pruning.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Spvit: Enabling faster vision transformers via latency-aware soft token pruning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.793398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.130928Z digest=sha256:17f3a7b0fa411cde138180e728966700d5045afddd6f5f3bd7a259863b79c533

Observation 7366d366-22a8-42ed-b489-fff6b0b1a64c · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:52.135759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:49:52.135759Z digest=sha256:1e5c94fd00f3d3da896814123fbbc6c761a3852b7565972684d85b8531db88fd

Observation bfda57c7-00f7-48a7-b788-59aab23c68da · outbound

This paper cites Videochat: Chat-centric video understanding.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Videochat: Chat-centric video understanding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.777893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.141193Z digest=sha256:8a0a2af6b700c8c25206fbac531d3d226e301d27192614ea317c7e4bfb542214

Observation 53e4dbbc-d041-417c-8a3f-fcd9abeda59b · outbound

This paper cites Unmasked teacher: Towards training-efficient video foundation models.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Unmasked teacher: Towards training-efficient video foundation models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.762358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.145954Z digest=sha256:4862673c4e06e4e8d2f5f4d19088204b20a71272f66c596397c9d5b76bed001d

Observation d4ac5416-df6d-4fd7-a6b0-bcab107d6709 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:52.150713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:49:52.150713Z digest=sha256:4070307ed55a26b626da1d69903e40aa85907ed82b96beb1545932a070606683

Observation 66e32273-bb97-4973-9e37-847d0845211b · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Llama-vid: An image is worth 2 tokens in large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.735877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.155387Z digest=sha256:2d0414abc713d2fe580229dac5adcd9151e6ba0987abfd4d852832a9ee90f9d7

Observation 88a193d4-94dc-420d-8cad-ec0a1bf467aa · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Video-llava: Learning united visual representation by alignment before projection

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.720410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.160095Z digest=sha256:93076ab4eb03e1dfc34afa8dff1b54630dd878273b80aad6b627de5fc1464688

Observation 01479777-00bf-465c-8921-da0de4ad68d5 · outbound

This paper cites Vila: On pre-training for visual language models.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Vila: On pre-training for visual language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.704059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.164601Z digest=sha256:07e63b3a2cca40ac3950a51dca3d716ec86486ee12b4ee7b1a330d0eb6ae54d6

Observation 3d365085-e043-45d3-81b4-72357ab214bc · outbound

This paper cites Visual instruction tuning.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Visual instruction tuning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.687922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.169265Z digest=sha256:91d62bbeb2edf82c4a2b00751fcf9e35fea907d3bb4dadadcd881c786799ac79

Observation c1dfc92f-58b0-4cad-acd9-d52448e12134 · outbound

This paper cites St-llm: Large language models are effective temporal learners.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models St-llm: Large language models are effective temporal learners

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.670456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.173802Z digest=sha256:6fdb2cb07fcce3f6f1c04f46e3c00dd7225bc07fb990a967b23d8cee85745643

Observation 444af41f-3654-4986-b6ad-dc803d627f54 · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.654884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.178782Z digest=sha256:23d979cd7e32c868c0c03d1d920f94ec0637eecace877d40848ef19953934360

Observation 7fdf3f4f-fb51-4c9a-8f97-5b040ac77ae1 · outbound

This paper cites Efficient inference of vision instruction-following models with elastic cache.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Efficient inference of vision instruction-following models with elastic cache

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.637596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.183524Z digest=sha256:ce0fbe74c4c3a0a4e67a73b9dd8418b573ad4a7f88a824edf65cc1a9a6c27fcd

Observation b3513bb3-1bba-4acf-bc75-9b61ca1dfdcf · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.620788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.188261Z digest=sha256:567c162d38af357e5473d7f18b1dad9da7683458ae6330dbdd8ce0fcfe1d5cb1

Observation d64665fc-a78a-4d1d-a015-e1d779b05a8d · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Egoschema: A diagnostic benchmark for very long-form video language understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.604400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.192760Z digest=sha256:800dd6930887c3ad147bc654759a72e3f8c2c2658318471cce25278cffab94b9

Observation 655f4a8a-6c93-4915-94a0-32464f0e6001 · outbound

This paper cites Learning transferable visual models from natural language supervision.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Learning transferable visual models from natural language supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:52.197356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:49:52.197356Z digest=sha256:32fe1fd7c73428d073e19bbec42f2babdac7f642ff49a0d8cded51b8e18edc61

Observation 88aee4a6-83c0-421f-bfdc-420889855522 · outbound

This paper cites Dynamicvit: Efficient vision transformers with dynamic token sparsification.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Dynamicvit: Efficient vision transformers with dynamic token sparsification

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.578089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.202129Z digest=sha256:39bd91fdfbbf4648a461ffc66b1ae4b7cdb1305c0860c5beb8818c2b4f33ec5f

Observation a8e3060c-1a14-422b-a9f8-33fec2064078 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.562640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.206740Z digest=sha256:2013e3f15a4ba5d416d0d8b3b8559a50d745b9b5452a6088cc36d00caf122eb1

Observation ed3450b3-cc7c-4344-87f6-d367817562d2 · outbound

This paper cites Llama: Open and efficient foundation language models.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Llama: Open and efficient foundation language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.546298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.211721Z digest=sha256:d31a542e62a2b6020411aed7549f44c3ad00f82a11697c97f19baadc977815de

Observation b6eb80ef-3fc5-4d0b-a3d0-31a93a4dcb93 · outbound

This paper cites Fastvit: A fast hybrid vision transformer using structural reparameterization.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Fastvit: A fast hybrid vision transformer using structural reparameterization

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.530444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.216971Z digest=sha256:c3efb6a09b149d7593bde872b337938be269b1e0e0c0b156f53317c5d3fbde86

Observation 5427a9af-9822-40b5-8e07-c113af179731 · outbound

This paper cites Look-m: Look-once optimization in kv cache for efficient multimodal long-context inference.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Look-m: Look-once optimization in kv cache for efficient multimodal long-context inference

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.513566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.221825Z digest=sha256:2e7393414410d5ddbfa01cfe83c69bb27a844ede6369af2c20d0be53bbb9ebbb

Observation b16479fe-e55b-4b5a-af37-f9777b044479 · outbound

This paper cites Tarsier: Recipes for training and evaluating large video description models.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Tarsier: Recipes for training and evaluating large video description models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.497397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.226571Z digest=sha256:b8d67e924c5a4b6da760d0716c6efc81c6f825c8e9e651590d98843e985ff24e

Observation cf22f3e6-e6de-4dcb-8634-df72158be49a · outbound

This paper cites Actionclip: A new paradigm for video action recognition.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Actionclip: A new paradigm for video action recognition

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.481738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.231305Z digest=sha256:11f919b9b232c52224e467ff19f71c39a67aa05f2e232207ccaffe179b56e759

Observation 8a6da8fc-f78e-4745-a4da-4d737627cc13 · outbound

This paper cites Internvideo2: Scaling video foundation models for multimodal video understanding.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Internvideo2: Scaling video foundation models for multimodal video understanding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.465594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.236108Z digest=sha256:7d045ac62ac4db9034545e2f8a931f4db6497152d73b716fe6735bb2c82f451d

Observation c16655bd-aa13-4471-99a6-9f99870feabb · outbound

This paper cites Freeva: Offline mllm as training-free video assistant.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Freeva: Offline mllm as training-free video assistant

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.448060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.240903Z digest=sha256:5abae1b2962bd27a978377cbd974e1f1dba8dc2fe4a6bf73e72a9549b5648ddd

Observation 31e1dacb-9f1a-46bd-8f4b-ecb579dbbf47 · outbound

This paper cites Pllava: Parameter-free llava extension from images to videos for video dense captioning.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Pllava: Parameter-free llava extension from images to videos for video dense captioning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.432007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.245737Z digest=sha256:bca83ccf3621d1a3e5dadb0b27416cc514f4b584300519809dcc2b3c1d21ad5b

Observation b2381bee-dfb4-4b3d-bbe3-d858b3a27169 · outbound

This paper cites Slowfast-llava: A strong training-free baseline for video large language models.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Slowfast-llava: A strong training-free baseline for video large language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.415047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.250883Z digest=sha256:bd5fabd66933c10a3d69b6da420b9d010d748759290b6f7937d0acf101e66a5f

Observation b5078aeb-6a01-4692-9ec9-08b0ff3b3258 · outbound

This paper cites Qwen2 technical report.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Qwen2 technical report

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.398911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.256159Z digest=sha256:2e51bc132219680ed7341acc0f7e65395e6f7722660a6dd7fa04597763ef7514

Observation 3dbe2622-126b-46c3-adc4-dd4a1693e392 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video understanding.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Video-llama: An instruction-tuned audio-visual language model for video understanding

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:52.380924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T10:49:52.261050Z digest=sha256:39169631a07ed348180417f9cebdc58d8fc841931db4ed2e50df9c4ade9489a9

Observation 6d371b21-56f9-475d-bedf-703558c711fc · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models H2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:52.265871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:49:52.265871Z digest=sha256:545f3dc2f93ebc841f78ccef87d77dd8de0c49263bc5061d05a1f2add4f12242

Observation ac1d57e7-0fd6-4e11-b1b8-66fc4376ca5e · outbound

This paper cites @esa (Ref.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models @esa (Ref

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:52.270808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:49:52.270808Z digest=sha256:6ffd865ced5be0c25dbbc7f54a901c749df176fe8f9d0726974fdc2bbd38778d

Observation 77aa930a-754e-4b4d-8eee-8428bbef4323 · outbound

This paper cites an unresolved cited work.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:52.277452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:49:52.277452Z digest=sha256:ca651424f1b79db3606fedf5e58abc51e7b477f2092013cf2467bb7063e581d0

Observation 8357fddf-4eb3-483c-8862-0e1b702a7478 · outbound

This paper cites an unresolved cited work.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:52.282672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:49:52.282672Z digest=sha256:747d8ce20f68d3432bdc44aab82217cacd5f29675ff5bbcd8cb61f86d9dca123

Pith citing papers

Observation d1096a89-6bc4-4e70-b5cf-9ab2e5fed281 · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:27.055696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:27.055696Z digest=sha256:2402f1ffbc606a094c0b11a4083515d79a66985977863a09dbf7697ab042c864

Observation d544c777-c601-45fb-a23e-d49639e58b38 · inbound

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs cites this paper.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.995554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.995554Z digest=sha256:f4719cde00df27d642c6bffd3908131eb9830d07955ebcd8b534cba44c368960

Observation f29c4805-84c0-4a34-92c7-6508dc49d3ec · inbound

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs cites this paper.

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T14:41:15.535604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:41:15.535604Z digest=sha256:84a8c5d5a56c00e4886c257fe5c11bec2221d64b1a53d90340d16c84f3f2105b

Observation 747571fc-405e-4119-b757-877208d82e2e · inbound

TrajTok: Learning Trajectory Tokens enables better Video Understanding cites this paper.

TrajTok: Learning Trajectory Tokens enables better Video Understanding PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:16:31.888882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T19:11:52.694778Z digest=sha256:4eda8d5fa0c98788ff25499a2804418061a8975e18f22b1f59482723a406a5e4

Observation 59c2e359-e839-46af-bc6a-e64c09bff64c · inbound

TrajTok: Learning Trajectory Tokens enables better Video Understanding cites this paper.

TrajTok: Learning Trajectory Tokens enables better Video Understanding PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T20:38:38.325046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:38:38.325046Z digest=sha256:57efe0d18c347e31d1832ad2c35f69f4d76395c9030e8c2633b0ec1ced7a5e29

Observation 352f6555-5169-49e5-ab23-c28f34d52a3b · inbound

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models cites this paper.

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:26:26.984952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T18:25:21.621268Z digest=sha256:e2e84de35f542969ec1c726adee65b565831fc9b8a5be5c5349237dd7ca14bee

Observation bf580084-5077-475c-b446-9c4c0ed3c70b · inbound

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs cites this paper.

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:41:03.942336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T15:23:08.671342Z digest=sha256:62c04513212cb7d6489f82230aefb7222b308863aa83eb4c73d245d30c5b70ad

Observation 391ab039-e668-49f0-92e4-cdc6a8d5e4a1 · inbound

Spectral Origins of the Self-Correction Blind Spot in Autoregressive Generation cites this paper.

Spectral Origins of the Self-Correction Blind Spot in Autoregressive Generation PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T15:30:23.485228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T15:30:23.485228Z digest=sha256:c7adae1f23f67c19cae006799f2b9e21b4d3a0cbad015d3e8ee711bff55b7b4f

Observation 6037976b-5fc2-4d7c-a168-9c15f12ab2bd · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.557165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.557165Z digest=sha256:a86a5016589461298d327abedbc1b87eddd7c8051e8c5753c175dd23d9dfd37e

Observation c93e8ec9-6253-41fa-a4c4-35e752a53ea7 · inbound

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs cites this paper.

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:57.501524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:43:57.501524Z digest=sha256:2f11131aab315c2ead7387bc08c7f367f4571a21f8e0b6a424870fc9751cedfa

Observation d9ce454e-6fdf-4a22-84c6-d681f1f16dce · inbound

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs cites this paper.

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T00:11:49.166898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:11:49.166898Z digest=sha256:d928474e2d1d753fc0a87369faefb56c43370f3a3560d636012471da54a3ac6c