Pith. sign in

Paper Citation Record · LEDGER

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler

As of 13 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2501.15513.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.15513 v2

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:18:18.203170Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T04:39:26.597389Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0594859a-02ad-4274-b5cb-675f791c02be · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736, 2022.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736, 2022

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.356749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.356749Z digest=sha256:7bd2eb2f6b050585751464d12bd01bf073e90b0fcb6ffd645d437c4be24ccd1b

Observation 8b723536-4ad9-4cf2-a9a7-dd830a5299c9 · outbound

This paper cites MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.386325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.386325Z digest=sha256:91e5f7c7a2b57640f6b1e1b5b1f8f2f614b19e3c2c3c50b9cd654c96b6bfcb4c

Observation b8072065-f2ca-4289-aeef-d172b1250d6f · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.411410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.411410Z digest=sha256:e858278b2cf39b58448db066c8419669e57fd922efe5ed50f14d10822710d760

Observation da7685aa-0d17-435c-8389-f3141c41ff2a · outbound

This paper cites On attention redundancy: A comprehensive study.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler On attention redundancy: A comprehensive study

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:18:18.819859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T14:18:17.446172Z digest=sha256:41e030d197256d56e7333a2738c1328a1cc9c5c62cfb6c653a5e1aeef59883cf

Observation d0610ac8-318e-447b-a2a3-8986869d718d · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.464986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.464986Z digest=sha256:17f23c471d392f2a285996e45c12336d8035ecd91a22f524a2f92de9f5b0304c

Observation 5fde84b7-b003-4cf6-8e4e-75596a287cc8 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.488069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.488069Z digest=sha256:bab3a8d2a2573554c8551440648daf792f38a343aaff7ee260e5676b84f05376

Observation c371f93b-dd3b-4003-8e41-3bdd221f361a · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.511758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.511758Z digest=sha256:95fc58cb3496fad4062f7ae9b3e2ac792eb75413685307cf4d7c5c6ec83a0315

Observation 94764db4-f09f-4834-9d11-812e5a47907d · outbound

This paper cites Efficient Multimodal Learning from Data-centric Perspective.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Efficient Multimodal Learning from Data-centric Perspective

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.531656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.531656Z digest=sha256:bdd9010cb06e35d28292093cf2ad843d398f141ebcde695ac93c0bbb07c3ee86

Observation 1228130b-8da1-477f-8bf7-98b9e7adb483 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.549431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.549431Z digest=sha256:705c03711e1b5dd94ea175c6f4b1a5e834f48ed2ea1d14a4c0dbff9dbd05cf57

Observation 2ea93c03-052e-44e2-a882-852655171b02 · outbound

This paper cites Qwen2.5-Coder Technical Report.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Qwen2.5-Coder Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.558446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.558446Z digest=sha256:e86cc2a6862ab74f792e8a553d738e404ac221ce892f0804c9518efa8aa06099

Observation ee1e81e6-e1d9-427f-89f6-59ac41686970 · outbound

This paper cites Phi-2: The surprising power of small language models.Microsoft Research Blog, 1(3):3, 2023.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Phi-2: The surprising power of small language models.Microsoft Research Blog, 1(3):3, 2023

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:18:18.727706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T14:18:17.570633Z digest=sha256:6c5c03e239e9f726f35f2104b9539f5be7e9ed667ece97193bf9a9fcb43ab4e0

Observation e7f2384a-558d-4916-b860-04b6f227d507 · outbound

This paper cites Chat-univi: Unified visual representation empowers large language models with image and video understanding.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Chat-univi: Unified visual representation empowers large language models with image and video understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.580777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.580777Z digest=sha256:18ce03f4bb74ab3629abe43d2b9e5b2ec905f068a1b0f6234bf89f4abec92df8

Observation 5035d900-4a80-48fe-97fb-804f35289731 · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.592863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.592863Z digest=sha256:2e72384ea5bf8d87a6e2176176b16cc0fb06c96b2d4211d32f7b0f756ddd6baf

Observation 39898d34-2355-4ed0-82d8-51ca383af8ed · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler VideoChat: Chat-Centric Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.605226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.605226Z digest=sha256:c858d4749d70e18171bfaea420da23e4610d202e4b40f98e34b359cf236ee21b

Observation 07ab9965-e9d8-471e-94ff-916d0f37064c · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.618502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.618502Z digest=sha256:300ce6787455fdebadcb53bc6694af7fc8a40bac044c4390f6a4d5a50487eaf9

Observation 584475cf-7c6c-47b5-9226-14e34cbec53e · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Llama-vid: An image is worth 2 tokens in large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:18:18.618280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T14:18:17.645852Z digest=sha256:5c516d8371bb1eb371a63c5ca062792d2b6c815b09e539f9cadb83026397c57a

Observation 6334ebe2-a2e4-40c8-a56f-ac9683e0c18d · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.694223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.694223Z digest=sha256:e40aa15fddc3f1702cf0d57acba1f1e075d136fd63aec306c5dcb0ffcc9f985b

Observation d0a71bda-cba0-4041-965d-e16e645a9f62 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.735844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.735844Z digest=sha256:39efffbdd5bd004bad333eaa6b87456b10ecd6a409f689c385e8d6aa0b8cface

Observation ee1fd6ce-9700-4652-9bbc-78cbd819c634 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, 2024.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Llava-next: Improved reasoning, ocr, and world knowledge, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.787631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.787631Z digest=sha256:4144deee29066c4c45a9cd111cd2ba88a668a616a92c5beb9aeff88709e2bf8b

Observation 87cf6f5f-bda8-418f-8086-cc2bfa6f68b9 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.816391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.816391Z digest=sha256:9746e8519bf6d322c88efc67118e48703bb4480f4c88acb74af961ff706eea57

Observation 72a5ebb0-6080-4aa6-8c4f-b2863db0fc86 · outbound

This paper cites St-llm: Large language models are effective temporal learners.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler St-llm: Large language models are effective temporal learners

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:18:18.580148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T14:18:17.845024Z digest=sha256:64856dce3807cf6eba3d8ca5ff322b0815817c42021a5c926b7674364111625b

Observation 38287e0a-7220-4daf-9cbd-e6fcfdf0004d · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.857839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.857839Z digest=sha256:ca4d629d9dd31eca565a1436ee149c3521260ea37bf5b7cfd5ad7c0c7ddd5c91

Observation 1ee46f66-a036-46a2-be2b-a8e42416aa54 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.898432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.898432Z digest=sha256:a24312139f5296b3855e24d12d458184d4b7ea2c922be8a6c7bec48b44353e95

Observation c07d4a9d-4c35-428a-866d-5cca2a9b9ade · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler DINOv2: Learning Robust Visual Features without Supervision

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.916023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.916023Z digest=sha256:a0a6b84336575a2bb237d0ba8234215100cdaad9024c907883afd3fab8f25b93

Observation 94812177-582d-490a-9e94-0bbc74402419 · outbound

This paper cites Learning transferable visual models from natural language supervision.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Learning transferable visual models from natural language supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.929562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.929562Z digest=sha256:894db9f5019abd36859a9bba4fa63cc94a2dbc67a7e0d80a72817bd16baedfed

Observation 294d4da3-d181-4240-89ef-097f885f3295 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Moviechat: From dense token to sparse memory for long video understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.934235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.934235Z digest=sha256:4b50f9b782722bd70f295ae06433e5a30427a0dd51bc1220ade2b07b3ac9f675

Observation 722f96a5-9ef2-4191-b2a7-581d452959de · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Gemma: Open Models Based on Gemini Research and Technology

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.947719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.947719Z digest=sha256:89a7556f1f27a21ff724b6fd32c9963e5dfbd390a0b946b799a5f0fe77ae2203

Observation d49048db-9381-46d5-b43a-c05e9d6b6cf5 · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.964511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.964511Z digest=sha256:8719711c1a65dbca5bc35d82e8abc56ecde5283c8c08b25c0c2e389e37ade73e

Observation a0a81bc9-a7e6-4d8f-9d69-ee685f2fb426 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.977599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.977599Z digest=sha256:722b9d8f509aa6f28f586fad6219ae5279e1eeb08fc0ff223116fdbd1b504d72

Observation aa8b79b0-698f-4177-b0f7-7f4bc1b232d7 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:18:18.557231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T14:18:17.988770Z digest=sha256:0f0bb81afcf9fe055597162609edecdae7a84aae9ee49323109d3f78add915dc

Observation 35d0e1c7-193d-47c4-8c84-d15690c79cfa · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.992793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.992793Z digest=sha256:359aa5ba2f65f41bdf230bf129e35b6720e05287ec2788388acb50b442691f32

Observation a8d7feb4-d9bf-4e6f-bcd9-228401fa8cf8 · outbound

This paper cites Sigmoid loss for language image pre-training.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Sigmoid loss for language image pre-training

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.998128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.998128Z digest=sha256:b00eae019967e1be0fc7e4f754819929f7333762f58d1c77ef8c43953c834714

Observation 4b280291-3a34-49c4-b416-3ed501db93fa · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:18.004242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:18.004242Z digest=sha256:60c163b916d2a69f841f343c8844d1a6b914fb6687e6fc2b2dd5fa67a981badb

Observation 8ef89105-5643-4fc6-b894-96a13c1bc000 · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler TinyLlama: An Open-Source Small Language Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:18.008511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:18.008511Z digest=sha256:c4dcdd2c2036bca93d8d25b2a84eca930f1f7fa1409db632d31d5c5787a6ee06

Observation 86aeb704-3adc-4170-afc5-5eac48021193 · outbound

This paper cites LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:18.012399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:18.012399Z digest=sha256:f64c3f3df843d350ee0d931e896b0cbb219b7b43c1718d1fe08ed7048095690a

Observation fe5d7b0f-5675-4012-8c79-da07e7ebffe9 · outbound

This paper cites Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:18.042072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:18.042072Z digest=sha256:4af10a4ef53c2fe38aa523b4cef3139456a06804c6b4d4cb89271d94fe192ee1

Observation 7ce12156-d8cc-4ea6-ba88-ee60aab2a8c3 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:18.086229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:18.086229Z digest=sha256:9d5d63475da199de5845bf286840de42db76e4a5988931cf085c34811fa99cc6

Observation ede505cb-e091-4b28-a534-6a91bffb337e · outbound

This paper cites TinyLLaVA: A Framework of Small-scale Large Multimodal Models.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler TinyLLaVA: A Framework of Small-scale Large Multimodal Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:18.139242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:18.139242Z digest=sha256:db482da999ed09e886c8042b511cc6acd9f9ceec867c0dfc1d3ce37f0e30f162

Observation 4eba6fca-bc13-46c2-9fb9-6984553edddb · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler MLVU: Benchmarking Multi-task Long Video Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:18.168610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:18.168610Z digest=sha256:a8a98cbe141e61c97f735205c175681a3362dd5a8ed0ab2ca3648b8f0fcb6311

Observation 2216e17b-a605-4ba7-9dc2-b2692de0cbf6 · outbound

This paper cites The best results are indicated byboldface.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler The best results are indicated byboldface

Reference 512

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:18:18.533114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T14:18:18.203170Z digest=sha256:5868fb00d31d87674d8ccfe57009b337ba6439e43d567ca06b95cfd50780c57a

Pith citing papers

Observation ca6bb9bc-a892-4943-a3f2-fc70019a7e86 · inbound

GMoT: Gated Motion-Aware Tokenization for Fine-Grained Micro-Gesture Video Reasoning with Multimodal LLMs cites this paper.

GMoT: Gated Motion-Aware Tokenization for Fine-Grained Micro-Gesture Video Reasoning with Multimodal LLMs TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T04:39:26.597389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:39:26.597389Z digest=sha256:7e4f7a90a318ca53e314d4341a9e61649c06a274e35755fc8fe4a997e02e3bd7