Pith. sign in

Paper Citation Record · LEDGER

Tarsier: Recipes for Training and Evaluating Large Video Description Models

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 45 inbound Pith citation observations for arXiv:2407.00634.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.00634 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 45 of 45 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:58:52.983532Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T17:40:00.906584Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7868b652-9b62-4519-befa-8237bb57d600 · inbound

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs cites this paper.

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:44:53.686898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T02:44:53.284345Z digest=sha256:c92d6784c107d38b009f8154605fa0d9d2eb47361431db2b29111e02824fb371

Observation 0be139c6-460b-4b87-ae5c-c34e807f64df · inbound

LLaVA-Video: Video Instruction Tuning With Synthetic Data cites this paper.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:20:32.962559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:a9575e19cd3539f66ae0c2c3a94f43b2d6aa989c69bfb1e1a980e263e2969762

Observation f6822db5-7cbd-4ebd-8023-6aa5e8df054d · inbound

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models cites this paper.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.734088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.734088Z digest=sha256:57256741b0648a245dc62dde1e3b53dcad19f89bac4c7b50eb8469ef808449be

Observation e108ea2e-3b48-4b8c-91e9-463ea9edbf50 · inbound

Open-Sora Plan: Open-Source Large Video Generation Model cites this paper.

Open-Sora Plan: Open-Source Large Video Generation Model Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-23T08:42:45.194961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T08:38:27.946746Z digest=sha256:28e9e8dbda8ffd94a4d624a904183c5053ffcd29d4835431b354949683943321

Observation 8b4458b4-060b-4a33-bdfa-057b2f6d6052 · inbound

PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation cites this paper.

PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T05:16:59.601418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:16:59.601418Z digest=sha256:10849f8fa23dc15ef9cf1430e0477a2dc2161b8ce7ea7c59e830ee0c6e6db9f2

Observation 3ee4e1f2-cd67-4798-a8a9-e436db706eac · inbound

Progress-Aware Video Frame Captioning cites this paper.

Progress-Aware Video Frame Captioning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:58.545529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:58.545529Z digest=sha256:f2d9ace21ba4ee58c66f9687b7dafe8392f6c3e0af72771a3e82faaf4c04efaa

Observation 1f9c87a1-a224-42c3-87d2-2b5769e407f0 · inbound

LinVT: Empower Your Image-level Large Language Model to Understand Videos cites this paper.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.477303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.477303Z digest=sha256:06cd930fa8c213c1bd07476aba9176271b9486775231cb7c776901a5eda03ff2

Observation c0db57b7-51b8-4e78-8b50-f50461a269ac · inbound

PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models cites this paper.

PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 121

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:42.664942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:42.664942Z digest=sha256:f85f3e62544f668d17f138965414ee9270a3309bf0b3efb6c83e1ffa4106b749

Observation 0b4c1064-cd9d-4bfc-afc1-2fc19c77a335 · inbound

CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval cites this paper.

CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:00.945147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:00.945147Z digest=sha256:08302eae5b995ec9c593a5ecd28a88b034edee4f7196d34f00bb666e71ee4d86

Observation 29df199b-4fd9-4a7e-9473-d21570d2bad3 · inbound

BounTCHA: A CAPTCHA Utilizing Boundary Identification in Guided Generative AI-extended Videos cites this paper.

BounTCHA: A CAPTCHA Utilizing Boundary Identification in Guided Generative AI-extended Videos Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-09T23:03:34.652174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T23:03:34.652174Z digest=sha256:8ca30ff8631a4571e5fdafb0821f379c8ed1c88a26d44d251637178ce28ec5ed

Observation 63f8375a-2a2e-47a9-a9bf-2f2200751a76 · inbound

Goku: Flow Based Video Generative Foundation Models cites this paper.

Goku: Flow Based Video Generative Foundation Models Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-08T21:07:32.440706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T21:07:32.440706Z digest=sha256:772179036a497d888235eb3b04ceab93dd181b6be086836edfd75c4487b9f468

Observation 1ba35c61-2595-4c3e-81cc-798898ead23d · inbound

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning cites this paper.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.869285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:8a2db3f710f209306992c832a1043111c48215c5334c05c2c31ebe876b9607c5

Observation d31b4ecc-a26f-4da3-b04f-050dde267969 · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 140

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.594061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:454df758b5de7d40e52ad2d223156c4fc8d0b39b7e16ba744a6ed7a05377d823

Observation 4165373e-c41e-4f73-bb44-aa18251c6109 · inbound

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval cites this paper.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:44.083171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:44.083171Z digest=sha256:5bd78795b0d55acc94720b04adf7ea706cfd7557b88715089842e2f298eea5d9

Observation e152143e-f3e1-440b-be7d-f896473bbd0b · inbound

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos cites this paper.

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:04.111822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:03:04.111822Z digest=sha256:2dad2836a9e5b5d4dd9491d6be2b335ef139934ec0e653b9d4eb428ac613aa56

Observation 8f47666e-7b6c-4623-bfbb-7765955299fa · inbound

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding cites this paper.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.490044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:334bfb5ea56e3a1977fc4260bf6111c84676fcb14c24ccf5aad38176bafeeefa

Observation c314e2d9-ad47-490c-a7b0-549e292d25e3 · inbound

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking cites this paper.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:35.885400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:35.885400Z digest=sha256:37a08e1e56f8d81c305106dac81453c7612b9050ff2427871f32357875dea673

Observation 543f8e17-0ffc-4fd6-aac0-f39935557265 · inbound

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs cites this paper.

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:50.622545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:50.622545Z digest=sha256:1484f7dc9125635283fdc17d357f10fcb16100c5265ca270b354ab560aeef69e

Observation b82aeb2f-ca6e-4fb8-a96a-3af13f1bd4a5 · inbound

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? cites this paper.

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:10.194652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:10.194652Z digest=sha256:e041e174783aa09e16b755a0158649f360d6373b6b4613f3c0b0f8d9e981c077

Observation 0cb707a5-6917-4e9f-b00d-1c255f096643 · inbound

How Important are Videos for Training Video LLMs? cites this paper.

How Important are Videos for Training Video LLMs? Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.884760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.884760Z digest=sha256:97041d66644ad814b9f53db934478566c5d637ab6ec14d4c87cfdb1818f90c2a

Observation c455dc02-2b25-4543-a95a-d0536125090a · inbound

LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering cites this paper.

LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T15:53:33.251744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:53:33.251744Z digest=sha256:3cf11a471a7a23c3b30e147ca762317f3d8477e1ca29e502298650c3fe643ab2

Observation 37757e5e-2466-411a-9b05-09d7f95e4a25 · inbound

OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence Reward cites this paper.

OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence Reward Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T16:58:52.983532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:58:52.983532Z digest=sha256:b6608435c1c41684f7ffc14b0802458ac320a146733cac47566e76b1d5a74f66

Observation b62dbf61-4e24-4aa2-be39-3f4e90809c3f · inbound

Adapting MLLMs for Nuanced Video Retrieval cites this paper.

Adapting MLLMs for Nuanced Video Retrieval Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:18.933713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T22:20:09.051957Z digest=sha256:fdbefa86f054a0386741ddfc79aab44ce8a1a24388e81fcbde1eaff44bf3485f

Observation 245512fd-c346-4587-9d79-985335bf2157 · inbound

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning cites this paper.

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:02:42.441350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:02:20.477517Z digest=sha256:2e2d684944a5bd29d7981a6e5bfad22b91b673da7ddc2eb2e79dcc25dae4afdf

Observation acb55947-e818-425b-b074-60982699ea2b · inbound

SCP: Spatial Causal Prediction in Video cites this paper.

SCP: Spatial Causal Prediction in Video Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:50:11.211833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T16:47:44.523606Z digest=sha256:87b3b5d1b5d58ccb2b36a30728c6403e8c8f16b69d73321f7840ddf8b4ddb4a1

Observation 31ef9471-cb5c-43dd-924d-6b79638221e1 · inbound

Progressive Video Condensation with MLLM Agent for Long-form Video Understanding cites this paper.

Progressive Video Condensation with MLLM Agent for Long-form Video Understanding Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:43:14.858948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T20:40:41.380829Z digest=sha256:97ae385d496d196d88dcd0fac8a22026e6ee216b867885ac5cc5f3e5bd2015bf

Observation 3b876ced-1ea4-4dad-ac02-ba66176a193c · inbound

Building a Precise Video Language with Human-AI Oversight cites this paper.

Building a Precise Video Language with Human-AI Oversight Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.401243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:8a87573f5e5ba2cb26d3e596968de7754ca132651ea085cde5dc752e3b9a96ff

Observation 97a2f1b6-2f1e-4d05-bab6-e6c6324a2baa · inbound

Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstruction cites this paper.

Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstruction Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:53:28.934400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T01:51:57.018809Z digest=sha256:bc4fec2ed464646176b8aa35dbc0247aca9f8ec620a304584a8e3fa0f32ea263

Observation 38714149-295f-4abe-9ab5-d23f4216ba24 · inbound

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models cites this paper.

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T05:28:04.465276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-20T05:27:30.938311Z digest=sha256:25e11a451395b3afe18e7ae3206293e9bfbc5ac36d7d7add87e1a420313b1aac

Observation c55f7ec3-6104-41fc-870c-a9ab35371b3e · inbound

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning cites this paper.

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:23:28.434241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T13:13:57.599970Z digest=sha256:dc088950ce11568333ff63562ccb78d8b0cc204c4b730e251e9aa12bd8f620ea

Observation 779d8d9e-b6f3-4051-86a8-a7c4d19b3f41 · inbound

UNIVID: Unified Vision-Language Model for Video Moderation cites this paper.

UNIVID: Unified Vision-Language Model for Video Moderation Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:07:09.272703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T22:56:26.674841Z digest=sha256:ea3ee20b354db5e69e21307f94535d01d25f54c62c734f7553db17a54fa88094

Observation d58a1299-f53b-4cbb-8f78-1f7c9e333eba · inbound

MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models cites this paper.

MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:17:09.548729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:49:18.420491Z digest=sha256:12d74ca2a4d71b3ef8aa52e9583c83e0fbab50923eadb794194fa794344796b4

Observation 6909c904-fb97-4a95-aee1-5bbac09e37f0 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.770571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:dca06219ea8235e0c9d8f5dd27d83bad1c21b7fde6f159d09b1c3093e96a9c04

Observation f129b174-1b38-463b-a386-c40262bf78de · inbound

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs cites this paper.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 80

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.312897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:8a1bad3452c0871126f1a962281479b58cd7e6ecf153bd6c764d6713e560f1c5

Observation e33d93dd-136a-4769-8026-647da8b3bed4 · inbound

Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA cites this paper.

Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:57:30.176143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T16:55:35.743040Z digest=sha256:e749c8fe2ab71919395076503b0eb4a48c933360b2ac46efb0cbfad5ed8d834f

Observation 9b7429ba-3931-4175-861f-4f270f630a92 · inbound

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning cites this paper.

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:29.024209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T17:21:38.543724Z digest=sha256:6af4c468e87719bfe2228497e59a30c3775c083402eb204c924ff396449bf139

Observation 2902fe35-5cb6-4ef3-83a3-04289e077e83 · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 297

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:03.121610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:99d53a7b52a43b1cce8d0f3886ee5d1b8336f85af97828e468fdaf17fbac8f46

Observation 1efce4af-dece-45af-8191-c58b9bd5045a · inbound

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning cites this paper.

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:40:00.907967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-25T23:29:24.520537Z digest=sha256:c1179fc99b90c4376d5d6f9d23d33eeddce04d4460ee31a28623eb2fc59f849e

Observation 87c325a0-12db-4fe8-b67c-73b20cc29f7e · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.256819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:7e1f6d88fbc743df84818ba9a6df533ac593a4bf23cf628a4098c939bb818c87

Observation de2f74fc-0519-4b49-ba64-ce5e88f3af4a · inbound

MotionAtlas: Detailed Region Captioning for Motion-Centric Videos cites this paper.

MotionAtlas: Detailed Region Captioning for Motion-Centric Videos Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:24:21.158963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T07:21:46.783970Z digest=sha256:fd73c104d06a857ae06801cd1141ecb9b480ed98ee18ae122b3f0174b18066f4

Observation a0274c15-76c9-48f4-a8cb-a533aa800a29 · inbound

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning cites this paper.

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:38:39.648627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-03T16:37:06.384435Z digest=sha256:0f679d81f29750d8b5aeea8d4c979114d4eb6585f00f5463a54d3671cfeaaad0

Observation 1a17a28f-d18a-4b89-92f5-7f5b45e0daa2 · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 167

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:5874cd6ec1e248064108401a5867825d9199c0c8a5e37e99e84a0a29a5bb62d9

Observation 0f4661c7-a4c4-4aeb-8791-76af60bd4357 · inbound

PercepCap: Video Captioner with Structured Spatio-Temporal Perception cites this paper.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.243631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.243631Z digest=sha256:bfa41d5fe64e64e3a8716924d336aa951bb7a3f8459078a2321fef8875b6ec30

Observation 9643093e-bfed-4251-9083-b15b50c20ec5 · inbound

RefCaptioner: Multi-Reference Image-Grounded Video Captioning cites this paper.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:20.012837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:20.012837Z digest=sha256:a4efbacf8308dd32ad406d3d8ceab8c717d4d04dfb816cc0fd63551743766002

Observation 90db9274-bbf3-46ad-8514-45a615392547 · inbound

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward cites this paper.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.059127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.059127Z digest=sha256:d818c1fc0416e370e80dc578db2d8b772712323329fae1ff2bd9c8e70fe168ed