Pith. sign in

Paper Citation Record · LEDGER

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward

As of 14 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 0 inbound Pith citation observations for arXiv:2608.06930.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06930 v1

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T18:13:44.573450Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

75 of 75 outbound references displayed

  • verified exact2
  • verified fuzzy26
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4eaed5f7-1f2a-4ff8-8f8f-92924844fd08 · outbound

This paper cites Avocado: An audiovisual video captioner driven by temporal orchestration.arXiv preprint arXiv:2510.10395, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Avocado: An audiovisual video captioner driven by temporal orchestration.arXiv preprint arXiv:2510.10395, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.001303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.001303Z digest=sha256:ad421d5cd3cd54022e9ae649c98e4924194ca38afacb5cf055555b1f975dd98c

Observation b695bdef-a788-4514-949d-8065a0a98e1f · outbound

This paper cites Ugc-videocaptioner: An omni ugc video detail caption model and new benchmarks.arXiv preprint arXiv:2507.11336, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Ugc-videocaptioner: An omni ugc video detail caption model and new benchmarks.arXiv preprint arXiv:2507.11336, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.012944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.012944Z digest=sha256:8d50b940b28dfd336b745bbcdf0f290e7fc9d0ea8ea0a25be6f74a02cb88b35f

Observation ad1b155c-49ef-45af-b7b2-78dea5511969 · outbound

This paper cites Daily-omni: Towards audio-visual reasoning with temporal alignment across modalities.arXiv preprint arXiv:2505.17862, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Daily-omni: Towards audio-visual reasoning with temporal alignment across modalities.arXiv preprint arXiv:2505.17862, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.020242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.020242Z digest=sha256:bf5622e627c8c913652002106dc2acd2cc6e6a17d09254ed08732e65a14f45ac

Observation f454879b-db8d-4894-b6a7-98084bef545c · outbound

This paper cites VidCapBench: A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VidCapBench: A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.027540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.027540Z digest=sha256:69743742ff07a4b154b42410e28b365ff7ec0b1b55cef6e9063f78029b3a0937

Observation 8f6e4ae1-519f-4569-be74-61cf0eed85a9 · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37: 19472–19495, 2024.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37: 19472–19495, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:47.162755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.035846Z digest=sha256:2d55f601828c7be75a5987d12742dc39953ffc6d2fc17fdf7d81b107557c85a5

Observation 55a344b0-2078-444f-9837-3581bfcb57db · outbound

This paper cites Mavors: Multi-granularity video representation for multimodal large language model.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Mavors: Multi-granularity video representation for multimodal large language model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.043255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.043255Z digest=sha256:64576cce2ab4b4d4c66787f50ef7c24c2135ab61c489a24aeff8a31c4c777cc0

Observation 5c73a9e0-e967-446c-a679-d61cfd231a08 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.051144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.051144Z digest=sha256:515a4866fbbf0996c8089baad3e377afb11cf838009c60cb16386483b4e1cb7f

Observation 90db9274-bbf3-46ad-8514-45a615392547 · outbound

This paper cites Tarsier: Recipes for Training and Evaluating Large Video Description Models.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.059127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.059127Z digest=sha256:9a1b70ea7337d314a34ad2757ce2623352640176084cb1ee4758a1a9d297a5d9

Observation b84cf9f9-f405-441b-9a80-a8a0db9a6853 · outbound

This paper cites Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.066103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.066103Z digest=sha256:d99099ac53f4a380acbf8867b7c4cfb34e880c7563242524e816239529404b17

Observation ac750a63-35b8-467a-9019-cfa9d9a1d905 · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.074176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.074176Z digest=sha256:b969c0828e6545bb85aa9a58923071337dbadecc8f8bb99d49630d95acb64c75

Observation 727dae58-6d8b-4478-bb13-886fcd349312 · outbound

This paper cites Advancing high-resolution video-language representation with large-scale video transcriptions.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Advancing high-resolution video-language representation with large-scale video transcriptions

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:47.108399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.080659Z digest=sha256:33c72c83035eb850d4b9b4ae1dbe7515e419625f142aa61f7f83c9b2c7c68f52

Observation 6995e3e9-3e81-4806-af86-eef40d176bdb · outbound

This paper cites video-salmonn 2: Captioning-enhanced audio-visual large language models.arXiv preprint arXiv:2506.15220, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward video-salmonn 2: Captioning-enhanced audio-visual large language models.arXiv preprint arXiv:2506.15220, 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.087226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.087226Z digest=sha256:f28cf8b6dd0f4ae66426b9781a40e4f524fc7785ddb0bbbeb129ea0013a1bdcb

Observation 7f752d0b-4aad-460d-9aba-c565c4d88bb7 · outbound

This paper cites AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.093438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.093438Z digest=sha256:4d6aebaf44dd1c701c24ee69b07f2b347dbbffcb8554dd7bc2d0d1464372ffcd

Observation 70ff3cbf-52e2-4513-b5d7-e931862558b8 · outbound

This paper cites Longvale: Vision-audio-language-event benchmark towards time-aware omni-modal perception of long videos.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Longvale: Vision-audio-language-event benchmark towards time-aware omni-modal perception of long videos

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:47.086591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.099582Z digest=sha256:d34677a46e84947b78af2dc612bf8769e44c7179c072680f07e33d43e91ca9c2

Observation 8db4405e-fa54-4f54-95d9-ff440f26235c · outbound

This paper cites Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.106889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.106889Z digest=sha256:92adba695ef0edb7a7900b694a6e34951c9851d547008d5d885677bcc13cbf60

Observation a402971b-163d-41f1-bf6a-5779e77e4a33 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.113320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.113320Z digest=sha256:4768c411c2daebb22034301af3d67c15981abaf829106fbf8915a305fef4e51e

Observation fafedc81-20fb-4608-872c-4b5035d3bde9 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.120212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.120212Z digest=sha256:8f84b3cf809bd39de151dca4f63c071cdd2a2410ac5436adf18aa002d02a990f

Observation 87b8dd1b-90a8-4b20-bc27-6bb85e24b4a2 · outbound

This paper cites Watch, listen, and describe: Globally and locally aligned cross-modal attentions for video captioning.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Watch, listen, and describe: Globally and locally aligned cross-modal attentions for video captioning

Reference 18

Resolution
verified exact
doi, observed 2026-08-10T18:13:44.635267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.126805Z digest=sha256:5c5835c704efdd0109b6cb3c349ca72b9875af2579f5ccecb67404b67f6cf514

Observation e10c871d-faef-49c8-bac8-ad08c41aec11 · outbound

This paper cites VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.133036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.133036Z digest=sha256:ea9b95b1488d821c64531d29250be575f39ee3d2835ae0a22b21ca7e1a4e30f8

Observation 3a17029f-7952-4eac-a5f8-50f77b513d43 · outbound

This paper cites Qwen2.5-VL Technical Report.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Qwen2.5-VL Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.140160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.140160Z digest=sha256:921532fe77e9024a3a82bf74adaf44074ef853a53f9832905c8732fcce984cbd

Observation 2d47876c-b084-4857-a211-94c6eba7760b · outbound

This paper cites Glave-cap: Global-local aligned video captioning with vision expert integration.arXiv preprint arXiv:2509.11360, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Glave-cap: Global-local aligned video captioning with vision expert integration.arXiv preprint arXiv:2509.11360, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.146443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.146443Z digest=sha256:e7d743820c3293535f0e28fc11af42e497cabcc453b9ac50e5c542cddc200607

Observation f8f1a29e-aef0-4ac8-9db6-e8e868898a3a · outbound

This paper cites Qwen2.5-Omni Technical Report.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Qwen2.5-Omni Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.151110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.151110Z digest=sha256:6976459d9f62e2005e3ee7f64abf7323e9d91ab12cdd59b1357a4b4bb507d523

Observation ead60c4d-46c1-4950-977f-efa478de054f · outbound

This paper cites OmniCaptioner: One Captioner to Rule Them All.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward OmniCaptioner: One Captioner to Rule Them All

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.157674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.157674Z digest=sha256:92358d9cc3cb0ff807e9e71e6e422a2c7042842a436ff23913c785c3fe87c163

Observation d1c38d65-50d3-4629-b59f-5b52c5a203e4 · outbound

This paper cites Omni-captioner: Data pipeline, models, and benchmark for omni detailed perception.arXiv preprint arXiv:2510.12720, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Omni-captioner: Data pipeline, models, and benchmark for omni detailed perception.arXiv preprint arXiv:2510.12720, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.163332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.163332Z digest=sha256:f9728594101b1c4f8966900965f4fd733a2e4edd96abebc6eb6fd93504b859ee

Observation aab17aee-273f-41a8-9111-9ea95809ed00 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.169838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.169838Z digest=sha256:597b4537e59968fcb129fc58d0c1cc7009efc1ff9f6ee13a675eb3ec62c7505b

Observation 91ab4e8d-8701-43e7-a6f1-180c2e0142ef · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.176917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.176917Z digest=sha256:30aa89d0953453d3c0012a086ddc2e7b0e3b619a313cf09f6e8001498e0f95ab

Observation 1dfeba4d-990f-4466-961f-7991cdc82a8b · outbound

This paper cites AdaTooler-V: Adaptive Tool-Use for Images and Videos.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward AdaTooler-V: Adaptive Tool-Use for Images and Videos

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.183615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.183615Z digest=sha256:f2cb84c435dad6b225858ce2016df8f197b9c33f639afe5adf520541b0d90923

Observation 0e486e27-48ba-4008-87db-37a63dcf1e18 · outbound

This paper cites Deepvideo-r1: Video rein- forcement fine-tuning via difficulty-aware regressive grpo.arXiv preprint arXiv:2506.07464, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Deepvideo-r1: Video rein- forcement fine-tuning via difficulty-aware regressive grpo.arXiv preprint arXiv:2506.07464, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.190682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.190682Z digest=sha256:ccb16d2cc32a2a655aa684827986dd66510afab1d88e22b888b9b36bfd64a721

Observation 89671b3f-be56-4556-bfb1-d94ce96e8412 · outbound

This paper cites Editthinker: Unlocking iterative reasoning for any image editor.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Editthinker: Unlocking iterative reasoning for any image editor

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.197366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.197366Z digest=sha256:5414b25f2d772f53c016fb7e06f1ed1a92862aa64b452c5e64493a1ed5d1409e

Observation 7e3a51c8-22b5-4d13-a606-cc7a0e510022 · outbound

This paper cites Sophiavl-r1: Reinforcing mllms reasoning with thinking reward.arXiv preprint arXiv:2505.17018, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Sophiavl-r1: Reinforcing mllms reasoning with thinking reward.arXiv preprint arXiv:2505.17018, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.204027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.204027Z digest=sha256:d0642ac9717b3d89c0e2aa6b3a035de9b2873ea14a95e322b15af89c545a47e0

Observation 3ec3c7c0-87f4-447b-885a-edc9ca334b31 · outbound

This paper cites Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.209331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.209331Z digest=sha256:d46ab1bbcc9c0dbb8fc51f56b35a45ba2b852a0c46864f4fc5a2405c9a3e8c2d

Observation cf7d5b55-0444-40c4-aacb-a628a9681f73 · outbound

This paper cites Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.215349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.215349Z digest=sha256:f36461b3ee4b9bcf03b9824ba8c86ae871c03dfdb93d08443a4e5ca21bee7e0c

Observation 65539722-4660-4024-be93-223a12fa8d76 · outbound

This paper cites OneThinker: All-in-one Reasoning Model for Image and Video.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward OneThinker: All-in-one Reasoning Model for Image and Video

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.221576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.221576Z digest=sha256:e2d54fe9ee802eef4711bb9e5908ee59379cb7c9fbad36f4fc3624428ae3b7c1

Observation 2eb3994f-034a-49ca-a654-13259944ce9d · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.228889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.228889Z digest=sha256:cfaae8aadeba5bf653c25d065ff8753bac84891a34037e3fdbaeabe1672c68fc

Observation 53fcffc3-aa5f-49b5-ac92-20dd6ee28176 · outbound

This paper cites VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.235558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.235558Z digest=sha256:6b13217435c4192edb970ed25181623a277a6b3494d121275c30db25f2925989

Observation d031b374-8983-4f7e-a325-cc8741da424e · outbound

This paper cites Exploring the role of audio in video captioning.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Exploring the role of audio in video captioning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:47.067484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.241601Z digest=sha256:2c35ff66794451a5c65e02c856c0cffae6d5e2cf6c563547481c7bf38c8dea35

Observation 81e8b0ea-9029-44e0-8c6e-06c0df6bdacb · outbound

This paper cites Hybrid transformers for music source separation.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Hybrid transformers for music source separation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:47.049145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.247461Z digest=sha256:7dc5088a03a1420cb3241c548292be837e9fde978d2032047d661a5dfc65fa02

Observation cc0ed116-12d4-4b94-929f-bb14edf82074 · outbound

This paper cites Audio-visual event localization in unconstrained videos.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Audio-visual event localization in unconstrained videos

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.255613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.255613Z digest=sha256:79c01ce7f184183dda3a1924ffaafc12d2a9d2fff65eb68913f1b9d7696fb73d

Observation f3944060-20a1-4e43-ac79-d894f5e3ef54 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Vggsound: A large-scale audio-visual dataset

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:47.018344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.261995Z digest=sha256:e6832f73020b01f9fc1428e0a4517d86ac1e2c9bfeeea163fa0b4cb7242ecba3

Observation 3ca62a8c-3c7f-419f-b375-76f43901be93 · outbound

This paper cites Condensed movies: Story based retrieval with contextual embeddings.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Condensed movies: Story based retrieval with contextual embeddings

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.268021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.268021Z digest=sha256:45f1d90ddebd090d547d467217c4dd2488032616bdbab3ae74a18a0eb640b3c7

Observation a4d9dd4f-6e26-4315-92f0-725f21550be2 · outbound

This paper cites Avqa: A dataset for audio-visual question answering on videos.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Avqa: A dataset for audio-visual question answering on videos

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.276994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.276994Z digest=sha256:c1d724f62c4ffaee2143024b1615ca8b4199ca9e957056c0e71e72010e83b626

Observation fee2fdf0-5625-4f0b-a12f-fb4d2404f25a · outbound

This paper cites Movienet: A holistic dataset for movie understanding.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Movienet: A holistic dataset for movie understanding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.978977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.284058Z digest=sha256:6b275f9e6deb63202179881caa628cb866b26a93045dde0bf90b0ff45926c630

Observation c4d13c8f-2283-41b8-ba93-d06bd766d3e0 · outbound

This paper cites A dataset for movie description.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward A dataset for movie description

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.962720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.292387Z digest=sha256:5da06f0c6c9627af93cde3c0fe50502be5ff7ec57277e4eac4a217584d38b820

Observation 09d3fbeb-410e-4146-9ec5-ad05891eb23f · outbound

This paper cites Auroracap: Efficient, performant video detailed captioning and a new benchmark.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Auroracap: Efficient, performant video detailed captioning and a new benchmark

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.946632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.300777Z digest=sha256:fd0cd6bd096a1ff31eb7197cfc7fd6f3a8ef270b65c1d0ef0c4a351336f9ec32

Observation cdd33149-e7cc-4cc7-89a3-72bc5fc4f766 · outbound

This paper cites Time-r1: Post-training large vision language model for temporal video grounding.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Time-r1: Post-training large vision language model for temporal video grounding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.929111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.313655Z digest=sha256:b93fe8669c3cf0ebbd4f874b3a2969fff7f305562e335ed4793618915df5d2ec

Observation 41f95256-3d3a-4c1f-a884-73c828d10848 · outbound

This paper cites Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-10T18:13:44.864443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.320002Z digest=sha256:8efa150b6069eb0a575a3d23e27107d107c1bcfaa5dac18052fc563d1972d5a1

Observation d6ea50e2-06d0-4b15-981a-95988c63ed04 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.329322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.329322Z digest=sha256:d3ad097fa73269100fb02ce735b4f8833e250aaa505b0ccea757f2a8a6f51d2b

Observation 324bbfe1-68ca-488a-968a-de6c43b9bfcf · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1 (2):3, 2022.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Lora: Low-rank adaptation of large language models.ICLR, 1 (2):3, 2022

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.336856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.336856Z digest=sha256:88017e936b7e1da4c3c132b469617ebe50c32e6a6d1364e21bf7d3cee92cd126

Observation fa408bef-362d-41bb-b75f-70b2b82f91e1 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Efficient memory management for large language model serving with pagedattention

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.342904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.342904Z digest=sha256:389a785c8fed238e6b6ad09bc79eebd1a8db8fb5713e255c4c074f7f71cbd51e

Observation 455e53d5-8b79-462a-8263-e33b9bceb9df · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Vbench: Comprehensive benchmark suite for video generative models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.369645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.369645Z digest=sha256:d496105a0c66b74182eaa67bc009a828e4e8d43ba4294fb7a75f4302553d46fe

Observation 367cf444-21a9-45a3-bae8-ce99e94b6814 · outbound

This paper cites VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.376026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.376026Z digest=sha256:90b32e8b03731054d5e65640c16bd327212bf006ca55f29cebdbd888eec1767c

Observation 1074ca18-4c48-469d-a37d-db46b91920bc · outbound

This paper cites HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.385618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.385618Z digest=sha256:f60796ce436441faa138875cf5dfab6849b604d2c55770b7e1c5ffa0cf440917

Observation 8aeb4fe0-9bc3-4f5c-9714-ed4492a388df · outbound

This paper cites ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.392418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.392418Z digest=sha256:ea3755f138b37e169dcdd939f88f9d4f387b84e01f39a8ee58909a7c64fe354f

Observation d59819bb-9364-46b4-a338-0f3746d8b9a2 · outbound

This paper cites Minicpm-o 2.6: A gpt-4o level mllm for vision, speech, and multimodal live stream- ing on your phone.https://github.com/OpenBMB/MiniCPM-V, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Minicpm-o 2.6: A gpt-4o level mllm for vision, speech, and multimodal live stream- ing on your phone.https://github.com/OpenBMB/MiniCPM-V, 2025

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.876062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.399890Z digest=sha256:2db0f2055b6800f5d615217f6ca1974f74f8f8827e1f06bf05fbc84725235356

Observation dc9de645-8519-4180-9fae-51dfd07e6618 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.406841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.406841Z digest=sha256:15b8edb01e0959d23a1641b3b30f04c3f2abd4e2a954c8e10ac00379405e65b4

Observation be280df9-dcdb-4e6d-9846-4a4242ff3b46 · outbound

This paper cites Self- critical sequence training for image captioning.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Self- critical sequence training for image captioning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.414464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.414464Z digest=sha256:b3ab9ca553213f1b5943289456a22a2dea53475b4e46ce3a514b6425eaf76302

Observation c49af00d-fb79-4ee8-ac0e-6cd6f366ea28 · outbound

This paper cites id": "sample id.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward id": "sample id

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.848800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.422217Z digest=sha256:76ed9715f4cbe857f58d108ad8a87b1fb6154e6950839fc3f3fea7f77894e100

Observation da44a066-d0cc-4715-a367-9fcfba7c58c0 · outbound

This paper cites Use the exact terminology found in the text (e.g., if the caption says ”shatters”, use ”shatters”, not ”breaks”).

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Use the exact terminology found in the text (e.g., if the caption says ”shatters”, use ”shatters”, not ”breaks”)

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.829527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.429120Z digest=sha256:b0ba8da3178e58ab8a82400829398181f0b2384ff1ccb88a07a6f7f59eb7f0ce

Observation 31b2fb4b-b1be-4066-88df-0ed9d35bbe6f · outbound

This paper cites an unresolved cited work.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:13:46.809850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.441284Z digest=sha256:0bd4f8768b7a338735a70df45b8ab5a5f4d31f3ea84ac0f0cf597bad6d29b4f0

Observation 06ff5225-4b67-4407-a027-13511abd61a4 · outbound

This paper cites Start directly with the answer content.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Start directly with the answer content

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.788962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.450979Z digest=sha256:0dccf85a804586c9f6f97f9313fc71cb27572be0ff9326d8970799ea1384e817

Observation fa9669e6-ff5d-45ad-bb6c-b38be7b63981 · outbound

This paper cites strengths.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward strengths

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.764825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.457246Z digest=sha256:157e6906ffdef62fb52e46e8e0b27f37518ae7bdae7166b88579d2c84837c717

Observation 4a20e370-b5df-4187-8a70-e124cda1f609 · outbound

This paper cites an unresolved cited work.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:13:46.742112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.464484Z digest=sha256:8a0b57d807d7c8f0aba8b76a6ff62bc4654a546f5129448f488159188e190d2c

Observation bee123fd-afc5-43ed-bb1d-2bd4514bc827 · outbound

This paper cites evaluation.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward evaluation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.717454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.473015Z digest=sha256:2b55ac3d175afae87c259d8cbe9c315aed54075309154706cbb20bb5a406347f

Observation 6f54366b-2833-4bff-a4bb-38ad54d5fcc9 · outbound

This paper cites Do not add any artistic interpretation, subjective analysis, or infer any character’s internal 23 thoughts, emotions, or intentions beyond what is explicitly visible or audible.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Do not add any artistic interpretation, subjective analysis, or infer any character’s internal 23 thoughts, emotions, or intentions beyond what is explicitly visible or audible

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.692726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.480978Z digest=sha256:f1aa6e56d12754aec5340767c7b7298a07aa1b1aa4b8b70a203284371b91de47

Observation 90a1999f-d774-4123-a162-e1bbae85ba49 · outbound

This paper cites While Character A is speaking, Character B simultaneously turns their head.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward While Character A is speaking, Character B simultaneously turns their head

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.668462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.486950Z digest=sha256:b6a8b6f5744e245e720a56e7a9808616611ee5a9728ee3f81541ba65cf6d306c

Observation 36b615f7-5d65-48b6-b411-84544350326a · outbound

This paper cites He closed the door,.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward He closed the door,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.648814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.494599Z digest=sha256:daff8a2b2b9b8a01f3f013ba4eddba3fed855b08d0566c9c5b49ca9028b06283

Observation a39cd8a8-82c3-4890-ab62-d46deb18892b · outbound

This paper cites The camera cuts to a close-up of her face,.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward The camera cuts to a close-up of her face,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.628650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.502462Z digest=sha256:b1a13d5d7b7a99a7406e7d62126ee81ff763b56164228861051b83d1cfb3c43f

Observation 52b76ef2-21c0-4633-a1fe-87058085ec5f · outbound

This paper cites an unresolved cited work.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:13:46.608847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.516642Z digest=sha256:6c28f9076fea57da038f7ae7d700f3f105ec8882c77303a18f8dea1feca81c56

Observation e8eebfdd-7bb2-4176-9ed5-f3a690bb0f80 · outbound

This paper cites Do not use bullet points, headings, or line breaks within the narrative.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Do not use bullet points, headings, or line breaks within the narrative

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.577670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.523799Z digest=sha256:ff54d3478e70ac6c7716a03a90a4eef16dcf2627ce2f75627e5ce768ac056522

Observation 70b54954-8d96-4140-823a-14fdffbda128 · outbound

This paper cites His voice, thick with sarcasm, says.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward His voice, thick with sarcasm, says

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.553912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.531066Z digest=sha256:868e8e9d11d2f3a533610d3e8f196b9e7d1c0080a0f85eb5b00dfd6c74baacfe

Observation 90c315f3-909b-46c3-93ad-2d8bb8a5336c · outbound

This paper cites thud" as a book hits the table, a.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward thud" as a book hits the table, a

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.526057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.543586Z digest=sha256:fff6c570083f7d8ac654512c357f265f59d956bdd7b8d1d5e4ce1bc2d6f7a30a

Observation 64b2eb73-aa40-48d3-8df0-360d4fdca671 · outbound

This paper cites While Person A’s hand reaches for the glass, Person B’s eyes dart to the side, and the music swells.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward While Person A’s hand reaches for the glass, Person B’s eyes dart to the side, and the music swells

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.499941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.557179Z digest=sha256:f41f3d1450dfda88f8f3997fc713138b28f9fdd041166af82d78cddad7ac3056

Observation 0092d044-e526-435f-8cb0-7f3df378f2fb · outbound

This paper cites Avoid summarizing.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Avoid summarizing

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.471133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.563308Z digest=sha256:39dd892dacfbef5511872b568aa5f04e8aa705ed72a79f2c4fdb7d8671ee8800

Observation 8457dc9e-0c4b-4dcb-8fe2-d16d504a83a2 · outbound

This paper cites strengths.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward strengths

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.445294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:13:44.573450Z digest=sha256:3ce55b339b45de0431d85268d64a6257f2ecfb7d258c4b896ffeb986dd082086

Observation 271c8407-3349-4ac9-8db6-390573dd5ff2 · outbound

This paper cites WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.362349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.362349Z digest=sha256:61d647be2f2f79c425ca98085477b306881ea12d296b2673a551794459d7da04

Pith citing papers

No inbound Pith citation observations are available.