Pith. sign in

Paper Citation Record · LEDGER

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

As of 19 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 7 inbound Pith citation observations for arXiv:2506.05328.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05328 v2

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:27:05.708277Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T22:09:56.270516Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T05:45:56.102421Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact3
  • verified fuzzy15
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a684ffce-ddae-4dfa-b02a-13da027d9496 · outbound

This paper cites Qwen2.5-VL Technical Report.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.464312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.464312Z digest=sha256:4e352695f48591d77a64d59345cc0e26ef7d2dddfe0520d2cb7454d795006b91

Observation 6be6bd9a-0430-44a7-9f86-3859b73daa87 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.468595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.468595Z digest=sha256:bb3cb2a16071db4c8bdaec79896ea4186d0416742479370bdccd1d89e68cf0f7

Observation 22c657ba-6fc6-4910-9e5a-f65fe28c6634 · outbound

This paper cites Eagle 2.5: Boosting long-context post-training for frontier vision-language models.arXiv preprint arXiv:2504.15271, 2025.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Eagle 2.5: Boosting long-context post-training for frontier vision-language models.arXiv preprint arXiv:2504.15271, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.471940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.471940Z digest=sha256:94af82edd030e396ff453f2502cc4a1d0cac1338a8ae5b9da073e693cbe30dac

Observation e2307a22-5132-418a-bd8d-d6bf2f3373d7 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.475164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.475164Z digest=sha256:141ac09713a517a1460eda18234c5fa70c15e4e6b7d10fde18e12377c7199758

Observation 77bb1d4b-203d-486f-8adf-71fbb007e644 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video understanding.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Video-llama: An instruction-tuned audio-visual language model for video understanding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.565887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:27:05.478496Z digest=sha256:617cf71371dc3b35dd0e1275a99508c2b0cbf7b3304d80f78b388f73242b0f92

Observation aec4a5df-6f37-466a-955a-a43cfada4361 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.482122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.482122Z digest=sha256:cd0475d5a755b04c51c4bfcdceee3c44cf396c10334b236becd235c2520ad4db

Observation e816cc37-4f83-46ff-b32a-26b4eaafe715 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.486127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.486127Z digest=sha256:b4449fc459f4e6d97750d68d68713bab15ab280a24a5154490c153e5409b6a2e

Observation dbda99d1-540f-4f58-b15b-70c985135e06 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.489310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.489310Z digest=sha256:d6ebe8dc3a1fe5a16b32e45556585003480d7269b83e73b600337dcb61e2a568

Observation fbd5d1e7-aff8-4073-980f-3d9d450d7a83 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.492480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.492480Z digest=sha256:46ced7ceb5959435849070e55882c46dba48b362fcb3cc4650524f685d0fe8c6

Observation 0ba4e3ee-4df0-48c8-b7f6-ca213a40175a · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.495697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.495697Z digest=sha256:cfd155840dcea12e746de2938355593a79e3110c3a186a6a5fabd3c5bca8849d

Observation 5c096c6d-ac30-465c-a225-b79619168202 · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.499105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.499105Z digest=sha256:eac2e7949794fecf76d6287fe1212c5d9201f51568c1eb09ec0dd453be80661f

Observation 09444582-a7b0-46d5-9633-3341ba5b0202 · outbound

This paper cites Cogvlm: Visual expert for pretrained language models.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Cogvlm: Visual expert for pretrained language models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.502419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.502419Z digest=sha256:6eedb8279828ce27aea26e75027d5d9812683d412bfa4e4c91f969b6061e3e76

Observation 46c9b6b1-1bdf-47dc-9e13-ea4ac9dc1f34 · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs CogVLM2: Visual Language Models for Image and Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.505880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.505880Z digest=sha256:4a261e46fb724e18fcee3d6d6f6f633720f0aba62e9b3605dc0c87b2a4b086d7

Observation 5e2d9d77-a5ee-42a6-8377-2f5e94b38832 · outbound

This paper cites VideoLLM: Modeling Video Sequence with Large Language Models.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs VideoLLM: Modeling Video Sequence with Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.509262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.509262Z digest=sha256:df1cedfb833162cf615f53942a513a49f04291a0078d9b93caf19baaebfc406f

Observation 5e625726-1323-43d1-b2b6-16c47c50e13f · outbound

This paper cites OVR: A Dataset for Open Vocabulary Temporal Repetition Counting in Videos.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs OVR: A Dataset for Open Vocabulary Temporal Repetition Counting in Videos

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.512719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.512719Z digest=sha256:adecd2882ff696c8cf880fcd1141e3a49c836829ae512535ec9171e19de7678a

Observation 300bb223-9ce7-44fe-9ffb-674a621388d9 · outbound

This paper cites Dvd-counting, 2025.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Dvd-counting, 2025

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.548934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:27:05.516090Z digest=sha256:30a08615b3d2d434578e4fdaf7326e313894f53f9dda4198330f64bf5f37d66d

Observation 3a54d292-f766-44d4-a434-7cda7ddb0235 · outbound

This paper cites Counting out time: Class agnostic video repetition counting in the wild.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Counting out time: Class agnostic video repetition counting in the wild

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.519281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.519281Z digest=sha256:4410c82a06bf99151e5a2165a67bedbcd47d14202e131fce09606c54936f5d89

Observation 3fdb8348-6f9e-4a57-b55f-82e31989d5c2 · outbound

This paper cites Repetitive activity counting by sight and sound.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Repetitive activity counting by sight and sound

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.532936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:27:05.522916Z digest=sha256:3f932b30767750aa7aaf8a7d25113e868affd682cdbcec6afa5d5094eb851fe6

Observation 771c62c1-f2e7-4a53-aedc-b111cfc00a7f · outbound

This paper cites TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action Counting.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action Counting

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.526108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.526108Z digest=sha256:c09018047c3807f9ff1f9e0e9e88a8bd99d3ae34979a23a15661dfed6794ffc7

Observation 02223deb-a7f6-4379-8ff1-a37f4262b12d · outbound

This paper cites CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.529740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.529740Z digest=sha256:24fa69ab6b36d57cb34a66841beda99ac2ac75df52ec648a6e987897abbbd96e

Observation be396407-bbd2-4d64-bb62-6ae72aab8163 · outbound

This paper cites Ola: Pushing the Frontiers of Omni-Modal Language Model.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.533023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.533023Z digest=sha256:465101a1c8df19391c17293fc1d331c0152e3df743d5edc4fb78c5a2a9f3e3e5

Observation 806f6ff7-5b7b-4842-9630-86556c9f42b7 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.536387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.536387Z digest=sha256:b57f3fbd9f17c83700985ea42d0bace08fb9f06f16b62c4381c9dbd1c5694dbf

Observation d411759c-8069-4827-b3b0-eabe637af90a · outbound

This paper cites Curriculum learning for reinforcement learning domains: A framework and survey.Journal of Machine Learning Research, 21(181):1–50, 2020.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Curriculum learning for reinforcement learning domains: A framework and survey.Journal of Machine Learning Research, 21(181):1–50, 2020

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.539317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.539317Z digest=sha256:cecfc5f3334619fb111ae85ede481bbac5635d1aaeb0e07dd8458b5f67458dc8

Observation c7b9bcc2-ec63-4c1a-b970-4323e5492402 · outbound

This paper cites AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.542467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.542467Z digest=sha256:254ce8fc39c069ff3dfd41983b9c3aa34812ce4909ce298c8aec762f0fba8f2f

Observation 9e0e43b2-8ea1-4d4b-9536-cd3afc7f67f2 · outbound

This paper cites Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.545846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.545846Z digest=sha256:5c15958cb08cd794b22ccb326f6938fd449bd8c6e3c59b699ecdb6f73a4b6625

Observation 686c5244-6f31-443a-a4b1-34f22aab81aa · outbound

This paper cites Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.549332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.549332Z digest=sha256:351be0ba5c0eec47d2901065ce2a3dc99d818378625bbe5f0760506f0d7320d2

Observation addbeab7-af44-472e-9131-91bcf2215c69 · outbound

This paper cites OMCAT: Omni Context Aware Transformer.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs OMCAT: Omni Context Aware Transformer

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:27:06.046811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:27:05.553172Z digest=sha256:c32d34b5b15c12df0a60a4c136edf5d090d3183e383d33e2eb036f3d0a22e9ad

Observation 9ca97b63-d6e3-4293-87f8-684f8ed388ba · outbound

This paper cites Baichuan-omni-1.5 technical report.arXiv preprint arXiv:2501.15368, 2025.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Baichuan-omni-1.5 technical report.arXiv preprint arXiv:2501.15368, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.557034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.557034Z digest=sha256:d25eb1452e9a091d1f33e5983ab497df1f7c3b5c2642b0c9d1ebb71f37fa8571

Observation 18af2022-4246-4367-be8f-081bb7661fe2 · outbound

This paper cites Crab: A Unified Audio-Visual Scene Understanding Model with Explicit Cooperation.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Crab: A Unified Audio-Visual Scene Understanding Model with Explicit Cooperation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.560273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.560273Z digest=sha256:f1368198a2a46f6a10bdf9f6a2d8f7682d973675b4d484d48826d8f324781074

Observation 3a2e00d1-6dff-4ecc-9867-4cf0fc9abd08 · outbound

This paper cites video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.563770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.563770Z digest=sha256:dbcb9663c4f89547e4532b353e35af799fe06edf516d552ff1c748d6659f0b6c

Observation 5ecc1de9-e393-4d1f-bf8e-82aeb980cc0d · outbound

This paper cites Meerkat: Audio-visual large language model for grounding in space and time.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Meerkat: Audio-visual large language model for grounding in space and time

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.567346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.567346Z digest=sha256:bdd0f54f1cc79600ad6dfcd1e3f9475d4e826c9bb933213dd1a9f75cdf5ed3dc

Observation e178f6d4-3cc7-4795-9423-6e16a9a8f312 · outbound

This paper cites PAVE: Patching and Adapting Video Large Language Models.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs PAVE: Patching and Adapting Video Large Language Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:27:05.978456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:27:05.570467Z digest=sha256:fb658e8c5a5f7a37070937092b0271dc6a8ac048264744b4d6da20c84b29cdf0

Observation bf31b9ad-6fc1-401a-be87-937002b6cf52 · outbound

This paper cites Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.574186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.574186Z digest=sha256:6b8e2644d7c1d57f4bfeef11052f7bab1394fed66f9b51229234f7c474cd82be

Observation c6254dd5-267a-4a5b-a5c4-fc16f2043bbb · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.503625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:27:05.578385Z digest=sha256:aff4397da27159eea6168362bc5897f6f6358e65c44b16a9d29dc7caa0c7e3c1

Observation 84bd407e-a595-4e5a-a0a5-5fa754d236f5 · outbound

This paper cites WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.581493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.581493Z digest=sha256:ebe5f98618e387ca557c3f95979e7c59c80ff1654f9faabff61524d2782b2d1e

Observation 0b61ec07-d356-4e35-894e-09aab9193737 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Chain-of-thought prompting elicits reasoning in large language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.584562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.584562Z digest=sha256:6d2b8bd88c42f4642ef292dbb4f52112d34ee7afc272416150046656ed09b185

Observation 34bc1aaf-6807-4143-893e-f983d1033b79 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.587891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.587891Z digest=sha256:8341620aa893bdab22dc9c12c3514afc47fddf2b3aea1eb07159d8dfa1d0b644

Observation 7d06622a-75a5-4ce8-a4ce-e48d40e83409 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.591545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.591545Z digest=sha256:051b258c0a581acdb88e2c725271a738680db9503444cf090521db495803f5bc

Observation 79e302ca-b92a-4cb5-9391-af5c5397bfa4 · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.594716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.594716Z digest=sha256:3cde6f3ae456005298d9980bf35ca4e54d65fb32f5433979a1045cacecc9e4ca

Observation 8debd7c1-08a0-4b71-8179-253c635383c6 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Gemini: A Family of Highly Capable Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.598090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.598090Z digest=sha256:1d634c6ee6a8b721bad6372ec2fb1dc51cd32cd1a275660d3a4ce86977eedfef

Observation 6435e326-212d-41e0-b84e-7df8a253e2bb · outbound

This paper cites Introducing gpt-4.1 in the api, 2025.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Introducing gpt-4.1 in the api, 2025

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.487558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:27:05.601374Z digest=sha256:893d3258e13f8218fa5802d7ca75bce61b30dcc0d3a58631306103f16bffbfcd

Observation 6d514cce-0015-4127-85dc-032e853a8a3c · outbound

This paper cites GPT-4o System Card.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs GPT-4o System Card

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.604689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.604689Z digest=sha256:3c1753841c4a67a958466890412885ebad01725df0a10c0a1e8e9ccc07be3838

Observation 2d5d455d-553e-4098-aa0f-ffc0607bc6b7 · outbound

This paper cites Seed1.5-VL Technical Report.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Seed1.5-VL Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.608221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.608221Z digest=sha256:683974ba8a2c944ea0f1ac18c56aed1e3eb5ab9c68639b883056411bc08a59dd

Observation 8ee657c0-142e-46e5-abb7-0f73ee36fc1f · outbound

This paper cites Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.611836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.611836Z digest=sha256:8b044869d6d15302800d942d78916a37ce6f4821129686c1c49c5bd2fbdfd430

Observation 8cf9e3ab-02f8-44a7-a346-1c46941f6bbd · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.618482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.618482Z digest=sha256:6961f8eb65fd77de14242588ccc204e5c9edd455897fb6a0a3cd9e30393243e5

Observation 1746ded8-6795-49d0-871f-472260bf428f · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.621901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.621901Z digest=sha256:22085fcbed74ac27678a61a11c3060e4d77e4b7302ca19b4ec7840ec9b1eba22

Observation dc908336-520b-4e60-8078-7f6460e1aaf8 · outbound

This paper cites Qwen2.5-Omni Technical Report.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Qwen2.5-Omni Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.625899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.625899Z digest=sha256:c14dd9a868e412343643ecaab24191684660a6ea1a333b054cfe2535f92b84a1

Observation 27005343-43db-4f9d-864c-8a8fd57809ca · outbound

This paper cites Avqa: A dataset for audio-visual question answering on videos.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Avqa: A dataset for audio-visual question answering on videos

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.629557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.629557Z digest=sha256:514bf7e47d847c95af016520d6167734a86cf572783dea9c3614362898d58450

Observation 5a2c7e4d-0bfa-46c4-be84-db3f5867773a · outbound

This paper cites Learning to answer questions in dynamic audio-visual scenarios.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Learning to answer questions in dynamic audio-visual scenarios

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.633066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.633066Z digest=sha256:9df55c7aae5821da325b06bd810992cc52561cba88ea3c43533040694fc70dc2

Observation 872115b7-75bf-42ad-94ca-e787115101ca · outbound

This paper cites Audio-visual event localization in unconstrained videos.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Audio-visual event localization in unconstrained videos

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.463851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:27:05.636748Z digest=sha256:386a9ba0e8e8663b06dba5e992929bced5bfa8da53b0e019394b96212920bc82

Observation 2fcdf031-2400-4683-a35c-832a7b127ab2 · outbound

This paper cites Dense-localizing audio-visual events in untrimmed videos: A large-scale benchmark and baseline.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Dense-localizing audio-visual events in untrimmed videos: A large-scale benchmark and baseline

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.453872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:27:05.639954Z digest=sha256:a4b6b2e3168987a2abb94b7b903c9433a766082ed4363e154789189d1680d956

Observation 3ad08ab3-41c1-486e-8ab4-6402bbae9db1 · outbound

This paper cites Cross-Modal learning for Audio-Visual Video Parsing.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Cross-Modal learning for Audio-Visual Video Parsing

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:27:05.828372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:27:05.643589Z digest=sha256:5973321ab8bed73d5eaeb5377555f73754d56f0701eb00f71af00f90ee4586ba

Observation fe57144d-5fa1-42e1-8c75-b318be294a27 · outbound

This paper cites Audio-visual segmentation with semantics.International Journal of Computer Vision, pages 1–21, 2024.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Audio-visual segmentation with semantics.International Journal of Computer Vision, pages 1–21, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.443673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:27:05.647139Z digest=sha256:e801a3590d66f9af40359370844b1ec1a6420c17bb57a1de2a8df33543dd487f

Observation 100a54ff-3995-45cd-b92f-440fe0f4a08f · outbound

This paper cites Enhancing geometric factors in model learning and inference for object detection and instance segmentation.IEEE transactions on cybernetics, 52(8):8574–8586, 2021.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Enhancing geometric factors in model learning and inference for object detection and instance segmentation.IEEE transactions on cybernetics, 52(8):8574–8586, 2021

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.433369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:27:05.650416Z digest=sha256:abb5d1c65a91a0e3e7a86d77c46bf135b030d4ffd073bb734a61a7e922dd38dd

Observation cc351f6b-87d6-47af-a2fa-68c710048516 · outbound

This paper cites URL https: //www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_ Card_Claude_3.pdf.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs URL https: //www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_ Card_Claude_3.pdf

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.423501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:27:05.653489Z digest=sha256:0323021ac9d6420bc16a8f46821b63fa6677d220667f609dd2069050e33d385b

Observation ab692d4d-79fa-40b9-b859-571f1388fd5e · outbound

This paper cites Onellm: One framework to align all modalities with language.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Onellm: One framework to align all modalities with language

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.657817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.657817Z digest=sha256:298882f7a3019f80236de59910f0b96bea9d5d214332400b582475779bb558a8

Observation 71987732-166b-49b7-b8f0-00bddd142d99 · outbound

This paper cites GroundingGPT:Language Enhanced Multi-modal Grounding Model.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs GroundingGPT:Language Enhanced Multi-modal Grounding Model

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.661256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.661256Z digest=sha256:8d1a4f984107f969b3160087ad67e0c399cbeb0fd71074358d86e56a219140af

Observation 6005d5e9-493a-467d-9ed5-a8d3c08c7d5d · outbound

This paper cites Valor: Vision-audio-language omni-perception pretraining model and dataset.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Valor: Vision-audio-language omni-perception pretraining model and dataset.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.407310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:27:05.664610Z digest=sha256:cf0cfcb6368f63c8bc0a5ace30c0a74fac58d3fb949c23dcbe94eafcb6f05804

Observation a29cf447-a518-4291-b3eb-d092abc80390 · outbound

This paper cites X-instructblip: A framework for aligning image, 3d, audio, video to llms and its emergent cross-modal reasoning.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs X-instructblip: A framework for aligning image, 3d, audio, video to llms and its emergent cross-modal reasoning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.397918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:27:05.667783Z digest=sha256:03887827052398164f1650bc59116c2137d217c4761d231cc0e42548be95dd0d

Observation c550d55b-4f9b-4f2d-89a0-50e4a988721a · outbound

This paper cites Avicuna: Audio-visual llm with interleaver and context-boundary alignment for temporal referential dialogue.arXiv e-prints, pages arXiv–2403, 2024.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Avicuna: Audio-visual llm with interleaver and context-boundary alignment for temporal referential dialogue.arXiv e-prints, pages arXiv–2403, 2024

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.387217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:27:05.670950Z digest=sha256:2535cf9cc2c8c4c1936e302e9440b2888a080bf1084ba8dc0a1fa6dbfb53430f

Observation 484f6ce8-fdd9-4447-b52b-7e1f6f78b46d · outbound

This paper cites AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.674331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.674331Z digest=sha256:43d8e9de43ce0981a81e41634198739762796cb7495a8eed4472432b0a9d3044

Observation 0b5835ec-d2b5-4b49-8558-da2721948a9c · outbound

This paper cites Omnibench: Towards the future of universal omni-language models.arXiv preprint arXiv:2409.15272, 2024.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Omnibench: Towards the future of universal omni-language models.arXiv preprint arXiv:2409.15272, 2024

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.677898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.677898Z digest=sha256:7a1b7453edb4fc1283e1b5b294449bffb9010c97edb1a0421bf24b06d1491130

Observation 5dc4b245-3c20-459a-9398-0aedc32eeaa4 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.680865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.680865Z digest=sha256:2e7fed60c8bd4ce92fb86f4f3206eb478173116dfcb5cbe0862a8030724f4471

Observation 0b62d6ce-670d-4aac-9cc7-e14526b3520e · outbound

This paper cites How many people spoke in the scene showing the conference table?.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs How many people spoke in the scene showing the conference table?

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.377743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:27:05.684003Z digest=sha256:fff5730e043f3760839b54c4930f5377e111a2a51ceeacd557143cb6fd07d5e1

Observation b2e7becb-f94b-4077-a055-9a04c361c860 · outbound

This paper cites an unresolved cited work.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:27:06.367572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:27:05.687707Z digest=sha256:4ce54ff8cfb28f3abbe9c0cb9acfa1c8b061600db44abe0e445bf15c410934d8

Observation 466bba4f-c527-4377-93ee-848320097d79 · outbound

This paper cites an unresolved cited work.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:27:06.357940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:27:05.691318Z digest=sha256:061b2f4683ae5b749dbac07fedc5c67cd9dc183fa7a31d1634dc1ce26ff636d1

Observation 6181835e-8342-4651-8f75-1f781fa56004 · outbound

This paper cites an unresolved cited work.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:27:06.347782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:27:05.694655Z digest=sha256:10fd86c0b1517996c4d7ec1c369b95433dc7a7066952b18505d146197a1816cc

Observation 24fcc273-cade-443a-a195-8b8c300f5146 · outbound

This paper cites an unresolved cited work.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:27:06.338960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:27:05.697979Z digest=sha256:29fc1256c75767bd6389264236857e92706f9e52f5f3d488a952e5401e84e792

Observation 4cc85d94-89fc-4989-bc42-0e2dc4b9a040 · outbound

This paper cites an unresolved cited work.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:27:06.328979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:27:05.701456Z digest=sha256:c62b76778a9577758d65cc383ef21c9f5ef03fc0a2ee45c30328ddcf39ef8b65

Observation 30856085-49c4-4a07-912b-1ea40b179069 · outbound

This paper cites an unresolved cited work.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:27:06.319383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:27:05.705052Z digest=sha256:68764c54949f423830a4ea443af948707e9426b23142bbf06e0fdb0a61eeb7dc

Observation 476f044f-235e-4f69-b4dd-82d77e9b3a79 · outbound

This paper cites question.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs question

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.309524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:27:05.708277Z digest=sha256:f19a4f87cf62515c73551ae4694f165087bd4460eb15732930438b6dea95a2a7

Pith citing papers

Observation 9c2be27d-d82f-4e75-b00e-629425ae9f37 · inbound

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models cites this paper.

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:45:56.105299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T05:45:07.700571Z digest=sha256:f95706e15aa8453f737aa217350798333d54094d4206d14f14d1e756c2aff025

Observation e3189c56-f1bb-4f92-8354-b7a926ffdf6e · inbound

SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance cites this paper.

SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T22:09:56.270516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:09:56.270516Z digest=sha256:075da52b4e185d9a9002cf45c0370809fa17f5911cfd0628198ecf3f5f7a8395

Observation e4baaf95-ed9e-476e-b0ff-3904fcbc7312 · inbound

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs cites this paper.

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:22.142010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T12:05:54.551728Z digest=sha256:bba05ee2e47bb54c1129b6e66bd74e3da860a427154488d58a195030d2220dcb

Observation da73e082-7283-4c33-aaae-e153c39db0af · inbound

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers cites this paper.

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:18:32.125415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T08:02:53.574120Z digest=sha256:bb3dc94ba523d3c25448e74636dc07215a7eaa686a15c6c6a20285c8267e9aa6

Observation 11430591-9fb8-4725-9863-066575a13626 · inbound

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation cites this paper.

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:52:12.698673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T03:49:58.240883Z digest=sha256:6433caf973c131447bc0a0da8d68e9755c743576d03b2c6c5635f841cb6510e1

Observation 2449202f-3603-4a76-9f9c-6f41ce858e3d · inbound

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation cites this paper.

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:09:50.259437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:06:20.658030Z digest=sha256:49659c967b81a8834bad7d61490c35a16ccf5d21f9bf9ecb635297f0b313ac69

Observation 15ec56f8-173e-4780-b416-249258d469bf · inbound

Empowering Long-form Omni-modal Understanding with Robust Audio Perception cites this paper.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:ea7f9fe5dffb785153b240c57e7d77987c39a25d3f781951280844ea7f0ed149