Pith. sign in

Paper Citation Record · LEDGER

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

As of 7 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 7 inbound Pith citation observations for arXiv:2506.05328.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05328 v2

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:27:05.708277Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T22:09:56.270516Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T05:45:56.102421Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact3
  • verified fuzzy15
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a684ffce-ddae-4dfa-b02a-13da027d9496 · outbound

This paper cites Qwen2.5-VL Technical Report.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.464312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.464312Z digest=sha256:df24e71ce52884f416b66ab154927f26bbf8b1be123fef252664bccb0ce9e01a

Observation 6be6bd9a-0430-44a7-9f86-3859b73daa87 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.468595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.468595Z digest=sha256:99f8ea5c385c8963a649e4309f92d3ca52a86147713ff764b5799910fb44edca

Observation 22c657ba-6fc6-4910-9e5a-f65fe28c6634 · outbound

This paper cites Eagle 2.5: Boosting long-context post-training for frontier vision-language models.arXiv preprint arXiv:2504.15271, 2025.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Eagle 2.5: Boosting long-context post-training for frontier vision-language models.arXiv preprint arXiv:2504.15271, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.471940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.471940Z digest=sha256:5814b2092bfe5b40ded007eb02c3c06a6aa784fc326d1a862d04179dbc9d2c49

Observation e2307a22-5132-418a-bd8d-d6bf2f3373d7 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.475164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.475164Z digest=sha256:26722c56c153b0148bb869071530260544df6b17e23d1d6cad8fc38e24a890ec

Observation 77bb1d4b-203d-486f-8adf-71fbb007e644 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video understanding.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Video-llama: An instruction-tuned audio-visual language model for video understanding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.565887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:27:05.478496Z digest=sha256:301f03eb608a2b801bc73d6fafe1fb72e47d913e90e68860c03ae60b4f79b7bf

Observation aec4a5df-6f37-466a-955a-a43cfada4361 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.482122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.482122Z digest=sha256:84f43c365edf5adfdef0d707d77f57ef03a3918afabeaf957df2fd375df4c89a

Observation e816cc37-4f83-46ff-b32a-26b4eaafe715 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.486127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.486127Z digest=sha256:475b9db01b4171b8fe012e24ef44e20e66a9a3bcd84825e9d8c458ffbaff5fbf

Observation dbda99d1-540f-4f58-b15b-70c985135e06 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.489310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.489310Z digest=sha256:cfcb57e9af09852bb7257b625715f3ecd4805ab4c068b36212698fa9eaac0f0e

Observation fbd5d1e7-aff8-4073-980f-3d9d450d7a83 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.492480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.492480Z digest=sha256:9db19703825af9fd188802cb50b6cf11f5572f91d63343eb43aaef8437c6cf3f

Observation 0ba4e3ee-4df0-48c8-b7f6-ca213a40175a · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.495697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.495697Z digest=sha256:05014b1be337d960e70b3c35fb63e316e6d2c6ce9fa46cb3e8a86971689d8053

Observation 5c096c6d-ac30-465c-a225-b79619168202 · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.499105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.499105Z digest=sha256:858e9d3cc92b761111fbbf50ced8ff89c2404e91e073fe303a1e41486a7353b7

Observation 09444582-a7b0-46d5-9633-3341ba5b0202 · outbound

This paper cites Cogvlm: Visual expert for pretrained language models.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Cogvlm: Visual expert for pretrained language models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.502419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.502419Z digest=sha256:b4b8a571e32dc749e6f9a76658330c8632abbfc8ec99bf9681f8b6046fb4e2b3

Observation 46c9b6b1-1bdf-47dc-9e13-ea4ac9dc1f34 · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs CogVLM2: Visual Language Models for Image and Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.505880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.505880Z digest=sha256:3735f3b8444da4e6c21339d7b4ef7c55c4ebdcec8cbd2241386ed81259c2df80

Observation 5e2d9d77-a5ee-42a6-8377-2f5e94b38832 · outbound

This paper cites VideoLLM: Modeling Video Sequence with Large Language Models.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs VideoLLM: Modeling Video Sequence with Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.509262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.509262Z digest=sha256:26a8b1e2072640e75e4bd62f9761b85245e8d7faf89dcf867065488a68bf1da1

Observation 5e625726-1323-43d1-b2b6-16c47c50e13f · outbound

This paper cites OVR: A Dataset for Open Vocabulary Temporal Repetition Counting in Videos.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs OVR: A Dataset for Open Vocabulary Temporal Repetition Counting in Videos

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.512719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.512719Z digest=sha256:427d176442bb304b3509754dd7c62c572c51359c361e25f92d18ab227b852d31

Observation 300bb223-9ce7-44fe-9ffb-674a621388d9 · outbound

This paper cites Dvd-counting, 2025.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Dvd-counting, 2025

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.548934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:27:05.516090Z digest=sha256:2a2c5c96d616d20da7f4ba8b39427b612523850c8be4a1fd7fd7b017b7d49d59

Observation 3a54d292-f766-44d4-a434-7cda7ddb0235 · outbound

This paper cites Counting out time: Class agnostic video repetition counting in the wild.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Counting out time: Class agnostic video repetition counting in the wild

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.519281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.519281Z digest=sha256:6ab8c79fd329b6cf3b588f1610d6a257dd3b1afda584e72457860da42a1702bd

Observation 3fdb8348-6f9e-4a57-b55f-82e31989d5c2 · outbound

This paper cites Repetitive activity counting by sight and sound.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Repetitive activity counting by sight and sound

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.532936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:27:05.522916Z digest=sha256:4063a9a77ea1393a3060781aa7cbef5ac916d99fbbbb95d5eb59ae63157c0e6c

Observation 771c62c1-f2e7-4a53-aedc-b111cfc00a7f · outbound

This paper cites TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action Counting.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action Counting

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.526108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.526108Z digest=sha256:333f9d5cd1cb306432e965814bd46e458d383f755ad1a8db6eae7ef7816056b2

Observation 02223deb-a7f6-4379-8ff1-a37f4262b12d · outbound

This paper cites CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.529740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.529740Z digest=sha256:08eae6b1a2a73e7b3d8ab02eee585a7dd4710f89934f133ca17dcfc8b83ee1bd

Observation be396407-bbd2-4d64-bb62-6ae72aab8163 · outbound

This paper cites Ola: Pushing the Frontiers of Omni-Modal Language Model.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.533023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.533023Z digest=sha256:0b8b643d45dbd95565e9c588b8e99ca153c37d040c0d018a676fce2542f450b8

Observation 806f6ff7-5b7b-4842-9630-86556c9f42b7 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.536387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.536387Z digest=sha256:760531d08a6e47355fbf22a7334f5b822535abe4f040a1c86f4699441ba3e74c

Observation d411759c-8069-4827-b3b0-eabe637af90a · outbound

This paper cites Curriculum learning for reinforcement learning domains: A framework and survey.Journal of Machine Learning Research, 21(181):1–50, 2020.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Curriculum learning for reinforcement learning domains: A framework and survey.Journal of Machine Learning Research, 21(181):1–50, 2020

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.539317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.539317Z digest=sha256:b50c1eef311d3f083bd8ca1c99dbb35403daa81497d720a5c7b4eedbe61192ce

Observation c7b9bcc2-ec63-4c1a-b970-4323e5492402 · outbound

This paper cites AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.542467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.542467Z digest=sha256:ea538057af2c1bace56a78185a5b031357b869dee2c5dd86c542d09a07bacd1f

Observation 9e0e43b2-8ea1-4d4b-9536-cd3afc7f67f2 · outbound

This paper cites Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.545846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.545846Z digest=sha256:14750b887d1726268d23d27a39e31e758bb4520d4bc05ffc93f5e7f26388c139

Observation 686c5244-6f31-443a-a4b1-34f22aab81aa · outbound

This paper cites Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.549332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.549332Z digest=sha256:7843d3ac9361933094d4b94425fe9e1968ff1bfcbaa7f0c4af5e45c769f18b0b

Observation addbeab7-af44-472e-9131-91bcf2215c69 · outbound

This paper cites OMCAT: Omni Context Aware Transformer.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs OMCAT: Omni Context Aware Transformer

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:27:06.046811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:27:05.553172Z digest=sha256:8bc4cdecc1a00ee3ddf0000068d1003569620c8ac4e2b94f868e173bc6868ac9

Observation 9ca97b63-d6e3-4293-87f8-684f8ed388ba · outbound

This paper cites Baichuan-omni-1.5 technical report.arXiv preprint arXiv:2501.15368, 2025.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Baichuan-omni-1.5 technical report.arXiv preprint arXiv:2501.15368, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.557034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.557034Z digest=sha256:4c7b375358fd4e3d7e5b5fff0d35810479b221c96f0c1d99ec89d7a15597d18f

Observation 18af2022-4246-4367-be8f-081bb7661fe2 · outbound

This paper cites Crab: A Unified Audio-Visual Scene Understanding Model with Explicit Cooperation.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Crab: A Unified Audio-Visual Scene Understanding Model with Explicit Cooperation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.560273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.560273Z digest=sha256:436a835492f79c0981b42c7fbc792b5811d60c404c2a0eec0f62baae05b59445

Observation 3a2e00d1-6dff-4ecc-9867-4cf0fc9abd08 · outbound

This paper cites video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.563770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.563770Z digest=sha256:d2339cd9d415465299fdce6dc1e8ee2fc850b8af42b712518f3e4484de59748e

Observation 5ecc1de9-e393-4d1f-bf8e-82aeb980cc0d · outbound

This paper cites Meerkat: Audio-visual large language model for grounding in space and time.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Meerkat: Audio-visual large language model for grounding in space and time

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.567346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.567346Z digest=sha256:f106bf74a6804971698e2a716c1bf86bb9477ac6914ed6006faf57cac2d1521a

Observation e178f6d4-3cc7-4795-9423-6e16a9a8f312 · outbound

This paper cites PAVE: Patching and Adapting Video Large Language Models.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs PAVE: Patching and Adapting Video Large Language Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:27:05.978456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:27:05.570467Z digest=sha256:90a3cfda487602d1dba39ba9702bcf94ad0977e571fcf454080c5dff8f8414c0

Observation bf31b9ad-6fc1-401a-be87-937002b6cf52 · outbound

This paper cites Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.574186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.574186Z digest=sha256:61cc8c0e313069af935aa0219e05b3d620a4aa4bad01eb9036b607970816f1a1

Observation c6254dd5-267a-4a5b-a5c4-fc16f2043bbb · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.503625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:27:05.578385Z digest=sha256:99c10813e54b4ddba37a57cb2323b3efb019ed227a9a0a00d175ce19ee9c4755

Observation 84bd407e-a595-4e5a-a0a5-5fa754d236f5 · outbound

This paper cites WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.581493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.581493Z digest=sha256:892c397369017b82c22aff6b0f497aaca18c2e13019a27cd5d82c730ef952ca2

Observation 0b61ec07-d356-4e35-894e-09aab9193737 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Chain-of-thought prompting elicits reasoning in large language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.584562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.584562Z digest=sha256:54244492a04056afdbd648b30a90e657be6ce4755aa31d634416500f559e6b27

Observation 34bc1aaf-6807-4143-893e-f983d1033b79 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.587891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.587891Z digest=sha256:01ba20b24d590df11160ad150a368191460c40d62346ea61fa5cd41ef26dbeb7

Observation 7d06622a-75a5-4ce8-a4ce-e48d40e83409 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.591545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.591545Z digest=sha256:5f433b10734ae3b90a1fb5ebce45d9e24dcc61bea78da4e254e041c5c4b1aada

Observation 79e302ca-b92a-4cb5-9391-af5c5397bfa4 · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.594716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.594716Z digest=sha256:6832eecca50d3a15ac92fd3d5445892a89c04da13db98ce0c21b49b938ac691d

Observation 8debd7c1-08a0-4b71-8179-253c635383c6 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Gemini: A Family of Highly Capable Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.598090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.598090Z digest=sha256:5810c6c3265188f15b32b17d3d16b5b945bc311eac380cc229166299e1d7d1e2

Observation 6435e326-212d-41e0-b84e-7df8a253e2bb · outbound

This paper cites Introducing gpt-4.1 in the api, 2025.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Introducing gpt-4.1 in the api, 2025

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.487558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:27:05.601374Z digest=sha256:23d386d4ed8f152bd58c976fd874cf736c8003e4b6208b86c7b0117769b46331

Observation 6d514cce-0015-4127-85dc-032e853a8a3c · outbound

This paper cites GPT-4o System Card.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs GPT-4o System Card

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.604689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.604689Z digest=sha256:65893164894b0bbaa54fdd7876d6bf59d1383800ca1263188f4b136b93f96582

Observation 2d5d455d-553e-4098-aa0f-ffc0607bc6b7 · outbound

This paper cites Seed1.5-VL Technical Report.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Seed1.5-VL Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.608221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.608221Z digest=sha256:82fab1b55a8f04ec8353da26419ee5ba78454017d899c4c5dfe68a3b8dba8473

Observation 8ee657c0-142e-46e5-abb7-0f73ee36fc1f · outbound

This paper cites Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.611836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.611836Z digest=sha256:75534e5503bed24ad15070137d82cd65847cf617bf0f9154d84394263520ace2

Observation 8cf9e3ab-02f8-44a7-a346-1c46941f6bbd · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.618482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.618482Z digest=sha256:24a4502380e0583b617ebaf64f8357d8f5988900c875fbefd2aa521b3d2aa25e

Observation 1746ded8-6795-49d0-871f-472260bf428f · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.621901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.621901Z digest=sha256:a820158beed0f42c62c0df93b7f84a4ee7967db08a1198d7936d66b030a24731

Observation dc908336-520b-4e60-8078-7f6460e1aaf8 · outbound

This paper cites Qwen2.5-Omni Technical Report.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Qwen2.5-Omni Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.625899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.625899Z digest=sha256:ceb3d1db7195266d5b99f298d30be5c6ebe3d60395e4c2b4fa8b1935894333ea

Observation 27005343-43db-4f9d-864c-8a8fd57809ca · outbound

This paper cites Avqa: A dataset for audio-visual question answering on videos.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Avqa: A dataset for audio-visual question answering on videos

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.629557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.629557Z digest=sha256:1e216e95737eb691d48cc2956b35e67f4100aac461fd3153ab0cbde8c820c325

Observation 5a2c7e4d-0bfa-46c4-be84-db3f5867773a · outbound

This paper cites Learning to answer questions in dynamic audio-visual scenarios.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Learning to answer questions in dynamic audio-visual scenarios

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.633066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.633066Z digest=sha256:0cec6f461abe3459e94e9ab76bc4d3649274327b6f43e83d6fd0d3755406cbe1

Observation 872115b7-75bf-42ad-94ca-e787115101ca · outbound

This paper cites Audio-visual event localization in unconstrained videos.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Audio-visual event localization in unconstrained videos

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.463851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:27:05.636748Z digest=sha256:52c57f070e60db1ee9372c3faddd98b30ea6356a88d4284db15deeb3b0d03f43

Observation 2fcdf031-2400-4683-a35c-832a7b127ab2 · outbound

This paper cites Dense-localizing audio-visual events in untrimmed videos: A large-scale benchmark and baseline.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Dense-localizing audio-visual events in untrimmed videos: A large-scale benchmark and baseline

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.453872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:27:05.639954Z digest=sha256:1d58a632a0682b51211f525f8fdcfc8c05936bc88e824acc83bc4bde3b83cc6f

Observation 3ad08ab3-41c1-486e-8ab4-6402bbae9db1 · outbound

This paper cites Cross-Modal learning for Audio-Visual Video Parsing.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Cross-Modal learning for Audio-Visual Video Parsing

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:27:05.828372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:27:05.643589Z digest=sha256:4bedddaa3e0141f5500f3a9198f2376ee39e8d0bba02785fe166a9991f55a519

Observation fe57144d-5fa1-42e1-8c75-b318be294a27 · outbound

This paper cites Audio-visual segmentation with semantics.International Journal of Computer Vision, pages 1–21, 2024.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Audio-visual segmentation with semantics.International Journal of Computer Vision, pages 1–21, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.443673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:27:05.647139Z digest=sha256:03dc2c3f9c3b5e42e515c4da1c4898550755b1e222e50f4ff01b3ffe8dcf67ac

Observation 100a54ff-3995-45cd-b92f-440fe0f4a08f · outbound

This paper cites Enhancing geometric factors in model learning and inference for object detection and instance segmentation.IEEE transactions on cybernetics, 52(8):8574–8586, 2021.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Enhancing geometric factors in model learning and inference for object detection and instance segmentation.IEEE transactions on cybernetics, 52(8):8574–8586, 2021

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.433369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:27:05.650416Z digest=sha256:c209eaf808c7a118e69b9614e78a29213c688bfb5e644b010ff666d79be7771c

Observation cc351f6b-87d6-47af-a2fa-68c710048516 · outbound

This paper cites URL https: //www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_ Card_Claude_3.pdf.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs URL https: //www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_ Card_Claude_3.pdf

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.423501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:27:05.653489Z digest=sha256:ac637c54d8f2f6277c185259e41dfe05cd77e94d99d3cc9c96847ac3f732f0d2

Observation ab692d4d-79fa-40b9-b859-571f1388fd5e · outbound

This paper cites Onellm: One framework to align all modalities with language.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Onellm: One framework to align all modalities with language

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.657817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.657817Z digest=sha256:009614c20da187f666eaadbf95d28efea707dc88b322c458a8566397800bf0fd

Observation 71987732-166b-49b7-b8f0-00bddd142d99 · outbound

This paper cites GroundingGPT:Language Enhanced Multi-modal Grounding Model.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs GroundingGPT:Language Enhanced Multi-modal Grounding Model

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.661256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.661256Z digest=sha256:4a5dea9a5b2ebc9bdc9e2ce5cf52551a44ba11d61cf02e5ca76289b4f8477dad

Observation 6005d5e9-493a-467d-9ed5-a8d3c08c7d5d · outbound

This paper cites Valor: Vision-audio-language omni-perception pretraining model and dataset.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Valor: Vision-audio-language omni-perception pretraining model and dataset.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.407310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:27:05.664610Z digest=sha256:1231062224a8e3880031ebc77b1c6761d20a9d94dc52f7c4abfe4bb613772e8a

Observation a29cf447-a518-4291-b3eb-d092abc80390 · outbound

This paper cites X-instructblip: A framework for aligning image, 3d, audio, video to llms and its emergent cross-modal reasoning.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs X-instructblip: A framework for aligning image, 3d, audio, video to llms and its emergent cross-modal reasoning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.397918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:27:05.667783Z digest=sha256:406aa295b9235d3f8690186086b3f370c239f2e281269accf5bcb94ed3f1c240

Observation c550d55b-4f9b-4f2d-89a0-50e4a988721a · outbound

This paper cites Avicuna: Audio-visual llm with interleaver and context-boundary alignment for temporal referential dialogue.arXiv e-prints, pages arXiv–2403, 2024.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Avicuna: Audio-visual llm with interleaver and context-boundary alignment for temporal referential dialogue.arXiv e-prints, pages arXiv–2403, 2024

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.387217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:27:05.670950Z digest=sha256:01c6ec11c8087d0cd145cb7d335f5a26e11f7637dfac8a7e8c634e9ab19f84cb

Observation 484f6ce8-fdd9-4447-b52b-7e1f6f78b46d · outbound

This paper cites AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.674331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.674331Z digest=sha256:0ae0ea3c1a06be74cb346371a3bb1132ffe44940d47cb30ca9e2853192ef242a

Observation 0b5835ec-d2b5-4b49-8558-da2721948a9c · outbound

This paper cites Omnibench: Towards the future of universal omni-language models.arXiv preprint arXiv:2409.15272, 2024.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Omnibench: Towards the future of universal omni-language models.arXiv preprint arXiv:2409.15272, 2024

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.677898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.677898Z digest=sha256:52e8e63aee9226f202840f274a220fb3b2a501e4c5aeddc3426f6f83a095c807

Observation 5dc4b245-3c20-459a-9398-0aedc32eeaa4 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.680865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.680865Z digest=sha256:e3cad23e1c1c986e3964b95f0678c803c4819ae4b060cbaae1abe0d3113151f0

Observation 0b62d6ce-670d-4aac-9cc7-e14526b3520e · outbound

This paper cites How many people spoke in the scene showing the conference table?.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs How many people spoke in the scene showing the conference table?

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.377743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:27:05.684003Z digest=sha256:7516bcee18ab5faa056944b8cfd6f0be862782da4a84b4268175179f6440d973

Observation b2e7becb-f94b-4077-a055-9a04c361c860 · outbound

This paper cites an unresolved cited work.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:27:06.367572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:27:05.687707Z digest=sha256:787604d2dba0340f0156f8396c7959f40fece92058afdbe3c5be6424f920faee

Observation 466bba4f-c527-4377-93ee-848320097d79 · outbound

This paper cites an unresolved cited work.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:27:06.357940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:27:05.691318Z digest=sha256:154fea76f8c97b73202b105be947d94195166ad48333a1f40ea5008f4a97d8ee

Observation 6181835e-8342-4651-8f75-1f781fa56004 · outbound

This paper cites an unresolved cited work.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:27:06.347782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:27:05.694655Z digest=sha256:459e573042d135db1b35159ff759b7b59d2a910550708a3172d7ff310854305b

Observation 24fcc273-cade-443a-a195-8b8c300f5146 · outbound

This paper cites an unresolved cited work.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:27:06.338960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:27:05.697979Z digest=sha256:1db25cd9a7507ff6e1bb4f131663bc50e71881524d66633e6a5688a112df0acd

Observation 4cc85d94-89fc-4989-bc42-0e2dc4b9a040 · outbound

This paper cites an unresolved cited work.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:27:06.328979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:27:05.701456Z digest=sha256:62ac06019334409b9908e0b97fd4348faf90cc6d69f9040f78b0b73f9b04e0ec

Observation 30856085-49c4-4a07-912b-1ea40b179069 · outbound

This paper cites an unresolved cited work.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:27:06.319383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:27:05.705052Z digest=sha256:b484528b7a1693fee777898a77a3e10f8420ac7c173a85ddaff56f988de8f84b

Observation 476f044f-235e-4f69-b4dd-82d77e9b3a79 · outbound

This paper cites question.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs question

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.309524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:27:05.708277Z digest=sha256:9b5242e9ac5bed8b59dada79bf768cfb21ac2d27694ad53e75d3902da83329fc

Pith citing papers

Observation 9c2be27d-d82f-4e75-b00e-629425ae9f37 · inbound

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models cites this paper.

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:45:56.105299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:45:07.700571Z digest=sha256:82d64b9bd19b08af9bee49fc3dce7eee2ac988aeeec94491189a23516f8a72a3

Observation e3189c56-f1bb-4f92-8354-b7a926ffdf6e · inbound

SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance cites this paper.

SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T22:09:56.270516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:09:56.270516Z digest=sha256:4d4fec187a77c96e1092debfcda52dc0403b610bf824b5f59e6239b9e6803e38

Observation e4baaf95-ed9e-476e-b0ff-3904fcbc7312 · inbound

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs cites this paper.

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:22.142010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T12:05:54.551728Z digest=sha256:08461b9da9a7ddfe9407c4460d0b8409e3440161c599d42200a34facb93dd806

Observation da73e082-7283-4c33-aaae-e153c39db0af · inbound

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers cites this paper.

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:18:32.125415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T08:02:53.574120Z digest=sha256:8827d58126e6165b71be61e0f8eb71709429d843d74a6f1b160a5ea2b5ce361f

Observation 11430591-9fb8-4725-9863-066575a13626 · inbound

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation cites this paper.

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:52:12.698673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T03:49:58.240883Z digest=sha256:c37309d40f12d152f9efcad7006e6d79f351ec55147d9e4479d1fe096a16a445

Observation 2449202f-3603-4a76-9f9c-6f41ce858e3d · inbound

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation cites this paper.

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:09:50.259437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T06:06:20.658030Z digest=sha256:bb767704733f60274809da2a333be013a3b608f0505681f6635553627a365e07

Observation 15ec56f8-173e-4780-b416-249258d469bf · inbound

Empowering Long-form Omni-modal Understanding with Robust Audio Perception cites this paper.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:d2a8ddfa97a76bf2d5064d2ab2c8f6b3652f0ea94fdfda9bc2ae3cfbcfb08287