Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:27:05.708277Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 7 inbound Pith citation observations for arXiv:2506.05328.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:27:05.708277Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-14T22:09:56.270516Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-18T05:45:56.102421Z
71 of 71 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a684ffce-ddae-4dfa-b02a-13da027d9496 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Qwen2.5-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6be6bd9a-0430-44a7-9f86-3859b73daa87 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22c657ba-6fc6-4910-9e5a-f65fe28c6634 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Eagle 2.5: Boosting long-context post-training for frontier vision-language models.arXiv preprint arXiv:2504.15271, 2025
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2307a22-5132-418a-bd8d-d6bf2f3373d7 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77bb1d4b-203d-486f-8adf-71fbb007e644 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Video-llama: An instruction-tuned audio-visual language model for video understanding
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aec4a5df-6f37-466a-955a-a43cfada4361 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e816cc37-4f83-46ff-b32a-26b4eaafe715 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbda99d1-540f-4f58-b15b-70c985135e06 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbd5d1e7-aff8-4073-980f-3d9d450d7a83 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ba4e3ee-4df0-48c8-b7f6-ca213a40175a · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c096c6d-ac30-465c-a225-b79619168202 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09444582-a7b0-46d5-9633-3341ba5b0202 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Cogvlm: Visual expert for pretrained language models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46c9b6b1-1bdf-47dc-9e13-ea4ac9dc1f34 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs CogVLM2: Visual Language Models for Image and Video Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e2d9d77-a5ee-42a6-8377-2f5e94b38832 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs VideoLLM: Modeling Video Sequence with Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e625726-1323-43d1-b2b6-16c47c50e13f · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs OVR: A Dataset for Open Vocabulary Temporal Repetition Counting in Videos
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 300bb223-9ce7-44fe-9ffb-674a621388d9 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Dvd-counting, 2025
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3a54d292-f766-44d4-a434-7cda7ddb0235 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Counting out time: Class agnostic video repetition counting in the wild
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fdb8348-6f9e-4a57-b55f-82e31989d5c2 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Repetitive activity counting by sight and sound
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 771c62c1-f2e7-4a53-aedc-b111cfc00a7f · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action Counting
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02223deb-a7f6-4379-8ff1-a37f4262b12d · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be396407-bbd2-4d64-bb62-6ae72aab8163 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Ola: Pushing the Frontiers of Omni-Modal Language Model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 806f6ff7-5b7b-4842-9630-86556c9f42b7 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d411759c-8069-4827-b3b0-eabe637af90a · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Curriculum learning for reinforcement learning domains: A framework and survey.Journal of Machine Learning Research, 21(181):1–50, 2020
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7b9bcc2-ec63-4c1a-b970-4323e5492402 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e0e43b2-8ea1-4d4b-9536-cd3afc7f67f2 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 686c5244-6f31-443a-a4b1-34f22aab81aa · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation addbeab7-af44-472e-9131-91bcf2215c69 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs OMCAT: Omni Context Aware Transformer
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9ca97b63-d6e3-4293-87f8-684f8ed388ba · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Baichuan-omni-1.5 technical report.arXiv preprint arXiv:2501.15368, 2025
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18af2022-4246-4367-be8f-081bb7661fe2 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Crab: A Unified Audio-Visual Scene Understanding Model with Explicit Cooperation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a2e00d1-6dff-4ecc-9867-4cf0fc9abd08 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ecc1de9-e393-4d1f-bf8e-82aeb980cc0d · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Meerkat: Audio-visual large language model for grounding in space and time
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e178f6d4-3cc7-4795-9423-6e16a9a8f312 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs PAVE: Patching and Adapting Video Large Language Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bf31b9ad-6fc1-401a-be87-937002b6cf52 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6254dd5-267a-4a5b-a5c4-fc16f2043bbb · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 84bd407e-a595-4e5a-a0a5-5fa754d236f5 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b61ec07-d356-4e35-894e-09aab9193737 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Chain-of-thought prompting elicits reasoning in large language models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34bc1aaf-6807-4143-893e-f983d1033b79 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Visual-RFT: Visual Reinforcement Fine-Tuning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d06622a-75a5-4ce8-a4ce-e48d40e83409 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Video-R1: Reinforcing Video Reasoning in MLLMs
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79e302ca-b92a-4cb5-9391-af5c5397bfa4 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8debd7c1-08a0-4b71-8179-253c635383c6 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Gemini: A Family of Highly Capable Multimodal Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6435e326-212d-41e0-b84e-7df8a253e2bb · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Introducing gpt-4.1 in the api, 2025
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6d514cce-0015-4127-85dc-032e853a8a3c · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs GPT-4o System Card
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d5d455d-553e-4098-aa0f-ffc0607bc6b7 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Seed1.5-VL Technical Report
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ee657c0-142e-46e5-abb7-0f73ee36fc1f · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cf9e3ab-02f8-44a7-a346-1c46941f6bbd · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1746ded8-6795-49d0-871f-472260bf428f · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc908336-520b-4e60-8078-7f6460e1aaf8 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Qwen2.5-Omni Technical Report
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27005343-43db-4f9d-864c-8a8fd57809ca · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Avqa: A dataset for audio-visual question answering on videos
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a2c7e4d-0bfa-46c4-be84-db3f5867773a · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Learning to answer questions in dynamic audio-visual scenarios
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 872115b7-75bf-42ad-94ca-e787115101ca · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Audio-visual event localization in unconstrained videos
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2fcdf031-2400-4683-a35c-832a7b127ab2 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Dense-localizing audio-visual events in untrimmed videos: A large-scale benchmark and baseline
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3ad08ab3-41c1-486e-8ab4-6402bbae9db1 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Cross-Modal learning for Audio-Visual Video Parsing
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fe57144d-5fa1-42e1-8c75-b318be294a27 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Audio-visual segmentation with semantics.International Journal of Computer Vision, pages 1–21, 2024
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 100a54ff-3995-45cd-b92f-440fe0f4a08f · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Enhancing geometric factors in model learning and inference for object detection and instance segmentation.IEEE transactions on cybernetics, 52(8):8574–8586, 2021
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cc351f6b-87d6-47af-a2fa-68c710048516 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs URL https: //www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_ Card_Claude_3.pdf
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ab692d4d-79fa-40b9-b859-571f1388fd5e · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Onellm: One framework to align all modalities with language
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71987732-166b-49b7-b8f0-00bddd142d99 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs GroundingGPT:Language Enhanced Multi-modal Grounding Model
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6005d5e9-493a-467d-9ed5-a8d3c08c7d5d · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Valor: Vision-audio-language omni-perception pretraining model and dataset.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a29cf447-a518-4291-b3eb-d092abc80390 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs X-instructblip: A framework for aligning image, 3d, audio, video to llms and its emergent cross-modal reasoning
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c550d55b-4f9b-4f2d-89a0-50e4a988721a · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Avicuna: Audio-visual llm with interleaver and context-boundary alignment for temporal referential dialogue.arXiv e-prints, pages arXiv–2403, 2024
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 484f6ce8-fdd9-4447-b52b-7e1f6f78b46d · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b5835ec-d2b5-4b49-8558-da2721948a9c · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Omnibench: Towards the future of universal omni-language models.arXiv preprint arXiv:2409.15272, 2024
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dc4b245-3c20-459a-9398-0aedc32eeaa4 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b62d6ce-670d-4aac-9cc7-e14526b3520e · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs How many people spoke in the scene showing the conference table?
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b2e7becb-f94b-4077-a055-9a04c361c860 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Unresolved cited work
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 466bba4f-c527-4377-93ee-848320097d79 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Unresolved cited work
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6181835e-8342-4651-8f75-1f781fa56004 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Unresolved cited work
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 24fcc273-cade-443a-a195-8b8c300f5146 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Unresolved cited work
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4cc85d94-89fc-4989-bc42-0e2dc4b9a040 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Unresolved cited work
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 30856085-49c4-4a07-912b-1ea40b179069 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Unresolved cited work
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 476f044f-235e-4f69-b4dd-82d77e9b3a79 · outbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs question
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9c2be27d-d82f-4e75-b00e-629425ae9f37 · inbound
XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e3189c56-f1bb-4f92-8354-b7a926ffdf6e · inbound
SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4baaf95-ed9e-476e-b0ff-3904fcbc7312 · inbound
Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation da73e082-7283-4c33-aaae-e153c39db0af · inbound
AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 11430591-9fb8-4725-9863-066575a13626 · inbound
Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2449202f-3603-4a76-9f9c-6f41ce858e3d · inbound
Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 15ec56f8-173e-4780-b416-249258d469bf · inbound
Empowering Long-form Omni-modal Understanding with Robust Audio Perception AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.