Pith. sign in

Paper Citation Record · LEDGER

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning

As of 19 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 1 inbound Pith citation observation for arXiv:2508.07470.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.07470 v2

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T22:09:16.309676Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T21:04:02.263300Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T21:05:03.827879Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bf48b8a6-817a-4536-915f-b4ca919a677b · outbound

This paper cites EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T22:09:16.793393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T22:09:15.243340Z digest=sha256:456fd0b5db579bdecd56f14a7bbd56f2cb6311aad4fac751d036f416eb53e642

Observation 3de3c3d9-f400-40bc-8f61-ad4adf1bb35d · outbound

This paper cites Girdhar, R.; El-Nouby, A.; Liu, Z.; Singh, M.; Alwala, K.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning Girdhar, R.; El-Nouby, A.; Liu, Z.; Singh, M.; Alwala, K

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:15.343828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:09:15.343828Z digest=sha256:ab7bee2be94b3e4e05335de5a6763e4f5d69d12f627101da8992d0b4a91db72a

Observation 5dc66f6d-baec-4bf6-9559-c827d76c459a · outbound

This paper cites Multimodal Pretraining for Dense Video Captioning.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning Multimodal Pretraining for Dense Video Captioning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:15.603717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:09:15.603717Z digest=sha256:e9a1201e0cadd1ea695eb6c4b116a028cfdd3f514de940c2d9f42d79b3e74b3a

Observation 2b12457c-5b01-447d-ac71-1375b9ea22f4 · outbound

This paper cites Jina CLIP: Your CLIP Model Is Also Your Text Retriever.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:15.681915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:09:15.681915Z digest=sha256:932159c088a31d4e2ea60dd7574ae6b063c45185d45ec9fbd0bcbeabe32a659f

Observation 32fe63cb-5832-4136-91ac-6194506ec5c5 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:15.767519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:09:15.767519Z digest=sha256:8c2d4a293a047852840099246c12706787ff447bc28194e20d3ce0ebad9f2723

Observation 6ce1a547-b414-41aa-948d-4eef6ec84102 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:15.844231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:09:15.844231Z digest=sha256:47e8bcf93b771b11de83898921d68d3deaa09b2e3a135ffea3d9111a8babd828

Observation 04d8c846-5ee5-46db-86a6-f8a3f6ead5d3 · outbound

This paper cites Qwen2.5 Technical Report.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning Qwen2.5 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:16.061859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:09:16.061859Z digest=sha256:bdc6c424e9962f7670c3aad023bb419fdd457985ed885ac457c6e96f8b1c3ddb

Observation b8227ad1-5869-42b2-abee-d7d250b9f5d2 · outbound

This paper cites UniMD: Towards Unifying Moment Retrieval and Temporal Action Detection.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning UniMD: Towards Unifying Moment Retrieval and Temporal Action Detection

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T22:09:16.487647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T22:09:16.128976Z digest=sha256:1ba7e2cd96f4cdfb891cf95cd51f5065fac51dbe17244ed3f6cc8909f8e2fc57

Observation 0b428c00-5abe-47ad-9039-1b4c412a6a56 · outbound

This paper cites Video Question Answering: Datasets, Algorithms and Challenges.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning Video Question Answering: Datasets, Algorithms and Challenges

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:16.203189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:09:16.203189Z digest=sha256:f5c8e7b3f1090339492b8bd75dace4dc97837a897ccc010d8989fea052b3add6

Observation 3deea570-b5f0-447b-afb2-615e92829ac0 · outbound

This paper cites In Inter- national Conference on Learning Representations (ICLR).

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning In Inter- national Conference on Learning Representations (ICLR)

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:09:17.047351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T22:09:16.309676Z digest=sha256:d81438c06ae38db0e15c8cd5ede603afa6a3446d911e33e821d68fc34d99fc0b

Observation 9838723f-089b-415e-9c58-314d24b06794 · outbound

This paper cites Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:15.943570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:09:15.943570Z digest=sha256:9e25a54def9fd1025bcc8461804ebc3153cbbbbc08d7ab4b6f92ad3e7a91b1a9

Observation 92a9f263-beb7-438d-bfcf-52840dfc59a3 · outbound

This paper cites ImageBind-LLM: Multi-modality Instruction Tuning.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning ImageBind-LLM: Multi-modality Instruction Tuning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:15.507027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:09:15.507027Z digest=sha256:411ef15aeff50dacec3d84381098e7c50a10154c3c32014488430c134df86f16

Observation ad6c6200-5aad-49e1-8cbb-e7cb7e6f0e10 · outbound

This paper cites In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 11287–11297.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 11287–11297

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:09:17.257252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T22:09:15.424425Z digest=sha256:9753b18fe7eca47711caa71aa27f78a35d64f61e4c4c4456de029721aeb01c1b

Observation 15aa2a3b-5e70-45df-922d-dd1d97003e78 · outbound

This paper cites AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:14.941204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:09:14.941204Z digest=sha256:a3f043c198440c46966d80914303888d41aac7242819392f38e1a70b228f1a78

Observation ea2b809f-564d-4508-92df-bd75ce360936 · outbound

This paper cites See https://vicuna.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning See https://vicuna

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:09:17.434401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T22:09:15.153338Z digest=sha256:6486e076745db2956fc2e6119c92d909b28366e113aac6579db5818bfcbc81cb

Observation ce3dab9b-d7de-4e10-903b-f2b647d01ca7 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:15.067940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:09:15.067940Z digest=sha256:5ff306dcfe956e7e793bdf832816b3941fb36d78d69b5c1246d099ae5621bbe5

Observation de74c809-b302-4a26-9fd4-f05ed2a14e2f · outbound

This paper cites Non-invertible Symmetries in 2D from Type IIB String Theory.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning Non-invertible Symmetries in 2D from Type IIB String Theory

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:14.844607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:09:14.844607Z digest=sha256:715a0073f950d33aef49870b7098eb3ac6ef6a892d688cd62eaf2559970ee3c6

Pith citing papers

Observation 7d8ab9f7-1791-41e7-8467-231f65af8c49 · inbound

Uncovering the Representation Geometry of Minimal Cores in Overcomplete Reasoning Traces cites this paper.

Uncovering the Representation Geometry of Minimal Cores in Overcomplete Reasoning Traces AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning

Reference 79

Resolution
malformed identifier
arxiv_id, observed 2026-06-30T21:05:03.829322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T21:04:02.263300Z digest=sha256:bc0cdfb0fccbc67a80038c2e565c48f74ac19fb0136853f235ad65cb489bd443