Pith. sign in

Paper Citation Record · LEDGER

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing

As of 19 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 3 inbound Pith citation observations for arXiv:2506.05414.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05414 v1

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:50:53.648115Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T23:39:22.070629Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T23:45:08.222067Z

Reference resolution

72 of 72 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 51c70d91-d0d3-4485-9597-c4635738810b · outbound

This paper cites Shelton and Timothy P.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Shelton and Timothy P

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:55.218328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:50:52.562058Z digest=sha256:31f555dcaf70ca7a075bd4454938463bcc253c43dfb8fbdb3389c1797710dd04

Observation 54604aca-b920-4463-a66d-cc2af4ca5b7f · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.568788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.568788Z digest=sha256:7ede0457220721b450422fd23cbb853e52183467cf3ea766355287fd43778806

Observation 8d839f7c-d13a-4593-8741-4824627e8516 · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.575432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.575432Z digest=sha256:d56afe66b88d0fd13b82657669bed1d001dc10b2d92cab926871323425284381

Observation d9994da5-2214-47aa-967d-1499f0e1d1ef · outbound

This paper cites Tenenbaum, Celso Miguel de Melo, Madhava Krishna, Liam Paull, Florian Shkurti, and Antonio Torralba.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Tenenbaum, Celso Miguel de Melo, Madhava Krishna, Liam Paull, Florian Shkurti, and Antonio Torralba

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.581847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.581847Z digest=sha256:2da58e9aff01324835ae06c90bbea9136951a399b5c7aed520890d235e5dea3e

Observation 7ab239f0-022e-42d6-bc5d-4a37fbd8672b · outbound

This paper cites LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.588064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.588064Z digest=sha256:9309b1d7420c0b31e1261486ea4f73714013d30d6c0ec4ccd533178c1e45e5c9

Observation 30493e97-6e53-4e79-89a1-589441b5d889 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.Advances in Neural Information Processing Systems, 36:46212–46244, 2023.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Egoschema: A diagnostic benchmark for very long-form video language understanding.Advances in Neural Information Processing Systems, 36:46212–46244, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.596168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.596168Z digest=sha256:b2719720f157160ecace5bee50c4cd187c5f2a27040ee941ea603e88bfa0af8c

Observation c43455d5-fd59-416f-a947-1c9b60b35c83 · outbound

This paper cites Openeqa: Embodied question answering in the era of foundation models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Openeqa: Embodied question answering in the era of foundation models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.603103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.603103Z digest=sha256:18f8a21b74249511627b7952f783c7b103e58590766a05b6aebb6f81d607cd20

Observation d85400df-0ead-4fef-bad4-b78914fb1e5b · outbound

This paper cites Learning to answer questions in dynamic audio-visual scenarios.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Learning to answer questions in dynamic audio-visual scenarios

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.610231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.610231Z digest=sha256:8687724b8a91c24e826630ff2c0b98cf4e73857a7e619d8347a77ca2af21136a

Observation a4d67649-81ae-4855-8570-07441948562b · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Ego4d: Around the world in 3,000 hours of egocentric video

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.617297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.617297Z digest=sha256:3f53f3d835cce60eada6bca1c2bc7260ad1f6336820fe877f18a4e754d5e4bc1

Observation e126f361-8d8a-4cd4-9614-0a94ea7d0776 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.624322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.624322Z digest=sha256:4016736340307c78d33bd6145c40d2d91a8cd6dbf3412e38de278d7b88a8123e

Observation 5815ef23-37f8-41bd-a5f7-b5a5bbffcd90 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.630717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.630717Z digest=sha256:4b0642e036b515cbcca3b654fef8fd6b4af8836dda614abca36fc75efa0e1408

Observation 1b379fc0-4ef3-45e3-93b6-51201ac64726 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.638441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.638441Z digest=sha256:b18b919fb3f441a14942969bb06d708d635301a303796e86e2a9a6afbe43bcef

Observation d7856dab-999d-4b83-a02d-48a4ba13866f · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing VideoChat: Chat-Centric Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.645520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.645520Z digest=sha256:07433f6f6eba275ce4dddb6226179ec92bc6e33acf63fc176c372160dc0e732c

Observation 961cccdb-6df9-40ec-8327-66d0ad9ea2b4 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.652513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.652513Z digest=sha256:8cee4aa7d6114c30163c31b81a780bbca3b0b32ed72d2eaa5904b579ad7b5623

Observation 64518349-8e1d-4ed3-8a5c-6eb46e3d765a · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.658655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.658655Z digest=sha256:23a85e507ad9f10eafa3770258bc44b808c84a48a7eeeac7f7fdcd1752c0c44f

Observation 3d34628c-381a-41de-ae2a-00fadf19fe1e · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing LLaVA-OneVision: Easy Visual Task Transfer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.664757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.664757Z digest=sha256:5ba2158ee371aa2d76e707c644cc3f2c30ec123b42b4ecce1779e3fd25515d12

Observation b19e6aaa-7840-46aa-9110-6e641672ca3c · outbound

This paper cites Listen, Think, and Understand.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Listen, Think, and Understand

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.670777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.670777Z digest=sha256:447cb216ada24be60927b3e20ea36b40cef73af53ddf3b7754e019e12db0ffbd

Observation fbef4dad-59d5-4be8-ba15-309b10b28eb1 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.677692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.677692Z digest=sha256:d913ed4da7d4c6ce6a1e1fa4094328f3fc5ea8a785e22333d65d6abfdee5b29a

Observation 2f83e2ec-4967-4a44-8c46-197267c26a52 · outbound

This paper cites Kimi-Audio Technical Report.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Kimi-Audio Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.684460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.684460Z digest=sha256:51faf8f41b59a072362bcff821dc102675278995d198a4aa2d5302291ad36003

Observation 95790d61-6845-4e0e-9e17-22c7694b5a54 · outbound

This paper cites Audio-reasoner: Improving reasoning capability in large audio language models, 2025.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Audio-reasoner: Improving reasoning capability in large audio language models, 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:55.135905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:50:52.691260Z digest=sha256:dd8a5f1924043d480bf55a6efb3447eae218bf95d7cdd818499f73556d6fd1c4

Observation f773e9ef-4b51-40a0-b55a-e8f37e6c6c5d · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.698238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.698238Z digest=sha256:7c300deb32625dec478d8408aa2220b8697dedcce8526d99e18495a601e426f2

Observation f5dcc519-42f4-4923-9c3a-02a13a5e89ac · outbound

This paper cites LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.705026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.705026Z digest=sha256:dac46a7e6718dd53c8f619a90f0477614ef6430b6a89115fbca81f30119fb34e

Observation 1286094a-ef5e-48ac-9c06-67421896a221 · outbound

This paper cites Egolife: Towards egocentric life assistant, 2025.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Egolife: Towards egocentric life assistant, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.712377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.712377Z digest=sha256:c18bf91df92dfb3303b34bd5498d985f149ce5dc5f05f02dcd902a46d868638c

Observation 757897fa-3c40-4072-bca9-bb4530bc2848 · outbound

This paper cites Ola: Pushing the Frontiers of Omni-Modal Language Model.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.719441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.719441Z digest=sha256:fe9108352810d50ffe995ecf0f6804c338cd9e0fb4de571357c41e14b47ef252

Observation a410c64e-8898-4152-9fd8-a7bd08e38da9 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.728028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.728028Z digest=sha256:92b03083270a0c4c265c17301cbe1eafb104dd9f555eece270ceab60e020b005

Observation e8db80ae-0bf3-485d-b398-3fc0753f61e4 · outbound

This paper cites video-SALMONN: Speech-enhanced audio-visual large language models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing video-SALMONN: Speech-enhanced audio-visual large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:55.110689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:50:52.734466Z digest=sha256:b196a2ea015e759efda4a95d5d42c9c1ed57b059fac7620f4c02c82becd929bc

Observation d2d7f427-c064-4870-b456-31241f56a0db · outbound

This paper cites Onellm: One framework to align all modalities with language.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Onellm: One framework to align all modalities with language

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:55.094678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:50:52.741622Z digest=sha256:42850f83fa8d7165b18d801a28dc2ea3ba27fc651eab1b62ace5b579de55b160

Observation b7dc4d74-c4f0-4614-b6be-86224e0a7b4e · outbound

This paper cites X-instructblip: A framework for aligning x-modal instruction-aware representations to llms and emergent cross-modal reasoning, 2023.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing X-instructblip: A framework for aligning x-modal instruction-aware representations to llms and emergent cross-modal reasoning, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:55.076047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:50:52.749258Z digest=sha256:f24b50ea2d6c84139fc8f454d94afd7064cd4296ef3df2fd9883eed398cdac90

Observation 20a70e97-adf7-4fe3-8876-f3fb30f6b31e · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.756287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.756287Z digest=sha256:3c047df2237d1ec4eb503b93d74085495702bf84edc64b531187d62a3ffc7e70

Observation c633c6df-ccb1-4ba4-9085-88e1fd92b35f · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.761992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.761992Z digest=sha256:f767f6f085748e1660989412f7375fb561140e252558f8889b866ed3deeff879

Observation 061cff8b-33b5-4e83-a9cd-f71616a5ae0e · outbound

This paper cites Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.768322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.768322Z digest=sha256:ec7076069d316d7dd2e41b2fdedeab4e940c211e5848c51c7b8424057fa02e68

Observation c14d7703-8bd6-4d2b-88ae-881a8999a0dd · outbound

This paper cites Robohop: Segment-based topological map representation for open-world visual navigation.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Robohop: Segment-based topological map representation for open-world visual navigation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:55.054845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:50:52.775435Z digest=sha256:451dc09c7dd03df4a2d296f85388c3c4cbcb09feabe76598bd6fa7285c437eb1

Observation 26c96cbc-7529-41bd-b782-88ab2c7aae1d · outbound

This paper cites Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.782291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.782291Z digest=sha256:e18fdc164a744072dc547752abd9fb1ab617fdd9ce68f8afeed7cef15f59d4f4

Observation 5615034d-7b85-4e48-83af-89adf22e580d · outbound

This paper cites 3d-mem: 3d scene memory for embodied exploration and reasoning, 2024.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing 3d-mem: 3d scene memory for embodied exploration and reasoning, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:55.021240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:50:52.789671Z digest=sha256:a5a2d468516fd63555650410eb4f8993ac3bd82e939b7a08fcb7a830094a1536

Observation cfa85fcf-2cdf-4e1d-8934-7ba0409f6ef4 · outbound

This paper cites 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.796197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.796197Z digest=sha256:eb4bb44ccaae4d98cffc841a0c2c1dff5283b1d9b1a15b17f64232698cfed7aa

Observation 3eabc801-a4d0-4b7b-ae40-b2bc915861bf · outbound

This paper cites Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.802988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.802988Z digest=sha256:e20ae7a5fc593be0eff4da3782e13a547ef4e01fcd180ed9b8a33883d38ce1fa

Observation 279161f6-8053-4e84-a3bf-af5e8044a284 · outbound

This paper cites Gridmm: Grid memory map for vision-and-language navigation.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Gridmm: Grid memory map for vision-and-language navigation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.809764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.809764Z digest=sha256:cc2172c3bdbd797412c240e46de56695240126a4a5fa75fed813f066872d6633

Observation 1f6e8217-2a87-4753-8dee-63be6fbd7fbc · outbound

This paper cites Visual language maps for robot navigation.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Visual language maps for robot navigation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.985179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:50:52.816719Z digest=sha256:32ce282f4ed3857fd9fc411c752ecc3795367c5e13d04c5870e83afff502cda4

Observation 88f90bd5-fb71-45df-912b-dd854a51244c · outbound

This paper cites ChatSplat: 3D Conversational Gaussian Splatting.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing ChatSplat: 3D Conversational Gaussian Splatting

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.823519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.823519Z digest=sha256:04d18e077630b97ef17a9bf72ffdf789aee5af89d4ec09adf253401468966837

Observation b1d1f0eb-5cd9-4f36-aa76-5afbb6c206db · outbound

This paper cites Language embedded 3d gaus- sians for open-vocabulary scene understanding.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Language embedded 3d gaus- sians for open-vocabulary scene understanding

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.965124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:50:52.831954Z digest=sha256:04423a2b573c46875ecbfc88484ae205ffaf3c9e272c6e2c4f6d8ae64287b111

Observation c43b3d02-3cd2-42ba-af71-3d520818c96b · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.840774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.840774Z digest=sha256:27afec708e1b7dc53ffbbb4dcf36443d74bfb8da2bc57a7c740eb5ffcdea9c97

Observation fff3a9a3-9b5c-44a7-98dd-53a5720c8082 · outbound

This paper cites Sound event localization and detection of overlapping sources using convolutional recurrent neural networks.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Sound event localization and detection of overlapping sources using convolutional recurrent neural networks

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.945823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:50:52.848124Z digest=sha256:a3fb6444dbec6c20f2f091a5df300433555f75d5ad0cfc1ab0672567f124cdde

Observation 604093e5-88f7-44b7-8c2a-be60f54dca3f · outbound

This paper cites Robust sound source tracking using srp-phat and 3d convolutional neural networks.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:300–311, 2020.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Robust sound source tracking using srp-phat and 3d convolutional neural networks.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:300–311, 2020

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.927560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:50:52.854282Z digest=sha256:c6c3a5bb072986666f3901d9521ecbe7c37fe5e39f1132909aac20dab2975109

Observation efc636f5-7775-4511-97f9-5a2ddf7bfe92 · outbound

This paper cites BAT: Learning to Reason about Spatial Sounds with Large Language Models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing BAT: Learning to Reason about Spatial Sounds with Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.861092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.861092Z digest=sha256:11649a523c68f18c0605ae8d9190ee4785477a7ebde6a7b46c076318b4d6386f

Observation bb8ad452-b6a5-4306-b951-3c05b1870867 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.869518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.869518Z digest=sha256:f8431d7320abdc7e0a1398fc51c3701644d1961e913f4d0fba44d5f171032bb9

Observation bc12e5ea-1e8b-4171-9691-9bf612a8fdfb · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.909983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:50:52.877319Z digest=sha256:04c9f19d359f5f1b314029dd7914cac23162b810663e4837e9b68fdec2476688

Observation 635378ff-4dab-49c3-9e8d-0bf0d600d337 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.882782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.882782Z digest=sha256:a1b9c1867b2f68c62b18631448d4dad56241860988ea2d233e302d5948a9fce7

Observation 7015e766-957a-4caa-ae1a-f5ee80c334fc · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.887571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.887571Z digest=sha256:02b14d329e83ad8e03a975322385add2aec5d456f1114f249ef0d04c0a9ffe0f

Observation b6c4f11c-2b5b-40bf-8e1d-df83df2d7f6e · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.895136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.895136Z digest=sha256:345c80c3db92c82064df46fc16c586770e920260165d6927eb53f60e8f69c01b

Observation 1b58dd93-e748-4d80-b495-5a2a619c530e · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.925613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.925613Z digest=sha256:e122f345fe1a6125fb3816c2b22cd4fb9cd9375e3605a8332da6d7510ee141b8

Observation 310fc69f-3eaf-45d3-aa84-59b71bcd55fb · outbound

This paper cites Egotaskqa: Understanding human tasks in egocentric videos.Advances in Neural Information Processing Systems, 35:3343–3360, 2022.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Egotaskqa: Understanding human tasks in egocentric videos.Advances in Neural Information Processing Systems, 35:3343–3360, 2022

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.725195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:50:53.049207Z digest=sha256:eb19f84b9a2bdc8ff0e245cea63040a245e6ebe6e049af64cf4d9f1396e0e635

Observation 19258998-a866-4f9e-8f3f-d97f04b0a596 · outbound

This paper cites Vitatecs: A diagnostic dataset for temporal concept understanding of video-language models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Vitatecs: A diagnostic dataset for temporal concept understanding of video-language models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.186552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.186552Z digest=sha256:342a3294b763e4f2b542f26d10135828531bb97b97b3b23469fe24b13a6ea229

Observation 9c23d341-727b-45fe-a305-5e5ebaa9831e · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing TempCompass: Do Video LLMs Really Understand Videos?

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.344906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.344906Z digest=sha256:07ed8860a2a7ffafcd0af64cb7118e418b9be6abc84b5dcfcb1fcd279a228f74

Observation de2fd68e-6b4e-4389-82a8-f67a378ad5e7 · outbound

This paper cites Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.504475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.504475Z digest=sha256:f8f44191880f299069ed8435fe4d4c2b146bc02027d55a1a4ddf4ca93c33be20

Observation 5bb61c32-e1aa-46ba-831d-2129db380b73 · outbound

This paper cites Tackling data bias in music-avqa: Crafting a balanced dataset for unbiased question-answering.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Tackling data bias in music-avqa: Crafting a balanced dataset for unbiased question-answering

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.669457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:50:53.539032Z digest=sha256:c2ef82a7b7acb80d2f318a29686f756167eb909ed7bd22add669993869026c07

Observation def4e580-b561-4b44-a3c7-3a91abb81f92 · outbound

This paper cites Avqa: A dataset for audio-visual question answering on videos.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Avqa: A dataset for audio-visual question answering on videos

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.543428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.543428Z digest=sha256:1da0c4c883988df6d1110bcafc563dbc55c78f380e27961fa4e8ce63340488c6

Observation 49b099fd-7b1b-4ebf-95bc-c124ef5a7b27 · outbound

This paper cites Pano-avqa: Grounded audio-visual question answering on 360deg videos.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Pano-avqa: Grounded audio-visual question answering on 360deg videos

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.416844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:50:53.548005Z digest=sha256:4d8fa9e9f12458fab9a4bc3cacb12125a8c0a7fbeb176d6d7385a6be766650c4

Observation 729992a3-cf9f-4247-93aa-6672fe9925b3 · outbound

This paper cites Egocentric audio-visual object localization.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Egocentric audio-visual object localization

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.377983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:50:53.553335Z digest=sha256:58270231e41b1d2680428a743e50128190b2d145ea7fb75e5fdb08b8e829f09f

Observation aa926751-849e-4ffe-be82-aada78b03281 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Scanqa: 3d question answering for spatial scene understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.558636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.558636Z digest=sha256:939b0895fa0bf9a95fe67a02481b8e34a50720a41710923ce1baed5bac7fc77d

Observation cdab30bd-6328-42ec-a59f-65cd428c537e · outbound

This paper cites Aria Everyday Activities Dataset.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Aria Everyday Activities Dataset

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.578305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.578305Z digest=sha256:fcfde4b593848336cb1a304ccddcb60507214e6191a3e65c66d5b338c87fae1f

Observation 535037db-2991-4d05-9af7-42db1e362124 · outbound

This paper cites EFM3D: A Benchmark for Measuring Progress Towards 3D Egocentric Foundation Models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing EFM3D: A Benchmark for Measuring Progress Towards 3D Egocentric Foundation Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.582860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.582860Z digest=sha256:22a20a93046bdd5564c65f1cca5b7792d80059c01a06bba016f959d87f668a28

Observation 807f7348-d4b2-45c7-ba36-a2e8112298f8 · outbound

This paper cites Photo-slam: Real-time simul- taneous localization and photorealistic mapping for monocular stereo and rgb-d cameras.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Photo-slam: Real-time simul- taneous localization and photorealistic mapping for monocular stereo and rgb-d cameras

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.344003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:50:53.589037Z digest=sha256:875ed12632d86ce771a34bb8b92be065b9361d956b76682a11093bfe20a4dbd2

Observation 1fc0ec9a-3838-4591-b4fa-0b8096f1dfec · outbound

This paper cites Image segmentation using text and image prompts.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Image segmentation using text and image prompts

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.322048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:50:53.594292Z digest=sha256:607b50f220472b9d9ea742e60e0955581404172d5f5862a43fd43e1c1b1530fd

Observation 3bca59fb-6576-4c28-8f48-23d558462aba · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing SAM 2: Segment Anything in Images and Videos

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.598703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.598703Z digest=sha256:5cf6f8990f9fde84e346450662ba32748bb5a98e8bf9d46b8dece6d405fcd04d

Observation d4850f78-4ad4-4792-a478-502cc15fb36c · outbound

This paper cites ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.604010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.604010Z digest=sha256:5a8a437b303c059976e21c87424692f590be93bf87ca3946e794ea02f4ee03bf

Observation 9f425254-a64c-4200-8c14-fc01a2b8f709 · outbound

This paper cites Brown University, 2000.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Brown University, 2000

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.300866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:50:53.610157Z digest=sha256:f5f2bb14732645048d627c1efd09d75f54d17937b2dceca4488beeac9d5ef1e6

Observation 5eb5c881-4c66-4058-9e95-398f0fe18c18 · outbound

This paper cites Coherent-to-diffuse power ratio estimation for dere- verberation.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 23(6):1006– 1018, 2015.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Coherent-to-diffuse power ratio estimation for dere- verberation.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 23(6):1006– 1018, 2015

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.282138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:50:53.616226Z digest=sha256:4b65d6e8a5cbf6459776fe05b3f0892e11f4606914e46f7af6e7fc24cf78f4a1

Observation 91ed6964-f625-4b2b-8f82-fc619039816a · outbound

This paper cites A density-based algorithm for discovering clusters in large spatial databases with noise.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing A density-based algorithm for discovering clusters in large spatial databases with noise

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.621308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.621308Z digest=sha256:d81b419f7cf868236e496436f5bd30dbc585e2f754a98f8a53738c751249ebfe

Observation 02188b04-effb-453d-b0b7-83a0edf3fc75 · outbound

This paper cites A new approach to linear filtering and prediction problems.Transactions of the ASME–Journal of Basic Engineering, 82(Series D):35–45, 1960.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing A new approach to linear filtering and prediction problems.Transactions of the ASME–Journal of Basic Engineering, 82(Series D):35–45, 1960

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.254676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:50:53.628005Z digest=sha256:92dd3d151c4203fc80e0a70bba908ff97c1aec0a58fc798161b23c680043a6b6

Observation c8411998-4f37-4049-accb-4eef08f76a8e · outbound

This paper cites The pascal visual object classes (voc) challenge.International Journal of Computer Vision, 88:303–338, 06 2010.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing The pascal visual object classes (voc) challenge.International Journal of Computer Vision, 88:303–338, 06 2010

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.635225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.635225Z digest=sha256:09ee403b14efeb22827f4095cb1a353a6a0a78bcc7bbe7bf1d757242602f19ea

Observation aa89350e-5ea6-4487-b41b-6fccf2ed5fd6 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.641537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.641537Z digest=sha256:d497624c6fdad0bfab44bbdf3d6705a48c61b9c5f3d56b660c494e5047462c09

Observation 5d13a14a-8e9a-4e43-adc4-bf1ea64118bd · outbound

This paper cites an unresolved cited work.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:50:54.224692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:50:53.648115Z digest=sha256:aaaa30523179d20d9393dd4dd3ee219979a348b0edfd0aac01316dd9c2280485

Pith citing papers

Observation ccdd4417-f720-4d73-a485-7bcd61c5b45d · inbound

EgoSound: Benchmarking Sound Understanding in Egocentric Videos cites this paper.

EgoSound: Benchmarking Sound Understanding in Egocentric Videos SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:46:42.583892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T21:44:47.636912Z digest=sha256:abb69f7272c2e330c9295868c974764ad7c9498a24f5e5678dbda4cea029e96b

Observation 162511d4-657a-4d71-8e94-51887078ef76 · inbound

Do Joint Audio-Video Generation Models Understand Physics? cites this paper.

Do Joint Audio-Video Generation Models Understand Physics? SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.223574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:ddb76c8086dc8bbbd476e9c2427041586e42665bb177254ab9e1c521e99cf33c

Observation 1105d00a-a7c5-41f2-b56c-abe588f4d4b2 · inbound

Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks cites this paper.

Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:13:11.900409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:10:31.059095Z digest=sha256:6d7a489ca630cea28c0884017592b4d2f0abdd29c10cb77cbaefc68fbc96e0cc