Pith. sign in

REVIEW 3 major objections 6 minor 104 references

Sonic Stage: Auto-Generating Interactive Spatial Soundscapes to Facilitate Dialogue Video Comprehension for Blind Viewers

T0 review · 3 major / 6 minor · reviewed 2026-08-01 · deepseek-v4-flash

Pith's one-line read Sonic Stage shows that a coherent 3D soundscape—spatialized dialogue, diegetic sounds, and tap-for details—can carry the visual information that audio description cannot fit into speech-dense scenes, improving blind viewers' comprehension.

desk verdict A solid systems contribution with a real advance in scene-space audio; the comparative claim leans on an unvalidated baseline, but the core result is likely to hold. read the letter →

arxiv 2607.20835 v2 pith:KYUQJ7V4 submitted 2026-07-23 cs.HC

classification cs.HC
keywords blindandlowvisionvideoaccessibilityspatialaudiodescriptiondialoguediegeticsound3Dscenereconstructioninteractivedescriptions
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Audio description cannot narrate actions during dialogue because it must not overlap speech, so dialogue-heavy films leave blind and low-vision viewers blind to character movement. Sonic Stage is the paper's answer: a fully automated pipeline that turns a dialogue clip into an interactive 3D soundscape, placing each voice at the character's position in the room, adding generated sound effects for on-screen actions, and offering short descriptions on tap. The paper's central claim is that these cues, anchored in a reconstructed 3D scene rather than the shifting screen image, let viewers keep a stable mental map across camera cuts. In a within-subject study of 12 blind and low-vision viewers against an interactive touch-exploration baseline, the authors report significantly higher accuracy on character position, movement, action, and visual-detail questions, plus higher spatial presence and narrative engagement. If right, the work shows a route to accessible video that works during speech rather than only in the gaps between speech.

What carries the argument

The machinery is the coherent 3D soundscape: a single reconstructed scene, built from sampled frames of the clip, into which all auditory cues are placed. The pipeline samples full and medium shots with moving characters, reconstructs a 3D point cloud of the space with a feed-forward model, tracks characters to recover their trajectories, and then optimizes the soundscape—listener at the characters' geometric center, left–right axis along the direction of greatest positional variance, logarithmic volume roll-off with distance, two-second trajectory smoothing, and a 70% spatial / 30% mono blend at speaker transitions. This shared reference frame is what keeps spatialized dialogue and diegetic

What would settle it

Take the same six user-study videos and re-run Sonic Stage with the dialogue artificially overlapped or with a scene change inserted mid-clip, then measure trajectory accuracy and recall. The paper's own stated assumptions predict a sharp drop: speaker labels would fail to link to on-screen characters, and the reconstructed 3D anchors would not exist. If comprehension accuracy for position and movement stayed near the reported 89% and 86% under those conditions, the boundary condition would be wrong; if it falls, the central claim is confirmed to hold only within the stated scope.

Watch

Extended reading notes

Core claim

On the authors' own terms, Sonic Stage's discovery is that the visual information audio description is forced to omit can be carried by the soundtrack itself. Three techniques work together: spatialized dialogue places each speaker's voice at a reconstructed 3D position so layout and movement are heard rather than described; diegetic sound effects render actions from the action's location; and interactive descriptions, invoked by a tap, supply dialogue-relevant details without pausing playback for long. The load-bearing design choice is scene-space anchoring: character trajectories and sound sources live in one reconstructed 3D scene, so camera cuts no longer jerk the audio around the way sc

Load-bearing premise

The system's benefit depends on the clip being a single physical space with non-overlapping speech and enough visual anchors for 3D reconstruction—the paper says so in its limitations—so scenes with overlapping dialogue, mid-scene cuts to new locations, or open environments would break the speaker linkage and spatial coherence that the user-study results rely on.

Editorial extensions

If this is right

  • Spatialized dialogue alone appears to carry layout and movement: participants scored 89% on position and 86% on movement, versus 44% and 18% with the touch-exploration baseline, suggesting audio description could offload spatial information to the soundtrack.
  • Diegetic sound with a half-second vibration made exploration self-timed: users explored more often with shorter pauses and yet watched no longer overall, evidence that cues can guide attention without stalling the narrative.
  • Because cues live in scene space rather than screen space, the same mechanism should survive camera cuts, which the paper identifies as the main cause of confusion in screen-space interactive systems.
  • The pipeline is compatible with existing audio description: it fills speech segments while audio description fills non-speech segments, and descriptions could be pruned to avoid redundancy with audio description.
  • The system's stated scope is single-space dialogue scenes with non-overlapping speech; results do not yet speak to overlapping dialogue, scene changes, or open environments.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the 45-point and 68-point gaps on position and movement replicate, the result suggests a design principle beyond accessibility: spatial consistency is an information channel of its own, and interactive systems that present spatial data frame-by-frame in screen space are structurally disadvantaged.
  • A testable extension: the same scene-space soundscape could be evaluated on live performances, where no camera cuts exist; the paper's participants already named concerts, opera, and dance as targets, but the system would need body-movement sonification that it does not currently attempt.
  • The paper's own constraints imply a falsifiable boundary: on clips with overlapping speech or open scenes, the speaker-to-character linkage and anchor-based reconstruction should degrade, and comprehension gains should shrink; measuring that degradation would clarify how much of the benefit comes from the soundscape versus the selection of easy videos.
  • Interactive description selection by dialogue relevance could generalize as a just-in-time audio design pattern: instead of describing everything, systems that surface the one detail tied to the current speech may reduce cognitive load in any audio-first interface.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper presents Sonic Stage, an automated system that converts dialogue-heavy videos into interactive 3D spatial soundscapes for blind and low-vision (BLV) viewers. Three techniques are combined: spatialized dialogue, diegetic sound effects, and on-demand interactive descriptions, all rendered in a stable 3D scene-space so that spatial cues remain coherent across camera cuts. The system is evaluated in a within-subject user study with 12 BLV participants against a baseline that the authors built and describe as 'modeled after SPICA.' The reported results show significantly higher objective recall for character position, movement, action, and visual detail, as well as higher spatial presence and narrative engagement, with no significant difference in dialogue comprehension. A separate technical evaluation reports 91.9% trajectory accuracy and 94.3% description accuracy on a 16-video dataset.

Significance. If the results hold, Sonic Stage addresses a real and important gap: audio description cannot convey visual actions during dense dialogue, and the paper's proposed auditory techniques are a plausible complement. The user study is a genuine contribution: it uses BLV participants, objective recall questions, a counterbalanced within-subject design, and both quantitative and qualitative analysis. The paper also ships a fully automated pipeline with concrete technical choices, including VGGT-based 3D reconstruction, which is a useful step toward scalable accessible video systems. However, the central comparative claim—that Sonic Stage improves comprehension over the state-of-the-art SPICA-style interaction—is only as strong as the fidelity of the author-implemented baseline, which the manuscript does not validate. The scope of the claim is also narrower than the abstract suggests, because the pipeline and evaluation are restricted to non-overlapping speech, single physical spaces, and dialogue-dominated scenes.

major comments (3)
  1. [§5.1.2, Appendix A.4, §2.1.4] The comparative claim is load-bearing and depends on an unvalidated baseline. The manuscript says the baseline is 'modeled after SPICA,' but no pilot study, expert audit, or feature-parity validation is reported to show that this implementation reproduces SPICA's spatial/temporal exploration behavior. Section 2.1.4 unqualifiedly refers to a 'comparative study between Sonic Stage and SPICA.' The baseline's movement recall of 18% is below the 33% chance level, which suggests that participants were not merely uninformed but actively confused by the screen-space interaction. This makes it difficult to attribute the large objective gaps (position 44% vs. 89%, movement 18% vs. 86%) to Sonic Stage's intrinsic benefits. The authors should either validate the baseline against SPICA, compare with the original SPICA system, or explicitly reframe the comparison as 'a touch-based screen-space baselin
  2. [§3.2, §7.5, Abstract] The central claim that Sonic Stage 'transforms dialogue videos' and significantly improves comprehension is bounded by assumptions that are stated only later: non-overlapping speech, a single physical space, sufficient visual anchors for 3D reconstruction, and high dialogue ratios (the user-study videos have 87–97% speech). These conditions hold for the selected clips but fail for much real film and TV, which includes overlapping dialogue, scene changes, and open or moving-camera scenes. The paper acknowledges these limits in §7.5, but the abstract and introduction present the system without these qualifications. The authors should state these boundary conditions prominently in the abstract and introduction, and ideally include an analysis of how the pipeline degrades when each assumption is violated, rather than leaving this to future work.
  3. [§4.2, §5.2.3] The statistical analysis is acceptable for a UIST-style user study, but the authors should address the multiple-comparison issue. Five per-category paired t-tests are reported without correction; while the effects for position, movement, action, and visual detail are individually significant at p<.01, a correction such as Bonferroni or Holm would make the evidence more robust and is standard for this number of tests. The technical evaluation also lacks an inter-rater reliability metric: Section 4.2 says one researcher labeled and a second 'reviewed the labels,' but no agreement score (e.g., Cohen's kappa) is reported. This should be added or explicitly justified.
minor comments (6)
  1. [§4.2] The claim that 'most inaccuracies were minor positional shifts below human auditory resolution' is not substantiated with measurements. If quantitative error magnitudes are available, reporting them would strengthen the argument that 91.9% trajectory accuracy is sufficient for the user experience.
  2. [§3.3.3] The soundscape parameters (V_max, V_min, D_near/D_far, spatial blend factor, smoothing window) are tuned with two BLV sound designers but no systematic sensitivity analysis is provided. A brief exploration of how the results change with these parameters would help establish that the user-study outcomes are not artifacts of a single hand-picked configuration.
  3. [§3.3.2] The character position is estimated as the mean of ten randomly sampled points from the projected segmentation mask. This is an arbitrary choice; a small sensitivity check (e.g., 5 vs. 20 points) or a rationale based on mask noise would be helpful.
  4. [Appendix A.4] The baseline description generation is said to use 'the method in the SPICA system' but no details are given about the object description model or prompts. Providing the exact generation pipeline would improve reproducibility.
  5. [§6.1, Figure 5] The figure labels for significance levels could be clearer: the text reports t-values and p-values, but the figure would benefit from explicit significance markers (e.g., asterisks) and error bars showing within-subject variability, not just between-subject standard errors.
  6. [§7.5] Minor typos: 'this approach that does not generalize well' should read 'this approach does not generalize well'; 'Future work could how to sonify' should read 'Future work could explore how to sonify.'

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the central claim is an empirical comparison, not a derivation from fitted inputs or self-citation.

full rationale

The paper's central assertion is empirical: a within-subject user study compares Sonic Stage with a baseline 'modeled after SPICA' and reports significantly better comprehension, spatial presence, and engagement. There is no mathematical derivation that reduces the reported outcome to the system's inputs. The technical evaluation uses manual labels and pretrained models (VGGT, YOLOv11, TalkNet, Gemini) against ground-truth video content; this is an external benchmark, not a fitted prediction. Hand-tuned parameters (Vmax, Vmin, Dnear, Dfar, smoothing window, spatial blend factor) were validated with two blind sound designers, but they were not fit to the comprehension outcome, so the headline improvements are not forced by those choices. The paper explicitly acknowledges scope limits—non-overlapping speech, single physical spaces, visual anchors for 3D reconstruction—which constrain when the system works but do not constitute circularity. The phrase in Section 2.1.4 calling the evaluation a 'comparative study between Sonic Stage and SPICA' is a loose description of a baseline modeled after SPICA, and the unvalidated fidelity of that baseline is a correctness/construct-validity risk, not a circularity pattern under the definitions used here. No equation is shown to equal another by construction, and no fitted parameter is renamed as a prediction. Self-citations appear but are not load-bearing for the main result.

Assumptions & free parameters 5 free parameters · 7 assumptions · 0 invented entities

The central claim rests on a pipeline of black-box models and on domain constraints that the authors explicitly delimit in Section 7.5. No new physical or mathematical entities are introduced; the free parameters are design constants tuned with two sound designers. The most structurally important assumptions are the non-overlapping-speech and single-scene constraints, and the reliability of the manual technical-evaluation labels.

free parameters (5)
  • Volume roll-off endpoints V_max, V_min = 1.0, 0.5
    Set by iterative testing with two BLV sound designers (Section 3.3.3); directly controls perceived distance changes in spatialized dialogue.
  • Spatial blend factor = 0.7
    70% spatial / 30% mono, validated by sound designers (Section 3.3.3) to smooth speaker transitions; affects the core spatialization experience.
  • Trajectory smoothing window = 2 s
    Two-second moving average applied to character trajectories (Section 3.3.3); affects the stability and responsiveness of spatial cues.
  • Motion sampling threshold = 25% of bounding-box width/height per second
    Frames sampled once per second for moving characters, otherwise only middle frame (Section 3.3.1); influences 3D reconstruction input and trajectory quality.
  • Distance roll-off percentiles D_near/D_far = median / 90th percentile of scene speaker-listener distances
    Chosen per-scene to limit outlier influence (Section 3.3.3); no ablation or independent validation is provided.
assumptions (7)
  • domain assumption Dialogue speech does not overlap between speakers
    Stated in Section 4.1 dataset description and Section 7.5: 'Currently, Sonic Stage assumes non-overlapping speech.' Speaker separation and TalkNet-based linking require this.
  • domain assumption The scene is contained in a single physical space with enough visual anchors for 3D reconstruction
    Section 7.5: open-ended scenes 'often lack visual anchors for 3D reconstruction'; the pipeline only reconstructs a fixed, stable soundscape.
  • domain assumption Stereo spatial audio with HRTF is an adequate channel for conveying spatial layout and movement to BLV viewers
    Section 3.6 uses Unity's Steam Audio HRTF; Section 7.4 acknowledges front-back ambiguity limits this channel.
  • domain assumption Pretrained models (VGGT, YOLOv11, BoT-SORT, TalkNet, Gemini-2.5, ElevenLabs) are sufficiently accurate for detection, reconstruction, description, and sound generation
    Sections 3.3–3.5 route every stage through black-box commercial or research models without independent validation of each stage.
  • ad hoc to paper A character's 3D position can be approximated by the mean of ten random points in the projected segmentation mask
    Section 3.3.2 defines this estimator; no ablation or accuracy analysis of the approximation is provided.
  • domain assumption The author-built baseline is a faithful implementation of SPICA
    Section 5.1.2 describes a 'baseline modeled after SPICA'; no validation against the original SPICA system is reported, so relative gains may depend on implementation quality.
  • domain assumption Manual labeling in the technical evaluation is reliable
    Sections 4.2–4.3 report one researcher labeling with a second reviewer; no inter-rater reliability, ground-truth benchmark, or confidence intervals are given.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Sonic Stage: Auto-Generating Interactive Spatial Soundscapes to Facilitate Dialogue Video Comprehension for Blind Viewers." pith.science (2026). https://pith.science/paper/KYUQJ7V4

@misc{pith2026260720835,
  author       = {Pith},
  title        = {Pith review of: Sonic Stage: Auto-Generating Interactive Spatial Soundscapes to Facilitate Dialogue Video Comprehension for Blind Viewers},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KYUQJ7V4}},
  note         = {Machine review of arXiv:2607.20835}
}
read the original abstract

Audio description (AD) makes film and television accessible to blind and low-vision (BLV) audiences by narrating characters' actions. However, in scenes with lots of dialogue, AD often omits important actions because it is constrained not to overlap with speech. It is not yet known how to convey characters' actions during dialogue. We present Sonic Stage, a system that transforms dialogue videos into interactive spatial soundscapes, enabling BLV audiences to intuitively understand characters' actions and movements through immersive auditory cues. Sonic Stage conveys essential visual information during dialogue through three auditory techniques: (1) spatialized dialogue to represent spatial layout, (2) diegetic sound to convey character actions, and (3) interactive descriptions to provide context-specific visual details. Evaluation with 12 BLV viewers showed that Sonic Stage significantly improved video comprehension, spatial presence, and narrative engagement. We highlight opportunities for enhancing video accessibility across diverse genres through immersive, interactive audio representations.

Figures

Figures reproduced from arXiv: 2607.20835 by the authors.

Figure 1
Figure 1. (Left) In videos with lots of dialogue, there is little opportunity to insert audio descriptions. As a result, blind viewers [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. A walkthrough of Sonic Stage. (A) Throughout the video, users hear spatialized dialogue coming from the speakers’ [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Sonic Stage constructs a coherent spatial soundscape [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (11 more)
Figure 4
Figure 4. Figure 4: Examples of videos used in the technical evaluation. [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Accuracy rates on recall questions by category. Each category included 72 responses per system. Statistical significance [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Usage patterns for each system. Sonic Stage led to more frequent exploration with shorter pauses, without increasing [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: The approach for integrating Sonic Stage with AD. 7.4 Toward Immersive Audio Representations Our study identifies several directions for advancing the immer￾sive audio representations of Sonic Stage. First, future systems could support alternative listening perspective…
Figure 8
Figure 8. Figure 8: Sonic Stage can be extended to diverse video genres. [PITH_FULL_IMAGE:figures/full_fig_p011_8.png]
Figure 9
Figure 9. Figure 9: Participant rating distributions for both systems (1 = strongly negative, 7 = strongly positive). Asterisks indicate [PITH_FULL_IMAGE:figures/full_fig_p014_9.png]
Figure 10
Figure 10. Figure 10: Participant ratings of the audio quality in Sonic Stage (1 = strongly negative, 7 = strongly positive). [PITH_FULL_IMAGE:figures/full_fig_p014_10.png]
Figure 11
Figure 11. Figure 11: Screenshots from the 16-video dataset used for system evaluation. [PITH_FULL_IMAGE:figures/full_fig_p015_11.png]
Figure 12
Figure 12. Figure 12: The baseline system supports touch-based spatial exploration. Users can (A) pause the video, (B) move their fingers [PITH_FULL_IMAGE:figures/full_fig_p016_12.png]
Figure 13
Figure 13. Figure 13: The Sonic Stage pipeline comprises three core modules. The first module applies 3D reconstruction to recover [PITH_FULL_IMAGE:figures/full_fig_p017_13.png]
Figure 14
Figure 14. Figure 14: An example of the 3D point cloud and character trajectories reconstructed by Sonic Stage. The trajectory variations [PITH_FULL_IMAGE:figures/full_fig_p017_14.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

104 extracted references · 7 linked inside Pith

  1. [1]

    Scene Detect

    2024. Scene Detect. https://www.scenedetect.com/. Accessed: 2025-06-11

  2. [2]

    Unity: Spatial Blend

    2025. Unity: Spatial Blend. https://docs.unity3d.com/ScriptReference/ AudioSource-spatialBlend.html. Accessed: 2025-06-11

  3. [3]

    Nir Aharon, Roy Orfaig, and Ben-Zion Bobrovsky. 2022. BoT-SORT: Robust associations multi-pedestrian tracking.arXiv preprint arXiv:2206.14651(2022)

  4. [4]

    Sercan Arik, Jitong Chen, Kainan Peng, Wei Ping, and Yanqi Zhou. 2018. Neural voice cloning with a few samples.Advances in neural information processing systems31 (2018)

  5. [5]

    Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al . 2025. Qwen2. 5-vl technical report.arXiv preprint arXiv:2502.13923(2025)

  6. [6]

    Bela Balazs. 1985. Theory of the film: Sound. 116–125 pages

  7. [7]

    Maxime Bleau, Camille van Acker, Natalina Martiniello, Joseph Paul Nemargut, and Maurice Ptito. 2023. Cognitive map formation in the blind is enhanced by three-dimensional tactile information.Scientific Reports13, 1 (2023), 9736

  8. [8]

    James C Bliss, Michael H Katcher, Charles H Rogers, and Raymond P Shepard

Show all 104 references
  1. [9]

    2010.Immersed in media: Telepres- ence in everyday life

    Cheryl Campanella Bracken and Paul Skalski. 2010.Immersed in media: Telepres- ence in everyday life. Routledge

  2. [10]

    Virginia Braun and Victoria Clarke. 2019. Reflecting on reflexive thematic analysis. Qualitative research in sport, exercise and health11, 4 (2019), 589–597

  3. [11]

    Rick Busselle and Helena Bilandzic. 2009. Measuring narrative engagement. Media psychology12, 4 (2009), 321–347

  4. [12]

    Matthew Butler, Leona M Holloway, Samuel Reinders, Cagatay Goncu, and Kim Marriott. 2021. Technology developments in touch-based accessible graphics: A systematic review of research 2010-2020. InProceedings of the 2021 CHI conference on human factors in computing systems. 1–15...

  5. [13]

    Ben Caldwell, Michael Cooper, Loretta Guarino Reid, Gregg Vanderheiden, Wendy Chisholm, John Slatin, and Jason White. 2008. Web content accessibility guide- lines (WCAG) 2.0.WWW Consortium (W3C)290, 1-34 (2008), 5–12

  6. [14]

    Anil Çamcı, Kristine Lee, Cody J Roberts, and Angus G Forbes. 2017. INVISO: a cross-platform user interface for creating virtual sonic environments. InProceed- ings of the 30th Annual ACM Symposium on User Interface Software and Technology. 507–518

  7. [15]

    Ruei-Che Chang, Chao-Hsien Ting, Chia-Sheng Hung, Wan-Chen Lee, Liang- Jin Chen, Yu-Tzu Chao, Bing-Yu Chen, and Anhong Guo. 2022. OmniScribe: Authoring Immersive Audio Descriptions for 360°Videos. InProceedings of the 35th Annual ACM Symposium on User Interface Software and Te...

  8. [16]

    Maryam Cheema, Sina Elahimanesh, Samuel Martin, Pooyan Fazli, and Hasti Seifi

  9. [17]

    Maryam Cheema, Hasti Seifi, and Pooyan Fazli. 2025. Describe Now: User-Driven Audio Description for Blind and Low Vision Individuals. InProceedings of the 2025 ACM Designing Interactive Systems Conference. 458–474

  10. [18]

    Boyuan Chen, Zhuo Xu, Sean Kirmani, Brain Ichter, Dorsa Sadigh, Leonidas Guibas, and Fei Xia. 2024. Spatialvlm: Endowing vision-language models with spatial reasoning capabilities. InProceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition. 14455–14465

  11. [19]

    Chang Chen, Sicheng Song, Shuchang Xu, Zhicheng Li, Huamin Qu, and Yanna Lin. 2025. RhythmTA: A Visual-Aided Interactive System for ESL Rhythm Train- ing via Dubbing Practice. InProceedings of the 38th Annual ACM Symposium on User Interface Software and Technology. 1–15

  12. [20]

    Shi Chen, Jingao Zhang, Suqi Lou, Xiaodong Wang, Wei Xiang, and Lingyun Sun. 2025. Voice by the Non-sighted: Practices and Challenges of Audiobook Voice Actors with Blind and Low Vision in China. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–19

  13. [21]

    Arnavi Chheda-Kothary, Ather Sharif, David Angel Rios, and Brian A Smith

  14. [22]

    2019.Audio-vision: sound on screen

    Michel Chion. 2019.Audio-vision: sound on screen. Columbia University Press

  15. [23]

    Hyunsung Cho, Alexander Wang, Divya Kartik, Emily Liying Xie, Yukang Yan, and David Lindlbauer. 2024. Auptimize: Optimal placement of spatial audio cues for extended reality. InProceedings of the 37th Annual ACM Symposium on User Interface Software and Technology. 1–14

  16. [24]

    It Brought Me Joy

    " It Brought Me Joy": Opportunities for Spatial Browsing in Desktop Screen Readers. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–18

  17. [25]

    Hainan Cui, Xiang Gao, Shuhan Shen, and Zhanyi Hu. 2017. HSfM: Hybrid structure-from-motion. InProceedings of the IEEE conference on computer vision and pattern recognition. 1212–1221

  18. [26]

    Ginger Delmas, Philippe Weinzaepfel, Thomas Lucas, Francesc Moreno-Noguer, and Grégory Rogez. 2022. Posescript: 3d human poses from natural language. In European Conference on Computer Vision. Springer, 346–362

  19. [27]

    Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, et al. 2025. Gemini 2.5: Pushing the frontier with advanced reasoning, multi- modality, long context, and next generation agentic...

  20. [28]

    Alessandro Favero, Luca Zancato, Matthew Trager, Siddharth Choudhary, Pra- muditha Perera, Alessandro Achille, Ashwin Swaminathan, and Stefano Soatto

  21. [29]

    Gaudio. 2025. AI stem splitter. https://www.gaudiolab.com/gaudio-studio/

  22. [30]

    Hilko Donker, Palle Klante, and Peter Gorny. 2002. The design of auditory user interfaces for blind users. InProceedings of the second Nordic conference on Human-computer interaction. 149–156

  23. [31]

    Google. 2025. Gemini Audio Understanding. https://ai.google.dev/gemini-api/ docs/audio

  24. [32]

    Google. 2025. Gemini Video Understanding. https://ai.google.dev/gemini-api/ docs/video-understanding

  25. [33]

    João Guerreiro, Yujin Kim, Rodrigo Nogueira, SeungA Chung, André Rodrigues, and Uran Oh. 2023. The design space of the auditory representation of objects and their behaviours in virtual reality for blind people.IEEE Transactions on visualization and computer graphics29, 5 (202...

  26. [34]

    Rohit Girmaji, Bhav Beri, Ramanathan Subramanian, and Vineet Gandhi. 2025. EditIQ: Automated Cinematic Editing of Static Wide-Angle Videos via Dialogue Interpretation and Saliency Cues. InProceedings of the 30th International Confer- ence on Intelligent User Interfaces. 609–623

  27. [35]

    Tilo Hartmann, Werner Wirth, Holger Schramm, Christoph Klimmt, Peter Vorderer, André Gysbers, Saskia Böcking, Niklas Ravaja, Jari Laarni, Timo Saari, et al. 2015. The spatial presence experience scale (SPES).Journal of Media Psychology(2015)

  28. [36]

    Leona Holloway, Kim Marriott, and Matthew Butler. 2018. Accessible maps for the blind: Comparing 3D printed models with tactile graphics. InProceedings of the 2018 CHI conference on human factors in computing systems. 1–13

  29. [37]

    Gaurav Jain, Basel Hindi, Connor Courtien, Xin Yi Therese Xu, Conrad Wyrick, Michael Malcolm, and Brian A. Smith. 2023. Front Row: Automatically Generating Immersive Audio Representations of Tennis Broadcasts for Blind Viewers. In Proceedings of the 36th Annual ACM Symposium o...

  30. [38]

    2003.Multiple view geometry in computer vision

    Richard Hartley and Andrew Zisserman. 2003.Multiple view geometry in computer vision. Cambridge university press

  31. [39]

    Lucy Jiang, Crescentia Jung, Mahika Phutane, Abigale Stangl, and Shiri Azenkot

  32. [40]

    Lucy Jiang, Mahika Phutane, and Shiri Azenkot. 2023. Beyond Audio Description: Exploring 360°Video Accessibility with Blind and Low Vision Users Through Collaborative Creation. InProceedings of the 25th International ACM SIGACCESS Conference on Computers and Accessibility(New ...

  33. [41]

    Ziyi Jiang, Mengjie Jian, Jiajia Hu, Hongze Zhao, Huamin Qu, Shuchang Xu, and Guanhong Liu. 2025. From Audio Description to Movement: Challenges and Design Principles in Video-based Body Exercise for Blind and Low Vision Users. InCompanion Publication of the 2025 Conference on...

  34. [42]

    Chutian Jiang, Emily Kuang, and Mingming Fan. 2025. How can haptic feedback assist people with blind and low vision (BLV): A systematic literature review. ACM Transactions on Accessible Computing18, 1 (2025), 1–57

  35. [43]

    Rahima Khanam and Muhammad Hussain. 2024. Yolov11: An overview of the key architectural enhancements.arXiv preprint arXiv:2410.17725(2024)

  36. [44]

    It’s Kind of Context Dependent

    “It’s Kind of Context Dependent”: Understanding Blind and Low Vision People’s Video Accessibility Preferences Across Viewing Scenarios. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems(Honolulu, HI, USA)(CHI ’24). Association for Computing Machine...

  37. [45]

    Mackenzie Leake, Abe Davis, Anh Truong, and Maneesh Agrawala. 2017. Com- putational video editing for dialogue-driven scenes.ACM Trans. Graph.36, 4 (2017), 130–1

  38. [46]

    James R Lewis. 2018. The system usability scale: past, present, and future.Inter- national Journal of Human–Computer Interaction34, 7 (2018), 577–590

  39. [47]

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis

  40. [48]

    David Chuan-En Lin, Anastasis Germanidis, Cristóbal Valenzuela, Yining Shi, and Nikolas Martelaro. 2023. Soundify: Matching sound effects to video. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 1–13

  41. [49]

    Guanhong Liu, Tianyu Yu, Chun Yu, Haiqing Xu, Shuchang Xu, Ciyuan Yang, Feng Wang, Haipeng Mi, and Yuanchun Shi. 2021. Tactile compass: Enabling visually impaired people to follow a path with continuous directional feedback. InProceedings of the 2021 CHI Conference on Human Fa...

  42. [50]

    Daniel Killough and Amy Pavel. 2023. Exploring Community-Driven Descriptions for Making Livestreams Accessible. InProceedings of the 25th International ACM SIGACCESS Conference on Computers and Accessibility. 1–13

  43. [51]

    Sheng Liu, Xiaohan Nie, and Raffay Hamid. 2022. Depth-guided sparse structure- from-motion for movies and tv shows. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 15980–15989

  44. [52]

    Xingyu Liu, Patrick Carrington, Xiang’Anthony’ Chen, and Amy Pavel. 2021. What makes videos accessible to blind and visually impaired people?. InPro- ceedings of the 2021 CHI Conference on Human Factors in Computing Systems. 1–14

  45. [53]

    Chaoyu Li, Sid Padmanabhuni, Maryam S Cheema, Hasti Seifi, and Pooyan Fazli

  46. [54]

    In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25)

    VideoA11y: Method and Dataset for Accessible Video Description. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Association for Computing Machinery, New York, NY, USA, Article 1055, 29 pages. doi:10.1145/3706598.3714096

  47. [55]

    Mariana Lopez, Gavin Kearney, and Krisztián Hofstädter. 2022. Seeing films through sound: Sound design, spatial audio, and accessibility for visually impaired audiences.British Journal of Visual Impairment40, 2 (2022), 117–144. Sonic Stage UIST ’26, November 02–05, 2026, Detro...

  48. [56]

    Mariana Julieta Lopez, Gavin Kearney, and Krisztian Hofstadter. 2021. Enhancing audio description: Inclusive cinematic experiences through sound design.Journal of Audiovisual Translation4, 1 (2021), 157–182

  49. [57]

    Hongbo Liu, Jingwen He, Yi Jin, Dian Zheng, Yuhao Dong, Fan Zhang, Ziqi Huang, Yinan He, Yangguang Li, Weichao Chen, et al. 2025. ShotBench: Expert- Level Cinematic Understanding in Vision-Language Models.arXiv preprint arXiv:2506.21356(2025)

  50. [58]

    Troy McDaniel, Lakshmie Narayan Viswanathan, and Sethuraman Panchanathan

  51. [59]

    2012.Acoustics: sound fields and transducers

    Tim Mellow. 2012.Acoustics: sound fields and transducers. Academic Press

  52. [60]

    Xingyu Liu, Biao Wang, Wayne Zhang, Ziqian Liao, Ziwen Li, Amy Pavel, Xi- ang’Anthony’ Chen, et al. 2025. CoSight: Exploring Viewer Contributions to Online Video Accessibility Through Descriptive Commenting.arXiv preprint arXiv:2508.08582(2025)

  53. [61]

    Xingyu" Bruce" Liu, Ruolin Wang, Dingzeyu Li, Xiang Anthony Chen, and Amy Pavel. 2022. CrossA11y: Identifying Video Accessibility Issues via Cross-modal Grounding. InProceedings of the 35th Annual ACM Symposium on User Interface Software and Technology. 1–14

  54. [62]

    Rosiana Natalie, Joshua Tseng, Hernisa Kacorri, and Kotaro Hara. 2023. Support- ing novices author audio descriptions via automatic feedback. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–18

  55. [63]

    Netflix. 2024. Audio Description Style Guide v2.5. https://partnerhelp. netflixstudios.com/hc/en-us/articles/215510667-Audio-Description-Style- Guide-v2-5. Accessed: 2025-06-29

  56. [64]

    Michał Maćkowski, Piotr Brzoza, Mateusz Kawulok, Rafał Meisel, and Dominik Spinczyk. 2023. Multimodal presentation of interactive audio-tactile graphics supporting the perception of visual information by blind people.ACM Transac- tions on Multimedia Computing, Communications a...

  57. [65]

    Zheng Ning, Zheng Zhang, Jerrick Ban, Kaiwen Jiang, Ruohong Gan, Yapeng Tian, and Toby Jia-Jun Li. 2024. MIMOSA: Human-AI Co-Creation of Computational Spatial Audio Effects on Videos. InProceedings of the 16th Conference on Creativity & Cognition(Chicago, IL, USA)(C&C ’24). As...

  58. [66]

    Ofcom. 2024. Guidelines on Providing TV and On-Demand Access Services. https://www.ofcom.org.uk/siteassets/resources/documents/tv-radio-and-on- demand/broadcast-codes/other-codes/ofcoms-guidelines-on-providing-tv- and-on-demand-access-services.pdf. Accessed: 2025-06-29

  59. [67]

    Amy Pavel, Gabriel Reyes, and Jeffrey P. Bigham. 2020. Rescribe: Authoring and Automatically Editing Audio Descriptions. InProceedings of the 33rd Annual ACM Symposium on User Interface Software and Technology(Virtual Event, USA) (UIST ’20). Association for Computing Machinery...

  60. [68]

    Pedro Morgado, Nuno Nvasconcelos, Timothy Langlois, and Oliver Wang. 2018. Self-supervised generation of spatial audio for 360 video.Advances in neural information processing systems31 (2018)

  61. [69]

    Rosiana Natalie, Ruei-Che Chang, Smitha Sheshadri, Anhong Guo, and Kotaro Hara. 2024. Audio Description Customization. InProceedings of the 26th Inter- national ACM SIGACCESS Conference on Computers and Accessibility(St. John’s, NL, Canada)(ASSETS ’24). Association for Computi...

  62. [70]

    SC Lannom. 2025. Guide to Camera Shots: Every Shot Size Explained. https: //www.studiobinder.com/blog/types-of-camera-shots-sizes-in-film/

  63. [71]

    Johannes L Schonberger and Jan-Michael Frahm. 2016. Structure-from-motion revisited. InProceedings of the IEEE conference on computer vision and pattern recognition. 4104–4113

  64. [72]

    Zheng Ning, Brianna L Wimer, Kaiwen Jiang, Keyi Chen, Jerrick Ban, Yapeng Tian, Yuhang Zhao, and Toby Jia-Jun Li. 2024. SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers. InProceedings of the 2024 CHI Conference o...

  65. [73]

    Ruijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian, Mike Zheng Shou, and Haizhou Li. 2021. Is someone speaking? exploring long-term temporal features for audio-visual active speaker detection. InProceedings of the 29th ACM international conference on multimedia. 3927–3935

  66. [74]

    Tencent. 2025. Tencent Transcription. https://cloud.tencent.com/product/asr/

  67. [75]

    Dimitrios Tzovaras, Georgios Nikolakis, Georgios Fergadis, Stratos Malasiotis, and Modestos Stavrakis. 2004. Design and implementation of haptic virtual environments for the training of the visually impaired.IEEE Transactions on Neural Systems and Rehabilitation Engineering12,...

  68. [76]

    Yi-Hao Peng, Jeffrey P Bigham, and Amy Pavel. 2021. Slidecho: Flexible non- visual exploration of presentation videos. InProceedings of the 23rd International ACM SIGACCESS Conference on Computers and Accessibility. 1–12

  69. [77]

    Minghan Qin, Wanhua Li, Jiawei Zhou, Haoqian Wang, and Hanspeter Pfister

  70. [78]

    InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Langsplat: 3d language gaussian splatting. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 20051–20060

  71. [79]

    Lakshmie Narayan Viswanathan, Troy McDaniel, Sreekar Krishna, and Sethu- raman Panchanathan. 2010. Haptics in audio described movies. In2010 IEEE International Symposium on Haptic Audio Visual Environments and Games. IEEE, 1–2

  72. [80]

    Volcengine. 2025. Volcengine Text-to-Speech. https://www.volcengine.com/ product/tts

  73. [81]

    Abigale Stangl, Shasta Ihorn, Yue-Ting Siu, Aditya Bodi, Mar Castanon, Lothar D Narins, and Ilmi Yoon. 2023. The potential of a visual dialogue agent in a tan- dem automated audio description system for videos. InProceedings of the 25th International ACM SIGACCESS Conference o...

  74. [82]

    Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rup- precht, and David Novotny. 2025. Vggt: Visual geometry grounded transformer. InProceedings of the Computer Vision and Pattern Recognition Conference. 5294– 5306

  75. [83]

    Yujia Wang, Wei Liang, Haikun Huang, Yongqi Zhang, Dingzeyu Li, and Lap-Fai Yu. 2021. Toward Automatic Audio Description Generation for Accessible Videos. InProceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama, Japan)(CHI ’21). Association for...

  76. [84]

    Frank Wilcoxon, S Katti, Roberta A Wilcox, et al . 1970. Critical values and probability levels for the Wilcoxon rank sum test and the Wilcoxon signed rank test.Selected tables in mathematical statistics1 (1970), 171–259

  77. [85]

    Valve Corporation. 2025. Steam Audio. https://valvesoftware.github.io/steam- audio

  78. [86]

    Tess Van Daele, Akhil Iyer, Yuning Zhang, Jalyn C Derry, Mina Huh, and Amy Pavel. 2024. Making Short-Form Videos Accessible with Hierarchical Video Summaries. InProceedings of the CHI Conference on Human Factors in Computing Systems. 1–17

  79. [87]

    Valentijn T Visch, Ed S Tan, and Dylan Molenaar. 2010. The emotional and cognitive effect of immersion in film viewing.Cognition and Emotion24, 8 (2010), 1439–1445

  80. [88]

    Shuchang Xu, Xiaofu Jin, Huamin Qu, and Yukang Yan. 2025. DanmuA11y: Making Time-Synced On-Screen Video Comments (Danmu) Accessible to Blind and Low Vision Users via Multi-Viewer Audio Discussions. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems ...

  81. [89]

    Shuchang Xu, Xiaofu Jin, Wenshuo Zhang, Huamin Qu, and Yukang Yan. 2025. Branch Explorer: Leveraging Branching Narratives to Support Interactive 360° Video Viewing for Blind and Low Vision Users.arXiv preprint arXiv:2507.09959 (2025)

  82. [90]

    Jindu Wang, Runze Cai, Shuchang Xu, Tianrui Hu, Huamin Qu, Shengdong Zhao, and Lin-Ping Yuan. 2026. Wearable AR for Restorative Breaks: How Interactive Narrative Experiences Support Relaxation for Young People. InProceedings of the 2026 CHI Conference on Human Factors in Compu...

  83. [91]

    Xudong Xu, Hang Zhou, Ziwei Liu, Bo Dai, Xiaogang Wang, and Dahua Lin

  84. [92]

    Baiqiao Yin, Qineng Wang, Pingyue Zhang, Jianshu Zhang, Kangrui Wang, Zihan Wang, Jieyu Zhang, Keshigeyan Chandrasegaran, Han Liu, Ranjay Kr- ishna, et al. 2025. Spatial Mental Modeling from Limited Views.arXiv preprint arXiv:2506.21458(2025)

  85. [93]

    Yuhang Zhao, Cynthia L Bennett, Hrvoje Benko, Edward Cutrell, Christian Holz, Meredith Ringel Morris, and Mike Sinclair. 2018. Enabling people with visual impairments to navigate virtual reality with a haptic and auditory cane simulation. InProceedings of the 2018 CHI conferen...

  86. [94]

    2013.Head-related transfer function and virtual auditory display

    Bosun Xie. 2013.Head-related transfer function and virtual auditory display. J. Ross Publishing

  87. [95]

    Shuchang Xu, Chang Chen, Zichen Liu, Xiaofu Jin, Lin-Ping Yuan, Yukang Yan, and Huamin Qu. 2024. Memory reviver: supporting photo-collection reminiscence for people with visual impairment via a proactive Chatbot. InProceedings of the 37th Annual ACM Symposium on User Interface...

  88. [96]

    Smith, and Yukang Yan

    Shuchang Xu, Xiaofu Jin, Gaurav Jain, Wenshuo Zhang, Huamin Qu, Brian A. Smith, and Yukang Yan. 2026. Sonic Stage: Automatically Generating an In- teractive Spatial Soundscape to Facilitate Dialogue Video Comprehension for Blind and Low Vision Viewers. InProceedings of the Ext...

  89. [99]

    Shuchang Xu, Ciyuan Yang, Wenhao Ge, Chun Yu, and Yuanchun Shi. 2020. Virtual Paving: Rendering a smooth path for people with visual impairment through vibrotactile and audio feedback.Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies4, 3 (2020), 1–25

  90. [104]

    Hang Zhou, Xudong Xu, Dahua Lin, Xiaogang Wang, and Ziwei Liu. 2020. Sep- stereo: Visually guided stereophonic audio generation by associating source separation. InEuropean Conference on Computer Vision. Springer, 52–69. UIST ’26, November 02–05, 2026, Detroit, MI, USA Xu et a...

  91. [2007]

    Optical-to-tactile image conversion for the blind.IEEE Transactions on Man-Machine Systems11, 1 (2007), 58–65

  92. [2013]

    In2013 IEEE International Conference on Multimedia and Expo (ICME)

    An evaluation of haptic descriptions for audio described films for individuals who are blind. In2013 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 1–6

  93. [2021]

    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Visually informed binaural audio generation without binaural audios. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 15485–15494

  94. [2023]

    Graph.42, 4 (2023), 139–1

    3D Gaussian splatting for real-time radiance field rendering.ACM Trans. Graph.42, 4 (2023), 139–1

  95. [2024]

    InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Multi-modal hallucination control by visual information grounding. InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 14303–14312

  96. [2025]

    arXiv preprint arXiv:2508.01092(2025)

    DescribePro: Collaborative Audio Description with Human-AI Interaction. arXiv preprint arXiv:2508.01092(2025)

Pith tools

Reviewed August 1, 2026 · model on record in the stance chip above.