Pith. sign in

Paper Citation Record · LEDGER

Learning to Highlight Audio by Watching Movies

As of 24 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 0 inbound Pith citation observations for arXiv:2505.12154.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12154 v1

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:43:56.792270Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

80 of 80 outbound references displayed

  • verified exact3
  • verified fuzzy55
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a51cb12f-4da1-434e-b5f3-bc35db7a6d7a · outbound

This paper cites Self-supervised learning of audio-visual objects from video.European Conference on Computer Vision (ECCV), 2020.

Learning to Highlight Audio by Watching Movies Self-supervised learning of audio-visual objects from video.European Conference on Computer Vision (ECCV), 2020

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.648890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.514472Z digest=sha256:e425ce009b4b9e32adb4c174a7d153b25570be22d9a9825c3583d15ca1a6db9d

Observation efd189c0-596a-480b-8655-4baec0512960 · outbound

This paper cites Condensed movies: Story based retrieval with con- textual embeddings, 2020.

Learning to Highlight Audio by Watching Movies Condensed movies: Story based retrieval with con- textual embeddings, 2020

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.637795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.519297Z digest=sha256:28ef5e524d29456c7a8c8a4bdd4619bf69c0c59f692076042ca4436abafb97b0

Observation 4ae7a440-8f83-4193-97db-7fa0c7546888 · outbound

This paper cites Ac- tion2sound: Ambient-aware generation of action sounds from egocentric videos.

Learning to Highlight Audio by Watching Movies Ac- tion2sound: Ambient-aware generation of action sounds from egocentric videos

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.626421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.522620Z digest=sha256:f2636268ef0dc53d7e81fdc9d21c60a38858a4c07d252381b6c3903c0d2e51e3

Observation 53f02eec-1e20-4b44-90aa-df0abfd222a2 · outbound

This paper cites iquery: Instruments as queries for audio-visual sound separation.

Learning to Highlight Audio by Watching Movies iquery: Instruments as queries for audio-visual sound separation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.614154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.526501Z digest=sha256:73cc9310fe613ec068ed6d4a14b51eb806ebb8dbf547ef8f401dd394efe59511

Observation 0734e2fc-d2f9-4d0e-aff6-241ae8fb73b1 · outbound

This paper cites Deep cross-modal audio-visual generation.

Learning to Highlight Audio by Watching Movies Deep cross-modal audio-visual generation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.602687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.530288Z digest=sha256:bae3d11432412811ea095c611178e86553b17b67ea2ec5727ae5508a87e6e38f

Observation b6a8e257-6df4-4329-ad6b-3aac5c0fbaf0 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Learning to Highlight Audio by Watching Movies How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:43:56.533483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:43:56.533483Z digest=sha256:222527131b347c683da22d53be14073c4b1aaff031881d6bead7d5b0331cca23

Observation f7e39f33-5a47-464d-96c6-546ecd93302e · outbound

This paper cites Hybrid spectrogram and waveform source separation.

Learning to Highlight Audio by Watching Movies Hybrid spectrogram and waveform source separation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.592847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.537629Z digest=sha256:e796b8c060ded55f02fef09cf68f1f2f409be5232b99a9c87afd4673af3117d0

Observation 707a6090-138d-412f-a9a4-d3a40e48138b · outbound

This paper cites an unresolved cited work.

Learning to Highlight Audio by Watching Movies Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:43:57.582419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.540935Z digest=sha256:32277f773cfb856d980e2a4fafbc1f4c7a2af760ff6b2313909e884cabf4232e

Observation 69ff8f18-e499-423c-aef9-456a83a10399 · outbound

This paper cites 2.5 d visual sound.

Learning to Highlight Audio by Watching Movies 2.5 d visual sound

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.571789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.544175Z digest=sha256:b504d3528112a1c4224743d4a0eac843d4052d4d7044099f361a2a9bfa341c89

Observation b6e9256d-c195-4a72-9fc8-23af8cf15207 · outbound

This paper cites Visualvoice: Audio- visual speech separation with cross-modal consistency.

Learning to Highlight Audio by Watching Movies Visualvoice: Audio- visual speech separation with cross-modal consistency

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.560090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.548256Z digest=sha256:93a4fb53f6b2260bd6db6f939fc19715e18f49bfa5af4f600600cd6b91de0582

Observation e6c7b969-1af2-45d8-b1eb-f5019b8a7e45 · outbound

This paper cites Learning to separate object sounds by watching unlabeled video.

Learning to Highlight Audio by Watching Movies Learning to separate object sounds by watching unlabeled video

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.549047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.551644Z digest=sha256:ed0785f270772f1ed408cec8f22b2e86934e3fe84ed27c80beae0c9a02aa22da

Observation f89806ec-22ee-4f94-a61a-9446c289d34e · outbound

This paper cites Visualechoes: Spatial visual represen- tation learning through echolocation.

Learning to Highlight Audio by Watching Movies Visualechoes: Spatial visual represen- tation learning through echolocation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.537485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.555313Z digest=sha256:046e2e81d69dc67fe5d24eef76f58cbd8dbe78ae6dd70ccd8f413c4e7bf0aee7

Observation c6f0c507-2f36-4c9d-9462-e1e5352bf898 · outbound

This paper cites Geometry- aware multi-task learning for binaural audio generation from video.

Learning to Highlight Audio by Watching Movies Geometry- aware multi-task learning for binaural audio generation from video

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.527645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.559422Z digest=sha256:ba1c9b4e6abf5ad6f7bbb726f9c1d53d363ea0907b96cf6a8c236e2bd75eeef7

Observation dbca678e-ff69-478e-a506-bcbb0b3a19da · outbound

This paper cites Imagebind: One embedding space to bind them all.

Learning to Highlight Audio by Watching Movies Imagebind: One embedding space to bind them all

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.517094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.562285Z digest=sha256:64f0f2f1f7a41e598a01ab2def5bc511a0e28f7b3d6949b4caf7f6a4a6902328

Observation 724abbb1-111d-4851-a1d4-656b623be8b5 · outbound

This paper cites Semiautomatic visual-attention modeling and its application to video compression.

Learning to Highlight Audio by Watching Movies Semiautomatic visual-attention modeling and its application to video compression

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.504850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.565249Z digest=sha256:06e970dbccbf1bc7e2ff552a1a20d1a9799062070bac865c309a60c22af1bc6b

Observation 6f000980-fe6e-4f15-903f-a6b14612f71e · outbound

This paper cites Graph- based visual saliency.Advances in neural information pro- cessing systems, 19, 2006.

Learning to Highlight Audio by Watching Movies Graph- based visual saliency.Advances in neural information pro- cessing systems, 19, 2006

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.492813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.568846Z digest=sha256:c55a720746c13a6a4ac30726fe50d2cc9c9e3667576cce4ac44faf1755f39008

Observation d41504b6-8ad8-4cea-9780-c53e07f9cb2b · outbound

This paper cites Discriminative sounding objects localization via self-supervised audiovisual matching.

Learning to Highlight Audio by Watching Movies Discriminative sounding objects localization via self-supervised audiovisual matching

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.482793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.571693Z digest=sha256:ac61af8d5eb965a60e6b4ad01ebff780228cfac40fc1d27155ae46e4d6f2f5bc

Observation 7be17340-7ffd-4a40-82f7-391657c0bb44 · outbound

This paper cites High-Quality Visually-Guided Sound Separation from Diverse Categories.

Learning to Highlight Audio by Watching Movies High-Quality Visually-Guided Sound Separation from Diverse Categories

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:43:56.574566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:43:56.574566Z digest=sha256:b759cc15ae3aa85a14997c305d0dd10315bdd377eb99e0718014c75c51941a6d

Observation 7d1955e9-c3f1-402e-a6a4-3f8493bd6dbb · outbound

This paper cites Egocentric Audio-Visual Object Localization.

Learning to Highlight Audio by Watching Movies Egocentric Audio-Visual Object Localization

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:43:56.945592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.578357Z digest=sha256:4548169f9ed4b2c9dc5bd50e52ab08a11a43a0c83c348db40f7a93dea50b58e5

Observation 85336f2f-9a80-42a8-9ea3-f6db42de32cf · outbound

This paper cites Scaling Concept With Text-Guided Diffusion Models.

Learning to Highlight Audio by Watching Movies Scaling Concept With Text-Guided Diffusion Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:43:56.582004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:43:56.582004Z digest=sha256:5b5d2751a9c352c8eca79110e8edb943fdb6ab228013f6b2d41bdf1d77831ed8

Observation 741c074a-9015-4618-ac23-770861379861 · outbound

This paper cites Modeling and Driving Human Body Soundfields through Acoustic Primitives.

Learning to Highlight Audio by Watching Movies Modeling and Driving Human Body Soundfields through Acoustic Primitives

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:43:56.585883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:43:56.585883Z digest=sha256:8208abc79ae0474a279e5b0d5e96c843d2d1959d3489427464b4e9cf9c31c43b

Observation 23ea06e0-87ed-482c-9860-65a8fd1bd170 · outbound

This paper cites FreSca: Scaling in Frequency Space Enhances Diffusion Models.

Learning to Highlight Audio by Watching Movies FreSca: Scaling in Frequency Space Enhances Diffusion Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:43:56.589829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:43:56.589829Z digest=sha256:161b189217e99bc12735b9fa33087982ffb263187b579905172a4cd4353b0d77

Observation fbda7ce5-5822-44b2-ac39-1928ac38e497 · outbound

This paper cites A model of saliency-based visual attention for rapid scene analysis.IEEE Transactions on pattern analysis and machine intelligence, 20(11):1254–1259, 1998.

Learning to Highlight Audio by Watching Movies A model of saliency-based visual attention for rapid scene analysis.IEEE Transactions on pattern analysis and machine intelligence, 20(11):1254–1259, 1998

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.470855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.592902Z digest=sha256:7dcf95e0cd9d5f79abbcad2e9728476b2ceedb54b5a74b7018a782db2350875c

Observation 94133073-34b5-48c0-afb4-5977fef7089c · outbound

This paper cites Vinet: Pushing the limits of visual modality for audio-visual saliency prediction.

Learning to Highlight Audio by Watching Movies Vinet: Pushing the limits of visual modality for audio-visual saliency prediction

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.459500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.596601Z digest=sha256:bd9fddfeb4cbdd232426a37054f4254392565086ab4106a9fb3a21e9143e4adf

Observation 7dd20d2a-4a2d-4c1f-b53e-fe320916bd52 · outbound

This paper cites Deepvs: A deep learning based video saliency prediction approach.

Learning to Highlight Audio by Watching Movies Deepvs: A deep learning based video saliency prediction approach

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.449445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.600611Z digest=sha256:81370a04d075c8ec5aa5b546fa415df20699589a713973dbe50fd4765c020a93

Observation 80a96d51-b988-49a0-822d-3c9664759b63 · outbound

This paper cites Listen, Chat, and Remix: Text-Guided Soundscape Remixing for Enhanced Auditory Experience.

Learning to Highlight Audio by Watching Movies Listen, Chat, and Remix: Text-Guided Soundscape Remixing for Enhanced Auditory Experience

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:43:56.604651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:43:56.604651Z digest=sha256:4fa654819e9424c7a49a77e3f2c65436385bfd5e6f82216fdabf0086f2d9d2d8

Observation 23f984b5-2772-4dd2-ab09-1d2effee0c11 · outbound

This paper cites Mu- sic mixing style transfer: A contrastive learning approach to disentangle audio effects.

Learning to Highlight Audio by Watching Movies Mu- sic mixing style transfer: A contrastive learning approach to disentangle audio effects

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.438056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.609564Z digest=sha256:6ccf191e628aba4477bed185c1fa27ed5c6a83f6b4aa5936a61f9cd3121d2248

Observation 3287e27c-0d82-490f-ad63-dd5acad755c3 · outbound

This paper cites Efficient training of audio transformers with patchout.

Learning to Highlight Audio by Watching Movies Efficient training of audio transformers with patchout

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.427306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.612796Z digest=sha256:1b0341dc0e4ee50c0943c60b3f00790bedb033f692028f2ec2d6b050dee8e947

Observation 07bba70a-f0b2-4f80-ba63-7abe270ccec2 · outbound

This paper cites Contextual encoder–decoder network for visual saliency prediction.Neural Networks, 129:261–270, 2020.

Learning to Highlight Audio by Watching Movies Contextual encoder–decoder network for visual saliency prediction.Neural Networks, 129:261–270, 2020

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.415971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.615732Z digest=sha256:e40627fdfbf5bb67b986a3990dd8357c64fc1f5e10333606f875f3e1185bfe93

Observation edf28029-4bb3-484f-aa78-11f2af63b8c4 · outbound

This paper cites Detecting mo- ments and highlights in videos via natural language queries.

Learning to Highlight Audio by Watching Movies Detecting mo- ments and highlights in videos via natural language queries

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:43:56.618846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:43:56.618846Z digest=sha256:b664c971d7f50e71db6fc1b02af126c2b9166ebbbc2992f1755a2394e4cdfb33

Observation 9c41f0c1-cf73-4cad-a3cd-bdd9bdd1b52b · outbound

This paper cites Neural Acoustic Context Field: Rendering Realistic Room Impulse Response With Neural Fields.

Learning to Highlight Audio by Watching Movies Neural Acoustic Context Field: Rendering Realistic Room Impulse Response With Neural Fields

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:43:56.621828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:43:56.621828Z digest=sha256:439dc81ba0b11aca71cc36d91317023fba3f1a9e67fdbdb165a129772078a12b

Observation 8534be2d-cf15-4240-b9f2-50dca2d259b8 · outbound

This paper cites Av-nerf: Learning neural fields for real- world audio-visual scene synthesis.

Learning to Highlight Audio by Watching Movies Av-nerf: Learning neural fields for real- world audio-visual scene synthesis

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.398688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.625686Z digest=sha256:055fa11cb6112c059b782bc7425c44bb5db62b0a2803ebcbc035b59758d505c6

Observation fbd16d28-3950-4686-ad02-2abd546af6c4 · outbound

This paper cites Language-guided joint audio-visual editing via one-shot adaptation.

Learning to Highlight Audio by Watching Movies Language-guided joint audio-visual editing via one-shot adaptation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.387931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.629422Z digest=sha256:517aea407be648e319d916fdc85642390ae0fc7f65543c42cc71017fd195d7e0

Observation 0d561870-2553-4441-baea-3fbe1ab9aa2c · outbound

This paper cites Univtg: Towards unified video-language temporal grounding.

Learning to Highlight Audio by Watching Movies Univtg: Towards unified video-language temporal grounding

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.377122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.632947Z digest=sha256:ec5e82541214b9f0ec9f8e6ec2ab2a52e75aa0da2fd9d8f4ca889db103641af6

Observation c8ac7ba1-5261-47cc-a9c6-a888539a8327 · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

Learning to Highlight Audio by Watching Movies AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:43:56.636913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:43:56.636913Z digest=sha256:3c5c22d4f79634661d68fedfbf5fe0d502f26dd54a695bbce6420c47b835f230

Observation 49cbf520-a9a8-43d5-96ec-777ff50b9ea6 · outbound

This paper cites Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection.

Learning to Highlight Audio by Watching Movies Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.365652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.640283Z digest=sha256:2207a495e55e98434651fa14e2a3abe6e6f002640110e7e0797d8139d1b72bac

Observation 0d3900ed-99b6-4b67-88da-b3530d474552 · outbound

This paper cites Play as you like: Timbre-enhanced multi-modal music style transfer.

Learning to Highlight Audio by Watching Movies Play as you like: Timbre-enhanced multi-modal music style transfer

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.352953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.643348Z digest=sha256:72a6543f3292eca24743b876763c59635007979fd689e4d31461a335bab062a7

Observation 431016d6-3d3d-42da-823d-2e78a30ad435 · outbound

This paper cites Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation.

Learning to Highlight Audio by Watching Movies Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.342275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.646793Z digest=sha256:13506953fbce0a83b5b9a97c71853745748d52b0744f7cdf177e9bfd1a94fe38

Observation bdab03a7-f680-4d18-96ef-6e955dc90fe6 · outbound

This paper cites Deep learning for black-box modeling of audio effects.Applied Sciences, 10(2):638, 2020.

Learning to Highlight Audio by Watching Movies Deep learning for black-box modeling of audio effects.Applied Sciences, 10(2):638, 2020

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.331961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.649716Z digest=sha256:cb07e070874ad903bf9a7e3535b8331af844a0185f72b21843c9ed8762065b59

Observation 1c1838d5-dfd6-4d6b-ad7f-76553942cffe · outbound

This paper cites Localizing Visual Sounds the Easy Way.

Learning to Highlight Audio by Watching Movies Localizing Visual Sounds the Easy Way

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:43:56.875627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.652693Z digest=sha256:be71070d3c48a766424a2a5c1e8cef45555b1db6748b6413800d9a361d82f968

Observation d1c98d3c-13eb-4573-993a-3fad69e861ae · outbound

This paper cites Query-dependent video representation for moment retrieval and highlight detection.

Learning to Highlight Audio by Watching Movies Query-dependent video representation for moment retrieval and highlight detection

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.321856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.656659Z digest=sha256:19b40d8b75eb9e420175b35114402baa3f22562288c2f5205a2f532220bf8753

Observation d1933679-1a7b-47be-901e-1fbf8c8d2b7a · outbound

This paper cites Self-supervised generation of spatial audio for 360 video.Advances in neural information processing systems, 2018.

Learning to Highlight Audio by Watching Movies Self-supervised generation of spatial audio for 360 video.Advances in neural information processing systems, 2018

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.311931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.659796Z digest=sha256:f367058752821229fe2d5337ca67b238f8af20ee963cd1ba429e939103ac9c95

Observation 5a9f839e-4a35-4ebc-a2f1-6429a22aecb7 · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

Learning to Highlight Audio by Watching Movies Audio-visual scene analysis with self-supervised multisensory features

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:43:56.662916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:43:56.662916Z digest=sha256:469e98317f4163add8c6e612c3f53fd5e55bd137618dbe3e1df5d1bff4da44e1

Observation e328d08e-43df-4140-8eb0-3c65150d10b9 · outbound

This paper cites Visually indicated sounds.

Learning to Highlight Audio by Watching Movies Visually indicated sounds

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.295557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.666206Z digest=sha256:02eb27293d872a8468fdc25eac9016547599556a5f6a4eae4a40de454d6e2073

Observation 168c549d-fd24-4996-8c6c-170a804288db · outbound

This paper cites Physical Modeling using Recurrent Neural Networks with Fast Convolutional Layers.

Learning to Highlight Audio by Watching Movies Physical Modeling using Recurrent Neural Networks with Fast Convolutional Layers

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:43:56.669576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:43:56.669576Z digest=sha256:4157b95c2d196effe3ed4cde043b3a5c6c97aa26b848957b77088e0afd0ee2a0

Observation 07c0f6e1-903b-4a55-a0d8-89dfa53deb06 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Learning to Highlight Audio by Watching Movies Movie Gen: A Cast of Media Foundation Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:43:56.672711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:43:56.672711Z digest=sha256:d750da01ddb687bb34231237441114914a1ff2706fa3c9f00b2f1b5c2cb0a84d

Observation 1d8d05b0-b0d6-4ce7-986f-a076a02267b5 · outbound

This paper cites Multiple sound sources localization from coarse to fine.

Learning to Highlight Audio by Watching Movies Multiple sound sources localization from coarse to fine

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.284465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.676491Z digest=sha256:fbd0911e8bff0c84e331c60fcfe8c37386b876583386adca9a50b283380190e1

Observation 063563da-5171-437a-9c77-cc836439d59f · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Learning to Highlight Audio by Watching Movies Learning transferable visual models from natural language supervi- sion

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:43:56.680364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:43:56.680364Z digest=sha256:43a7938298368df152aa81b5cbc6c50f88c0c013509705c65cd7ec0baf05137b

Observation f05f1a3c-6f88-40ab-8328-6fef90a5d6f4 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67, 2020.

Learning to Highlight Audio by Watching Movies Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67, 2020

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T20:43:56.683920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:43:56.683920Z digest=sha256:e48804b6d0071345668adb671d69219e9083f03eed872dc6b641ff7b111a88cf

Observation 7638f41c-ab81-432f-b841-a17ed6e74334 · outbound

This paper cites Modeling nonlinear audio effects with end-to-end deep neural networks.

Learning to Highlight Audio by Watching Movies Modeling nonlinear audio effects with end-to-end deep neural networks

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.261368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.687835Z digest=sha256:e7d932b1622381b62e063a049f1ba93f97b3e732edddfc443e9e929eebe764e6

Observation f9226e98-9c7a-418b-b13f-196d4a7b7510 · outbound

This paper cites Dynamic storyboard generation in an engine-based virtual environment for video production.

Learning to Highlight Audio by Watching Movies Dynamic storyboard generation in an engine-based virtual environment for video production

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.250108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.691332Z digest=sha256:c17a7b86cffcb57d4a844a5395681ba51001702725ad5cb0dd677f28bc9d6c8b

Observation 06ffbbbc-e9f2-4ab6-8bd5-0fce5b7d3964 · outbound

This paper cites U- net: Convolutional networks for biomedical image segmenta- tion.

Learning to Highlight Audio by Watching Movies U- net: Convolutional networks for biomedical image segmenta- tion

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.237273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.694435Z digest=sha256:770d92e42ecaa6c55cc4faf2dbea699888ce2ff23f14a496d6b7af80f0dfb4c9

Observation c93b2b5f-62e9-4ed8-b979-29b9b4f3d239 · outbound

This paper cites Separate and reconstruct: Asymmetric encoder- decoder for speech separation.

Learning to Highlight Audio by Watching Movies Separate and reconstruct: Asymmetric encoder- decoder for speech separation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.226711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.698015Z digest=sha256:e154d45d67d5d2f3cb6a9c2353cc2d066baf59ac4018587c34012d45041af104

Observation 6f206bb7-7f26-4109-bee8-fabfb42f6dfa · outbound

This paper cites Freeu: Free lunch in diffusion u-net.

Learning to Highlight Audio by Watching Movies Freeu: Free lunch in diffusion u-net

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:43:56.701754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:43:56.701754Z digest=sha256:5a47d1400faba7fd73e164f492e61c7af8f5888de971efb99c860aef8908088c

Observation 36502121-6b6d-4db9-ab22-6d7caab824fc · outbound

This paper cites Steinmetz and Joshua D.

Learning to Highlight Audio by Watching Movies Steinmetz and Joshua D

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.208937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.705407Z digest=sha256:65815ee201ae2681d3332c553c38d22c7db624c614311f5f08035544a9178d58

Observation 26208ba6-5455-4d66-8fd0-39b920f92be0 · outbound

This paper cites Style Transfer of Audio Effects with Differentiable Signal Processing.

Learning to Highlight Audio by Watching Movies Style Transfer of Audio Effects with Differentiable Signal Processing

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T20:43:56.709150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:43:56.709150Z digest=sha256:c78ff76a461ffa298296e70c296d83df61b2d4612fea9f151c456a2ee088279f

Observation 0350546e-c7a0-4146-bd0a-8a9dbf807237 · outbound

This paper cites Attention is all you need in speech separation.

Learning to Highlight Audio by Watching Movies Attention is all you need in speech separation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.197769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.712576Z digest=sha256:df1874c38042038cc09efc2da4bdc7b025543b6dd733a91e1e62afe4b80aab5f

Observation 21aec616-cc3f-4502-87a9-638554ad89f3 · outbound

This paper cites Audio-visual event localization in unconstrained videos.

Learning to Highlight Audio by Watching Movies Audio-visual event localization in unconstrained videos

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.187126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.716308Z digest=sha256:560d4bd8f783eefc3d409894da8e54603d3d407fc60f07ba2568921ac2b637b7

Observation 956420ac-a5d3-43de-be8a-05a7a83109b2 · outbound

This paper cites Unified mul- tisensory perception: Weakly-supervised audio-visual video parsing.

Learning to Highlight Audio by Watching Movies Unified mul- tisensory perception: Weakly-supervised audio-visual video parsing

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.177004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.720114Z digest=sha256:5cd28b45d4a8a17aa981303448a8c545de3b915cd005bbb9012ddc2870b5fcb4

Observation b1998293-151d-4492-a605-d5e121ae2904 · outbound

This paper cites Cyclic co-learning of sounding object visual grounding and sound separation.

Learning to Highlight Audio by Watching Movies Cyclic co-learning of sounding object visual grounding and sound separation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.166104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.723508Z digest=sha256:5da4351b114bdbc1ddbfcd28d007a5b1aa38926e929ad2dd317d19477ba82d98

Observation 240678ba-d242-42be-a186-a6b14e95e1c3 · outbound

This paper cites Diff-MST: Differentiable Mixing Style Transfer.

Learning to Highlight Audio by Watching Movies Diff-MST: Differentiable Mixing Style Transfer

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:43:56.834097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.726488Z digest=sha256:b14d3fec4e3e97124f0623b90ab4a69315440ed0412edfd2fcdc666eb1e1528f

Observation 73a6042c-3a3b-4f8e-9f2b-a6bfe96d4d65 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

Learning to Highlight Audio by Watching Movies Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T20:43:56.730758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:43:56.730758Z digest=sha256:d68b967a4ea44196269d26e0e4c2b9fe02be8547e4b957319c129e76c8f7b85e

Observation e82070b1-dc75-47cd-8e42-a328bcf9458e · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

Learning to Highlight Audio by Watching Movies Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T20:43:56.733736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:43:56.733736Z digest=sha256:b452ac9d62640e4ab6cd851d96771337a02ab3b5b49c624097426450de6ec622

Observation 588209f5-3e04-4bd8-969f-c3d9dbd5a46a · outbound

This paper cites Lave: Llm-powered agent assistance and lan- guage augmentation for video editing.

Learning to Highlight Audio by Watching Movies Lave: Llm-powered agent assistance and lan- guage augmentation for video editing

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.148137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.736927Z digest=sha256:b753846a2b38d73708fac7e7a674d7ade5fa8307ab1c44b932004fcef4c76ec5

Observation 2c7b2960-a070-4896-873e-d5ce0e1b6b3c · outbound

This paper cites Re- mastering divide and remaster: A cinematic audio source separation dataset with multilingual support.

Learning to Highlight Audio by Watching Movies Re- mastering divide and remaster: A cinematic audio source separation dataset with multilingual support

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.138080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.739997Z digest=sha256:8a13545a599e7fd6de519acbee7a85bcb06e9e13b1ed69784d26cda0aa671346

Observation 5f9369f9-e204-4661-b1de-1db0971b3ade · outbound

This paper cites Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners.

Learning to Highlight Audio by Watching Movies Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.126835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.742881Z digest=sha256:abcddb76652d47fe1c07759816a0b4686f989825313660f8c913e059d3045a25

Observation 39e53fed-cf3b-47f2-bb7d-3ea808ac5f13 · outbound

This paper cites Visually informed binaural audio generation without binaural audios.

Learning to Highlight Audio by Watching Movies Visually informed binaural audio generation without binaural audios

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.116273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.746460Z digest=sha256:95c556f1c1fac5ac0daa39985a5f0ca38180d40cb3fad75575e6efa6c5e50101

Observation ba62b237-6e36-4e89-a921-933173a54bc0 · outbound

This paper cites Parallel waveform synthesis based on generative adversarial networks with voicing-aware conditional discrimi- nators.

Learning to Highlight Audio by Watching Movies Parallel waveform synthesis based on generative adversarial networks with voicing-aware conditional discrimi- nators

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.103980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.750477Z digest=sha256:460d2126055a46f83ba5a5e8b771297185f353339edc233b1d01d9f530a3b577

Observation 58100920-6946-4136-8f85-3a21f59bbebb · outbound

This paper cites Don’t separate, learn to remix: End-to-end neural remixing with joint optimization.

Learning to Highlight Audio by Watching Movies Don’t separate, learn to remix: End-to-end neural remixing with joint optimization

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.091200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.754003Z digest=sha256:05724d729c9cd84caba3b39d05da28e34aa3e2e64c3f2950ebde31e79a7e389d

Observation 1fdaf460-d2cd-4561-816d-67b5374b9ab1 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Learning to Highlight Audio by Watching Movies Adding conditional control to text-to-image diffusion models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T20:43:56.757729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:43:56.757729Z digest=sha256:854ce44aac251e627592837d0c9a5f92132d876439070cb4c2ad6d39e22cbc17

Observation 76347d0f-2d62-4dd0-bf79-94fbd0648e83 · outbound

This paper cites The sound of pixels.

Learning to Highlight Audio by Watching Movies The sound of pixels

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.075481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.761397Z digest=sha256:aed002e424e7dcedc15b32e210775557e176f85f1bd9328ffa6963a52007b533

Observation 415e7205-2712-4e0b-82de-4fba7e2ef488 · outbound

This paper cites github.io/VisAH/) to illustrate our method and show- case our results.We strongly encourage readers to visit this webpage and use headphones.

Learning to Highlight Audio by Watching Movies github.io/VisAH/) to illustrate our method and show- case our results.We strongly encourage readers to visit this webpage and use headphones

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.064682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.764301Z digest=sha256:21c96a8cc528234642f60fe27c1df2673abc2b5ed3a9fb5d14967644d9307b65

Observation 0a7cf33a-e181-407f-b6ec-d94281377909 · outbound

This paper cites Here, we provide case studies to illustrate the conditions under which such failures occur.

Learning to Highlight Audio by Watching Movies Here, we provide case studies to illustrate the conditions under which such failures occur

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.054799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.768303Z digest=sha256:d4928606ff92b3d0ac50ebcbdd7249d01655ffb76642cbb2d5e4997f2a682250

Observation 94548f8d-8eca-4427-99a0-eee2e9a3e3da · outbound

This paper cites an unresolved cited work.

Learning to Highlight Audio by Watching Movies Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:43:57.044369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.771662Z digest=sha256:05732497116df83063993fdce0e25aafcb18e8d5b4deecd52170eb09a1be887f

Observation cabe2c01-cf69-4541-9555-05bdb35bd001 · outbound

This paper cites As illustrated in Fig.

Learning to Highlight Audio by Watching Movies As illustrated in Fig

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.032536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.774632Z digest=sha256:b2c6f9466834e4b67fc933f70f21561746d85003df282dc189385b6ca3a39671

Observation 2ef2ea07-7f12-4c1f-98b1-4422d3d36c6f · outbound

This paper cites The MR-STFT loss is implemented by computing theℓ1 distance between the am- plitude spectrograms of the predicted signalˆsand the ground truth signals.

Learning to Highlight Audio by Watching Movies The MR-STFT loss is implemented by computing theℓ1 distance between the am- plitude spectrograms of the predicted signalˆsand the ground truth signals

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.021947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.777714Z digest=sha256:ea00732425841aabf30c077f93bfbc926afe033c6f839c1a78b011b09359084c

Observation c3c7a137-52bd-4382-ae78-2c65293b09dc · outbound

This paper cites a dark, elegant outfit.

Learning to Highlight Audio by Watching Movies a dark, elegant outfit

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:57.010808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.781104Z digest=sha256:6890784a995e592b9534b73bca8e0706be23b1981a03ef6f5ba641c72b2bd300

Observation f4be45b9-2dbe-4033-8b72-8c5670f5ae98 · outbound

This paper cites While our method requires more time, it remains efficient for practical applications.

Learning to Highlight Audio by Watching Movies While our method requires more time, it remains efficient for practical applications

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:56.999109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.784741Z digest=sha256:f258bd2600322a3c057cd5c83e85afc0ee156d1407ea252fb92a5c374779e18a

Observation ad2d3e47-8658-4e39-8e88-954984990296 · outbound

This paper cites 12 across differ- ent levels of dataset difficulty, as discussed in Sec 5.3.2 and shown in Tab.

Learning to Highlight Audio by Watching Movies 12 across differ- ent levels of dataset difficulty, as discussed in Sec 5.3.2 and shown in Tab

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:43:56.988438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.788803Z digest=sha256:1f7f02cd1f3434052bbb5e132229df749ff971b57f29e6716c836aba91f2270c

Observation 61a0208a-99d0-4bc1-984b-f29713def338 · outbound

This paper cites an unresolved cited work.

Learning to Highlight Audio by Watching Movies Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:43:56.977948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:43:56.792270Z digest=sha256:b81d6085a1ccf146fa27168bdbc8b66dfc1f50755d89fc04207a724f49da09e8

Pith citing papers

No inbound Pith citation observations are available.