REVIEW 2 major objections 5 minor 102 references
The Sound of Water: Inferring Physical Properties from Pouring Liquids
T0 review · 2 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read From the sound of pouring alone, the paper recovers the air-column length, container height and radius, flow rate, and time to fill, by tracking the fundamental axial resonance whose wavelength is linear in the air-column length.
desk verdict A genuinely novel physics-grounded pipeline for pouring analysis, with a real but fixable circularity issue in the headline dynamic metric. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the axial-resonance wavelength identity $\lambda(t)=4(l(t)+\beta R)$, with $\beta=0.62$ fixed; it converts pitch into a linear metric for the air column. Around this sit four derived formulas: $l(t)=(\lambda(t)-\lambda(T))/4$, $H=(\lambda(0)-\lambda(T))/4$, $R=\lambda(T)/(4\beta)$, $Q(t)=-(\pi R^2/4)\,d\lambda/dt$, plus the early-pour approximation $\tau(t)\approx-\lambda(t)/(d\lambda/dt)$ for time to fill. The detector is an audio transformer trained in two stages: synthetic pre-training on simulated pours, then visual co-supervision on real pours, where a scale factor $\alpha$ links metric wavelengths to pixel air-column lengths and radii. The wavelength curve carries the entire argument, because every physical property is a boundary value, an intercept, or a slope of that curve.
What would settle it
Pour into a transparent cylinder with a ruler beside it at a known constant rate, record audio and video, extract the fundamental frequency from the spectrogram, and compare $l(t)$ from Eq. (5) against the visually tracked water level; if the discrepancy grows with $R$ or with flow rate, or if the apparent pitch is not single-valued, the fixed-$\beta$ linear relation fails.
Extended reading notes
Core claim
The central claim is that the pitch of pouring water is the fundamental axial resonance of a pipe closed at one end, with a fixed end correction: $\lambda(t)=4(l(t)+0.62R)$, where $\lambda$ is the wavelength of the fundamental, $l(t)$ the air-column length, and $R$ the container radius. If this identity holds over the pour, then the wavelength curve is enough to read off $l(t)$ at every instant, $H$ and $R$ from the boundary at start and end, $Q(t)$ from its slope, and the time to fill from its early behavior. The paper supports the claim by building a transformer-based pitch detector that outputs a wavelength distribution per time step, pre-training it on synthetic pours generated with a differentiable synthesizer, and then fine-tuning it on real videos using the video stream as a weak teacher through the scale-aware equation $\alpha\lambda(t)/4 = l_{\text{px}}(t)+\beta R_{\text{px}}$. Tested on a new dataset of 805 real pouring videos, the co-supervised model outperforms classical and learned pitch estimators and estimates physical properties with the errors reported above, while its features also support container-shape classification and liquid-mass regression on a previous dataset.
Load-bearing premise
The load-bearing premise is that real pouring audio contains one clean, dominant axial-resonance pitch that follows $\lambda(t)=4(l(t)+0.62R)$ at every instant, so every later measurement is a boundary value or slope of that curve; the paper itself shows hemispherical containers and some bottleneck containers violate this.
Editorial extensions
If this is right
- A single smartphone recording of a pour into a cylinder-like container yields absolute metric estimates of the container and the pour with no manual measurement, with reported mean absolute errors of 0.60 cm in air-column length, 2.27 cm in height, 1.39 cm in radius, and 22.5 ml/s in flow rate on Test set I.
- Visual co-supervision improves estimates most near the end of the pour, where the audio signal is weak, and that is exactly where radius and flow-rate errors are dominated by $\lambda(T)$ and the slope of $\lambda$, so co-supervision translates into better static and dynamic properties.
- Pitch-detection features encode shape: an unseen three-way shape classification reaches 90.91% sample accuracy and 92.47% mean class accuracy, and linear probing on a prior pouring dataset gives 1.20 oz mean absolute error for liquid mass.
- The detector generalizes beyond cylinders to semi-conical, bottleneck, cup, teapot, and wine-glass containers, across glass, plastic, steel, ceramic, and cardboard, and to in-the-wild YouTube pours, although hemispherical containers and multi-modal bottleneck resonances are acknowledged failure cases.
Reading between the lines
- The paper leaves implicit that the same pitch-to-wavelength pipeline could serve as a contact-free calibration device: if the container is known, the formula gives flow rate and total poured volume, and if the flow rate is known, the same audio gives the container's dimensions; a direct test would compare audio-inferred volume against a scale.
- A natural extension is multi-pitch tracking, since the failure on bottleneck containers with two simultaneous frequency modes suggests treating axial and radial resonances as separate tracks, or using a harmonic-aware architecture, rather than a single fundamental wavelength.
- Because the relation is stated with a constant $\beta=0.62$, the approach makes a testable prediction about how inferred size errors should scale with container radius; measuring that scaling would indicate whether the fixed end correction is an adequate approximation across the dataset's size range.
- The result suggests that the metric-ruler claim may extend to other container-filling sounds, such as grains or viscous liquids, as long as an air column with a single dominant resonance exists; this is an extrapolation the paper does not make.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper claims that, from only audio of liquid pouring, physical properties of the container-liquid system can be recovered: air-column length l(t), container height H, radius R, volume flow rate Q(t), and time-to-fill τ. The authors derive these properties from the axial-resonance relation λ(t)=4(l(t)+βR) in Section 3, train a wav2vec2-based wavelength-prediction network using simulated pouring sounds and visual co-supervision from a DINO-based video teacher in Section 4, introduce a new 805-video dataset in Section 5, and report quantitative results for these properties, plus shape classification and liquid-mass estimation on an external dataset in Section 6. The central theoretical derivation is mathematically straightforward, and the main empirical concern is the provenance of the air-column ground truth used for the headline result.
Significance. If the empirical claims are taken at face value, the paper makes a meaningful advance: it gives a simple and largely parameter-free physical mapping from pitch to metric properties, demonstrates that a learned pitch detector can outperform classical pitch trackers by a large margin, and introduces a valuable dataset for audio-visual physical inference. The independent manual-ruler evaluation of H and R in Table 3 and the external-dataset mass-estimation results are concrete strengths, and the failure cases in Section 6.4 are honestly disclosed. However, the strongest dynamic-property claim (0.60 cm air-column MAE) currently lacks a demonstrated independent ground-truth source, which is essential before the 'audio as metric ruler' and 'human-like capabilities' claims can be accepted.
major comments (2)
- [Section 6.1, Table 2; Appendix A.3; Eq. (10)] The manuscript does not state the source of the air-column ground truth used in Table 2. The co-supervised audio model is fine-tuned with the MSE objective in Eq. (10), where the target is the video network's prediction l_px(t)+βR_px, and the video network is trained on pseudo-labels obtained from temporal-difference heatmaps with a RANSAC polynomial fit (Appendix A.3). If Table 2 evaluates against the same pseudo-label pipeline, the reported 0.60 cm error is not a measurement against physical liquid level but an agreement score with the video teacher, and the apparent gain of co-supervision over the audio-only variant (0.60 vs 0.78 cm) could be largely an artifact of fitting that teacher. This is load-bearing because the 'audio is effectively a metric ruler' claim (Section 3.1) and the claimed 'human-like capabilities' (Section 1) rest on this number. The manual ruler measurements for H and R in Table 3 provide independent support for the static-property chain, but they do not validate the dynamic l(t) curve. Please specify the ground-truth source and, if it is the pseudo-label pipeline, add independent manual or sensor-based liquid-level annotations for at least a subset of videos, reporting both pseudo-label-based and independent errors.
- [Section 4.3, Eq. (10)-(11); Appendix A.3] The per-video scale factor α is estimated from the ratio between the audio network's own wavelength predictions and the video network's pixel measurements, weighted by RMS energy, and is then fixed while the same audio network is fine-tuned toward Eq. (10). This makes the co-supervision loop partially self-referential: systematic errors in the pre-fine-tune audio model can be absorbed into α and are then not penalized by the objective. The verification in Appendix A.3 (α in [30,80], inverse relation with container size) is only a sanity check. Please provide a sensitivity analysis, e.g., compute α from ground-truth wavelengths on a subset of videos or from an independent metric-to-pixel calibration, and show how the final property errors in Table 3 change. This matters because α is the bridge that converts audio wavelengths into metric quantities used during training.
minor comments (5)
- [Section 5, Table 1] The split arithmetic is unclear: Table 1 reports 18 train containers/195 videos, 13 Test I containers/54 videos, 19 Test II containers/327 videos, and 25 Test III containers/434 videos, while the text says the totals are 18 + 25 = 43 containers and 195 + 54 + 434 = 683 videos, omitting Test II and the overlap between Test II and Test III; please clarify the unique-container counts and how the 122 remaining videos are defined.
- [Section 6.4, Figure 15] The caption lists '(a) Hemispherical container (cup)' and '(b) Bottle-neck container', but the body text describes (a) as a bottleneck case and (b) as a hemispherical case; the caption and text should be made consistent.
- [Table 3] The column header 'Synthetic ↓' is confusing because the text refers to this model as 'audio-only'; using one consistent name would improve readability.
- [Section 6.1, Table 3] The rows for time-to-fill use the notation 'τ 1 4 (t)', 'τ 1 2 (t)', and 'τ 3 4 (t)'; please define this notation explicitly, since it is not obvious that these denote the fraction of the original audio given to the model.
- [Appendix A.3] The verification that empirical scale factors are in [30,80] relies on 'generic values' of f, s, and Z, but those values are not stated; please provide the assumed values and units so the range can be reproduced.
Circularity Check
Partial circularity: the headline air-column MAE likely targets the same video pseudo-labels used to co-supervise the audio model; other property estimates rest on independent measurements.
-
fitted input called prediction
[Sec. 4.3 / Eq. (10); Sec. 6.1 / Table 2; App. A.3]
"To compute length of air column from wavelengths, we rely on Eq. (5). We compare our models with the baselines in estimating l(t) and report the mean absolute error averaged over all time points. ... The (psuedo) labels to train this network are obtained using temporal difference between adjacent frames and classical image processing techniques (Derivative of Gaussian on temporal difference heatmaps) to obtain clean ground truths. ..."
Eq. (10) is the fine-tuning loss: α·λ_audio/4 = l_px+βR_px, with l_px,R_px from the video teacher. The video teacher is trained on A.3's pseudo-labels (temporal-difference heatmaps + RANSAC). Table 2 reports MAE of l(t) but never states an independent source; the only l(t) 'ground truths' the paper describes are those pseudo-labels. Hence the headline 0.60 cm error can measure agreement with the same curve the audio net was fine-tuned to match, sharing smoothing/lag biases, rather than physical truth; the gain over the audio-only 0.78 cm is partly teacher-fitting. H/R, flow-rate and time-to-fill evaluations use manual rulers and actual fill time, so the circularity is partial.
full rationale
The physics chain from Eq. (3) to Eqs. (5)-(8) is a self-contained derivation from standard organ-pipe acoustics with an externally cited end-correction β=0.62; it does not borrow its conclusion from the learned model. Pitch detection is pre-trained on synthetic pouring sounds generated by the same physics equation, but the model is then tested on real recordings, and the static properties are compared against manual ruler measurements, flow rate against volume/time, and time-to-fill against actual fill duration. The only material circularity concern is the dynamic air-column metric: the audio network is fine-tuned with Eq. (10) to match a video teacher trained on the pseudo-label pipeline of Appendix A.3, and the paper does not state that Table 2's l(t) ground truth is independent of that pipeline. If it is the same pipeline, the 0.60 cm figure is a teacher-consistency score rather than an independent physical measurement. This warrants a moderate score, but it does not invalidate the static-property or cross-dataset results, which are externally grounded; there is no load-bearing self-citation or appeal to a uniqueness theorem by the authors.
Assumptions & free parameters
free parameters (2)
- scale factor alpha =
per video, estimated in [30, 80]
- end-correction factor beta =
0.62 (fixed)
assumptions (5)
- domain assumption Axial resonance in a cylindrical container is described by f(t) = c/(4(l(t) + beta*R)) with beta = 0.62.
- domain assumption The observed fundamental pitch of real pouring audio corresponds to this axial resonance mode.
- domain assumption Volume flow rate is approximately constant within a single pouring video.
- domain assumption The end-correction term beta*R is negligible compared to H at the start of pouring.
- domain assumption The scale factor alpha is constant per video and equals (1/Z)(f/s) for a perspective camera.
Cite this review
Pith. "Pith review of The Sound of Water: Inferring Physical Properties from Pouring Liquids." pith.science (2026). https://pith.science/paper/BVF3ET5M
@misc{pith2026241111222,
author = {Pith},
title = {Pith review of: The Sound of Water: Inferring Physical Properties from Pouring Liquids},
year = {2026},
howpublished = {\url{https://pith.science/paper/BVF3ET5M}},
note = {Machine review of arXiv:2411.11222}
}
read the original abstract
We study the connection between audio-visual observations and the underlying physics of a mundane yet intriguing everyday activity: pouring liquids. Given only the sound of liquid pouring into a container, our objective is to automatically infer physical properties such as the liquid level, the shape and size of the container, the pouring rate and the time to fill. To this end, we: (i) show in theory that these properties can be determined from the fundamental frequency (pitch); (ii) train a pitch detection model with supervision from simulated data and visual data with a physics-inspired objective; (iii) introduce a new large dataset of real pouring videos for a systematic study; (iv) show that the trained model can indeed infer these physical properties for real data; and finally, (v) we demonstrate strong generalization to various container shapes, other datasets, and in-the-wild YouTube videos. Our work presents a keen understanding of a narrow yet rich problem at the intersection of acoustics, physics, and learning. It opens up applications to enhance multisensory perception in robotic pouring.
Figures
Figures from the paper (17 more)
Reference graph
Works this paper leans on
-
[1]
Afouras, J
T. Afouras, J. S. Chung, and A. Zisserman. Deep lip reading: a comparison of models and an online application. In Conference of the International Speech Communication Association (INTERSPEECH),
-
[2]
Deep audio-visual speech recognition
Triantafyllos Afouras, Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman. Deep audio-visual speech recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 44(12):8717–8727, 2018. 3
2018
-
[3]
Hearing water temperature: Characterizing the development of nuanced perception of auditory events
Tanushree Agrawal, Michelle Lee, Amanda Calcetas, Danielle Clarke, Naomi Lin, and Adena Schachner. Hearing water temperature: Characterizing the development of nuanced perception of auditory events. In Annual Meeting of the Cognitive Science Society, 2020. URL https://api.semanticscholar.org/ CorpusID:231792766. 2
2020
-
[4]
Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text
Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong. Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text. Advances in Neural Information Processing Systems (NeurIPS), 34:24206–24221, 2021. 3
2021
-
[5]
Self- supervised learning by cross-modal audio-video clustering
Humam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani, Bernard Ghanem, and Du Tran. Self- supervised learning by cross-modal audio-video clustering. Advances in Neural Information Processing Systems (NeurIPS), 33:9758–9770, 2020. 3
2020
-
[6]
Herbert Anderson and Floyd C
S. Herbert Anderson and Floyd C. Ostensen. Effect of frequency on the end correction of pipes. Physical Review, 1928. URL https://api.semanticscholar.org/CorpusID:121139086. 4
1928
-
[7]
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman. Look, listen and learn. In International Conference on Computer Vision (ICCV), pages 609–617, 2017. 3
2017
-
[8]
Objects that sound
Relja Arandjelovic and Andrew Zisserman. Objects that sound. In European Conference on Computer Vision (ECCV), pages 435–451, 2018. 3
2018
Show all 102 references
-
[9]
Labelling unlabelled videos from scratch with multi-modal self-supervision
Yuki Asano, Mandela Patrick, Christian Rupprecht, and Andrea Vedaldi. Labelling unlabelled videos from scratch with multi-modal self-supervision. Advances in Neural Information Processing Systems (NeurIPS), 33:4660–4671, 2020. 3
2020
-
[10]
Soundnet: Learning sound representations from unlabeled video
Yusuf Aytar, Carl V ondrick, and Antonio Torralba. Soundnet: Learning sound representations from unlabeled video. Advances in Neural Information Processing Systems (NeurIPS), 29, 2016. 3, 14
2016
-
[11]
Pournet: Robust robotic pouring through curriculum and curiosity-based reinforcement learning
Edwin Babaians, Tapan Sharma, Mojtaba Karimi, Sahand Sharifzadeh, and Eckehard Steinbach. Pournet: Robust robotic pouring through curriculum and curiosity-based reinforcement learning. In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pages 9332–9...
2022
-
[12]
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in Neural Information Processing Systems (NeurIPS), 33:12449–12460, 2020. 2, 6, 22
2020
-
[13]
Piyush Bagad, Makarand Tapaswi, Cees G. M. Snoek, and Andrew Zisserman. The Sound of Water: Inferring Physical Properties from Pouring Liquids. In ICASSP, 2025. 1
2025
-
[14]
On the vibrations of elastic shells partly filled with liquid
Sudhansukumar Banerji. On the vibrations of elastic shells partly filled with liquid. Phys. Rev., Mar 1919. doi: 10.1103/PhysRev.13.171. URL https://link.aps.org/doi/10.1103/PhysRev.13.171. 2, 4
1919 doi
-
[15]
The physics of sound
Richard E Berg and David G Stork. The physics of sound. Pearson Education India, 1982. 2
1982
-
[16]
Cabe and John B
Patrick A. Cabe and John B. Pittenger. Human sensitivity to acoustic information from vessel filling. Journal of experimental psychology. Human perception and performance, 2000. 2, 3, 5, 11
2000
-
[17]
Perception of object length by sound
Claudia Carello, Krista L Anderson, and Andrew J Kunkler-Peck. Perception of object length by sound. Psychological science, 9(3):211–214, 1998. 2
1998
-
[18]
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In International Conference on Computer Vision (ICCV), pages 9650–9660, 2021. 8, 22, 23
2021
-
[19]
Multimodal clustering networks for self-supervised learning from unlabeled videos
Brian Chen, Andrew Rouditchenko, Kevin Duarte, Hilde Kuehne, Samuel Thomas, Angie Boggust, Rameswar Panda, Brian Kingsbury, Rogerio Feris, David Harwath, et al. Multimodal clustering networks for self-supervised learning from unlabeled videos. In International Conference on Co...
2021
-
[20]
Vggsound: A large-scale audio-visual dataset
Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman. Vggsound: A large-scale audio-visual dataset. In International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020. 2, 3
2020
-
[21]
Localizing visual sounds the hard way
Honglie Chen, Weidi Xie, Triantafyllos Afouras, Arsha Nagrani, Andrea Vedaldi, and Andrew Zisserman. Localizing visual sounds the hard way. In Conference on Computer Vision and Pattern Recognition (CVPR), pages 16867–16876, 2021. 3 17
2021
-
[23]
Sound localization by self-supervised time delay estimation
Ziyang Chen, David F Fouhey, and Andrew Owens. Sound localization by self-supervised time delay estimation. In European Conference on Computer Vision (ECCV), pages 489–508. Springer, 2022. 3
2022
-
[24]
T-foley: A controllable waveform-domain diffusion model for temporal-event-guided foley sound synthesis
Yoonjin Chung, Junwon Lee, and Juhan Nam. T-foley: A controllable waveform-domain diffusion model for temporal-event-guided foley sound synthesis. In International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 6820–6824. IEEE, 2024. 3
2024
-
[25]
Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Jian Ma, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100. Internation...
2022 doi
-
[26]
Yin, a fundamental frequency estimator for speech and music
Alain De Cheveigné and Hideki Kawahara. Yin, a fundamental frequency estimator for speech and music. The Journal of the Acoustical Society of America, 111(4):1917–1930, 2002. 10, 12, 25
1917
-
[27]
A probabilistic approach to liquid level detection in cups using an rgb-d camera
Chau Do, Tobias Schubert, and Wolfram Burgard. A probabilistic approach to liquid level detection in cups using an rgb-d camera. In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 2075–2080, 2016. doi: 10.1109/IROS.2016.7759326. 3
2016
-
[28]
Learning to pour using deep deterministic policy gradients
Chau Do, Camilo Gordillo, and Wolfram Burgard. Learning to pour using deep deterministic policy gradients. In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pages 3074–3079, 2018. doi: 10.1109/IROS.2018.8593654. 3
2018
-
[29]
Precision pouring into unknown containers by service robots
Chenyu Dong, Masaru Takizawa, Shunsuke Kudoh, and Takashi Suehiro. Precision pouring into unknown containers by service robots. In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 5875–5882, 2019. doi: 10.1109/IROS40897.2019.8967911. 3
2019
-
[30]
Conditional generation of audio from video via foley analogies
Yuexi Du, Ziyang Chen, Justin Salamon, Bryan Russell, and Andrew Owens. Conditional generation of audio from video via foley analogies. In Conference on Computer Vision and Pattern Recognition (CVPR), pages 2426–2436, 2023. 3
2023
-
[31]
Ddsp: Differentiable digital signal processing
Jesse Engel, Lamtharn (Hanoi) Hantrakul, Chenjie Gu, and Adam Roberts. Ddsp: Differentiable digital signal processing. In International Conference on Learning Representations (ICLR), 2020. URL https://openreview.net/forum?id=B1x1ma4tDr. 6
2020
-
[32]
Predicting 3d shapes, masks, and properties of materials, liquids, and objects inside transparent containers, using the transproteus cgi dataset, 2021
Sagi Eppel, Haoping Xu, Yi Ru Wang, and Alan Aspuru-Guzik. Predicting 3d shapes, masks, and properties of materials, liquids, and objects inside transparent containers, using the transproteus cgi dataset, 2021. 3
2021
-
[33]
Anthony P. French. In vino veritas: A study of wineglass acoustics. American Journal of Physics, 1983. URL https://api.semanticscholar.org/CorpusID:120875058. 2, 4
1983
-
[34]
Gemmeke, Daniel P
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter. Audio set: An ontology and human-labeled dataset for audio events. In International Conference on Acoustics, Speech and Signal Processing (ICASS...
2017
-
[35]
Audiovisual masked autoencoders
Mariana-Iuliana Georgescu, Eduardo Fonseca, Radu Tudor Ionescu, Mario Lucic, Cordelia Schmid, and Anurag Arnab. Audiovisual masked autoencoders. In International Conference on Computer Vision (ICCV), pages 16144–16154, 2023. 3
2023
-
[36]
The ecological approach to visual perception: classic edition
James J Gibson. The ecological approach to visual perception: classic edition. Psychology press, 2014. 2
2014
-
[37]
Imagebind: One embedding space to bind them all
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra. Imagebind: One embedding space to bind them all. In Conference on Computer Vision and Pattern Recognition (CVPR), pages 15180–15190, 2023. 3
2023
-
[38]
Contrastive audio-visual masked autoencoder
Yuan Gong, Andrew Rouditchenko, Alexander H Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, and James Glass. Contrastive audio-visual masked autoencoder. arXiv:2210.07839, 2022. 3
2022 arXiv
-
[39]
Listen, think, and understand
Yuan Gong, Hongyin Luo, Alexander H Liu, Leonid Karlinsky, and James Glass. Listen, think, and understand. arXiv:2305.10790, 2023. 3
2023 arXiv
-
[40]
Spectral information for detection of acoustic time to arrival
Michael S Gordon, Frank A Russo, and Ewen MacDonald. Spectral information for detection of acoustic time to arrival. Attention, Perception, & Psychophysics, 75:738–750, 2013. 2
2013
-
[41]
Tuning of musical glasses
Thomas Guignard. Tuning of musical glasses. In Master’s Thesis, ETH Zurich, 2003. URL https: //api.semanticscholar.org/CorpusID:137983171. 4
2003
-
[42]
Audioclip: Extending clip to image, text and audio
Andrey Guzhov, Federico Raue, Jörn Hees, and Andreas Dengel. Audioclip: Extending clip to image, text and audio. In International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 976–980. IEEE, 2022. 3 18
2022
-
[43]
Hermann L. F. Helmholtz and Alexander John Ellis. On the sensations of tone as a physiological basis for the theory of music. Nature, 12:449–452, 2005. URL https://api.semanticscholar.org/ CorpusID:119511156. 16, 22, 23
2005
-
[44]
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus). arXiv:1606.08415, 2016. 22
2016 arXiv
-
[45]
Mavil: Masked audio-video learners
Po-Yao Huang, Vasu Sharma, Hu Xu, Chaitanya Ryali, Yanghao Li, Shang-Wen Li, Gargi Ghosh, Jitendra Malik, Christoph Feichtenhofer, et al. Mavil: Masked audio-video learners. Advances in Neural Information Processing Systems (NeurIPS), 36, 2024. 3
2024
-
[46]
Learning to pour, 2017
Yongqiang Huang and Yu Sun. Learning to pour, 2017. 3
2017
-
[47]
Robot gaining accurate pouring skills through self- supervised learning and generalization
Yongqiang Huang, Juan Wilches, and Yu Sun. Robot gaining accurate pouring skills through self- supervised learning and generalization. Robotics and Autonomous Systems, 136:103692, 2021. 3
2021
-
[48]
End corrections of organ pipes
Arthur Taber Jones. End corrections of organ pipes. Journal of the Acoustical Society of America, 1941. URL https://api.semanticscholar.org/CorpusID:120564889. 4
1941
-
[49]
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al. The kinetics human action video dataset. arXiv:1705.06950, 2017. 2
2017 arXiv
-
[50]
Crepe: A convolutional representation for pitch estimation
Jong Wook Kim, Justin Salamon, Peter Li, and Juan Pablo Bello. Crepe: A convolutional representation for pitch estimation. In International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 161–165. IEEE, 2018. 6, 7, 10, 12, 25
2018
-
[51]
Adam: a method for stochastic optimization
DP Kingma. Adam: a method for stochastic optimization. arXiv:1412.6980, 2014. 7, 9
2014 arXiv
-
[52]
Segment anything
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. InInternational Conference on Computer Vision (ICCV), pages 4015–4026, 2023. 8
2023
-
[53]
Perception of material from contact sounds
Roberta L Klatzky, Dinesh K Pai, and Eric P Krotkov. Perception of material from contact sounds. Presence, 9(4):399–410, 2000. 2
2000
-
[54]
Cooperative learning of audio and video models from self-supervised synchronization
Bruno Korbar, Du Tran, and Lorenzo Torresani. Cooperative learning of audio and video models from self-supervised synchronization. Advances in Neural Information Processing Systems (NeurIPS) , 31,
-
[55]
Hearing shape
Andrew J Kunkler-Peck and Michael T Turvey. Hearing shape. Journal of Experimental psychology: human perception and performance, 26(1):279, 2000. 2
2000
-
[56]
Temporal convolu- tional networks for action segmentation and detection
Colin Lea, Michael D Flynn, Rene Vidal, Austin Reiter, and Gregory D Hager. Temporal convolu- tional networks for action segmentation and detection. In Conference on Computer Vision and Pattern Recognition (CVPR), pages 156–165, 2017. 14
2017
-
[57]
Sound-guided semantic video generation
Seung Hyun Lee, Gyeongrok Oh, Wonmin Byeon, Chanyoung Kim, Won Jeong Ryoo, Sang Ho Yoon, Hyunjun Cho, Jihyun Bae, Jinkyu Kim, and Sangpil Kim. Sound-guided semantic video generation. In European Conference on Computer Vision (ECCV), pages 34–50. Springer, 2022. 3
2022
-
[58]
Making sense of audio vibration for liquid height estimation in robotic pouring
Hongzhuo Liang, Shuang Li, Xiaojian Ma, Norman Hendrich, Timo Gerkmann, Fuchun Sun, and Jianwei Zhang. Making sense of audio vibration for liquid height estimation in robotic pouring. In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), November 2019. 3
2019
-
[59]
Making sense of audio vibration for liquid height estimation in robotic pouring
Hongzhuo Liang, Shuang Li, Xiaojian Ma, Norman Hendrich, Timo Gerkmann, Fuchun Sun, and Jianwei Zhang. Making sense of audio vibration for liquid height estimation in robotic pouring. In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 5333–533...
2019
-
[60]
Robust robotic pouring using audition and haptics
Hongzhuo Liang, Chuangchuang Zhou, Shuang Li, Xiaojian Ma, Norman Hendrich, Timo Gerkmann, Fuchun Sun, Marcus Stoffel, and Jianwei Zhang. Robust robotic pouring using audition and haptics. In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, Oct...
2020
-
[61]
Robust robotic pouring using audition and haptics
Hongzhuo Liang, Chuangchuang Zhou, Shuang Li, Xiaojian Ma, Norman Hendrich, Timo Gerkmann, Fuchun Sun, Marcus Stoffel, and Jianwei Zhang. Robust robotic pouring using audition and haptics. In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 108...
2020
-
[62]
Pourit!: Weakly-supervised liquid perception from a single image for visual closed-loop robotic pouring
Haitao Lin, Yanwei Fu, and Xiangyang Xue. Pourit!: Weakly-supervised liquid perception from a single image for visual closed-loop robotic pouring. In International Conference on Computer Vision (ICCV), pages 241–251, October 2023. 3
2023
-
[63]
Qi Liu, Fan Feng, Chuanlin Lan, and Rosa H. M. Chan. Va2mass: Towards the fluid filling mass estimation via integration of vision and audio learning. In ICPR Workshops, 2020. URL https: //api.semanticscholar.org/CorpusID:232023100. 3
2020
-
[64]
Active contrastive learning of audio-visual video representations
Shuang Ma, Zhaoyang Zeng, Daniel McDuff, and Yale Song. Active contrastive learning of audio-visual video representations. arXiv:2009.09805, 2020. 3 19
2009 arXiv
-
[65]
Language segment-anything
Luca Medeiros. Language segment-anything. Github, 2024. 22
2024
-
[66]
Absolute and relative cues for the auditory perception of egocentric distance
Donald H Mershon and John N Bowers. Absolute and relative cues for the auditory perception of egocentric distance. Perception, 8(3):311–322, 1979. 2
1979
-
[67]
Audio-visual instance discrimination with cross- modal agreement
Pedro Morgado, Nuno Vasconcelos, and Ishan Misra. Audio-visual instance discrimination with cross- modal agreement. In Conference on Computer Vision and Pattern Recognition (CVPR), pages 12475– 12486, 2021. 3
2021
-
[68]
Harvest: A high-performance fundamental frequency estimator from speech signals
Masanori Morise et al. Harvest: A high-performance fundamental frequency estimator from speech signals. In Conference of the International Speech Communication Association (INTERSPEECH), pages 2321–2325, 2017. 2
2017
-
[69]
See the glass half full: Reasoning about liquid containers, their volume and content
Roozbeh Mottaghi, Connor Schenck, Dieter Fox, and Ali Farhadi. See the glass half full: Reasoning about liquid containers, their volume and content. International Conference on Computer Vision (ICCV), pages 1889–1898, 2017. URL https://api.semanticscholar.org/CorpusID:7410030. 3
2017
-
[70]
Self-supervised transparent liquid segmentation for robotic pouring
Gautham Narayan Narasimhan, Kai Zhang, Ben Eisner, Xingyu Lin, and David Held. Self-supervised transparent liquid segmentation for robotic pouring. IEEE International Conference on Robotics and Automation (ICRA), pages 4555–4561, 2022. URL https://api.semanticscholar.org/Corpu...
2022
-
[71]
Audio-visual scene analysis with self-supervised multisensory features
Andrew Owens and Alexei A Efros. Audio-visual scene analysis with self-supervised multisensory features. In European Conference on Computer Vision (ECCV), pages 631–648, 2018. 3
2018
-
[72]
Ambient sound provides supervision for visual learning
Andrew Owens, Jiajun Wu, Josh H McDermott, William T Freeman, and Antonio Torralba. Ambient sound provides supervision for visual learning. In European Conference on Computer Vision (ECCV), pages 801–816. Springer, 2016. 3
2016
-
[73]
Robot motion planning for pouring liquids
Zherong Pan, Chonhyon Park, and Dinesh Manocha. Robot motion planning for pouring liquids. Proceedings of the International Conference on Automated Planning and Scheduling, 26(1):518–526, Mar
-
[74]
Why can you hear a difference between pouring hot and cold water? an investigation of temperature dependence in psychoacoustics
He Peng and Joshua D Reiss. Why can you hear a difference between pouring hot and cold water? an investigation of temperature dependence in psychoacoustics. In Audio Engineering Society Convention
-
[75]
Critcher
Hannah Perfecto, Kristin Donnelly, and Clayton R. Critcher. V olume estimation through mental simulation. Psychological Science, 30, 2019. 2
2019
-
[76]
V olume estimation through mental simulation
Hannah Perfecto, Kristin Donnelly, and Clayton R Critcher. V olume estimation through mental simulation. Psychological science, 30(1):80–91, 2019. 3
2019
-
[77]
Pouring by feel: An analysis of tactile and proprioceptive sensing for accurate pouring
Pedro Piacenza, Daewon Lee, and V olkan Isler. Pouring by feel: An analysis of tactile and proprioceptive sensing for accurate pouring. In IEEE International Conference on Robotics and Automation (ICRA), pages 10248–10254, 2022. doi: 10.1109/ICRA46639.2022.9811898. 3
2022
-
[78]
C. E. Pykett. End corrections, natural frequencies, tone colour and physical modelling of organ pipes,
-
[79]
The theory of sound, volume 2
John William Strutt Baron Rayleigh. The theory of sound, volume 2. Macmillan, 1896. 2, 4
-
[80]
Pesto: Pitch estimation with self- supervised transposition-equivariant objective
Alain Riou, Stefan Lattner, Gaëtan Hadjeres, and Geoffroy Peeters. Pesto: Pitch estimation with self- supervised transposition-equivariant objective. In International Society for Music Information Retrieval Conference (ISMIR), 2023. 10, 12, 25
2023
-
[81]
Size, shape, and material properties of sound models
Davide Rocchesso, Laura Ottaviani, Federico Fontana, and Federico Avanzini. Size, shape, and material properties of sound models. The sounding object, pages 95–110, 2003. 2
2003
-
[82]
Identifying a sound-producing object’s direction of motion and change in speed
Michael K Russell. Identifying a sound-producing object’s direction of motion and change in speed. Auditory Perception & Cognition, 6(3-4):353–368, 2023. 2
2023
-
[83]
Detection and tracking of liquids with fully convolutional networks,
Connor Schenck and Dieter Fox. Detection and tracking of liquids with fully convolutional networks,
-
[84]
Towards learning to perceive and reason about liquids
Connor Schenck and Dieter Fox. Towards learning to perceive and reason about liquids. In International Symposium on Experimental Robotics, 2016. URL https://api.semanticscholar.org/CorpusID: 12918749
2016
-
[85]
Perceiving and reasoning about liquids using fully convolutional networks
Connor Schenck and Dieter Fox. Perceiving and reasoning about liquids using fully convolutional networks. The International Journal of Robotics Research, 37:452 – 471, 2017. URL https://api. semanticscholar.org/CorpusID:6123383. 3
2017
-
[86]
Visual closed-loop control for pouring liquids
Connor Schenck and Dieter Fox. Visual closed-loop control for pouring liquids. In IEEE International Conference on Robotics and Automation (ICRA) , pages 2629–2636, 2017. doi: 10.1109/ICRA.2017. 7989307. 3 20
2017 doi
-
[87]
You can hear the difference between hot and cold water, March 2017
Tom Scott. You can hear the difference between hot and cold water, March 2017. URL https: //www.youtube.com/watch?v=Ri_4dDvcZeM. 2
2017
-
[88]
Time-contrastive networks: Self-supervised learning from video
Pierre Sermanet, Corey Lynch, Yevgen Chebotar, Jasmine Hsu, Eric Jang, Stefan Schaal, Sergey Levine, and Google Brain. Time-contrastive networks: Self-supervised learning from video. InIEEE International Conference on Robotics and Automation (ICRA), pages 1134–1141. IEEE, 2018. 3
2018
-
[89]
Learning audio-visual speech representation by masked multimodal cluster prediction
Bowen Shi, Wei-Ning Hsu, Kushal Lakhotia, and Abdelrahman Mohamed. Learning audio-visual speech representation by masked multimodal cluster prediction. arXiv:2201.02184, 2022. 3
2022 arXiv
-
[90]
Intuitive physical inference from sound
James Traer and J McDermott. Intuitive physical inference from sound. In 2018 Conference on Cognitive Computational Neuroscience, pages 2018–1057, 2018. 2
2018
-
[91]
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of Machine Learning Research (JMLR), 9(11), 2008. 12
2008
-
[92]
Attention is all you need
A Vaswani. Attention is all you need. Advances in Neural Information Processing Systems (NeurIPS),
-
[93]
The sound of temperature: What information do pouring sounds convey concerning the temperature of a beverage
Carlos Velasco, Russ Jones, Scott King, and Charles Spence. The sound of temperature: What information do pouring sounds convey concerning the temperature of a beverage. Journal of Sensory Studies, 28,
-
[94]
The sound of temperature: What information do pouring sounds convey concerning the temperature of a beverage
Carlos Velasco, Russ Jones, Scott King, and Charles Spence. The sound of temperature: What information do pouring sounds convey concerning the temperature of a beverage. Journal of Sensory Studies, 28(5): 335–345, 2013. 3
2013
-
[95]
The use of helmholtz resonance for measuring the volume of liquids and solids
Emile S Webster and Clive E Davies. The use of helmholtz resonance for measuring the volume of liquids and solids. Sensors, 10(12):10663–10672, 2010. 2, 11
2010
-
[96]
Analyzing liquid pouring sequences via audio-visual neural networks
Justin Wilson, Auston Sterling, and Ming Lin. Analyzing liquid pouring sequences via audio-visual neural networks. In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2019. 3, 12, 13, 14, 16, 25, 26
2019
-
[97]
Physics 101: Learning physical object properties from unlabeled videos
Jiajun Wu, Joseph J Lim, Hongyi Zhang, Joshua B Tenenbaum, and William T Freeman. Physics 101: Learning physical object properties from unlabeled videos. InBritish Machine Vision Conference (BMVC),
-
[98]
Liquid pouring monitoring via rich sensory inputs, 2018
Tz-Ying Wu, Juan-Ting Lin, Tsun-Hsuang Wang, Chan-Wei Hu, Juan Carlos Niebles, and Min Sun. Liquid pouring monitoring via rich sensory inputs, 2018. 3
2018
-
[99]
Merlot reserve: Neural script knowledge through vision and language and sound
Rowan Zellers, Jiasen Lu, Ximing Lu, Youngjae Yu, Yanpeng Zhao, Mohammadreza Salehi, Aditya Kusupati, Jack Hessel, Ali Farhadi, and Yejin Choi. Merlot reserve: Neural script knowledge through vision and language and sound. In Conference on Computer Vision and Pattern Recogniti...
2022
-
[100]
Pouring dynamics estimation using gated recurrent units, 2021
Qi Zheng. Pouring dynamics estimation using gated recurrent units, 2021. 3
2021
-
[101]
Psychoacoustics: Facts and models, volume 22
Eberhard Zwicker and Hugo Fastl. Psychoacoustics: Facts and models, volume 22. Springer Science & Business Media, 2013. 2 21 A Appendix / supplemental material A.1 Dataset Recording setup. We use the OnePlus Nord CE 5G phone with its in-built microphone to record videos of liq...
2013
-
[145]
Audio Engineering Society, 2018. 16
2018
-
[2016]
URL https://ojs.aaai.org/index.php/ICAPS/article/ view/13787
doi: 10.1609/icaps.v26i1.13787. URL https://ojs.aaai.org/index.php/ICAPS/article/ view/13787. 3
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.