Pith. sign in

Paper Citation Record · LEDGER

Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2407.04416.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.04416 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:36:19.799975Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.536675Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 00306dbf-1ccc-40bd-9615-44556ec9a28a · inbound

Audio-Language Models for Audio-Centric Tasks: A Systematic Survey cites this paper.

Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions

Reference 135

Resolution
unresolved
no resolver link, observed 2026-08-10T14:36:19.799975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:36:19.799975Z digest=sha256:4de39eb33cba300198f1d3cc371de6bf19693bd74595122200b38f2391354f73

Observation f491f5a3-11f0-4614-94de-1a71b8fca4f4 · inbound

Sounding that Object: Interactive Object-Aware Image to Audio Generation cites this paper.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:17.755251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:17.755251Z digest=sha256:568740c22f469eca5df641acdd46f937572778b7bdc076f4a7750879aacb9a4e

Observation 400238fd-6039-465d-a1a4-3597ab50cd9d · inbound

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos cites this paper.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:02:46.490467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:02:46.490467Z digest=sha256:6726875bbe4dae4bb9a39ea45be7af56f1e66966966a20419ae847d3e3b28e00

Observation 42eee5fb-72fb-4fed-b827-dc747763a45f · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions

Reference 67

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.538025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:589a577838b246e15a9868b77c031320d32656fbed780a073a9edc8cf16117cd

Observation 9b7c5262-0aff-4d9f-bd34-95b37b4fdf84 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:f747a00f3683e05d8e3b599c059cdf940b702d4f8435a0c086cbeaf82fd238fb