Pith. sign in

Paper Citation Record · LEDGER

Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization

As of 9 August 2026, this Paper Citation Record lists 7 of 7 outbound references and 3 inbound Pith citation observations for arXiv:2602.10230.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.10230 v2

Coverage vector

measured 7 of 7 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T01:17:47.362069Z

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T08:43:32.422846Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

7 of 7 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved6
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 739419b8-9968-432a-b95a-06f712fcb33e · outbound

This paper cites Sequence Transduction with Recurrent Neural Networks.

Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization Sequence Transduction with Recurrent Neural Networks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T01:17:46.970140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:17:46.970140Z digest=sha256:311a6eb3b1f19eba25be887f12c8b493f3461f760564d951c8e94a11d4d03878

Observation 756e342d-7bf4-4be8-b1ed-6d74b418cf49 · outbound

This paper cites Schick, T., Dwivedi-Yu, J., Dessi, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., and Scialom, T.

Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization Schick, T., Dwivedi-Yu, J., Dessi, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., and Scialom, T

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T01:17:47.237653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:17:47.237653Z digest=sha256:525bd62c7b7aae02b4380af811a0b0dbaa0bd3db7dcead17d5c78fc8e4ce534a

Observation 6f92203b-6391-4f61-aa9f-9cc912f31ddc · outbound

This paper cites word": "much.

Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization word": "much

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-03T01:17:47.362069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:17:47.362069Z digest=sha256:93f42d7e6e8b63c7c1907bce253c467a7422a0bc3d76819a52943bd4c65d7b38

Observation 2b8c1b47-2992-4ede-99ee-0329898dec86 · outbound

This paper cites GroundingGPT:Language Enhanced Multi-modal Grounding Model.

Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization GroundingGPT:Language Enhanced Multi-modal Grounding Model

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T01:17:47.065638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:17:47.065638Z digest=sha256:3476a9a4e1aa1fb2763d6f83d24b42cb17bc4ebf850b8aebf1b26a11af72504a

Observation bc4856d6-0694-4807-94b7-1f02e9a70a05 · outbound

This paper cites cc/paper_files/paper/2023/file/ d842425e4bf79ba039352da0f658a906-Paper-Conference.

Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization cc/paper_files/paper/2023/file/ d842425e4bf79ba039352da0f658a906-Paper-Conference

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T01:17:47.268297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:17:47.268297Z digest=sha256:30a6d8e96b0b2705e176a4327af6899ef2199fb552b14be0ec1e36673c9316d2

Observation 6f18f1dd-31ad-4a52-824c-bc3794f08c4b · outbound

This paper cites Voxtral.

Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization Voxtral

Reference 2024

Resolution
malformed identifier
no resolver link, observed 2026-08-03T01:17:47.177241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:17:47.177241Z digest=sha256:8d038cc5d6ae166dbbc3e45242d06b1fa277466909dcb36372fc217d3dc87d7e

Observation 7df5fd3d-4755-4212-87af-cdc4cbb74137 · outbound

This paper cites Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding.

Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T01:17:46.910143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:17:46.910143Z digest=sha256:e263315abf69b5d8ba6b6e8f84ef5eab592ce9abe9d4f13e333bcace49c794a4

Pith citing papers

Observation 4ad3990f-7093-488d-86af-0f21858899a4 · inbound

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding cites this paper.

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T08:43:32.422846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:43:32.422846Z digest=sha256:c2de014493d057cbf53fd267200af44fe564092bef6c58686181631540a63ba2

Observation 5109a179-b385-46fc-b755-48b774a3b227 · inbound

SoundscapeAgent: Agentic Soundscape Construction for Controllable Synthesis and Scalable Audio-Language Supervision cites this paper.

SoundscapeAgent: Agentic Soundscape Construction for Controllable Synthesis and Scalable Audio-Language Supervision Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T06:28:47.970632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:28:47.970632Z digest=sha256:b2bd869ab2fc0fe9242c0f57909c0f8bc2574ee566ddd46c3057b9a34557f7a6

Observation a227c080-866e-436d-b888-9ae5bb68591b · inbound

From Semantics to Readout: Mechanistic Understanding of Audio Tokens after Fine-Tuning for Temporal Audio Grounding cites this paper.

From Semantics to Readout: Mechanistic Understanding of Audio Tokens after Fine-Tuning for Temporal Audio Grounding Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T02:46:57.963783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:46:57.963783Z digest=sha256:a67f2548b0cfda902ecf05f4001845a505bf5e8452316a70533d2db19d5449e0