Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2503.11197.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.11197 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:31:13.461419Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T20:07:21.344787Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9f43c0ce-d66c-4ef9-a17f-0fc807f4c018 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 292

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:24.508679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:f22ce8dd24c569e79c16c56ccb261dbe2f2845fdbbd4280a8766c0029bf5bed5

Observation 7b9ad463-7fbf-40ee-8a5f-8390c59237f0 · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.461419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.461419Z digest=sha256:90b826ae242e2fd7239b8256f335f83580a16456cd42aaa0fe3b13debdcefa03

Observation 02f72257-f736-4ef4-9ac9-4788cfbaf892 · inbound

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models cites this paper.

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:42:45.293345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T03:42:44.523919Z digest=sha256:67489fb92f9e6944d744ca62b362ae4a07b3f63a5cd3d3858829440430dddce9

Observation 08b7bb0e-fa30-455a-9a89-729280f973ba · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 171

Resolution
unresolved
no resolver link, observed 2026-08-05T20:28:59.977640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:28:59.977640Z digest=sha256:8e9d60b9d196c2de1702e5ae0bfacb3b358c0f4c6e2b09408f08107f50a134ab

Observation 6a3b6bae-50c7-4904-b015-a960435a0c0e · inbound

Group Relative Policy Optimization for Speech Recognition cites this paper.

Group Relative Policy Optimization for Speech Recognition Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T12:07:21.394981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:07:21.394981Z digest=sha256:b7536633107a2791c5f4bd25943fc019833fb295a28585467785766e50043bfc

Observation 9b56048b-e90a-4405-b888-06f9d163bd80 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 262

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.171376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:87407a92b48d955b88170f5cf1a124aab1f33747eaa0ee1fff019e17a78d2047

Observation 513d9acc-d37d-431a-9187-d31604016403 · inbound

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning cites this paper.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:18.440292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:18.440292Z digest=sha256:f1ab4c8aedd1e3ea721ddc22bc9fe6a1f5fadaf36074fffea7be917c7f7ad326

Observation 1e30c90c-79f5-444e-a87c-a4d8b13e6018 · inbound

Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models cites this paper.

Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:46:03.575240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T07:43:23.913399Z digest=sha256:ea04ca7f40c2763b5a2cbb7aa6a3b20be7c66540641028793cfed4ee46a8f4f1

Observation 8f292c1d-1083-4035-b4a8-109329d1dec2 · inbound

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling cites this paper.

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:31:32.069649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T12:26:35.347190Z digest=sha256:5aa15baaebe38887f254dd67d0f555e5a3dffe6a78593d16ca577daee414c6ab

Observation 946b3ddd-5e8a-452d-8d87-9c6b7ded12a0 · inbound

VERITAS: A Multi-Agent Co-Scientist for Verifiable Image-Derived Hypothesis Testing cites this paper.

VERITAS: A Multi-Agent Co-Scientist for Verifiable Image-Derived Hypothesis Testing Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T21:34:30.078813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:34:30.078813Z digest=sha256:a41967bd9c88f2b81eb648902f6fb252a7ce0d7b572069ae36068869f248c7d8

Observation 8293b213-4cda-4579-b637-2a6d80e6fd61 · inbound

Why Your Tokenizer Fails in Information Fusion: A Timing-Aware Pre-Quantization Fusion for Video-Enhanced Audio Tokenization cites this paper.

Why Your Tokenizer Fails in Information Fusion: A Timing-Aware Pre-Quantization Fusion for Video-Enhanced Audio Tokenization Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:31:01.814073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T14:49:51.481873Z digest=sha256:a00bc48f524d27fc095a5d5fba44259b3e1220ffa65fddbaaac882c8c05cdea7

Observation aceb0ce8-818d-4344-8794-de122493689c · inbound

TinyMU: A Compact Audio-Language Model for Music Understanding cites this paper.

TinyMU: A Compact Audio-Language Model for Music Understanding Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:17:36.954358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T08:17:23.740979Z digest=sha256:5aa7e74e2926169dbdadca9acc4b4d6c43eb31436a818e749357061e2ebb0dd5

Observation e4d01878-c041-45ed-a53e-f627b8b95a83 · inbound

Audio-DeepThinker: Progressive Reasoning-Aware Reinforcement Learning for High-Quality Chain-of-Thought Emergence in Audio Language Models cites this paper.

Audio-DeepThinker: Progressive Reasoning-Aware Reinforcement Learning for High-Quality Chain-of-Thought Emergence in Audio Language Models Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:26:04.051187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T03:41:27.425176Z digest=sha256:7badbcb02393285b76e37afdcbf20d3c93405496ce545d5e09919e620005386f

Observation 044e02f1-1000-4e55-9cbb-922572d725cf · inbound

Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization cites this paper.

Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:46:25.869198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:01:23.060562Z digest=sha256:e8dd91f0a5914b53e6fe315802c3b87b701bd9e6ecadd24f88e1f6befa277caa

Observation cf1683a5-dbf9-4766-9b64-1d903fbb5bdf · inbound

Beyond Content: A Comprehensive Speech Toxicity Dataset and Detection Framework Incorporating Paralinguistic Cues cites this paper.

Beyond Content: A Comprehensive Speech Toxicity Dataset and Detection Framework Incorporating Paralinguistic Cues Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T18:37:42.814418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T18:36:34.351533Z digest=sha256:8f04fe13d18003c65869a20f1bba193f48c5665b624c0d44fea95057cc30eda8

Observation 7e39656b-afa1-4cab-afec-3ed85f069a99 · inbound

A Survey of Audio Reasoning in Multimodal Foundation Models cites this paper.

A Survey of Audio Reasoning in Multimodal Foundation Models Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:09:24.292449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T02:08:06.976461Z digest=sha256:b7f9e545848027eb83289546c93b2fb50a9bfd9f9f65add42d98579372897afd

Observation ff16a0b4-21dd-4882-99d2-2b437f8372ca · inbound

Learning When to Think While Listening in Large Audio-Language Models cites this paper.

Learning When to Think While Listening in Large Audio-Language Models Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:43:50.737960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T18:37:34.409802Z digest=sha256:4a81c90081366ec156b9c533a7b05fe55ebda4aa731bd1da08dfe516bb7845b3

Observation 4399527a-8e7c-4c25-82d8-6525afc60902 · inbound

LaSR: Context-Aware Speech Recognition via Latent Reasoning cites this paper.

LaSR: Context-Aware Speech Recognition via Latent Reasoning Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:22:34.992549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T19:12:59.843096Z digest=sha256:0d29f70300c029934d06ec819d937ae9c4ba2f06c43212ed7da474214f6e23ab

Observation 908f10dc-06eb-4515-9197-97752fbe235a · inbound

FSA-GRPO: Teaching Auditory LLMs to Use Few-shot Demonstrations cites this paper.

FSA-GRPO: Teaching Auditory LLMs to Use Few-shot Demonstrations Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T16:05:49.453256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-01T16:02:03.730543Z digest=sha256:3992ce120c53c4b465620d2efd11426a2b9595436b5fd90e7213d1c1b63756d9

Observation 7cfff068-1d25-4f50-8ef7-0f8da95b5d03 · inbound

VISA: A Visual Information Strengthened Audio-Reasoning System for the Interspeech 2026 ARC Agent Track cites this paper.

VISA: A Visual Information Strengthened Audio-Reasoning System for the Interspeech 2026 ARC Agent Track Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:07:21.346714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T21:02:19.300441Z digest=sha256:f1dfb7bcfa6dc3541cdeb998c8dcedf464fff9125f28eb5741cbe1717ab18464

Observation 1b580945-e9cd-485e-94c7-0b3bb827eaf8 · inbound

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos cites this paper.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:42.917716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:42.917716Z digest=sha256:e6adfec79b92427372800223b260a55f23dd448c4fa4d4d9539139e7320722ef

Observation 5924d362-6311-4e55-906c-f7e5c3a8e008 · inbound

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning cites this paper.

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.897788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.897788Z digest=sha256:f44bef2f7f03b3fc8d82a7f3de2c96fe5b346df0ae4a59c88a1ca92545784cff

Observation 18352f4d-a462-4945-a5a9-133ff75fc6eb · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:16.661745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:16.661745Z digest=sha256:f5f7405cdfd8ed32a01295183c3884dc5940a62a370282a8c800ca208ca03773

Observation 6f8e5bc0-3f3e-42a5-aa65-05b8214630aa · inbound

Weak-to-Strong On-Policy Distillation cites this paper.

Weak-to-Strong On-Policy Distillation Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 129

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:30.612742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:30.612742Z digest=sha256:38780fb2c450ce9322073d90a373d4cb419f8c3480e9756c10e912af3f12ebba