Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning

As of 17 August 2026, this Paper Citation Record lists 100 of 100 outbound references and 0 inbound Pith citation observations for arXiv:2608.02831.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02831 v1

Coverage vector

measured 100 of 100 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:03:20.975265Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 100 outbound references displayed

  • verified exact11
  • verified fuzzy2
  • unresolved87
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 72216fe6-495e-419e-904f-05bda35caf94 · outbound

This paper cites 2023 , url =.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning 2023 , url =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.581426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.581426Z digest=sha256:ae5ec150c34b037d5055033764df4e8c68cc6abc28278b9d0a546f97b788b36d

Observation 71ca77b9-c836-42c1-87dd-248d0c611006 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Advances in Neural Information Processing Systems , volume=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.625181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.625181Z digest=sha256:082c424ceacfa325a3b72b83fff8b47aea2c3c314c655edc5545c68b675a95f1

Observation 792fb518-733e-4133-bf11-e4b9fb2ddb5d · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Movie Gen: A Cast of Media Foundation Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.649697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.649697Z digest=sha256:c9ae2b27673e119c8dd4d5c438d73450c22120bba046c9eee068eb2658c265a5

Observation 6b351283-fb42-4f08-9ae0-69dce7cdb245 · outbound

This paper cites Towards Conversational Medical AI with Eyes, Ears and a Voice.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Towards Conversational Medical AI with Eyes, Ears and a Voice

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.661572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.661572Z digest=sha256:bec42e15e6cbcfe7a84f84b9bbdfbaf60abd1d7a767ab8b4627594cbe5e27fa2

Observation b192ec25-585c-4b2a-98b4-41192c1b3902 · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2025 , pages=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Findings of the Association for Computational Linguistics: ACL 2025 , pages=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.666574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.666574Z digest=sha256:0713f8e043ac12906528d5c89c5cfcfaed2654c163a131b1b74806455d4c006c

Observation 7381af69-928f-4fd3-8a28-94607cb518dd · outbound

This paper cites International Conference on Learning Representations , volume=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning International Conference on Learning Representations , volume=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.671105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.671105Z digest=sha256:7be2a55d588349e20bacfaa9d9905fe14285c2f0fbf5eaef8ced745dea65ac8e

Observation bc2b904c-0fd8-4a08-8563-f4a70426a18c · outbound

This paper cites Reward Hacking in Rubric-Based Reinforcement Learning.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Reward Hacking in Rubric-Based Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.709576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.709576Z digest=sha256:9c4a91d1176ad2721896647e86d45dc9d6f634307d090a80786b647e611f2ef6

Observation d8664f22-1f09-40cd-ac1d-461a833847cb · outbound

This paper cites arXiv preprint arXiv:2602.05125 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2602.05125 , year=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.779868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.779868Z digest=sha256:1772b9e0d26a55579949fd098243c8e4edb86175b52f4737f4374c71e2e16b12

Observation 916fd730-9271-4eb6-870c-6a209d65ee9c · outbound

This paper cites Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.825053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.825053Z digest=sha256:0b658136a7f1dca237ac18be553f170cc187105fc61c57bd5b79d15c0de4c298

Observation b8a86bb6-d29c-45c5-b3e6-ccfa8dd20dac · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.844112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.844112Z digest=sha256:363b78d31e43d2c3c59b83953f002254585886acb7e6469c2d5766e8fece4f52

Observation 53f5c82c-5342-4939-976b-994ae60b48b5 · outbound

This paper cites Weak-to-Strong On-Policy Distillation.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Weak-to-Strong On-Policy Distillation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.871173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.871173Z digest=sha256:e986fceb35e9c9dbf23c06fb8d8f64dd461f96ece4deff175a7fa48d1f3b556f

Observation c5d69191-b1f5-4aa5-9a5b-2f7c6f16fc0e · outbound

This paper cites arXiv preprint arXiv:2602.21628 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2602.21628 , year=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.876157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.876157Z digest=sha256:3ad831851bc83278b8fe9eea6428504e402660ae1816e37580cb3d1426b7e460

Observation 6cf13027-e7d0-4867-ad6c-464d6450b724 · outbound

This paper cites Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.880654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.880654Z digest=sha256:15fc415c9cc532bf7af3686a8f1e922ad5c3d5cb66794e3e40b78a5e7c927c9c

Observation 0508a662-86ec-45ba-b422-fbde282f5187 · outbound

This paper cites arXiv preprint arXiv:2603.16600 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2603.16600 , year=

Reference 14

Resolution
verified exact
raw_fallback, observed 2026-08-15T15:03:23.732358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T15:03:18.885146Z digest=sha256:caca8a45d57ef8627783c12334597892a5256a1fd8529157d3a55eb33a0440c8

Observation bf9335c5-335b-4e01-9de7-8a82f72280bc · outbound

This paper cites Visual Preference Optimization with Rubric Rewards.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Visual Preference Optimization with Rubric Rewards

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.914108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.914108Z digest=sha256:91212647f647630f2437d23323b6b68afe61fc9b430310f899d9c17878f5dbd1

Observation e3bb04da-0188-4c5e-81ba-9e91f6e90658 · outbound

This paper cites AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.008040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.008040Z digest=sha256:38365434a512df38d329297a46260a6e11491d3b99e4f62281cf926463bf2c60

Observation a5403e7e-d75e-4a25-ba2c-3ccadb4a6018 · outbound

This paper cites arXiv preprint arXiv:2602.04649 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2602.04649 , year=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.084775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.084775Z digest=sha256:e8b9b81cf7f1da43beb1faa48f57b64b456f49e5976681b69534d7bcd299b21b

Observation a2bce307-639f-4de9-b3f6-da494af5e87b · outbound

This paper cites arXiv preprint arXiv:2602.01511 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2602.01511 , year=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.110231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.110231Z digest=sha256:8fec9767c1cc13aa97885690fd51efe296f3ee048da1cc62235820841b043702

Observation 02d019f0-b7a9-4c4b-87a5-92cb6fe569b8 · outbound

This paper cites arXiv preprint arXiv:2510.07284 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2510.07284 , year=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.114984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.114984Z digest=sha256:3f999ca19360121470f1b045f28df89f3a7340c62e9ee3361c2d22b588a8aff4

Observation b46170ed-2603-4e1a-b81a-36d0e91c1504 · outbound

This paper cites arXiv preprint arXiv:2602.10885 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2602.10885 , year=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.119412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.119412Z digest=sha256:9117eb786d2925318a878de2f51fe67715327c77376fcdd8747cc2b5e452d2d0

Observation 9d8fc2a4-822b-4370-95d0-d656a24c9c3c · outbound

This paper cites DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.123830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.123830Z digest=sha256:22b7ad78486585fd38c0aef214498c24fdf5bf7c4da196df08ccb0f4e56a54ea

Observation f8766714-db81-404c-bcc1-0bb3da494bfd · outbound

This paper cites Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.128691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.128691Z digest=sha256:919c9671405f2b60d4e89ddb96253a069c28a4697c1f2d4db0061a196545ed8c

Observation e08e7ca9-3d77-45e1-9500-41b955f68030 · outbound

This paper cites arXiv e-prints , pages=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv e-prints , pages=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.134675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.134675Z digest=sha256:976e941893be5b9bd0353e1e5c979807c7d37262a0b0429d9885b3eead560425

Observation a4458c19-3996-47b3-b819-7272e8cdea8f · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2025 , pages=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Findings of the Association for Computational Linguistics: ACL 2025 , pages=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.209889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.209889Z digest=sha256:ab5037e71f40ae705e9c778e6ccb0706cd5ab9fa797ae1330bd100ff334295ba

Observation 755ce286-98ea-412e-8f11-4c2f041745b4 · outbound

This paper cites arXiv preprint arXiv:2510.07743 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2510.07743 , year=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.286838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.286838Z digest=sha256:93e50680f1a9f6c3e2bdd9a4e6e4f10d6e0b79fc400bf39bb6d598c99a89dc8d

Observation 10f29666-f0fb-4867-9216-f43af663fb32 · outbound

This paper cites International Conference on Learning Representations , volume=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning International Conference on Learning Representations , volume=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.312041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.312041Z digest=sha256:bcfd8af2aa013d2c59c2b63a452efcc3ce594f18b6be583861d44765677ba419

Observation 84cf2d6b-28d9-4685-8548-b2c5901ac301 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Advances in Neural Information Processing Systems , volume=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.338833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.338833Z digest=sha256:d9567ee881727992986eabe98219804c6efad8bbc42b0c7a37eee86825ea0bec

Observation e79e831e-a13f-4653-abf1-eabac051cf77 · outbound

This paper cites arXiv preprint arXiv:2512.20061 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2512.20061 , year=

Reference 28

Resolution
verified exact
raw_fallback, observed 2026-08-15T15:03:23.289149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T15:03:19.351013Z digest=sha256:f781351149490bf134e25a77ad947ce45f353a1b988d8926dc1d6ee258034ca7

Observation 805ab792-6c08-4804-baff-6260c54b2b9c · outbound

This paper cites 2025 , note =.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning 2025 , note =

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.355489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.355489Z digest=sha256:3168515bbb4ac09053c632cc3ce8acb08fc05b0d859ff03f29512d67ad125a91

Observation 91cf2320-3323-4c02-ab90-991d9d9f2bfd · outbound

This paper cites Qwen2-Audio Technical Report.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Qwen2-Audio Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.360251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.360251Z digest=sha256:d6393e7861d8b45809a0f87227c3c19f46fd30bf0f9f374c17033b43230b7a73

Observation 83f50cdd-9422-48ef-9bf1-25b52589f42d · outbound

This paper cites arXiv preprint arXiv:2512.23808 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2512.23808 , year=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.366027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.366027Z digest=sha256:5ddf6857f2fd85d306b04c7a6ea36b0087e6155c3e7fe1a0360a8091618ac576

Observation 1368ded1-c337-493f-bfda-201846e9122c · outbound

This paper cites GitHub Repository , howpublished =.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning GitHub Repository , howpublished =

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.370901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.370901Z digest=sha256:ad2e1dddb78c47776569d973a03d89de0f349daa51c09e64d18693bcaf7900eb

Observation 96e52be7-700f-40ed-b7bf-8da55657667b · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Advances in Neural Information Processing Systems , volume=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.451588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.451588Z digest=sha256:d96adcc86d8bf368bb3b3a50f1848814403ead0247f928a1f3d2ea7168cfa0ec

Observation f57fddee-fb5b-4140-9948-206b046ed244 · outbound

This paper cites Phi-4 Technical Report.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Phi-4 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.523791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.523791Z digest=sha256:0dccbace689c63347d38e32593624511e78718be28dcecb8bd404e4b717f830b

Observation dbc53e42-853c-4d2b-9912-440442962e9f · outbound

This paper cites Step-Audio 2 Technical Report.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Step-Audio 2 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.543459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.543459Z digest=sha256:159d9987c7a0bd66594b61c6aa4d56c2a5795c9b03c9dad8538bf79feba4b44f

Observation fcd7c6fe-712e-492e-bfe7-f094a964f728 · outbound

This paper cites 2025 , eprint=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning 2025 , eprint=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.556237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.556237Z digest=sha256:d2aba428dd3a148b92133163fcab126c45749c3365b7692025a1f814cff67f10

Observation 3fb62d8f-43a0-49c1-939d-9d3c9cd7c99d · outbound

This paper cites Kimi-Audio Technical Report.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Kimi-Audio Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.561459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.561459Z digest=sha256:5c3ad3b39dd37c0354fc64bf07765ce7b7967a4b46db5f1b413e03f503f6898d

Observation 892b46d4-76e5-4e00-8810-a2539cbca3f9 · outbound

This paper cites Kimi-Audio Technical Report.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Kimi-Audio Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.566176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.566176Z digest=sha256:45b720d86b99d7823984ec42565db40a232e60b1f78df4744d9d428fc74b4553

Observation 062b5297-4454-4015-8984-b51b5b9e74ac · outbound

This paper cites 2023 , url =.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning 2023 , url =

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.571016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.571016Z digest=sha256:4046bd1f682db0f04844b36b31b8a5f3d611a5ddc0f775114cfaf833fa7ddb6b

Observation ed269b75-4fc8-4ee2-8213-14e5013f8b44 · outbound

This paper cites Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.576273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.576273Z digest=sha256:4bdd4608f3827889f482ede56e65a89fe20deab6b81a691a52780240d7d1c11b

Observation 408d63c8-a2f3-40f7-984f-14471f6baf20 · outbound

This paper cites Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.649811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.649811Z digest=sha256:b3a51a4fe046e6ac9db5ccf783d19558c42bf25fcd9aae8af9389b99555a1b47

Observation bf21f3db-143e-4426-9af6-e154cca7e7f6 · outbound

This paper cites Proceedings of the 30th ACM International Conference on Multimedia , pages=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Proceedings of the 30th ACM International Conference on Multimedia , pages=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.700472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.700472Z digest=sha256:b91d61903533b0aeca87ad0162915e30a850f8a6327d1d96cf75f83a0ebb9ca2

Observation 4979d106-6962-4cce-a85d-f05c22290330 · outbound

This paper cites GPT-4o System Card.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning GPT-4o System Card

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.705103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.705103Z digest=sha256:b90a5d7adc5be18c4f746c560b69a8f03420a2388c7545bfaead6e866d6bbc33

Observation bce34aef-7560-453d-b0c9-683dc715a272 · outbound

This paper cites GPT-4 Technical Report.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning GPT-4 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.710688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.710688Z digest=sha256:1d45fb24d4f5b8542ca380067ba9cf830c7caef9865911ee4249fd50a6ed8064

Observation 30e0b09b-1c71-4b5b-a9c7-a616aa2077bc · outbound

This paper cites Wav2CLIP: Learning Robust Audio Representations from Clip , booktitle =.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Wav2CLIP: Learning Robust Audio Representations from Clip , booktitle =

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.758595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.758595Z digest=sha256:789991325c104e93fdac67300e39091c577865fab73e2aefafd178c2cdf592d6

Observation b1a30ca8-d537-4c9c-8ff5-81f697b2f874 · outbound

This paper cites Pengi: An Audio Language Model for Audio Tasks , booktitle =.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Pengi: An Audio Language Model for Audio Tasks , booktitle =

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.852290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.852290Z digest=sha256:aeec64349dfe6132144111ba2e55780bfb006c3ebb3e63e727fa9974be568123

Observation 9828d1af-c9e9-40fd-8c0c-e72a46e73ad2 · outbound

This paper cites Liu and Leonid Karlinsky and James R.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Liu and Leonid Karlinsky and James R

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.892520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.892520Z digest=sha256:3b6badd0ec64601bad7d14a2a68830882094a77b8959836b2e8e9fc91672f073

Observation b78d0eb0-5e1a-4b70-b1e3-1801971d62a0 · outbound

This paper cites Optimal Transport for Treatment Effect Estimation , booktitle =.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Optimal Transport for Treatment Effect Estimation , booktitle =

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.919121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.919121Z digest=sha256:2b434fc4649d5dc11a99434ec5cd401ff547226e70e61fd5ce69fcfd3de5e050

Observation 728c110c-c656-4ac6-aead-94f440a4a37c · outbound

This paper cites A Review for Deep Reinforcement Learning in Atari:Benchmarks, Challenges, and Solutions.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning A Review for Deep Reinforcement Learning in Atari:Benchmarks, Challenges, and Solutions

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-15T15:03:22.979713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T15:03:19.935357Z digest=sha256:2c0d99b61035c6093a895f7de70eb71e13f5ba79e6ec8587a2400540ac6627ff

Observation e9e73edb-7dfe-40b3-ba74-c91d9523a45f · outbound

This paper cites Learnable Behavior Control: Breaking Atari Human World Records via Sample-Efficient Behavior Selection , booktitle =.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Learnable Behavior Control: Breaking Atari Human World Records via Sample-Efficient Behavior Selection , booktitle =

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.948878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.948878Z digest=sha256:0bf53825295314cbdfea686ba04cbcd9af4f16b5055c4b519dcac8b10b98d10f

Observation b07f9220-dfda-4365-92c6-e414a912928a · outbound

This paper cites Generalized Data Distribution Iteration , booktitle =.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Generalized Data Distribution Iteration , booktitle =

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.953538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.953538Z digest=sha256:06298458a6cfc7a5edea31c23a16d97677f4df2a8b37b29e37cbb130ca47b412

Observation 1f805ee3-1663-457a-90ad-8389f43503e5 · outbound

This paper cites GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-15T15:03:22.960912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T15:03:19.958819Z digest=sha256:794d4a2a9dbdb6a0918a14a142bb84da53fbb1c042bd389f09ea8e2990c961ed

Observation ffaa1268-31b3-4a35-a753-2b83caf1f6b4 · outbound

This paper cites CoRR , volume =.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning CoRR , volume =

Reference 53

Resolution
verified exact
doi, observed 2026-08-15T15:03:21.712228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T15:03:19.963566Z digest=sha256:551e3f5f357819c0ccf6cd8dcc6c041ba14d888fb1e1e804fd7544e12ec451ec

Observation 138fa139-ddc6-468a-8ed4-59caedd7aac4 · outbound

This paper cites ConvFormer: Revisiting Transformer for Sequential User Modeling.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning ConvFormer: Revisiting Transformer for Sequential User Modeling

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.968134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.968134Z digest=sha256:6481883305805117b102c1369b817e642b2a305cb0397ee71c8432b92f4ce72e

Observation d41bf1c5-dac8-4837-b86d-d1fa01317ec9 · outbound

This paper cites An Entropy Regularization Free Mechanism for Policy-based Reinforcement Learning.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning An Entropy Regularization Free Mechanism for Policy-based Reinforcement Learning

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-15T15:03:22.865970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T15:03:19.988866Z digest=sha256:b7426bf610970679d7a35c8c12f4c59e7c4131fd13ec596327e2754dae86cb36

Observation 11f073e4-12f2-4547-8378-34affe85b0c5 · outbound

This paper cites CoRR , volume =.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning CoRR , volume =

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.058803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.058803Z digest=sha256:1515bc107c68ca62bd795d74562b904934051f3b3d4f8bdf16bed82e8d109df3

Observation 9ed0d688-2461-4f30-b7b5-801b438fa434 · outbound

This paper cites Critic PI2: Master Continuous Planning via Policy Improvement with Path Integrals and Deep Actor-Critic Reinforcement Learning.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Critic PI2: Master Continuous Planning via Policy Improvement with Path Integrals and Deep Actor-Critic Reinforcement Learning

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-15T15:03:22.728094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T15:03:20.105498Z digest=sha256:3933f6ef3e289c1efc93268556871acf5e0b05c3c09316ed3903f2a4857bdb15

Observation 42050181-ba8e-4490-8b73-1e98df06e938 · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning The Thirteenth International Conference on Learning Representations,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.115850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.115850Z digest=sha256:9df238a06a12b463ab46cd8f642e262a1c775171543cd56b54c38b13816a97a9

Observation 387a6d8f-8aa8-45f4-92a0-01c432cbf2c6 · outbound

This paper cites CoRR , volume =.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning CoRR , volume =

Reference 59

Resolution
verified exact
doi, observed 2026-08-15T15:03:21.519822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T15:03:20.135289Z digest=sha256:24f188ea6a10bcebf1dfffef2d8d63d268eb491e42250c57b79c5339db059612

Observation a87e8a5b-1927-4c1b-b1d2-a852b17fb217 · outbound

This paper cites CoRR , volume =.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning CoRR , volume =

Reference 60

Resolution
verified exact
doi, observed 2026-08-15T15:03:21.373181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T15:03:20.141645Z digest=sha256:c857e51a761a61030ad23187fe0c7199d553a7e532283568e8afc5343c849b71

Observation 756766ed-ab52-499c-a4eb-9979d6b2fe6b · outbound

This paper cites Proximal Policy Optimization Algorithms.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Proximal Policy Optimization Algorithms

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.146090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.146090Z digest=sha256:8e1518ef1a5387f22925e88052b512f6e0f48b791236106f0b700d99e3310df3

Observation f0ea545a-123f-4ed7-b3a7-c1aa43dac51f · outbound

This paper cites CASA: Bridging the Gap between Policy Improvement and Policy Evaluation with Conflict Averse Policy Iteration.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning CASA: Bridging the Gap between Policy Improvement and Policy Evaluation with Conflict Averse Policy Iteration

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-15T15:03:22.574300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T15:03:20.150236Z digest=sha256:35c0abc1a5cba109d557ad350fa86461f38ee3ecc3fe5777bf24a8dda39cca77

Observation b92d5a4c-34ec-4aa9-bae0-66dd32b5265f · outbound

This paper cites PRANCE: Joint Token-Optimization and Structural Channel-Pruning for Adaptive ViT Inference.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning PRANCE: Joint Token-Optimization and Structural Channel-Pruning for Adaptive ViT Inference

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.155097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.155097Z digest=sha256:bb71628f71dcf6e2559479e79beded5e9c49edc86f747697197ed878a2e510c5

Observation ed3c6e28-6c7b-4d0a-bdf1-61f853351175 · outbound

This paper cites The Twelfth International Conference on Learning Representations,.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning The Twelfth International Conference on Learning Representations,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.161486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.161486Z digest=sha256:88652b40091ade86fb5222bea1689f4a25ad261def12ba045c061d62a5ffb846

Observation c7951792-715b-4bd3-9bab-db3d33bbc5fa · outbound

This paper cites Forty-first International Conference on Machine Learning,.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Forty-first International Conference on Machine Learning,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.165992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.165992Z digest=sha256:09dcfc9e91258606e5bde30ae1a16f9ad7ec3cb061039fd20e90182c17b0cbe1

Observation 481d5bb3-0fc2-4202-8a8e-79832571323d · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.170110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.170110Z digest=sha256:37ed21c84e4f2e3ed9689683075148b7999263d0765acf7156bd1db34a0203e1

Observation 1f4173fc-edb0-4d53-8eac-efb12e9a8f99 · outbound

This paper cites Qwen2-Audio Technical Report.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Qwen2-Audio Technical Report

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.198940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.198940Z digest=sha256:37098617158ba5460278295f4e5b3cdfbca5b638e4836f54a6511ce4b6622913

Observation 660e61f1-abc9-44cc-b825-aab9f7c42f1d · outbound

This paper cites Qwen2.5-Omni Technical Report.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Qwen2.5-Omni Technical Report

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.279703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.279703Z digest=sha256:ff2a4e7960962a4e37907b59b687aced38359e447468d0a715fc80ba4e560d7a

Observation a3d8f8ca-0a76-46f2-984a-064371b741c4 · outbound

This paper cites 2025 , eprint=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning 2025 , eprint=

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.351600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.351600Z digest=sha256:023a85b88fb20f835b5b62d606b2f10d933c9fa8d9f41de1fe8b72b60c22bd20

Observation b0489be7-acc5-4149-8819-c40f686a191b · outbound

This paper cites Chi and Quoc V.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Chi and Quoc V

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.360298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.360298Z digest=sha256:ab58f9d95f7699dc8715eb2114d982fcd5d580e6e3a2197a8a9c1e10d3b423bc

Observation 9af4c5fc-1cac-44ca-b38d-c5c61af4733f · outbound

This paper cites VideoCoT: A Video Chain-of-Thought Dataset with Active Annotation Tool.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning VideoCoT: A Video Chain-of-Thought Dataset with Active Annotation Tool

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.377256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.377256Z digest=sha256:73820b3c4896dd6cb027fd01c70e1a8e2167f53df4483085aa8928b88b30f9e9

Observation e69da737-0048-43c4-87f9-0e5d530e99fa · outbound

This paper cites Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.385682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.385682Z digest=sha256:491ab572c32f494d510b68072c6e0fd2a4f57adab20df8ce98d7d48d62f6fe45

Observation ed36e967-d8b3-4e60-9359-74763bee8d0a · outbound

This paper cites arXiv preprint arXiv:2505.09439 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2505.09439 , year=

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.395521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.395521Z digest=sha256:db77f5d54bf210bb81208fa4b383d2c6777e431100841f23aefa9f96cf54b7f7

Observation 1c989b9f-481e-4eaf-8111-9fc5d7675a07 · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:03:24.089602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T15:03:20.400512Z digest=sha256:c4e279be726ca1a5f52473f58e7902a0fe46e033f91f0212c2784ba739882b94

Observation 877eeec4-b2c2-4518-af80-6161986805d6 · outbound

This paper cites CoRR , volume =.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning CoRR , volume =

Reference 76

Resolution
verified exact
doi, observed 2026-08-15T15:03:21.160237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T15:03:20.455039Z digest=sha256:761651ac1e1d3e24ba6532ef4fb83793e1e87f2077ad667807c4c3d9e9744381

Observation ce525799-8429-4255-a847-31f28ecd5dc1 · outbound

This paper cites OpenAI o1 System Card.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning OpenAI o1 System Card

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.510220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.510220Z digest=sha256:1a090bc1f925bcdf74bd56c6ff2bee056c324dc8c8f25f75abb41187a3ad011f

Observation 94cde8fa-d300-442e-a0de-67bbba8cd3f2 · outbound

This paper cites 2025 , eprint=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning 2025 , eprint=

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.555188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.555188Z digest=sha256:9254fc913cf8878ad7461d4fb2bfc22237f7b9e6bdcbac21b97ecc8d3f3151e1

Observation b083773d-80a2-43c4-86ea-b54cf54fa156 · outbound

This paper cites Flow of Reasoning: Training LLMs for Divergent Reasoning with Minimal Examples.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Flow of Reasoning: Training LLMs for Divergent Reasoning with Minimal Examples

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.562839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.562839Z digest=sha256:339e9fe51a959a3b41ef0019358a700c96e0c361dbf8ae7b57cba98c14b65b16

Observation 389babb4-8067-4623-94f0-e6a98e2176cc · outbound

This paper cites ArrowGEV: Grounding Events in Video via Learning the Arrow of Time.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning ArrowGEV: Grounding Events in Video via Learning the Arrow of Time

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.568224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.568224Z digest=sha256:916a83caba7141f1ecfa9cee5dc633a3fc25e8d419e3b3d59af293113165cae9

Observation 65d8f06a-9ef3-4083-85a6-64e54213076d · outbound

This paper cites arXiv preprint arXiv:2601.04171 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2601.04171 , year=

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.573032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.573032Z digest=sha256:fd39ca22af1b0b7cb6d9726c51bd9cfd427f80fcedcb943e7b4438c13e2b9491

Observation 0026c232-25bd-40df-9311-566675c39b47 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Advances in Neural Information Processing Systems , volume=

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.577377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.577377Z digest=sha256:c020c737d38f4a6a163e4a33bc11d2e631aa9cba6ee70d7c3cb1aa2bda9a40dd

Observation 652fe124-886a-49a2-a202-f4243add5ad9 · outbound

This paper cites arXiv preprint arXiv:2511.12344 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2511.12344 , year=

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.581542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.581542Z digest=sha256:b048a2b21949bdd8410299baf6968a14c64d7b20c568fa1ab3148d925c4dea47

Observation 602e71bc-99a6-4250-b101-c9e970599ba2 · outbound

This paper cites Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.621494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.621494Z digest=sha256:c240660d4d418eda5849a6bde849eb24d2574a8ec80bcedbd7623118f1e97d94

Observation a17686b4-faa4-4dbe-a94b-175b59ec8dd1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.696940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.696940Z digest=sha256:112c0fa250cc01b01dfb44abb779f0cbfe79f727ef78a3b795d478d8ce9641b2

Observation 0fd3169b-a755-4720-9a19-c79a79d28949 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.707594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.707594Z digest=sha256:1430a238c1c55600be4812185f2ef0a39b43e851b0b5a708b52e0db9009ab6d3

Observation 1827b252-7810-415f-8faa-4cc420b239e4 · outbound

This paper cites Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.714065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.714065Z digest=sha256:3759ff266effefcd309afaec99cba5bfa70e35d47f37b845a7da759e4e3952dc

Observation df1668f8-2120-411b-aaae-29d81836699a · outbound

This paper cites MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.718955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.718955Z digest=sha256:284fd58aa63843f3d6cb4a739e54a8464b505b625d7c837383e1c0ce41f60668

Observation b9eb6e47-3bb7-486b-8094-dbbdd94d5ef3 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:03:24.058821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T15:03:20.723150Z digest=sha256:ed8a1bd05fefc6c4c94ca7b2699e2bbe5816b56c26d3151d1ca079f58550efa7

Observation 748f3498-457e-4e52-b56b-997bedb389a7 · outbound

This paper cites Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.727700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.727700Z digest=sha256:34d1bab4f8a374ac4fb68129d13458f1af384cd754d3d5afc69fd657871cde51

Observation 4c0b6bb3-54e3-4b8e-91e8-fdde00dd4f61 · outbound

This paper cites arXiv preprint arXiv:2510.11454 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2510.11454 , year=

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.732762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.732762Z digest=sha256:60c64f40de5f3770b82ff956343ae0b8116a986c2c6b81a9ca44876486fbb374

Observation 237af61d-854e-4b40-8c41-d0306751dc8b · outbound

This paper cites arXiv preprint arXiv:2602.13685 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2602.13685 , year=

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.738254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.738254Z digest=sha256:fdf70a88532051530c6128a26fc254e304288ee6cf43cbf25d3e3428f5487c1d

Observation b1ed2f30-6bc0-41c1-8dea-40b86a3f9ffb · outbound

This paper cites arXiv preprint arXiv:2602.10439 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2602.10439 , year=

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.755835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.755835Z digest=sha256:6e0e2fe5ad0db28b18d69073e00cffd204071e3b54489ca8f8daee772a11ebd0

Observation e01833b1-7489-43d6-860f-65c49c36f36e · outbound

This paper cites arXiv preprint arXiv:2511.15848 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2511.15848 , year=

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.836254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.836254Z digest=sha256:ba959108805501ac5736ffb4968891d6868419f50f2853a66fa88cd4a17dd61e

Observation eb4a273d-1e96-4cd6-a3b7-815e029e6bac · outbound

This paper cites arXiv preprint arXiv:2510.20867 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2510.20867 , year=

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.940583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.940583Z digest=sha256:a977766d2ec9714e8ea5ffd81d7b9552b9ae6d014473c13498d5b680476aa167

Observation 5abab500-ae57-464d-a697-239e877fa343 · outbound

This paper cites ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.950109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.950109Z digest=sha256:4aa819c8c8bf6d9bf66c489095b65ccc058f509910a5573fd24b0b7ec1446ed2

Observation dd5cdea5-df55-44ed-b1d4-650180864af5 · outbound

This paper cites International conference on machine learning , pages=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning International conference on machine learning , pages=

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.956623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.956623Z digest=sha256:28b69e42406b3419163477211d0cbb0d79bebd1fb5d189aed092e1f02131247d

Observation 8c6403d2-eafe-41d5-b5b6-83de8429d7cb · outbound

This paper cites ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.960816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.960816Z digest=sha256:0266bf9956638e720c241659e79a7a66b0889bd49ff3c51346b972e209ed9bfe

Observation da89599d-4536-4118-9ac9-de4e36b77b62 · outbound

This paper cites arXiv preprint arXiv:2503.02318 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2503.02318 , year=

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.965844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.965844Z digest=sha256:e39f412e5f479f906f4d3723cfe6cf924f0344eb1a98a88610f61623e7fec6b5

Observation 19c6c492-c904-440e-8f5e-f348a5ff0837 · outbound

This paper cites MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.969779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.969779Z digest=sha256:64308184bc48581a0dff245a2a543c7e8224a3854466b70ab708d3f22f259545

Observation b80465ce-e474-44f0-a20b-6006416d3654 · outbound

This paper cites MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.975265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.975265Z digest=sha256:60b8ed2a6cfd73a210ce42575d9604744ca845b52f20b54ba51e772a494be239

Pith citing papers

No inbound Pith citation observations are available.