Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning

As of 16 August 2026, this Paper Citation Record lists 100 of 100 outbound references and 0 inbound Pith citation observations for arXiv:2608.02831.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02831 v1

Coverage vector

measured 100 of 100 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:03:20.975265Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 100 outbound references displayed

  • verified exact11
  • verified fuzzy2
  • unresolved87
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 72216fe6-495e-419e-904f-05bda35caf94 · outbound

This paper cites 2023 , url =.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning 2023 , url =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.581426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.581426Z digest=sha256:cbb2210d55dbba47c83afe236751f9716f50429b66130c41cc1f128d3cc19d90

Observation 71ca77b9-c836-42c1-87dd-248d0c611006 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Advances in Neural Information Processing Systems , volume=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.625181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.625181Z digest=sha256:b04af26c4f15931bf35e1b849610e721ba6de507dc527a0c0d290d54fcbf6786

Observation 792fb518-733e-4133-bf11-e4b9fb2ddb5d · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Movie Gen: A Cast of Media Foundation Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.649697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.649697Z digest=sha256:2c0eb24e52af4d98992392259dc98cce77317d67819a8f0736a27b01d544117a

Observation 6b351283-fb42-4f08-9ae0-69dce7cdb245 · outbound

This paper cites Towards Conversational Medical AI with Eyes, Ears and a Voice.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Towards Conversational Medical AI with Eyes, Ears and a Voice

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.661572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.661572Z digest=sha256:5434bbfea817aa7a2c6d58d1810f9b61cdf1500e646defffdaa158b349747255

Observation b192ec25-585c-4b2a-98b4-41192c1b3902 · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2025 , pages=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Findings of the Association for Computational Linguistics: ACL 2025 , pages=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.666574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.666574Z digest=sha256:2f4f194aa9f7d6bcc163e9f5463027e4abb34ad389e380c636e289e66a9438ae

Observation 7381af69-928f-4fd3-8a28-94607cb518dd · outbound

This paper cites International Conference on Learning Representations , volume=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning International Conference on Learning Representations , volume=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.671105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.671105Z digest=sha256:f323256b61c988a05f1971ce9be0dd21fadc4dbeed60edcb85bc0d1512212849

Observation bc2b904c-0fd8-4a08-8563-f4a70426a18c · outbound

This paper cites Reward Hacking in Rubric-Based Reinforcement Learning.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Reward Hacking in Rubric-Based Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.709576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.709576Z digest=sha256:5c6b0b479cccec760012f187c8af50497043e04d27049b271c0a63617b3a8dec

Observation d8664f22-1f09-40cd-ac1d-461a833847cb · outbound

This paper cites arXiv preprint arXiv:2602.05125 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2602.05125 , year=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.779868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.779868Z digest=sha256:813d68ce9688fac0267d37d1c2b4ea2aed43c19c01c4efabe46b2ba886cc0159

Observation 916fd730-9271-4eb6-870c-6a209d65ee9c · outbound

This paper cites Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.825053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.825053Z digest=sha256:77ce8189aff1158b95cfc87d55dc2429de7b60011126c0d067bcb858b3bb3bce

Observation b8a86bb6-d29c-45c5-b3e6-ccfa8dd20dac · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.844112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.844112Z digest=sha256:a8419fef1ca6843207806b1b733e88bf6d00bb442bf70372343b9e0c2ff7218b

Observation 53f5c82c-5342-4939-976b-994ae60b48b5 · outbound

This paper cites Weak-to-Strong On-Policy Distillation.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Weak-to-Strong On-Policy Distillation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.871173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.871173Z digest=sha256:93ced3dc83a510f0db22ed7af412a82b600b9cac7fa35c11d60086fde903d32d

Observation c5d69191-b1f5-4aa5-9a5b-2f7c6f16fc0e · outbound

This paper cites arXiv preprint arXiv:2602.21628 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2602.21628 , year=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.876157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.876157Z digest=sha256:f23af89f5b589aa72710031b01d3133b4d7669a1fa4ea5cde40eaa80a85f1c0d

Observation 6cf13027-e7d0-4867-ad6c-464d6450b724 · outbound

This paper cites Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.880654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.880654Z digest=sha256:6a15c529dbd0b50eb45a5a75facb003b2bb055d581ee89590494d6f2148752c2

Observation 0508a662-86ec-45ba-b422-fbde282f5187 · outbound

This paper cites arXiv preprint arXiv:2603.16600 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2603.16600 , year=

Reference 14

Resolution
verified exact
raw_fallback, observed 2026-08-15T15:03:23.732358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T15:03:18.885146Z digest=sha256:a3a63776e63483183d85cae8fe05a5e8fd80fbaf3415c0480dce19628bb6f0fa

Observation bf9335c5-335b-4e01-9de7-8a82f72280bc · outbound

This paper cites Visual Preference Optimization with Rubric Rewards.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Visual Preference Optimization with Rubric Rewards

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.914108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.914108Z digest=sha256:d04b08bb2c7f4b999462b9e706c5f61939d7fe04d5fdd88652863d3cd180a1ad

Observation e3bb04da-0188-4c5e-81ba-9e91f6e90658 · outbound

This paper cites AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.008040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.008040Z digest=sha256:fa19ff22592c188a645c39c4a18b3bacdc805ac0e34fbc6964329795e1f1aecb

Observation a5403e7e-d75e-4a25-ba2c-3ccadb4a6018 · outbound

This paper cites arXiv preprint arXiv:2602.04649 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2602.04649 , year=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.084775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.084775Z digest=sha256:0a7cbbe7a8c969489237b35ef6194c44a45b341d46a7008f93aaa89b2df6c117

Observation a2bce307-639f-4de9-b3f6-da494af5e87b · outbound

This paper cites arXiv preprint arXiv:2602.01511 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2602.01511 , year=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.110231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.110231Z digest=sha256:20744ad9078d65b9d2921229baefff1e80c882c3377cbbc04a090f3af5c1432b

Observation 02d019f0-b7a9-4c4b-87a5-92cb6fe569b8 · outbound

This paper cites arXiv preprint arXiv:2510.07284 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2510.07284 , year=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.114984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.114984Z digest=sha256:2429d2f1e46ecfd0178100301c551d4f119b5aafe01826002ca2ad7ba589d365

Observation b46170ed-2603-4e1a-b81a-36d0e91c1504 · outbound

This paper cites arXiv preprint arXiv:2602.10885 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2602.10885 , year=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.119412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.119412Z digest=sha256:bc20fa431f0bb41ed98e04403c960bc78acc5eda3c9f75c25b16a46e03247906

Observation 9d8fc2a4-822b-4370-95d0-d656a24c9c3c · outbound

This paper cites DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.123830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.123830Z digest=sha256:6209f489379e5f88e39aa101ac35deea89cfb87b13083818a0654fa02ad8bcbd

Observation f8766714-db81-404c-bcc1-0bb3da494bfd · outbound

This paper cites Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.128691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.128691Z digest=sha256:4df231c9b43de0e227e83661f3639cd845d8cf2efd295a5860f76c02cbec2ead

Observation e08e7ca9-3d77-45e1-9500-41b955f68030 · outbound

This paper cites arXiv e-prints , pages=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv e-prints , pages=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.134675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.134675Z digest=sha256:14365e56e34880d0208e80166cb4a03a5f4392fe1efff5191e40a2c6d7cfee88

Observation a4458c19-3996-47b3-b819-7272e8cdea8f · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2025 , pages=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Findings of the Association for Computational Linguistics: ACL 2025 , pages=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.209889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.209889Z digest=sha256:cdf97894b7bd9631373f732d068c1a0f781538cfd58ff93708f6a166fa2ad806

Observation 755ce286-98ea-412e-8f11-4c2f041745b4 · outbound

This paper cites arXiv preprint arXiv:2510.07743 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2510.07743 , year=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.286838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.286838Z digest=sha256:a425bbae8c9bd0e2abb27bef84536ea50e4ed0a5b4856ae3c0224eebac737d34

Observation 10f29666-f0fb-4867-9216-f43af663fb32 · outbound

This paper cites International Conference on Learning Representations , volume=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning International Conference on Learning Representations , volume=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.312041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.312041Z digest=sha256:a77d67194ca7aefabf2b9f5061ff16ceff0d36115d100fa95bf087b8fa3315b9

Observation 84cf2d6b-28d9-4685-8548-b2c5901ac301 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Advances in Neural Information Processing Systems , volume=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.338833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.338833Z digest=sha256:ea5b5d77753b268e5de1e4a7934ad837d90c6bfca9e51dd43264bdb4d9623dab

Observation e79e831e-a13f-4653-abf1-eabac051cf77 · outbound

This paper cites arXiv preprint arXiv:2512.20061 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2512.20061 , year=

Reference 28

Resolution
verified exact
raw_fallback, observed 2026-08-15T15:03:23.289149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T15:03:19.351013Z digest=sha256:2e4dc974ba645da754e64b81adb38f735cc435364cbe0e4b53e9fd4bb5879ec2

Observation 805ab792-6c08-4804-baff-6260c54b2b9c · outbound

This paper cites 2025 , note =.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning 2025 , note =

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.355489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.355489Z digest=sha256:c120aa82e4e61ee98998626b89946855752013b912d0bd8dc13f00aad276c8ba

Observation 91cf2320-3323-4c02-ab90-991d9d9f2bfd · outbound

This paper cites Qwen2-Audio Technical Report.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Qwen2-Audio Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.360251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.360251Z digest=sha256:02d7a08afdff9854c2ed5b91f8cf20cb9fe5e3689c78d67147d1cb3898e084d0

Observation 83f50cdd-9422-48ef-9bf1-25b52589f42d · outbound

This paper cites arXiv preprint arXiv:2512.23808 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2512.23808 , year=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.366027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.366027Z digest=sha256:b922f69cb6c0c2829661dc7b389cb283cf71c0ec4e27390dff3c5288c473f7c1

Observation 1368ded1-c337-493f-bfda-201846e9122c · outbound

This paper cites GitHub Repository , howpublished =.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning GitHub Repository , howpublished =

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.370901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.370901Z digest=sha256:fa701e494266dab8fe40c0b07041c468056e7f232923e1d20e714ad979a5a9ab

Observation 96e52be7-700f-40ed-b7bf-8da55657667b · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Advances in Neural Information Processing Systems , volume=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.451588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.451588Z digest=sha256:235f15b5a464bd0dbf3a90fad2025e9769a54a36269e72a77a22d6ff1ab1deef

Observation f57fddee-fb5b-4140-9948-206b046ed244 · outbound

This paper cites Phi-4 Technical Report.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Phi-4 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.523791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.523791Z digest=sha256:84cdbc4cc12502a3e62c8a117275c6a328f9b5034739e7dfdbc5d5b4fa2371be

Observation dbc53e42-853c-4d2b-9912-440442962e9f · outbound

This paper cites Step-Audio 2 Technical Report.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Step-Audio 2 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.543459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.543459Z digest=sha256:681f68362306fb88eb3d883f00213e26be660c68697ef24ec3b26d9759a7d307

Observation fcd7c6fe-712e-492e-bfe7-f094a964f728 · outbound

This paper cites 2025 , eprint=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning 2025 , eprint=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.556237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.556237Z digest=sha256:3e6d517bafaff9cba3025ddc695a0cddda689bbc683de34bd779234d82628e30

Observation 3fb62d8f-43a0-49c1-939d-9d3c9cd7c99d · outbound

This paper cites Kimi-Audio Technical Report.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Kimi-Audio Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.561459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.561459Z digest=sha256:cd89ef59b7e2fe4f1345407eb5d95de2be99e073d586e4d8cbfebb033561f351

Observation 892b46d4-76e5-4e00-8810-a2539cbca3f9 · outbound

This paper cites Kimi-Audio Technical Report.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Kimi-Audio Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.566176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.566176Z digest=sha256:b47fbd9cb16cf0c9654e028a5a7b00a500115df022008177f91b17a042830259

Observation 062b5297-4454-4015-8984-b51b5b9e74ac · outbound

This paper cites 2023 , url =.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning 2023 , url =

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.571016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.571016Z digest=sha256:ad489cdae20c2fc431cd906b32c5b974b3ef0751efcb038d2f62b1f2bab79e87

Observation ed269b75-4fc8-4ee2-8213-14e5013f8b44 · outbound

This paper cites Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.576273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.576273Z digest=sha256:ab1099395052c6570eb5503b4ae0bb86aaf9125604a494665ae23d93e75c62f5

Observation 408d63c8-a2f3-40f7-984f-14471f6baf20 · outbound

This paper cites Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.649811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.649811Z digest=sha256:3b8a03e586e933977f3005b177b901ff7dd09a7d1425effbc3c4aa86abc91a1b

Observation bf21f3db-143e-4426-9af6-e154cca7e7f6 · outbound

This paper cites Proceedings of the 30th ACM International Conference on Multimedia , pages=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Proceedings of the 30th ACM International Conference on Multimedia , pages=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.700472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.700472Z digest=sha256:b48da2d930b0e2bea64bac3a2e6926ae6911f35bebdbb1525f3fcf31df32d1ae

Observation 4979d106-6962-4cce-a85d-f05c22290330 · outbound

This paper cites GPT-4o System Card.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning GPT-4o System Card

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.705103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.705103Z digest=sha256:e6a3a7daa292c486141f9ccdb53caf679bf5f58689f744a89e05c09ae29f82ba

Observation bce34aef-7560-453d-b0c9-683dc715a272 · outbound

This paper cites GPT-4 Technical Report.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning GPT-4 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.710688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.710688Z digest=sha256:c9975250f76ec4ac75b985993d3da24305c13d0fa76208ae7f47f67f832158ee

Observation 30e0b09b-1c71-4b5b-a9c7-a616aa2077bc · outbound

This paper cites Wav2CLIP: Learning Robust Audio Representations from Clip , booktitle =.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Wav2CLIP: Learning Robust Audio Representations from Clip , booktitle =

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.758595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.758595Z digest=sha256:a96c674324ec2938c79ab5a6379d53f78e8c8443cb2ceecd66eed57325f83589

Observation b1a30ca8-d537-4c9c-8ff5-81f697b2f874 · outbound

This paper cites Pengi: An Audio Language Model for Audio Tasks , booktitle =.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Pengi: An Audio Language Model for Audio Tasks , booktitle =

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.852290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.852290Z digest=sha256:bb63c2be0ade6fe900d6cadcd4e49d51a826bd04510485873e5d889425ecd852

Observation 9828d1af-c9e9-40fd-8c0c-e72a46e73ad2 · outbound

This paper cites Liu and Leonid Karlinsky and James R.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Liu and Leonid Karlinsky and James R

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.892520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.892520Z digest=sha256:93f58b6c629329c4cf0797ffd96c846fce1f01c059e54ed6b7c71318cd007164

Observation b78d0eb0-5e1a-4b70-b1e3-1801971d62a0 · outbound

This paper cites Optimal Transport for Treatment Effect Estimation , booktitle =.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Optimal Transport for Treatment Effect Estimation , booktitle =

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.919121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.919121Z digest=sha256:282df5a4fb28f64d13cc50645d9c037d358fe4fdccdf827afa3bf3cbe775661d

Observation 728c110c-c656-4ac6-aead-94f440a4a37c · outbound

This paper cites A Review for Deep Reinforcement Learning in Atari:Benchmarks, Challenges, and Solutions.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning A Review for Deep Reinforcement Learning in Atari:Benchmarks, Challenges, and Solutions

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-15T15:03:22.979713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T15:03:19.935357Z digest=sha256:b882c034f5e95fc8d8c0df2841a1198fe846afa6f014610510aca1ef7e3872c2

Observation e9e73edb-7dfe-40b3-ba74-c91d9523a45f · outbound

This paper cites Learnable Behavior Control: Breaking Atari Human World Records via Sample-Efficient Behavior Selection , booktitle =.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Learnable Behavior Control: Breaking Atari Human World Records via Sample-Efficient Behavior Selection , booktitle =

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.948878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.948878Z digest=sha256:c0c5aef217335ebb8d050c608572c1edb87ef50ff6bf09e7b8a955fa3e5bca6a

Observation b07f9220-dfda-4365-92c6-e414a912928a · outbound

This paper cites Generalized Data Distribution Iteration , booktitle =.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Generalized Data Distribution Iteration , booktitle =

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.953538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.953538Z digest=sha256:f40d7b34e7b05730d4c8bba7ea599c19e70ce2642e3cf3c9b196716fed49dc8e

Observation 1f805ee3-1663-457a-90ad-8389f43503e5 · outbound

This paper cites GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-15T15:03:22.960912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T15:03:19.958819Z digest=sha256:6e89c484c3bd28ba99aaa988933577e9a93d66c9e31e14731810f999215d4d34

Observation ffaa1268-31b3-4a35-a753-2b83caf1f6b4 · outbound

This paper cites CoRR , volume =.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning CoRR , volume =

Reference 53

Resolution
verified exact
doi, observed 2026-08-15T15:03:21.712228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T15:03:19.963566Z digest=sha256:0453f1e4aade28f1ee248a858a7a1d6c17284e578a4cb232c5ea764143558828

Observation 138fa139-ddc6-468a-8ed4-59caedd7aac4 · outbound

This paper cites ConvFormer: Revisiting Transformer for Sequential User Modeling.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning ConvFormer: Revisiting Transformer for Sequential User Modeling

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.968134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.968134Z digest=sha256:8bc51ee0ba84a41ffd120fddb2899c0881c41b6fa826d41a7790528954ffb8c6

Observation d41bf1c5-dac8-4837-b86d-d1fa01317ec9 · outbound

This paper cites An Entropy Regularization Free Mechanism for Policy-based Reinforcement Learning.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning An Entropy Regularization Free Mechanism for Policy-based Reinforcement Learning

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-15T15:03:22.865970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T15:03:19.988866Z digest=sha256:efa68b2876aaa5a26c863963f4cd3a66cbbec49ffce75d2443ae0142782e2f13

Observation 11f073e4-12f2-4547-8378-34affe85b0c5 · outbound

This paper cites CoRR , volume =.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning CoRR , volume =

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.058803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.058803Z digest=sha256:6fc1972fce2b8eca62cf4990aa7c98942f1397aa3a885e1ad9ea105757e09074

Observation 9ed0d688-2461-4f30-b7b5-801b438fa434 · outbound

This paper cites Critic PI2: Master Continuous Planning via Policy Improvement with Path Integrals and Deep Actor-Critic Reinforcement Learning.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Critic PI2: Master Continuous Planning via Policy Improvement with Path Integrals and Deep Actor-Critic Reinforcement Learning

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-15T15:03:22.728094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T15:03:20.105498Z digest=sha256:dc771fcbdc8ddccf7f3fd75d8bcf70ec57c78c50d47fc64a5e0ca20fb70ab1b5

Observation 42050181-ba8e-4490-8b73-1e98df06e938 · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning The Thirteenth International Conference on Learning Representations,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.115850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.115850Z digest=sha256:6a0c9feb5b5d19a1922236959f390bb9a0a93855b2403718bf49b4eda9b95207

Observation 387a6d8f-8aa8-45f4-92a0-01c432cbf2c6 · outbound

This paper cites CoRR , volume =.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning CoRR , volume =

Reference 59

Resolution
verified exact
doi, observed 2026-08-15T15:03:21.519822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T15:03:20.135289Z digest=sha256:f9ff2de1eb5b29c22d722816bc28f1ae012e8f06691239472243183902e424f3

Observation a87e8a5b-1927-4c1b-b1d2-a852b17fb217 · outbound

This paper cites CoRR , volume =.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning CoRR , volume =

Reference 60

Resolution
verified exact
doi, observed 2026-08-15T15:03:21.373181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T15:03:20.141645Z digest=sha256:07dce4a2ac12e5a4b8f1c66bf699b137dc08053f4940aea3e4aaeaf4abce50af

Observation 756766ed-ab52-499c-a4eb-9979d6b2fe6b · outbound

This paper cites Proximal Policy Optimization Algorithms.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Proximal Policy Optimization Algorithms

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.146090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.146090Z digest=sha256:f9f61fb079b4f30e32b98db295c1c0908dd1a162c12860022a4bebc2c7136c13

Observation f0ea545a-123f-4ed7-b3a7-c1aa43dac51f · outbound

This paper cites CASA: Bridging the Gap between Policy Improvement and Policy Evaluation with Conflict Averse Policy Iteration.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning CASA: Bridging the Gap between Policy Improvement and Policy Evaluation with Conflict Averse Policy Iteration

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-15T15:03:22.574300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T15:03:20.150236Z digest=sha256:8c8f8f141a58379e4f013a0e0ec1bb631402d70bde58e3303e3773cfc1b5e743

Observation b92d5a4c-34ec-4aa9-bae0-66dd32b5265f · outbound

This paper cites PRANCE: Joint Token-Optimization and Structural Channel-Pruning for Adaptive ViT Inference.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning PRANCE: Joint Token-Optimization and Structural Channel-Pruning for Adaptive ViT Inference

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.155097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.155097Z digest=sha256:1da671d88dab245bd8245637cf2357f0df5c4012cd9f10beb91357e047e2aaa3

Observation ed3c6e28-6c7b-4d0a-bdf1-61f853351175 · outbound

This paper cites The Twelfth International Conference on Learning Representations,.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning The Twelfth International Conference on Learning Representations,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.161486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.161486Z digest=sha256:7f27c4d75260489ce5857187d527e5a6f3cd68fd10087fb3247d84c03820a06a

Observation c7951792-715b-4bd3-9bab-db3d33bbc5fa · outbound

This paper cites Forty-first International Conference on Machine Learning,.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Forty-first International Conference on Machine Learning,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.165992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.165992Z digest=sha256:e7ffec75b1fab5746a91ba1740c1422655fb65c9cc91cca6a3088728c10d5546

Observation 481d5bb3-0fc2-4202-8a8e-79832571323d · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.170110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.170110Z digest=sha256:9d0e9710218aecf9d08a85eef44d2d1fe994f1e4779b1433d991d204cbf408d5

Observation 1f4173fc-edb0-4d53-8eac-efb12e9a8f99 · outbound

This paper cites Qwen2-Audio Technical Report.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Qwen2-Audio Technical Report

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.198940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.198940Z digest=sha256:1527e70cd32899c5c869b3ff71ee0f2d56a2a822586e6c13e1436574c621492c

Observation 660e61f1-abc9-44cc-b825-aab9f7c42f1d · outbound

This paper cites Qwen2.5-Omni Technical Report.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Qwen2.5-Omni Technical Report

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.279703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.279703Z digest=sha256:0ea2301255e91be1950f61c0308f0ff7e5b825268843488e03d6b6aa36a984fa

Observation a3d8f8ca-0a76-46f2-984a-064371b741c4 · outbound

This paper cites 2025 , eprint=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning 2025 , eprint=

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.351600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.351600Z digest=sha256:85e7d33f953f598ea66a2e09ba8425f54cdecbf5a1cc5c55f9c654a15a1137d9

Observation b0489be7-acc5-4149-8819-c40f686a191b · outbound

This paper cites Chi and Quoc V.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Chi and Quoc V

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.360298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.360298Z digest=sha256:00c75687ec9b6acafdad7000748c544ea1390ab17be6b85ddc06830f5a0b906a

Observation 9af4c5fc-1cac-44ca-b38d-c5c61af4733f · outbound

This paper cites VideoCoT: A Video Chain-of-Thought Dataset with Active Annotation Tool.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning VideoCoT: A Video Chain-of-Thought Dataset with Active Annotation Tool

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.377256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.377256Z digest=sha256:54743cde492b71571185052cfed3af15755586c0acc1e518491e1a772f3dec8f

Observation e69da737-0048-43c4-87f9-0e5d530e99fa · outbound

This paper cites Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.385682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.385682Z digest=sha256:56d1b32e298ead0397f20ec747177a517c44da062e91bd0da6f1d45aacbbdce5

Observation ed36e967-d8b3-4e60-9359-74763bee8d0a · outbound

This paper cites arXiv preprint arXiv:2505.09439 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2505.09439 , year=

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.395521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.395521Z digest=sha256:1cf23749b5ced3f51d9c22bdc56ce3d08660cccac7f84a735a5f9f2fabdd3e86

Observation 1c989b9f-481e-4eaf-8111-9fc5d7675a07 · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:03:24.089602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T15:03:20.400512Z digest=sha256:1b396235f9f77e4455d81631a43a0d42dad0afe41c292ca4f010ab93b49b777e

Observation 877eeec4-b2c2-4518-af80-6161986805d6 · outbound

This paper cites CoRR , volume =.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning CoRR , volume =

Reference 76

Resolution
verified exact
doi, observed 2026-08-15T15:03:21.160237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T15:03:20.455039Z digest=sha256:d9f9892397550443056c55fc37bd5025df500e916642e529a54759788c2e4acb

Observation ce525799-8429-4255-a847-31f28ecd5dc1 · outbound

This paper cites OpenAI o1 System Card.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning OpenAI o1 System Card

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.510220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.510220Z digest=sha256:18b55c3a2fd4b82149729b6c3d69324ccce2b4bb1b7118e6a9d83b9de67ae038

Observation 94cde8fa-d300-442e-a0de-67bbba8cd3f2 · outbound

This paper cites 2025 , eprint=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning 2025 , eprint=

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.555188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.555188Z digest=sha256:e8af1486e5e49b375bd79ce3a677372eab7cc37f54a1f7af2cbba74146f6b0bd

Observation b083773d-80a2-43c4-86ea-b54cf54fa156 · outbound

This paper cites Flow of Reasoning: Training LLMs for Divergent Reasoning with Minimal Examples.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Flow of Reasoning: Training LLMs for Divergent Reasoning with Minimal Examples

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.562839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.562839Z digest=sha256:7dd684eaf5fac5530ece916b381410602e5b97ee88b5cf71721c6e36575922f4

Observation 389babb4-8067-4623-94f0-e6a98e2176cc · outbound

This paper cites ArrowGEV: Grounding Events in Video via Learning the Arrow of Time.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning ArrowGEV: Grounding Events in Video via Learning the Arrow of Time

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.568224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.568224Z digest=sha256:9d3fc8a5378fd8fdf89f74955d4f27126e16b826132cb6df901b6780aca322d9

Observation 65d8f06a-9ef3-4083-85a6-64e54213076d · outbound

This paper cites arXiv preprint arXiv:2601.04171 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2601.04171 , year=

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.573032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.573032Z digest=sha256:0ebd7c4be50e8e89dddff4077a31a29b3f7a0e337a5ed210691f3df7786aa226

Observation 0026c232-25bd-40df-9311-566675c39b47 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Advances in Neural Information Processing Systems , volume=

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.577377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.577377Z digest=sha256:e61dca1245ebc324479e6eb111661ba76db28cf388746a3bf2525ccd2ca3814f

Observation 652fe124-886a-49a2-a202-f4243add5ad9 · outbound

This paper cites arXiv preprint arXiv:2511.12344 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2511.12344 , year=

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.581542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.581542Z digest=sha256:8653b8e181bd9145ee34e0f8916f4d57cf2f8a35c51a9b860bdc0d5d2c74044a

Observation 602e71bc-99a6-4250-b101-c9e970599ba2 · outbound

This paper cites Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.621494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.621494Z digest=sha256:5d70f84c8bcc536dcd9ec784e9164ad252f129ad86de0db1dccdf6186f1fe34e

Observation a17686b4-faa4-4dbe-a94b-175b59ec8dd1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.696940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.696940Z digest=sha256:933223f8035a7b687c665c0f756535da1645ce800f4b2d9c0fad763f460dd75a

Observation 0fd3169b-a755-4720-9a19-c79a79d28949 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.707594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.707594Z digest=sha256:0e252b3eb2e255a7a74f091d35467f204a2816af78daace6417dba34387cac3b

Observation 1827b252-7810-415f-8faa-4cc420b239e4 · outbound

This paper cites Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.714065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.714065Z digest=sha256:4447130e7494dba1c2caa362695cd8234e157945f097fa611f8b9f275a946af8

Observation df1668f8-2120-411b-aaae-29d81836699a · outbound

This paper cites MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.718955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.718955Z digest=sha256:a703b234b61aea1a67174f126ddda2897ec833d92d52d19bf5857781085faefd

Observation b9eb6e47-3bb7-486b-8094-dbbdd94d5ef3 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:03:24.058821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T15:03:20.723150Z digest=sha256:d7f53e74666e4dc98012e495036da77c669f46b04ceed154ec948386571db87a

Observation 748f3498-457e-4e52-b56b-997bedb389a7 · outbound

This paper cites Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.727700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.727700Z digest=sha256:1a1c1f3bfee081c9dbbe1d7e24ecc36f5bdbf59fd41233a8996b1b0979435d41

Observation 4c0b6bb3-54e3-4b8e-91e8-fdde00dd4f61 · outbound

This paper cites arXiv preprint arXiv:2510.11454 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2510.11454 , year=

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.732762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.732762Z digest=sha256:d6f52efdfa22c9f40af5bc3978d72c7b11e776aa9186441249c42aa6ff4477ed

Observation 237af61d-854e-4b40-8c41-d0306751dc8b · outbound

This paper cites arXiv preprint arXiv:2602.13685 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2602.13685 , year=

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.738254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.738254Z digest=sha256:3aef31a0a18337604f600a461c2dc7e393d6a2eaf25dd40a4fdb1621bd71282f

Observation b1ed2f30-6bc0-41c1-8dea-40b86a3f9ffb · outbound

This paper cites arXiv preprint arXiv:2602.10439 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2602.10439 , year=

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.755835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.755835Z digest=sha256:1d63ecb72073ee969519c484554f3bd3b366005d537f002fb2eef86c8d74ef67

Observation e01833b1-7489-43d6-860f-65c49c36f36e · outbound

This paper cites arXiv preprint arXiv:2511.15848 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2511.15848 , year=

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.836254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.836254Z digest=sha256:82bac05277c68316fd6fe98dec4b8dfbbfe3bc03b36eb41fa5b19441b0f0e2e6

Observation eb4a273d-1e96-4cd6-a3b7-815e029e6bac · outbound

This paper cites arXiv preprint arXiv:2510.20867 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2510.20867 , year=

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.940583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.940583Z digest=sha256:408d17cf8079f2db68181374f0853d091486df5490a141b596530f89c0acd60c

Observation 5abab500-ae57-464d-a697-239e877fa343 · outbound

This paper cites ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.950109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.950109Z digest=sha256:4c082d25d301229b0e5995879070ec16de2476958280a00149b9673407cf704f

Observation dd5cdea5-df55-44ed-b1d4-650180864af5 · outbound

This paper cites International conference on machine learning , pages=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning International conference on machine learning , pages=

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.956623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.956623Z digest=sha256:e46fb59c94153f97dbd6c7fdbd15dfe1092439e16cc522d41623f39d8098979c

Observation 8c6403d2-eafe-41d5-b5b6-83de8429d7cb · outbound

This paper cites ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.960816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.960816Z digest=sha256:1cf1e46c1b1badff679c481670913b107bf12208f52b68f5eaaf666f40a7aecc

Observation da89599d-4536-4118-9ac9-de4e36b77b62 · outbound

This paper cites arXiv preprint arXiv:2503.02318 , year=.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv preprint arXiv:2503.02318 , year=

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.965844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.965844Z digest=sha256:fceaf3062fb0126abf0eea5fe5c155225c2f3138d478148a7aa2a509ae4141fa

Observation 19c6c492-c904-440e-8f5e-f348a5ff0837 · outbound

This paper cites MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.969779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.969779Z digest=sha256:6809a1d0dd38ff2f16ae7703807285127301ebb71d329ce426adb6b2bd2181da

Observation b80465ce-e474-44f0-a20b-6006416d3654 · outbound

This paper cites MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:20.975265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:20.975265Z digest=sha256:486d3d56aae4e0761c6919ac4020be990404bafce01b216b6c99da236df9be80

Pith citing papers

No inbound Pith citation observations are available.