Pith. sign in

Paper Citation Record · LEDGER

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy

As of 5 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 2 inbound Pith citation observations for arXiv:2603.15646.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.15646 v2

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T17:08:23.269275Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T15:03:42.266254Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-12T07:51:38.881425Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact17
  • verified fuzzy6
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f7ed8e7f-0c5a-4c00-9a81-8acffdac8a78 · outbound

This paper cites H., Gendler, A., Baruch, E.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy H., Gendler, A., Baruch, E

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T17:10:11.210472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:175a6dc2249a8929a4fdf3484ee84eb5f50abb5dcd49e02311890c3603a8f9dd

Observation bf8328d8-4c95-4557-a204-05683fd2b2ec · outbound

This paper cites HealthBench: Evaluating Large Language Models Towards Improved Human Health.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy HealthBench: Evaluating Large Language Models Towards Improved Human Health

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T17:10:10.350720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:ea7e3a2cd4ae7113677baaef5e9739465bfafb7898cc833e289e8517a3552404

Observation 02e6903c-9be5-42d8-bf6e-aef97cd6cd38 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Constitutional AI: Harmlessness from AI Feedback

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:10:10.332916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:0b42a0f37c0d75a13789d72be9ff51cab5ce40c5292c7d919c5714c02b487cce

Observation 1bacc138-fa68-4527-b4fe-2a77ad2b0d19 · outbound

This paper cites XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-26T03:04:06.160463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:f76d5275827b65f174f2273b2f00b0b1f4d348b3c558d529060d6ba27e7399f3

Observation 67c5f086-2392-481f-9936-c73202b9170d · outbound

This paper cites Language models that think, chat better.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Language models that think, chat better

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:10:10.403000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:00727008f7d965eb2063e7b9a596c76a08c8ffec79b2530fb58ffaae8c15675c

Observation 5d59b23f-9125-4152-91fd-b49ca58de998 · outbound

This paper cites Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:10:10.378588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:0cf72d786a4c03d9867fad4acd65ffe3973590c790efb543389a8b4b0c58207d

Observation babd34df-a8b7-480c-8daf-1b43780aa7f0 · outbound

This paper cites arXiv preprint arXiv:2511.10507 , year=.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy arXiv preprint arXiv:2511.10507 , year=

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:10:10.357083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:85fa7491376ee055539653f779969580f67ddf7694f8e26820f88d2c7efac3fb

Observation e85495b3-c760-4a36-93ab-334b15bc8a38 · outbound

This paper cites Reinforcement Learning with Rubric Anchors.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Reinforcement Learning with Rubric Anchors

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:10:10.378397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:65039522e6ed56b216d057044cff595342ce09834ecd9783e5da2353122771cb

Observation 926a2f97-4b4c-4901-b841-6f6c88568dad · outbound

This paper cites MaPPO: Maximum a Posteriori Preference Optimization with Prior Knowledge.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy MaPPO: Maximum a Posteriori Preference Optimization with Prior Knowledge

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T17:10:10.383743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:96ac4ef33e6f577dc9c639ccabf77f1ec9fd476a7658ebc7a8b8501b1b36b8c6

Observation 743c52f8-3add-4133-91d6-b5d6be436fbe · outbound

This paper cites Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language Models.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:10:10.403325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:f5bda0ee9e9e5a19caddb70e76d8c8e8606a2ff6cbbdf5bbdb6fff219b3866e1

Observation b8c6df27-566c-470d-9452-5dd60f4f12a1 · outbound

This paper cites Learning to optimize multi-objective alignment through dynamic reward weighting.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Learning to optimize multi-objective alignment through dynamic reward weighting

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:10:10.396976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:6000fb52bf89f6eda1ba4e867d8df4d83839f7d5dd9aeb0576f26e15289e9d5e

Observation 64ad7a06-22b8-4fe7-8cd6-be8316a5cda4 · outbound

This paper cites Online rubrics elicitation from pairwise comparisons.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Online rubrics elicitation from pairwise comparisons

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:10:10.372183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:aa47b04174a1fe08f7053320cc1a3094d8cb6cd9dd2a512c1f2b26a8bd35f8a3

Observation 230ddb86-54b7-48d0-85af-f7e6af4abd7e · outbound

This paper cites Proximal Policy Optimization Algorithms.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Proximal Policy Optimization Algorithms

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:10:10.339593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:ce70a44c1686f9d716a2c08530066d6fc0a8039c9abbc6e6151393a5e390cec6

Observation 4d72c8f2-1d8b-486d-aed7-81956b721a65 · outbound

This paper cites DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T00:00:28.487879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:9fa4ed3c36e8d54d851f95004fe2d1d3bc6260c0c6b10ae28fc0c18dd124f4f3

Observation b353b112-8ab4-4c92-b29c-0352a7f7fd9f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:10:10.373798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:ba5a437f94f53aa01e9b6f7eb544ac534ce3f5ccf59e527a5b0a6478abef2c59

Observation 98fc3c8c-3610-4314-9f0a-407b49f7958d · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy HybridFlow: A Flexible and Efficient RLHF Framework

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:10:10.390724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:fc2dd6ce0667d4407786ad220dd141c5a9e9c40d66919261e621a314d7c1d8e2

Observation 1544c5fe-57fb-47d7-8d66-bcc4ca1ba7ba · outbound

This paper cites Nemotron-cascade: Scaling cascaded reinforcement learning for general-purpose reasoning models.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Nemotron-cascade: Scaling cascaded reinforcement learning for general-purpose reasoning models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:10:10.384722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:ca72ba969fda0646f5cbd94505fb15189fd387a3468b6351a52491d70c6fae0e

Observation 58184558-409e-4387-8175-a31b123cae3c · outbound

This paper cites arXiv preprint arXiv:2602.01511 , year=.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy arXiv preprint arXiv:2602.01511 , year=

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:10:10.351260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:3a656658412222fc5fe5a8dda7e0410185751f30fcae8b8e11ad54db7ae288ed

Observation 03f65816-e7db-4a58-b780-04d4effd86f9 · outbound

This paper cites Qwen3 Technical Report.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Qwen3 Technical Report

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:10:10.389425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:e607a35d319a69c2d4a2d36caadf72754fedb5c008d37e98166cfe14fcaea971

Observation df6a2493-f2bf-462a-84cb-ac6ed9e751ae · outbound

This paper cites Does this image satisfy this rule?.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Does this image satisfy this rule?

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:10:10.409449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:a458cf6248e26baf19d5eb8cfe15ec6fa4d1253714d0730ff373e9bc989397e8

Observation a8683238-ad7e-4b0a-aac5-b3bfd90353d4 · outbound

This paper cites Group Sequence Policy Optimization.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Group Sequence Policy Optimization

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:10:10.360727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:9f43754c7430ccf37bbbeb2e62d0e82610ad741c67cfa52c9ab7d5ad364f52bf

Observation f49e56ad-0276-47f8-ac17-be9fbaa1550f · outbound

This paper cites Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-08-04T02:23:17.018487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:351da067fa603c7929844b3648fed4400315f272c31b64d45d893513e80be3e7

Observation d3a505f3-a108-4713-8047-b886cb783476 · outbound

This paper cites I’m an emergency medicine physician.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy I’m an emergency medicine physician

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T17:10:11.215521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:e810ccd165a870ca484d83bbedcd00dc112ac4fe987e620d3067c152a635d2e1

Observation 86af8d85-5403-4cc1-8bdd-d5597f3e90b2 · outbound

This paper cites accuracy.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy accuracy

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T17:10:11.209097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:a630a4868a66fb0684bae239e5dee2948d76ab63953736096be826072c6b6a36

Observation c880ebe9-ddf9-44eb-a386-db0d3aaeffc7 · outbound

This paper cites ‘list [ { “criterion.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy ‘list [ { “criterion

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T17:10:11.208091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:c37454f1b9eaec68f71571cc02d6e9f0766eef945752265c0ff560769e8ad390

Observation 858c46ae-06b5-4637-ae2b-1cad7bfae204 · outbound

This paper cites Llama-3.1-8B-Instruct starts at 0.34 and achieves 0.70, while Qwen3-8B starts at a higher score0.58and achieves0.76.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Llama-3.1-8B-Instruct starts at 0.34 and achieves 0.70, while Qwen3-8B starts at a higher score0.58and achieves0.76

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T17:10:11.213055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:5a6c6d9e1af95a1f924e6b46001cd9fc42d681b1f570efa42747b1530b152262

Observation 58b581e5-cd56-4160-a5e6-b9b8e55bdbfe · outbound

This paper cites In ARL, we use the fixed meta-class Order 0: [completeness, accuracy, instruction following, context awareness, communication quality].

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy In ARL, we use the fixed meta-class Order 0: [completeness, accuracy, instruction following, context awareness, communication quality]

Reference 27

Resolution
malformed identifier
raw_fallback, observed 2026-05-15T17:10:11.204202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:66bcef25bf5917b872b6debffb5268aac96e287372d1f0acef8e34342b1a29f0

Observation 28e83434-7346-41a8-b8c9-87c6ed66c940 · outbound

This paper cites The maximum response length is set to 2048, and the temperature in LLM sampling is set to 1.0 in the training process.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy The maximum response length is set to 2048, and the temperature in LLM sampling is set to 1.0 in the training process

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T17:10:11.211669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:1ad0b1dc1de2bbcaae7064b2f7b915648372d349f66a3511dbe87095c5461cd2

Pith citing papers

Observation 01682382-2014-48d5-872e-b6959ccacae6 · inbound

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems cites this paper.

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy

Reference 164

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:51:38.936160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T01:47:40.772146Z digest=sha256:3a0eae3a7ad9810cbae61254e8c9f15bab9699aafce388e3a8c4065e07c8a9ba

Observation 623215f7-1583-430c-8a30-443982b723de · inbound

Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions cites this paper.

Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T15:03:42.266254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:03:42.266254Z digest=sha256:580be3d3da83af3479643584468d262cd3a7902e18e91ea15767bfa36008e46e