Pith. sign in

Paper Citation Record · LEDGER

Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents

As of 22 August 2026, this Paper Citation Record lists 10 of 10 outbound references and 0 inbound Pith citation observations for arXiv:2511.10687.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.10687 v3

Coverage vector

measured 10 of 10 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T22:48:13.389900Z

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

10 of 10 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ebfb17e4-39e8-481a-8c56-62b9d8c4efc8 · outbound

This paper cites an unresolved cited work.

Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T22:48:13.182497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:48:13.182497Z digest=sha256:2fe633835a6ec4ea9f58d4fb62c0db08071392f4dd3d8b30b6692c6081c04a5a

Observation 7779f0d2-8a8a-461e-8760-1f5f4956eec8 · outbound

This paper cites an unresolved cited work.

Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T22:48:13.286579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:48:13.286579Z digest=sha256:28cb5afa03f8112da3e20583557f8bb09cb4e98950413ce4f2f0b252bd6d66a1

Observation 6d528103-0915-49a8-bde4-887ad2280f9a · outbound

This paper cites A.2.1 PLUGGING THE SIGNALS INTO POST-TRAINING Success route (RL-style).Use {ri,t} as per-message rewards for each agent policy πi.

Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents A.2.1 PLUGGING THE SIGNALS INTO POST-TRAINING Success route (RL-style).Use {ri,t} as per-message rewards for each agent policy πi

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T22:48:13.389900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:48:13.389900Z digest=sha256:6db9f12629cb919bec4f6249a57b2f931296638894f587279a5be68667e8fd83

Observation d8b051f7-8a88-452a-9f97-57fe3d7156d7 · outbound

This paper cites Hanhan Zhou, Tian Lan, and Vaneet Aggarwal.

Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents Hanhan Zhou, Tian Lan, and Vaneet Aggarwal

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T22:48:12.972733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:48:12.972733Z digest=sha256:6091c430f86c50f441c0aecbb7fa84c12ca4ccb61c7fa9438ba75d21b851d052

Observation 4dbc5025-f6a2-4325-b958-526dbb994880 · outbound

This paper cites an unresolved cited work.

Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T22:48:13.075374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:48:13.075374Z digest=sha256:efc9cccf0b0bc37adff2859a99a7dfcf9acc90cd932d2158959cbaba6d73ec12

Observation 9386696a-3c67-4049-beee-c373ae487399 · outbound

This paper cites Bounding the Estimation Error of Sampling-based Shapley Value Approximation.

Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents Bounding the Estimation Error of Sampling-based Shapley Value Approximation

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-03T22:48:12.451470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:48:12.451470Z digest=sha256:cb1e554aa887f646b0c93682e8bc497ee50d6990e8c379f43fc9de7a2414bf59

Observation fbdf2d18-0fcd-47dd-a69d-c5e64652322b · outbound

This paper cites STaR: Bootstrapping Reasoning With Reasoning.

Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents STaR: Bootstrapping Reasoning With Reasoning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-03T22:48:12.820859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:48:12.820859Z digest=sha256:5133458a115d6e06e1105cc811e6ddba142a1402f86e786b0090a86aa273de47

Observation 8bf93735-835e-413d-9cbc-2c78b0cbc75d · outbound

This paper cites Training language models to follow instructions with human feedback.

Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents Training language models to follow instructions with human feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T22:48:12.563763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:48:12.563763Z digest=sha256:67d3ddbf7905334a0269aba2d1a8038a23ce9d245625849abbbd13bd4c07c223

Observation 6536a929-c228-4cbf-8654-d7abe1360432 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T22:48:12.670792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:48:12.670792Z digest=sha256:a41d8621fafb3d741686b89393627fb452545917b79fc44471e7151027391c1b

Observation 6fa6716e-04e7-41c2-a61f-c2dcbd4576e4 · outbound

This paper cites SCAR: Shapley Credit Assignment for More Efficient RLHF.

Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents SCAR: Shapley Credit Assignment for More Efficient RLHF

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T22:48:12.321881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:48:12.321881Z digest=sha256:9e6233b22a01eed1e6877b2c5e54898f5432d7415591ac59dd2911fe776ffa93

Pith citing papers

No inbound Pith citation observations are available.