Pith. sign in

Paper Citation Record · LEDGER

Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents

As of 8 August 2026, this Paper Citation Record lists 10 of 10 outbound references and 0 inbound Pith citation observations for arXiv:2511.10687.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.10687 v3

Coverage vector

measured 10 of 10 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T22:48:13.389900Z

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

10 of 10 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ebfb17e4-39e8-481a-8c56-62b9d8c4efc8 · outbound

This paper cites an unresolved cited work.

Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T22:48:13.182497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:48:13.182497Z digest=sha256:92e7a9b85e55f20b4694c43252e727b768ed3daf8de08081060deeb94f9c1b66

Observation 7779f0d2-8a8a-461e-8760-1f5f4956eec8 · outbound

This paper cites an unresolved cited work.

Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T22:48:13.286579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:48:13.286579Z digest=sha256:b0453098b94f650434b81e083a4c974cad3b80c2a32f3a981c082f964b713426

Observation 6d528103-0915-49a8-bde4-887ad2280f9a · outbound

This paper cites A.2.1 PLUGGING THE SIGNALS INTO POST-TRAINING Success route (RL-style).Use {ri,t} as per-message rewards for each agent policy πi.

Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents A.2.1 PLUGGING THE SIGNALS INTO POST-TRAINING Success route (RL-style).Use {ri,t} as per-message rewards for each agent policy πi

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T22:48:13.389900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:48:13.389900Z digest=sha256:37f5003860043849b5b7c28bb79ba47147f08c886f6cee39f02562494239455d

Observation d8b051f7-8a88-452a-9f97-57fe3d7156d7 · outbound

This paper cites Hanhan Zhou, Tian Lan, and Vaneet Aggarwal.

Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents Hanhan Zhou, Tian Lan, and Vaneet Aggarwal

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T22:48:12.972733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:48:12.972733Z digest=sha256:a11245c28d6e6cf746477e0e2701f6795a2a709856904a7e54c7fb7a85147b93

Observation 4dbc5025-f6a2-4325-b958-526dbb994880 · outbound

This paper cites an unresolved cited work.

Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T22:48:13.075374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:48:13.075374Z digest=sha256:e4c0eca04101783a6175fade4c67133520635df93f369d685a588f7480a7ead5

Observation 9386696a-3c67-4049-beee-c373ae487399 · outbound

This paper cites Bounding the Estimation Error of Sampling-based Shapley Value Approximation.

Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents Bounding the Estimation Error of Sampling-based Shapley Value Approximation

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-03T22:48:12.451470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:48:12.451470Z digest=sha256:be640a09b4254e9bee9da698dd5fd48602664e144f4957fd0190fc3cbbda3694

Observation fbdf2d18-0fcd-47dd-a69d-c5e64652322b · outbound

This paper cites STaR: Bootstrapping Reasoning With Reasoning.

Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents STaR: Bootstrapping Reasoning With Reasoning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-03T22:48:12.820859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:48:12.820859Z digest=sha256:7216d54e3fa96c09101216ab62f8245fe8e5f817f7319376070bdfc755ec4c83

Observation 8bf93735-835e-413d-9cbc-2c78b0cbc75d · outbound

This paper cites Training language models to follow instructions with human feedback.

Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents Training language models to follow instructions with human feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T22:48:12.563763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:48:12.563763Z digest=sha256:f39b5c96c087952e57ab42900338874d7863e19f0be2895edbcc8bafab685151

Observation 6536a929-c228-4cbf-8654-d7abe1360432 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T22:48:12.670792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:48:12.670792Z digest=sha256:bc28571d11a99d85f1a2a75474f44e1d01695023bffe986780a69dd18868d0d2

Observation 6fa6716e-04e7-41c2-a61f-c2dcbd4576e4 · outbound

This paper cites SCAR: Shapley Credit Assignment for More Efficient RLHF.

Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents SCAR: Shapley Credit Assignment for More Efficient RLHF

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T22:48:12.321881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:48:12.321881Z digest=sha256:f26dfc56acf1cbe97861135bffa11d3b8a21846f7871c736942b65b3e6c76896

Pith citing papers

No inbound Pith citation observations are available.