Pith. sign in

Paper Citation Record · LEDGER

Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards

As of 6 August 2026, this Paper Citation Record lists 9 of 9 outbound references and 1 inbound Pith citation observation for arXiv:2604.09855.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.09855 v1

Coverage vector

measured 9 of 9 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T17:11:15.484711Z

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-08T22:26:44.052574Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T22:35:40.643385Z

Reference resolution

9 of 9 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved2
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 391fbdef-0764-4a7e-bb02-2fb2a960e055 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards Kimi K2: Open Agentic Intelligence

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T07:26:01.930259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:11:15.484711Z digest=sha256:244a9f36827111bb8205ed68f7bb0d9e9e870ee585adb04b7e9010732e86a39e

Observation 655c7202-9547-4e29-8c93-d009b353b531 · outbound

This paper cites codename_1.

Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards codename_1

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:41:46.652154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:11:15.484711Z digest=sha256:0c2170a7a7ac8e9c89dc8fba35db65042aa5e115c02527fe7ea2f5d482bccf04

Observation 8a53b19d-28cc-4b90-8155-9c68803a23b2 · outbound

This paper cites an unresolved cited work.

Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-05-17T11:41:46.655409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:11:15.484711Z digest=sha256:b587a0292db9874ebe6b5d2be1750071c47d2535a5fac3f60ca9389fbc626c30

Observation 2dcb90a2-f3c4-4a34-9862-d325c731cc66 · outbound

This paper cites $M (N codename_1) is a exact copy of seller’s previous offer.

Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards $M (N codename_1) is a exact copy of seller’s previous offer

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:41:46.644828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:11:15.484711Z digest=sha256:7cab95b4084b48ddadd4c848cd101dee9c62f20e144cfd0817f8cb1473447cdb

Observation de725330-876c-47d2-b242-9695dfdb5ebb · outbound

This paper cites Happy By Clinique For Men. Cologne Spray 1.7 Oz.

Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards Happy By Clinique For Men. Cologne Spray 1.7 Oz

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:41:46.648565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:11:15.484711Z digest=sha256:dd520f0ce0b4ef7c08f3048f9c11788dfd19ba09480e45f211e06a4ced2c830b

Observation a34d4894-f55c-4e2e-bd4d-257c3eab1c59 · outbound

This paper cites codename_1.

Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards codename_1

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:41:46.658549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:11:15.484711Z digest=sha256:23831cc49dc3d5178c26e2f9577da1bee306836fbe6a2c892a48e34a32ad6140

Observation 76033cc6-4f99-4fa7-8b9f-5cb5dc201beb · outbound

This paper cites an unresolved cited work.

Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-17T11:41:46.638146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:11:15.484711Z digest=sha256:c27b0734ba6e00715c61b629e97b141c6f7f62c0c3c2dcefaa837a01c1ea6ecc

Observation 3e1cd05e-56a4-4491-b652-4ee70e1371b5 · outbound

This paper cites codename_1.

Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards codename_1

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:41:46.634943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:11:15.484711Z digest=sha256:8543d6ec5b76093d35e97a26f25291862628fbe6a3d18b9c2f7989d57a7176a0

Observation 3edba574-afe3-4e30-96cb-1f1b72198cf9 · outbound

This paper cites Happy By Clinique For Men. Cologne Spray 1.7 Oz.

Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards Happy By Clinique For Men. Cologne Spray 1.7 Oz

Reference 9

Resolution
malformed identifier
raw_fallback, observed 2026-05-17T11:41:46.641568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:11:15.484711Z digest=sha256:42a020c1183d44e1129f77721a75cdb5663974aecc630f1550f918f57fdf0b0d

Pith citing papers

Observation d9257adb-84d8-4271-8db2-894109063678 · inbound

Strategic Bargaining in Multi-Buyer Markets: Reinforcement Learning from Verifiable Rewards for LLM Negotiations cites this paper.

Strategic Bargaining in Multi-Buyer Markets: Reinforcement Learning from Verifiable Rewards for LLM Negotiations Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards

Reference 55

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T22:35:40.645323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-08T22:26:44.052574Z digest=sha256:a594fe063f9b4efd8aac520d015a6477444334e63389316d83ab6d77fb98cc02