Pith. sign in

REVIEW 1 cited by

Towards General Negotiation Strategies with End-to-End Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.15096 v1 pith:VEKUORQR submitted 2024-06-21 cs.MA cs.LG

classification cs.MAcs.LG
keywords negotiationagentslearningproblemsreinforcementcausingnegotiateactions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The research field of automated negotiation has a long history of designing agents that can negotiate with other agents. Such negotiation strategies are traditionally based on manual design and heuristics. More recently, reinforcement learning approaches have also been used to train agents to negotiate. However, negotiation problems are diverse, causing observation and action dimensions to change, which cannot be handled by default linear policy networks. Previous work on this topic has circumvented this issue either by fixing the negotiation problem, causing policies to be non-transferable between negotiation problems or by abstracting the observations and actions into fixed-size representations, causing loss of information and expressiveness due to feature design. We developed an end-to-end reinforcement learning method for diverse negotiation problems by representing observations and actions as a graph and applying graph neural networks in the policy. With empirical evaluations, we show that our method is effective and that we can learn to negotiate with other agents on never-before-seen negotiation problems. Our result opens up new opportunities for reinforcement learning in negotiation agents.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Differentiable Normative Guidance for Nash Bargaining Solution Recovery

    cs.GT 2026-03 conditional novelty 6.0 of 10

    Guided graph diffusion with a differentiable Nash-product penalty recovers individually rational, near-Nash-bargaining utility splits without explicit Pareto frontier knowledge.

Pith tools