Pith. sign in

REVIEW 1 cited by

LADDER: A Human-Level Bidding Agent for Large-Scale Real-Time Online Auctions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1708.05565 v2 pith:PMPQ6P7B submitted 2017-08-18 cs.LG cs.AIcs.CLcs.GT

classification cs.LGcs.AIcs.CLcs.GT
keywords agentbiddingonlinereal-timeauctionsdeepinformationinputs
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present LADDER, the first deep reinforcement learning agent that can successfully learn control policies for large-scale real-world problems directly from raw inputs composed of high-level semantic information. The agent is based on an asynchronous stochastic variant of DQN (Deep Q Network) named DASQN. The inputs of the agent are plain-text descriptions of states of a game of incomplete information, i.e. real-time large scale online auctions, and the rewards are auction profits of very large scale. We apply the agent to an essential portion of JD's online RTB (real-time bidding) advertising business and find that it easily beats the former state-of-the-art bidding policy that had been carefully engineered and calibrated by human experts: during JD.com's June 18th anniversary sale, the agent increased the company's ads revenue from the portion by more than 50%, while the advertisers' ROI (return on investment) also improved significantly.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Generative Auto-Bidding in Large-Scale Competitive Auctions via Diffusion Completer-Aligner

    cs.GT 2025-09 conditional novelty 5.0 of 10

    Diffusion-completer training with a trajectory aligner makes diffusion-based auto-bidding work at scale, improving conversion value by 29.9% on a sparse public benchmark and by 2.0% in production at Kuaishou.

Pith tools