Pith. sign in

REVIEW 1 cited by

Future-Conditioned Recommendations with Multi-Objective Controllable Decision Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.07212 v1 pith:SBPLQSXZ submitted 2025-01-13 cs.IR

classification cs.IR
keywords objectivescontrollablerecommendationrecommendationsstrategiesfutureitemmulti-objective
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Securing long-term success is the ultimate aim of recommender systems, demanding strategies capable of foreseeing and shaping the impact of decisions on future user satisfaction. Current recommendation strategies grapple with two significant hurdles. Firstly, the future impacts of recommendation decisions remain obscured, rendering it impractical to evaluate them through direct optimization of immediate metrics. Secondly, conflicts often emerge between multiple objectives, like enhancing accuracy versus exploring diverse recommendations. Existing strategies, trapped in a "training, evaluation, and retraining" loop, grow more labor-intensive as objectives evolve. To address these challenges, we introduce a future-conditioned strategy for multi-objective controllable recommendations, allowing for the direct specification of future objectives and empowering the model to generate item sequences that align with these goals autoregressively. We present the Multi-Objective Controllable Decision Transformer (MocDT), an offline Reinforcement Learning (RL) model capable of autonomously learning the mapping from multiple objectives to item sequences, leveraging extensive offline data. Consequently, it can produce recommendations tailored to any specified objectives during the inference stage. Our empirical findings emphasize the controllable recommendation strategy's ability to produce item sequences according to different objectives while maintaining performance that is competitive with current recommendation strategies across various objectives.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Generative Auto-Bidding in Large-Scale Competitive Auctions via Diffusion Completer-Aligner

    cs.GT 2025-09 conditional novelty 5.0 of 10

    Diffusion-completer training with a trajectory aligner makes diffusion-based auto-bidding work at scale, improving conversion value by 29.9% on a sparse public benchmark and by 2.0% in production at Kuaishou.

Pith tools