Pith. sign in

REVIEW 7 cited by

RecSim: A Configurable Simulation Platform for Recommender Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.04847 v2 pith:K3ZEWERY submitted 2019-09-11 cs.LG cs.HCcs.IRstat.ML

classification cs.LGcs.HCcs.IRstat.ML
keywords recsimuserenvironmentsbehaviorconfigurableitemplatformrecommender
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose RecSim, a configurable platform for authoring simulation environments for recommender systems (RSs) that naturally supports sequential interaction with users. RecSim allows the creation of new environments that reflect particular aspects of user behavior and item structure at a level of abstraction well-suited to pushing the limits of current reinforcement learning (RL) and RS techniques in sequential interactive recommendation problems. Environments can be easily configured that vary assumptions about: user preferences and item familiarity; user latent state and its dynamics; and choice models and other user response behavior. We outline how RecSim offers value to RL and RS researchers and practitioners, and how it can serve as a vehicle for academic-industrial collaboration.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 51 citations worldwide. Full citation record

  1. Hierarchical Residual Policy Optimization for Generative Recommendations

    cs.IR 2026-08 conditional novelty 6.0 of 10

    HRPO decomposes item-level rewards into token-level 'residual credits' along semantic identifier hierarchies and optimizes the generator with a PPO-style objective, improving session utility in KuaiSim and production ...

  2. CoSimRec: Measuring Coordinated-Content Penetration in Recommender Feedback Loops

    cs.IR 2026-07 conditional novelty 6.0 of 10

    Under controlled simulation, coordinated content reaches non-bot recommendation slots when rankers reward popularity or feedback (APR-Lift up to 0.47 on LastFM), while random ranking shows none.

  3. How Does Empowering Users with Greater System Control Affect News Filter Bubbles?

    cs.IR 2026-06 conditional novelty 6.0 of 10

    Users who could adjust a news recommender's stance and topic sliders changed how extreme their feed became depending on their starting point, but did not consistently increase political diversity.

  4. Modeling Earth-Scale Human-Like Societies with One Billion Agents

    cs.MA 2025-06 conditional novelty 6.0 of 10

    Light Society scales LLM-agent social simulations to one billion agents by substituting most LLM interactions with a distilled surrogate model.

  5. PUB: An LLM-Enhanced Personality-Driven User Behaviour Simulator for Recommender System Evaluation

    cs.IR 2025-06 conditional novelty 6.0 of 10

    PUB uses LLM-inferred Big Five personality traits to generate synthetic recommender-system interactions that the authors say mimic real Amazon behavior and preserve algorithm performance rankings.

  6. CreAgent: Towards Long-Term Evaluation of Recommender System under Platform-Creator Information Asymmetry

    cs.IR 2025-02 conditional novelty 6.0 of 10

    CreAgent combines an LLM with game-theoretic beliefs and fast-slow thinking to reproduce creator behavior under information asymmetry, and it is used to evaluate recommender systems over long time horizons.

  7. Beyond Action Imitation: Learning a Decision-Aware User Simulator for Online Advertising

    cs.IR 2026-07 conditional novelty 5.5 of 10

    DASH folds cross-domain user histories, distills teacher thinking traces, and RL-tunes a small LLM with action plus rubric rewards to jointly predict ad actions and decision traces on Tencent data.

Pith tools