REVIEW 4 major objections 5 minor 299 references
Playstyle and Artificial Intelligence: An Initial Blueprint Through the Lens of Video Games
T0 review · 4 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read This dissertation argues that playstyle — an agent's consistent pattern of decision-making — can be measured by one label-free metric: the divergence of two agents' action distributions on the discrete states both have visited, and that thi
desk verdict The dissertation's playstyle metric is a real but not-new contribution; the 'general' claim outruns the evidence, yet the framework and ablations deserve refereeing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Hierarchical State Discretization (HSD) encoder, a vector-quantized variational autoencoder whose discrete codes are organized into a multiscale hierarchy and trained with two joint objectives: an observation reconstruction loss that preserves perceptual capacity, and a policy prediction loss that forces codes to retain decision-relevant semantics. On top of it, Playstyle Distance compares only the intersection of symbolic states two agents have both visited, aggregating local 2-Wasserstein distances between their action distributions with expected averaging. A multiscale composition of three granularities (coarse singleton, intermediate 2^20 codes, fine hierar
What would settle it
A dual test would settle it. First, self-distance: split one agent's trajectory in half and compute the metric between halves with the same encoder — it should be near zero and far below the distance between different policies; an equal or larger self-distance means the metric tracks sampling noise, not style. Second, oracle consistency: build three rule-based TORCS agents that human observers unanimously rate as A ≈ B in style and C clearly different; the paper's stated consistency criterion requires d(A,C) < d(B,C) to be reproduced, so a single inversion would falsify the semantic-alignment
Extended reading notes
Core claim
Playstyle is defined as the decision-making style of an agent expressed through interaction with a responsive environment, where multiple viable choices exist and external consequences give actions meaning. The central operational claim is that this latent construct can be measured by Playstyle Distance: project both agents' observations into a discrete symbolic state space via Hierarchical State Discretization (HSD), intersect the states both agents visited, and take the expected 2-Wasserstein divergence between their empirical action distributions on that intersection. Because style manifests in how agents choose among comparable situations, restricting comparison to genuinely shared conte
Load-bearing premise
The load-bearing assumption is that the discrete symbolic states learned from human gameplay mean the same decision context for every agent, so that comparing two agents' action choices only on the states both happened to visit is a fair and sufficient basis for measuring their stylistic difference.
Editorial extensions
If this is right
- Playstyle measurement becomes label-free and domain-agnostic: the same encoder-plus-intersection procedure works for rule-based bots, human players, and learned agents, across racing games and Atari.
- Strategic diversity and competitive balance become quantifiable on the same distance: the dissertation defines diverse trajectory counts, Top-D diversity, and Top-B balance directly from the playstyle metric.
- Style becomes a trainable target: the capacity and popularity dimensions can be folded into loss functions, letting reinforcement and imitation learning produce agents with specified stylistic tendencies.
- Imitation fidelity can be scored as playstyle distance between an expert and a learner, a process-level complement to outcome-based measures of success.
- If style is a core dimension of intelligent behavior, modeling playstyle becomes a candidate ingredient for agents built to express recognizable values, not just optimize rewards.
Reading between the lines
- The intersection strategy is a generic answer to the state-alignment problem in behavior comparison; it should transfer to any sequential decision domain (robotics, driving, human-computer interaction) where a shared discrete state space can be learned, though the dissertation trains a fresh encoder per environment, leaving single-encoder cross-environment transfer as an open, testable question.
- The exponential similarity kernel predicts a specific psychophysical signature: human similarity judgments between playstyle traces should track e^{-d} rather than raw distance; a paired user study could confirm or reject that cognitive grounding.
- The capacity-popularity axes suggest an operational definition of 'meta shifts' in live games: tracking how a population's styles move through this two-dimensional space over time would turn a qualitative community term into a measurable quantity.
- Because Playstyle Distance is a metric, it induces a geometry on the space of agents; interpolating between policies would yield 'midpoint styles,' a natural and untested route to style blending and novel style generation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This dissertation-style manuscript proposes 'playstyle' as a decision-making style exhibited by agents in interactive environments, and develops a conceptual blueprint spanning philosophy, measurement, expression, and applications. The central technical contribution is a label-free Playstyle Distance metric based on Hierarchical State Discretization (HSD): observations are mapped to discrete symbolic states, and two agents are compared by the divergence of their empirical action distributions over the intersection of visited discrete states. Experiments are reported across three domains: TORCS (rule-based agents), RGSK (six human participants), and seven Atari games (RL agents trained with DQN, C51, Rainbow, IQN). The metric is extended to multiscale states, a probabilistic similarity kernel, continuous playstyle spectra, and downstream diversity/balance indicators. The manuscript also discusses imitation learning and human-likeness, and concludes with speculative connections to AGI and 'soul' digitization.
Significance. If the proposed Playstyle Distance is valid, it would constitute a notable step toward a general, label-free, transferable measurement of decision-making style, with applications in game AI, player modeling, and agent diversity. The paper's strengths include a concrete algorithmic proposal (HSD with dual reconstruction/policy decoders), a relevant ablation showing that policy supervision is necessary for usable intersections (Table 5), and experiments spanning rule-based, human, and learning-based agents. However, the generality claim is considerably broader than the evidence: the human sample is very small, semantic alignment of discrete states across agents is assumed rather than tested, and several reported accuracy numbers are inconsistent across tables. The core construct-validity question is therefore left open.
major comments (4)
- [§6.5.1–6.5.2, Eqs. defining S_φ(A,B) and d_φ(A,B|s)] The Playstyle Distance is defined as the divergence of action distributions on the intersection of HSD discrete states. This only measures playstyle if the same symbolic state in two agents' datasets refers to the same decision context. The evidence offered—CIFAR-10 classification (Table 2), state utilization (Table 3), noise robustness (Table 4), and the policy-decoder ablation (Table 5)—shows that HSD codes preserve object-class information and are influenced by policy supervision, but it does not demonstrate cross-agent semantic alignment. Intersection sizes are reported (e.g., 187.53 with the policy decoder in Table 5), but no coverage statistics, per-state sample counts, or tests of state-conditional action correspondence are given. Since the metric's validity depends entirely on this alignment, the 'general metric' claim is under-supported. Please add direct evidence: distribution
- [§6.5.2, 'Evaluation on RGSK' and Table 8] The human-player evaluation uses only six participants, each instructed to maintain distinct playstyles. With such a small, instruction-conditioned sample, near-perfect accuracies (93–99%) largely reflect the experimental setup rather than generalizable human playstyle discrimination. The manuscript does not report per-participant variance, leave-one-participant-out results, or any statistical significance. This is load-bearing because the abstract claims a 'general' metric across human, rule-based, and RL agents, and the human evidence is a cornerstone of that claim. I recommend additional participants, cross-participant validation, and explicit reporting of variability.
- [Table 13 vs Tables 7–8] There is a substantial unexplained discrepancy in the central accuracy numbers. Table 13 reports single-scale 220 (t=2) accuracy of 73.3±8.2% for TORCS and 79.2±7.9% for RGSK, whereas Tables 7 and 8 report 91.20% (5 Speed) and 93.11% (Nitro) for what appears to be the same configuration. If the tasks differ (e.g., 25-class aggregate vs 5-class per-dimension), this must be stated explicitly. As written, a reader cannot determine which numbers support the main claim. Given that the accuracy tables are the primary experimental evidence, this inconsistency must be resolved.
- [§6.5.4, Eq. (6.3) and surrounding text] The exponential perceptual kernel P(d)=1/e^d is asserted on the basis of a proof 'provided in the Appendix' of a prior paper, but no derivation appears in this manuscript. The similarity measure P S∩ is a core extension, and the kernel's functional form and the normalizer D_M^Φ are not derived or validated here. Please include the derivation in full, or explicitly state the kernel as a modeling assumption and justify it with data. As written, this is a load-bearing gap in the similarity extension.
minor comments (5)
- [§11.2] The heading 'Video Game Industry' appears twice consecutively, disrupting the section numbering and flow.
- [§6.5.4] The notation P S∩ and D_M^Φ is visually cluttered and not clearly defined in the main text. Please clarify the indexing and the exact set over which the sum is taken.
- [§6.5.2, Table 12 and surrounding text] MKL is called 'approx. JS' in the table but the text states it is not equivalent to Jensen-Shannon divergence. Reconcile this label or rename the metric.
- [Throughout] There are minor typographical issues, such as 'V AEs' in §2.4 and irregular spacing around mathematical symbols. A copyedit pass would improve readability.
- [Reproducibility] No code, data, or trained encoders are made available. Given the centrality of the HSD training procedure and the transfer claims, a public release of the evaluation code and datasets would substantially strengthen the contribution.
Circularity Check
Core playstyle-distance derivation is self-contained; minor load-bearing self-citation for the perceptual similarity kernel.
-
self citation load bearing
[Section 6.5.4, 'Perception of Similarity' (around Eq. 6.3)]
"This perceptual relation is the only relation under our assumptions from human cognition and probability. The original paper by Lin et al. [181] (2024) provides a proof using differential equations in the Appendix."
The exponential similarity kernel P(d)=1/e^d is asserted to be the unique consequence of the stated assumptions, but the proof is not reproduced in this dissertation; it is deferred to the authors' own prior work (Lin et al., 2024). The uniqueness claim is therefore load-bearing on a self-citation rather than on a self-contained derivation. This is a minor circularity concern because the main Playstyle Distance accuracy results and the HSD-based comparisons do not depend on this perceptual kernel; the kernel primarily underlies the Playstyle Similarity extension.
full rationale
The central measurement claim—Playstyle Distance as the divergence between action distributions conditioned on HSD-discretized intersection states—is not circular. The HSD encoder is trained on human TORCS gameplay without style labels; the style dimensions and agent identities are used only as held-out evaluation targets. Accuracy is reported on rule-based TORCS agents, human RGSK players, and Atari RL agents, so the main results are not fitted to the target labels. The state-space size and intersection threshold are selected on TORCS and then transferred to RGSK and Atari; this is hyperparameter transfer rather than a fitted-input-called-prediction reduction. The only notable issue is the perceptual similarity kernel in Section 6.5.4, whose uniqueness proof is deferred to a self-citation; this is ancillary to the core Playstyle Distance validation. Overall, the paper shows no significant circularity in its central derivation.
Assumptions & free parameters
free parameters (4)
- HSD state space size (default) =
2^20 = 1,048,576 states
- Intersection threshold t =
2
- Playstyle similarity normalizer D_M^Phi =
Computed as the average of all observed pairwise distances in the comparison set
- Multiscale mixture weights =
Equal weights over encoders {1, 2^20 HSD, 256-res HSD}
assumptions (5)
- ad hoc to paper Playstyle Consistency: preference comparisons are anchored to a current belief state b via PF(b', b), avoiding global transitivity or ordering.
- domain assumption Beliefs are foundational, unprovable premises that initiate action, and the self is an internal system of beliefs.
- domain assumption Human perception of similarity follows a logarithmic or exponential Weber-Fechner relationship, leading to P(d) = exp(-d).
- domain assumption The environment is a 'perfect generator': states produced by environment interaction are inherently valid, so GAN-style fake/real evaluation is inapplicable.
- ad hoc to paper The dual-loop model (external interaction loop plus internal deliberation loop) underlies all agent decision-making and style formation.
Cite this review
Pith. "Pith review of Playstyle and Artificial Intelligence: An Initial Blueprint Through the Lens of Video Games." pith.science (2026). https://pith.science/paper/NRSSFURU
@misc{pith2026250819152,
author = {Pith},
title = {Pith review of: Playstyle and Artificial Intelligence: An Initial Blueprint Through the Lens of Video Games},
year = {2026},
howpublished = {\url{https://pith.science/paper/NRSSFURU}},
note = {Machine review of arXiv:2508.19152}
}
read the original abstract
Contemporary artificial intelligence (AI) development largely centers on rational decision-making, valued for its measurability and suitability for objective evaluation. Yet in real-world contexts, an intelligent agent's decisions are shaped not only by logic but also by deeper influences such as beliefs, values, and preferences. The diversity of human decision-making styles emerges from these differences, highlighting that "style" is an essential but often overlooked dimension of intelligence. This dissertation introduces playstyle as an alternative lens for observing and analyzing the decision-making behavior of intelligent agents, and examines its foundational meaning and historical context from a philosophical perspective. By analyzing how beliefs and values drive intentions and actions, we construct a two-tier framework for style formation: the external interaction loop with the environment and the internal cognitive loop of deliberation. On this basis, we formalize style-related characteristics and propose measurable indicators such as style capacity, style popularity, and evolutionary dynamics. The study focuses on three core research directions: (1) Defining and measuring playstyle, proposing a general playstyle metric based on discretized state spaces, and extending it to quantify strategic diversity and competitive balance; (2) Expressing and generating playstyle, exploring how reinforcement learning and imitation learning can be used to train agents exhibiting specific stylistic tendencies, and introducing a novel approach for human-like style learning and modeling; and (3) Practical applications, analyzing the potential of these techniques in domains such as game design and interactive entertainment. Finally, the dissertation outlines future extensions, including the role of style as a core element in building artificial general intelligence (AGI).
Figures
Figures from the paper (32 more)
Reference graph
Works this paper leans on
-
[1]
Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek Gordon Murray, Benoit Steiner, Paul A. Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. Tensorflow: A system fo...
2016
-
[2]
Pieter Abbeel and Andrew Y . Ng. Apprenticeship learning via inverse reinforcement learning. In International Conference on Machine Learning (ICML), 2004
2004
-
[3]
Machado, Pablo Samuel Castro, and Marc G
Rishabh Agarwal, Marlos C. Machado, Pablo Samuel Castro, and Marc G. Bellemare. Contrastive behavioral similarity embeddings for generalization in reinforcement learn- ing. In International Conference on Learning Representations (ICLR), 2021
2021
-
[4]
Christiano, John Schulman, and Dan Mané
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul F. Christiano, John Schulman, and Dan Mané. Concrete problems in ai safety. CoRR, abs/1606.06565, 2016
arXiv 2016
-
[5]
Summa Theologica , volume 5
Thomas Aquinas. Summa Theologica , volume 5. Guerin, 1274. Originally written 1265–1274. Printed edition used: Guerin, 1869
-
[6]
Metaphysics
Aristotle. Metaphysics. Clarendon Press, 350 BCE. Originally written circa 350 BCE; this edition translated by Christopher Kirwan and published in 1993
1993
-
[7]
Wasserstein gan
Martín Arjovsky, Soumith Chintala, and Léon Bottou. Wasserstein gan. CoRR, 2017
2017
-
[8]
Finite-time analysis of the multi- armed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer. Finite-time analysis of the multi- armed bandit problem. Machine Learning, 47:235–256, 2002. 258
2002
Show all 299 references
-
[9]
Bernard J. Baars. A Cognitive Theory of Consciousness . Cambridge University Press, 1988
1988
-
[10]
Agent57: Outperforming the atari human benchmark
Adrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo, and Charles Blundell. Agent57: Outperforming the atari human benchmark. In International Conference on Machine Learning (ICML) , 2020
2020
-
[11]
Never give up: Learning directed exploration strate- gies
Adrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo, Bilal Piot, Steven Kapturowski, Olivier Tieleman, Martín Arjovsky, Alexander Pritzel, Andrew Bolt, and Charles Blundell. Never give up: Learning directed exploration strate- gies. In International...
2020
-
[12]
Markov, Yi Wu, Glenn Powell, Bob Mc- Grew, and Igor Mordatch
Bowen Baker, Ingmar Kanitscheider, Todor M. Markov, Yi Wu, Glenn Powell, Bob Mc- Grew, and Igor Mordatch. Emergent tool use from multi-agent autocurricula. In Inter- national Conference on Learning Representations (ICLR), 2020
2020
-
[13]
Video pretraining (vpt): Learning to act by watching unlabeled online videos
Bowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga, Jie Tang, Adrien Ecoffet, Brandon Houghton, Raul Sampedro, and Jeff Clune. Video pretraining (vpt): Learning to act by watching unlabeled online videos. In Advances in Neural Information Processing Systems (NeurIPS), 2022
2022
-
[14]
Re-evaluating evaluation
David Balduzzi, Karl Tuyls, Julien Pérolat, and Thore Graepel. Re-evaluating evaluation. In Advances in Neural Information Processing Systems (NeurIPS), 2018
2018
-
[15]
Open-ended learning in symmetric zero-sum games
David Balduzzi, Marta Garnelo, Yoram Bachrach, Wojciech Czarnecki, Julien Pérolat, Max Jaderberg, and Thore Graepel. Open-ended learning in symmetric zero-sum games. In International Conference on Machine Learning (ICML), 2019
2019
-
[16]
Hunt, Tom Schaul, Hado P
André Barreto, Will Dabney, Rémi Munos, Jonathan J. Hunt, Tom Schaul, Hado P. Van Hasselt, and David Silver. Successor features for transfer in reinforcement learn- ing. In Advances in Neural Information Processing Systems (NeurIPS), 2017
2017
-
[17]
Hearts, clubs, diamonds, spades: Players who suit muds
Richard Bartle. Hearts, clubs, diamonds, spades: Players who suit muds. Journal of MUD Research, 1996
1996
-
[18]
What is game balancing? - an examination of concepts
Alexander Becker and Daniel Görlich. What is game balancing? - an examination of concepts. ParadigmPlus, 2020. 259
2020
-
[19]
M. G. Bellemare, W. Dabney, and R. Munos. A distributional perspective on reinforce- ment learning. In International Conference on Machine Learning (ICML), 2017
2017
-
[20]
Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling
Marc G. Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling. The arcade learn- ing environment: An evaluation platform for general agents. Journal of Artificial Intelli- gence Research, 2013
2013
-
[21]
Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Rémi Munos
Marc G. Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Rémi Munos. Unifying count-based exploration and intrinsic motivation. In Advances in Neural Information Processing Systems (NeurIPS), 2016
2016
-
[22]
Representation learning: A review and new perspectives
Yoshua Bengio, Aaron Courville, and Pascal Vincent. Representation learning: A review and new perspectives. IEEE Transactions on Pattern Analysis and Machine Intelligence, 35(8):1798–1828, 2013
2013
-
[23]
Culture, Class, Distinction
Tony Bennett, Mike Savage, Elizabeth Bortolaia Silva, Alan Warde, Modesto Gayo-Cal, and David Wright. Culture, Class, Distinction. Routledge, London, 2009
2009
-
[24]
Theano: A cpu and gpu math expression compiler
James Bergstra, Olivier Breuleux, Frédéric Bastien, Pascal Lamblin, Razvan Pascanu, Guillaume Desjardins, Joseph Turian, David Warde-Farley, and Yoshua Bengio. Theano: A cpu and gpu math expression compiler. In Proceedings of the Python for Scientific Computing Conference (Sci...
2010
-
[25]
A Treatise Concerning the Principles of Human Knowledge
George Berkeley. A Treatise Concerning the Principles of Human Knowledge. Clarendon Press, 1710. Reprinted in many editions; standard version from the 1710 first edition
-
[26]
Daniel E. Berlyne. Curiosity and exploration: Animals spend much of their time seeking stimuli whose significance raises problems for psychology. Science, 153(3731):25–33, 1966
1966
-
[27]
System neural diversity: Measuring behavioral heterogeneity in multi-agent learning
Matteo Bettini, Ajay Shankar, and Amanda Prorok. System neural diversity: Measuring behavioral heterogeneity in multi-agent learning. CoRR, abs/2305.02128, 2023
2023 arXiv
-
[28]
On a measure of divergence between two multinomial populations
Anil Bhattacharyya. On a measure of divergence between two multinomial populations. Sankhy¯a: The Indian Journal of Statistics, 1946
1946
-
[29]
The Holy Bible
Bible. The Holy Bible. Church of England, 600 BCE. Compilation of texts from circa 600 BCE to 100 CE; this edition is the King James Version (KJV), first published in 1611. 260
-
[30]
Christopher M. Bishop. Pattern Recognition and Machine Learning. Springer, 2006
2006
-
[31]
Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang, Xin Zhang, Jake Zhao, and Karol Zieba
Mariusz Bojarski, Davide Del Testa, Daniel Dworakowski, Bernhard Firner, Beat Flepp, Prasoon Goyal, Lawrence D. Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang, Xin Zhang, Jake Zhao, and Karol Zieba. End to end learning for self-driving cars. CoRR, abs/1604.07316, 2016
2016 arXiv
-
[32]
Are you living in a computer simulation? The Philosophical Quarterly, 53(211):243–255, 2003
Nick Bostrom. Are you living in a computer simulation? The Philosophical Quarterly, 53(211):243–255, 2003
2003
-
[33]
Distinction: A Social Critique of the Judgement of Taste
Pierre Bourdieu. Distinction: A Social Critique of the Judgement of Taste . Harvard University Press, Cambridge, MA, 1984. Originally published in French in 1979 as La distinction
1984
-
[34]
Heads-up limit hold’em poker is solved
Michael Bowling, Neil Burch, Michael Johanson, and Oskari Tammelin. Heads-up limit hold’em poker is solved. Science, 347(6218):145–149, 2015
2015
-
[35]
Ralph Allan Bradley and Milton E. Terry. Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika, 1952
1952
-
[36]
Disagreement-regularized imitation learn- ing
Kianté Brantley, Wen Sun, and Mikael Henaff. Disagreement-regularized imitation learn- ing. In International Conference on Learning Representations (ICLR), 2020
2020
-
[37]
Random forests
Leo Breiman. Random forests. Machine Learning, 45(1):5–32, 2001
2001
-
[38]
Openai gym
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. Openai gym. CoRR, abs/1606.01540, 2016
2016 arXiv
-
[39]
Driving event detection and driving style classification using artificial neural networks
Patrick Brombacher, Johannes Masino, Michael Frey, and Frank Gauterin. Driving event detection and driving style classification using artificial neural networks. In IEEE Inter- national Conference on Industrial Technology (ICIT), 2017
2017
-
[40]
Bentz, Léon Bottou, Isabelle Guyon, Yann LeCun, Cliff Moore, Eduard Säckinger, and Roopak Shah
Jane Bromley, James W. Bentz, Léon Bottou, Isabelle Guyon, Yann LeCun, Cliff Moore, Eduard Säckinger, and Roopak Shah. Signature verification using a "siamese" time delay neural network. International Journal of Pattern Recognition and Artificial Intelligence (IJPRAI), 7, 1993
1993
-
[41]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini 261 Agarwal, Ariel Herbert-V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, ...
2020
-
[42]
Jerome S. Bruner. Acts of Meaning. Harvard University Press, 1990
1990
-
[43]
Perception and the Representative Design of Psychological Experi- ments
Egon Brunswik. Perception and the Representative Design of Psychological Experi- ments. University of California Press, 1956
1956
-
[44]
Human Motivation and Emotion
Ross Buck. Human Motivation and Emotion. John Wiley & Sons, 1988
1988
-
[45]
Storkey, and Oleg Klimov
Yuri Burda, Harrison Edwards, Amos J. Storkey, and Oleg Klimov. Exploration by random network distillation. In International Conference on Learning Representations (ICLR), 2019
2019
-
[46]
Peter Burkholder, Donald Jay Grout, and Claude V
J. Peter Burkholder, Donald Jay Grout, and Claude V . Palisca. A History of Western Music: Tenth International Student Edition. W. W. Norton & Company, New York, 10th, international student edition edition, 2019
2019
-
[47]
Joseph Jr
Murray Campbell, A. Joseph Jr. Hoane, and Feng-Hsiung Hsu. Deep blue. Artificial Intelligence, 134(1–2):57–83, 2002
2002
-
[48]
William E. Caplin. Classical Form: A Theory of Formal Functions for the Instrumental Music of Haydn, Mozart, and Beethoven. Oxford University Press, New York, 1998
1998
-
[49]
Bellemare
Pablo Samuel Castro, Subhodeep Moitra, Carles Gelada, Saurabh Kumar, and Marc G. Bellemare. Dopamine: A research framework for deep reinforcement learning. CoRR, abs/1812.06110, 2018
2018 arXiv
-
[50]
scary robots
Stephen Cave, Kate Coughlan, and Kanta Dihal. "scary robots": Examining public re- sponses to AI. In AAAI/ACM Conference on AI, Ethics, and Society (AIES), 2019
2019
-
[51]
Vision-language models as a source of rewards
Harris Chan, V olodymyr Mnih, Feryal Behbahani, Michael Laskin, Luyu Wang, Fabio Pardo, Maxime Gazeau, Himanshu Sahni, Dan Horgan, Kate Baumli, Yannick 262 Schroecker, Stephen Spencer, Richie Steigerwald, John Quan, Gheorghe Comanici, Se- bastian Flennerhag, Alexander Neitz, L...
2023
-
[52]
Chang, Masatoshi Uehara, Dhruv Sreenivas, Rahul Kidambi, and Wen Sun
Jonathan D. Chang, Masatoshi Uehara, Dhruv Sreenivas, Rahul Kidambi, and Wen Sun. Mitigating covariate shift in imitation learning via offline data with partial coverage. In Advances in Neural Information Processing Systems (NeurIPS), 2021
2021
-
[53]
Similarity estimation techniques from rounding algorithms
Moses Charikar. Similarity estimation techniques from rounding algorithms. In ACM Symposium on Theory of Computing (STOC), 2002
2002
-
[54]
Detecting individual decision-making style in the game of go
Chun-Jung Chen. Detecting individual decision-making style in the game of go. Master’s thesis, National Yang Ming Chiao Tung University, 2023
2023
-
[55]
Learning to evaluate the artness of ai-generated images
Junyu Chen, Jie An, Hanjia Lyu, Christopher Kanan, and Jiebo Luo. Learning to evaluate the artness of ai-generated images. IEEE Transactions on Multimedia, 26, 2024
2024
-
[56]
Diffusion model-augmented behavioral cloning
Shang-Fu Chen, Hsiang-Chun Wang, Ming-Hao Hsu, Chun-Mao Lai, and Shao-Hua Sun. Diffusion model-augmented behavioral cloning. InInternational Conference on Machine Learning (ICML), 2024
2024
-
[57]
Modeling intransitivity in matchup and comparison data
Shuo Chen and Thorsten Joachims. Modeling intransitivity in matchup and comparison data. In International Conference on Web Search and Data Mining (WSDM), 2016
2016
-
[58]
Infogan: Interpretable representation learning by information maximizing generative ad- versarial nets
Xi Chen, Yan Duan, Rein Houthooft, John Schulman, Ilya Sutskever, and Pieter Abbeel. Infogan: Interpretable representation learning by information maximizing generative ad- versarial nets. In Advances in Neural Information Processing Systems (NeurIPS), 2016
2016
-
[59]
Open Innovation: The New Imperative for Creating and Profiting from Technology
Henry William Chesbrough. Open Innovation: The New Imperative for Creating and Profiting from Technology. Harvard Business Press, 2003
2003
-
[60]
Learning phrase representations us- ing rnn encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. Learning phrase representations us- ing rnn encoder–decoder for statistical machine translation. In Conference on Empirical Methods in Natural Language Pr...
2014
-
[61]
Christensen
Clayton M. Christensen. The Innovator’s Dilemma: When New Technologies Cause Great Firms to Fail. Harvard Business Review Press, 1997
1997
-
[62]
Maximizing cosine similarity between spatial features for unsupervised domain adaptation in semantic segmentation
Inseop Chung, Daesik Kim, and Nojun Kwak. Maximizing cosine similarity between spatial features for unsupervised domain adaptation in semantic segmentation. In IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2022
2022
-
[63]
Cuisine and Culture: A History of Food and People
Linda Civitello. Cuisine and Culture: A History of Food and People. John Wiley & Sons, Hoboken, NJ, 2011
2011
-
[64]
Surfing Uncertainty: Prediction, Action, and the Embodied Mind
Andy Clark. Surfing Uncertainty: Prediction, Action, and the Embodied Mind . Oxford University Press, 2015
2015
-
[65]
de Condorcet
M. de Condorcet. Essai sur l’application de l’analyse à la probabilité des décisions rendues à la pluralité des voix. Imprimerie Royale, Paris, 1785
-
[66]
The Analects
Confucius. The Analects. Ballantine Books, 475 BCE. Originally compiled in the 5th century BCE; this edition published in 2003
2003
-
[67]
Support-vector networks
Corinna Cortes and Vladimir Vapnik. Support-vector networks. Machine Learning, 20 (3):273–297, 1995
1995
-
[68]
Costa, Paul T
Jr. Costa, Paul T. and Robert R. McCrae. Domains and facets: Hierarchical personality assessment using the revised neo personality inventory. Journal of Personality Assess- ment, 1995
1995
-
[69]
Efficient selectivity and backup operators in monte-carlo tree search
Rémi Coulom. Efficient selectivity and backup operators in monte-carlo tree search. In International Conference on Computer and Games (CG), pp. 72–83, 2006
2006
-
[70]
elo ratings
Rémi Coulom. Computing “elo ratings” of move patterns in the game of go. Journal of the International Computer Games Association (ICGA Journal), 30(4):198–206, 2007
2007
-
[71]
Whole-history rating: A bayesian rating system for players of time- varying strength
Rémi Coulom. Whole-history rating: A bayesian rating system for players of time- varying strength. In International Conference on Computers and Games (CG), 2008
2008
-
[72]
Flow: The Psychology of Optimal Experience
Mihály Csíkszentmihályi. Flow: The Psychology of Optimal Experience. Harper & Row, 1990. 264
1990
-
[73]
Robots that can adapt like animals
Antoine Cully, Jeff Clune, Danesh Tarapore, and Jean-Baptiste Mouret. Robots that can adapt like animals. Nature, 521(7553):503–507, 2015
2015
-
[74]
Implicit quantile net- works for distributional reinforcement learning
Will Dabney, Georg Ostrovski, David Silver, and Rémi Munos. Implicit quantile net- works for distributional reinforcement learning. InInternational Conference on Machine Learning (ICML), 2018
2018
-
[75]
Primal wasser- stein imitation learning
Robert Dadashi, Léonard Hussenot, Matthieu Geist, and Olivier Pietquin. Primal wasser- stein imitation learning. In International Conference on Learning Representations (ICLR), 2021
2021
-
[76]
Rae, Andreas Glaese, Yujia Rist, Tom Goodside, Catherine Olsson, Jarrid Rae, Thomas Van Overveldt, Daniel Levy, Xinyun Li, Anton Bakhtin, Bryan McCann, Panos Mandilaras, Samuel R
Sumanth Dathathri, Abigail See, Sumedh Ghaisas, Po-Sen Huang, Rob McAdam, Jo- hannes Welbl, Vandana Bachani, Alex Kaskasoli, Robert Stanforth, Tatiana Matejovi- cova, James Sutherland, Chaitanya Malaviya, Nicholas Nunez, Luke Belrose, Samuel Humeau, Abhinav Sridhar, Tom Jones,...
2024
-
[77]
Actions, reasons, and causes
Donald Davidson. Actions, reasons, and causes. The Journal of Philosophy , 60(23): 685–700, 1963
1963
-
[78]
Gureckis, and Brenden M
Guy Davidson, Graham Todd, Julian Togelius, Todd M. Gureckis, and Brenden M. Lake. Goals as reward-producing programs. Nature Machine Intelligence, 2025
2025
-
[79]
Dempster, Nan M
Arthur P. Dempster, Nan M. Laird, and Donald B. Rubin. Maximum likelihood from incomplete data via the em algorithm. Journal of the Royal Statistical Society: Series B (Methodological), 39(1):1–22, 1977
1977
-
[80]
Daniel C. Dennett. The Intentional Stance. MIT Press, 1987
1987
-
[81]
Meditations on First Philosophy
René Descartes. Meditations on First Philosophy . Hackett Publishing, Indianapolis, 3rd edition, 1641. Originally published in Latin in 1641. This edition: 3rd ed., 1993, translated by Donald A. Cress. 265
1993
-
[82]
BERT: Pre- training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre- training of deep bidirectional transformers for language understanding. North American Chapter of the Association for Computational Linguistics (NAACL), 2019
2019
-
[83]
Jacob D. Dodson. The relation of strength of stimulus to rapidity of habit-formation in the kitten. Journal of Animal Behavior, 1915
1915
-
[84]
Flownet: Learning optical flow with convolutional networks
Alexey Dosovitskiy, Philipp Fischer, Eddy Ilg, Philip Häusser, Caner Hazirbas, Vladimir Golkov, Patrick van der Smagt, Daniel Cremers, and Thomas Brox. Flownet: Learning optical flow with convolutional networks. In IEEE International Conference on Com- puter Vision (ICCV), 2015
2015
-
[85]
Yannakakis
Anders Drachen, Alessandro Canossa, and Georgios N. Yannakakis. Player modeling us- ing self-organization in tomb raider: Underworld. InIEEE Conference on Computational Intelligence and Games (CIG), 2009
2009
-
[86]
Temporal cycle-consistency learning
Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson, Pierre Sermanet, and Andrew Zis- serman. Temporal cycle-consistency learning. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019
2019
-
[87]
Elo-mmr: A rating system for massive multiplayer compe- titions
Aram Ebtekar and Paul Liu. Elo-mmr: A rating system for massive multiplayer compe- titions. In The Web Conference (WWW), 2021
2021
-
[88]
Reason for Being: A Meditation on Ecclesiastes
Jacques Ellul. Reason for Being: A Meditation on Ecclesiastes . Wm. B. Eerdmans Publishing, Grand Rapids, MI, 1990
1990
-
[89]
Arpad E. Elo. The USCF Rating System: Its Development, Theory, and Applications . United States Chess Federation, 1966
1966
-
[90]
Impala: Scalable distributed deep-rl with importance weighted actor- learner architectures
Lasse Espeholt, Hubert Soyer, Rémi Munos, Karen Simonyan, V olodymyr Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, Shane Legg, and Koray Kavukcuoglu. Impala: Scalable distributed deep-rl with importance weighted actor- learner architectures. In Internatio...
2018
-
[91]
A density-based al- gorithm for discovering clusters in large spatial databases with noise
Martin Ester, Hans-Peter Kriegel, Jörg Sander, and Xiaowei Xu. A density-based al- gorithm for discovering clusters in large spatial databases with noise. In International Conference on Knowledge Discovery and Data Mining (KDD), 1996. 266
1996
-
[92]
Diversity is all you need: Learning skills without a reward function
Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine. Diversity is all you need: Learning skills without a reward function. In International Conference on Learning Representations (ICLR), 2019
2019
-
[93]
Generalized data distribution iteration
Jiajun Fan and Changnan Xiao. Generalized data distribution iteration. In International Conference on Machine Learning (ICML), 2022
2022
-
[94]
Learnable behavior control: Breaking atari human world records via sample-efficient behavior selection
Jiajun Fan, Yuzheng Zhuang, Yuecheng Liu, Jianye Hao, Bin Wang, Jiangcheng Zhu, Hao Wang, and Shu-Tao Xia. Learnable behavior control: Breaking atari human world records via sample-efficient behavior selection. InInternational Conference on Learning Representations (ICLR), 2023
2023
-
[95]
G. T. Fechner. Elements of Psychophysics. Holt, Rinehart and Winston, 1966. Originally published in 1860
1966
-
[96]
Testing the manifold hy- pothesis
Charles Fefferman, Sanjoy Mitter, and Hariharan Narayanan. Testing the manifold hy- pothesis. Journal of the American Mathematical Society, 29(4):983–1049, 2016
2016
-
[97]
Guided cost learning: Deep inverse op- timal control via policy optimization
Chelsea Finn, Sergey Levine, and Pieter Abbeel. Guided cost learning: Deep inverse op- timal control via policy optimization. In International Conference on Machine Learning (ICML), 2016
2016
-
[98]
Ramirez, Ayzaan Wahid, Laura Downs, Adrian Wong, Johnny Lee, Igor Mordatch, and Jonathan Tompson
Pete Florence, Corey Lynch, Andy Zeng, Oscar A. Ramirez, Ayzaan Wahid, Laura Downs, Adrian Wong, Johnny Lee, Igor Mordatch, and Jonathan Tompson. Implicit behavioral cloning. In Conference on Robot Learning (CoRL), 2021
2021
-
[99]
Noisy networks for exploration
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Matteo Hes- sel, Ian Osband, Alex Graves, V olodymyr Mnih, Rémi Munos, Demis Hassabis, Olivier Pietquin, Charles Blundell, and Shane Legg. Noisy networks for exploration. In Inter- national Conference on Lear...
2018
-
[100]
Fox go, 2024
Fox Go. Fox go, 2024. URL https://www.foxwq.com/
2024
-
[101]
Friedman
Jerome H. Friedman. Greedy function approximation: A gradient boosting machine. Annals of Statistics, 29(5):1189–1232, 2001
2001
-
[102]
Learning robust rewards with adversarial in- verse reinforcement learning
Justin Fu, Katie Luo, and Sergey Levine. Learning robust rewards with adversarial in- verse reinforcement learning. In International Conference on Learning Representations (ICLR), 2018. 267
2018
-
[103]
Evalu- ating human-like behaviors of video-game agents autonomously acquired with biological constraints
Nobuto Fujii, Yuichi Sato, Hironori Wakama, Koji Kazai, and Haruhiro Katayose. Evalu- ating human-like behaviors of video-game agents autonomously acquired with biological constraints. In Advances in Computer Entertainment Technology (ACE), 2013
2013
-
[104]
Extreme q-learning: Maxent rl without entropy
Divyansh Garg, Joey Hejna, Matthieu Geist, and Stefano Ermon. Extreme q-learning: Maxent rl without entropy. In International Conference on Learning Representations (ICLR), 2023
2023
-
[105]
Gatys, Alexander S
Leon A. Gatys, Alexander S. Ecker, and Matthias Bethge. Image style transfer using convolutional neural networks. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016
2016
-
[106]
Combining online and offline knowledge in uct
Sylvain Gelly and David Silver. Combining online and offline knowledge in uct. In International Conference on Machine Learning (ICML), pp. 273–280, 2007
2007
-
[107]
Gut Feelings: The Intelligence of the Unconscious
Gerd Gigerenzer. Gut Feelings: The Intelligence of the Unconscious. Penguin, 2007
2007
-
[108]
Rich feature hierar- chies for accurate object detection and semantic segmentation
Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. Rich feature hierar- chies for accurate object detection and semantic segmentation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2014
2014
-
[109]
Glickman
Mark E. Glickman. Parameter estimation in large dynamic paired comparison experi- ments. Journal of the Royal Statistical Society Series C: Applied Statistics, 1999
1999
-
[110]
Clément Godard, Oisin Mac Aodha, and Gabriel J. Brostow. Unsupervised monocular depth estimation with left-right consistency. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017
2017
-
[111]
Über formal unentscheidbare sätze der principia mathematica und ver- wandter systeme i
Kurt Gödel. Über formal unentscheidbare sätze der principia mathematica und ver- wandter systeme i. Monatshefte für Mathematik und Physik, 38:173–198, 1931
1931
-
[112]
Gonzalez and Richard E
Rafael C. Gonzalez and Richard E. Woods. Digital Image Processing . Pearson, 4th edition, 2017
2017
-
[113]
Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio. Generative adversarial nets. In Advances in Neural Information Processing Systems (NeurIPS), 2014
2014
-
[114]
Ways of Worldmaking
Nelson Goodman. Ways of Worldmaking. Hackett Publishing, Indianapolis, 1978. 268
1978
-
[115]
Yannakakis
Daniele Gravina, Antonios Liapis, and Georgios N. Yannakakis. Quality diversity through surprise. IEEE Transactions on Evolutionary Computation , 23(4):603–616, 2019
2019
-
[116]
Shuyue Guan and Murray H. Loew. A novel measure to evaluate generative adversarial networks based on direct analysis of generated images. Neural Computing and Applica- tions, 2021
2021
-
[117]
Reinforcement learn- ing with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine. Reinforcement learn- ing with deep energy-based policies. In International Conference on Machine Learning (ICML), 2017
2017
-
[118]
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International Conference on Machine Learning (ICML), 2018
2018
-
[119]
The Code of Hammurabi
Hammurabi. The Code of Hammurabi. University of Chicago Press, 1754 BCE. Orig- inally inscribed circa 1754 BCE; this edition translated from the original Akkadian and published in 1904
1904
-
[120]
The Elements of Statistical Learning
Trevor Hastie, Robert Tibshirani, and Jerome Friedman. The Elements of Statistical Learning. Springer, 2009
2009
-
[121]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016
2016
-
[122]
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. Mask r-cnn. IEEE International Conference on Computer Vision (ICCV), 2017
2017
-
[123]
Artistic styles: Revisiting the analysis of modern artists’ careers
Christiane Hellmanzik. Artistic styles: Revisiting the analysis of modern artists’ careers. Journal of Cultural Economics, 33:201–232, 2009
2009
-
[124]
Trueskill™: A bayesian skill rating system
Ralf Herbrich, Tom Minka, and Thore Graepel. Trueskill™: A bayesian skill rating system. In Advances in Neural Information Processing Systems (NIPS), 2006
2006
-
[125]
Clipscore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. Clipscore: A reference-free evaluation metric for image captioning. In Conference on Empirical Methods in Natural Language Processing (EMNLP), 2021. 269
2021
-
[126]
Hessel, J
M. Hessel, J. Modayil, H. van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. Gheshlaghi Azar, and D. Silver. Rainbow: Combining improvements in deep reinforcement learning. In AAAI Conference on Artificial Intelligence (AAAI), 2018
2018
-
[127]
Hester, M
T. Hester, M. Vecerik, O. Pietquin, M. Lanctot, T. Schaul, B. Piot, D. Horgan, J. Quan, A. Sendonaris, I. Osband, G. Dulac-Arnold, J. P. Agapiou, J. Z. Leibo, and A. Gruslys. Deep q-learning from demonstrations. In AAAI Conference on Artificial Intelligence (AAAI), 2018
2018
-
[128]
Agapiou, Joel Z
Todd Hester, Matej Vecerík, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Dan Horgan, John Quan, Andrew Sendonaris, Ian Osband, Gabriel Dulac-Arnold, John P. Agapiou, Joel Z. Leibo, and Audrunas Gruslys. Deep q-learning from demonstrations. In AAAI Conference on Arti...
2018
-
[129]
Gans trained by a two time-scale update rule converge to a local nash equi- librium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equi- librium. In Advances in Neural Information Processing Systems (NeurIPS), 2017
2017
-
[130]
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon. Generative adversarial imitation learning. In Advances in Neural Information Processing Systems (NeurIPS), 2016
2016
-
[131]
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems (NeurIPS), 2020
2020
-
[132]
Towards human-like rl: Taming non-naturalistic behavior in deep rl via adap- tive behavioral costs in 3d games
Kuo-Hao Ho, Ping-Chun Hsieh, Chiu-Chou Lin, You-Ren Luo, Feng-Jian Wang, and I- Chen Wu. Towards human-like rl: Taming non-naturalistic behavior in deep rl via adap- tive behavioral costs in 3d games. In Asian Conference on Machine Learning (ACML), 2023
2023
-
[133]
Leviathan
Thomas Hobbes. Leviathan. Penguin Classics, London, 1651. This edition: Penguin Classics, 1985. Originally published in 1651
1985
-
[134]
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural Computa- tion, 9:1735–1780, 1997
1997
-
[135]
Distributed prioritized experience replay
Dan Horgan, John Quan, David Budden, Gabriel Barth-Maron, Matteo Hessel, Hado van Hasselt, and David Silver. Distributed prioritized experience replay. In International Conference on Learning Representations (ICLR), 2018. 270
2018
-
[136]
Multilayer feedforward net- works are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White. Multilayer feedforward net- works are universal approximators. Neural Networks, 2(5):359–366, 1989
1989
-
[137]
Image quality metrics: Psnr vs
Alain Horé and Djemel Ziou. Image quality metrics: Psnr vs. ssim. In International Conference on Pattern Recognition (ICPR), 2010
2010
-
[138]
Measuring policy distance for multi-agent reinforcement learning
Tianyi Hu, Zhiqiang Pu, Xiaolin Ai, Tenghai Qiu, and Jianqiang Yi. Measuring policy distance for multi-agent reinforcement learning. In International Conference on Au- tonomous Agents and Multiagent Systems (AAMAS), 2024
2024
-
[139]
Efficient action-constrained reinforce- ment learning via acceptance-rejection method and augmented mdps
Wei Hung, Shao-Hua Sun, and Ping-Chun Hsieh. Efficient action-constrained reinforce- ment learning via acceptance-rejection method and augmented mdps. In International Conference on Learning Representations (ICLR), 2025
2025
-
[140]
The case for dynamic difficulty adjustment in games
Robin Hunicke. The case for dynamic difficulty adjustment in games. In ACM SIGCHI International Conference on Advances in Computer Entertainment Technology (ACE) , 2005
2005
-
[141]
Ideas Pertaining to a Pure Phenomenology and to a Phenomenologi- cal Philosophy: First Book: General Introduction to a Pure Phenomenology
Edmund Husserl. Ideas Pertaining to a Pure Phenomenology and to a Phenomenologi- cal Philosophy: First Book: General Introduction to a Pure Phenomenology . Springer Science & Business Media, Dordrecht, 1913. Originally published in German in 1913. This edition: English transla...
1913
-
[142]
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International Conference on Machine Learning (ICML), 2015
2015
-
[143]
The Art of Color: The Subjective Experience and Objective Rationale of Color
Johannes Itten. The Art of Color: The Subjective Experience and Objective Rationale of Color. Van Nostrand Reinhold Company, New York, 1973. Translated from German by Ernst van Haagen. Originally published as Kunst der Farbe
1973
-
[144]
Étude comparative de la distribution florale dans une portion des alpes et des jura
Paul Jaccard. Étude comparative de la distribution florale dans une portion des alpes et des jura. Bulletin de la Société Vaudoise des Sciences Naturelles, 37:547–579, 1901
1901
-
[145]
Czarnecki, Jeff Don- ahue, Ali Razavi, Oriol Vinyals, Tim Green, Iain Dunning, Karen Simonyan, Chrisantha 271 Fernando, and Koray Kavukcuoglu
Max Jaderberg, Valentin Dalibard, Simon Osindero, Wojciech M. Czarnecki, Jeff Don- ahue, Ali Razavi, Oriol Vinyals, Tim Green, Iain Dunning, Karen Simonyan, Chrisantha 271 Fernando, and Koray Kavukcuoglu. Population based training of neural networks.CoRR, abs/1711.09846, 2017
2017 arXiv
-
[146]
The Principles of Psychology, volume 1–2
William James. The Principles of Psychology, volume 1–2. Henry Holt and Company, New York, 1890. Originally published in two volumes
-
[147]
Lawrence Zitnick, and Ross Girshick
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Judy Hoffman, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick. Inferring and executing programs for visual reasoning. In IEEE International Conference on Computer Vision (ICCV), 2017
2017
-
[148]
Jones and Benjamin K
Cameron R. Jones and Benjamin K. Bergen. Large language models pass the turing test. CoRR, abs/2503.23674, 2025
2025 arXiv
-
[149]
Unity: A general platform for intelligent agents
Arthur Juliani, Vincent-Pierre Berges, Esh Vckay, Yuan Gao, Hunter Henry, Marwan Mattar, and Danny Lange. Unity: A general platform for intelligent agents. CoRR, abs/1809.02627, 2018
2018 arXiv
-
[150]
John M. Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, Alex Bridgland, Clemens Meyer, Simon A A Kohl, Andy Ballard, Andrew Cowie, Bernardino Romera-Paredes, Stanislav...
2021
-
[151]
Daniel Jurafsky and James H. Martin. Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition . Prentice Hall, 2 edition, 2008
2008
-
[152]
Littman, and Anthony R
Leslie Pack Kaelbling, Michael L. Littman, and Anthony R. Cassandra. Planning and acting in partially observable stochastic domains. Artificial Intelligence, 101(1–2):99– 134, 1998
1998
-
[153]
Kahneman and A
D. Kahneman and A. Tversky. Prospect theory: An analysis of decision under risk. Econometrica, 47(2):263–291, 1979
1979
-
[154]
Thinking, Fast and Slow
Daniel Kahneman. Thinking, Fast and Slow. Farrar, Straus and Giroux, New York, 2011. 272
2011
-
[155]
Benchmarking end-to-end behavioural cloning on video games
Anssi Kanervisto, Joonas Pussinen, and Ville Hautamäki. Benchmarking end-to-end behavioural cloning on video games. In IEEE Conference on Games (CoG), 2020
2020
-
[156]
Critique of Pure Reason
Immanuel Kant. Critique of Pure Reason . Cambridge University Press, Cambridge,
-
[157]
Re- current experience replay in distributed reinforcement learning
Steven Kapturowski, Georg Ostrovski, John Quan, Rémi Munos, and Will Dabney. Re- current experience replay in distributed reinforcement learning. In International Confer- ence on Learning Representations (ICLR), 2019
2019
-
[158]
A style-based generator architecture for generative adversarial networks
Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019
2019
-
[159]
The road to artificial superintelligence: A comprehen- sive survey of superalignment
HyunJin Kim, Xiaoyuan Yi, Jing Yao, Jianxun Lian, Muhua Huang, Shitong Duan, JinYeong Bak, and Xing Xie. The road to artificial superintelligence: A comprehen- sive survey of superalignment. CoRR, abs/2412.16468, 2024
2024 arXiv
-
[160]
Kingma and Max Welling
Diederik P. Kingma and Max Welling. Auto-encoding variational bayes. InInternational Conference on Learning Representations (ICLR), 2014
2014
-
[161]
The soviet scientific programme on ai: if a machine cannot ‘think’, can it ‘control’? BJHS Themes, 8:111–125, 2023
Olessia Kirtchik. The soviet scientific programme on ai: if a machine cannot ‘think’, can it ‘control’? BJHS Themes, 8:111–125, 2023
2023
-
[162]
Gaming industry report 2025: Market size & trends,
Andrea Knezovic. Gaming industry report 2025: Market size & trends,
2025
-
[163]
Bandit based monte-carlo planning
Levente Kocsis and Csaba Szepesvári. Bandit based monte-carlo planning. In European Conference on Machine Learning (ECML), pp. 282–293, 2006
2006
-
[164]
Deep neural decision forests
Peter Kontschieder, Madalina Fiterau, Antonio Criminisi, and Samuel Rota Bulò. Deep neural decision forests. In IEEE International Conference on Computer Vision (ICCV), 2015
2015
-
[165]
Offline reinforcement learning with implicit q-learning
Ilya Kostrikov, Ashvin Nair, and Sergey Levine. Offline reinforcement learning with implicit q-learning. In International Conference on Learning Representations (ICLR) , 2022. 273
2022
-
[166]
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton. Learning multiple layers of features from tiny images. Technical report, University of Toronto, 2009
2009
-
[167]
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Sys- tems (NeurIPS), 2012
2012
-
[168]
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine. Conservative q-learning for offline reinforcement learning. In Advances in Neural Information Processing Sys- tems (NeurIPS), 2020
2020
-
[169]
Diffusion-reward adversarial imitation learning
Chun-Mao Lai, Hsiang-Chun Wang, Ping-Chun Hsieh, Frank Wang, Min-Hung Chen, and Shao-Hua Sun. Diffusion-reward adversarial imitation learning. In Advances in Neural Information Processing Systems (NIPS), 2024
2024
-
[170]
Can agents run relay race with strangers? generalization of rl to out-of-distribution trajectories
Li-Cheng Lan, Huan Zhang, and Cho-Jui Hsieh. Can agents run relay race with strangers? generalization of rl to out-of-distribution trajectories. In International Con- ference on Learning Representations (ICLR), 2023
2023
-
[171]
A unified game-theoretic ap- proach to multiagent reinforcement learning
Marc Lanctot, Vinícius Flores Zambaldi, Audrunas Gruslys, Angeliki Lazaridou, Karl Tuyls, Julien Pérolet, David Silver, and Thore Graepel. A unified game-theoretic ap- proach to multiagent reinforcement learning. In Advances in Neural Information Pro- cessing Systems (NeurIPS), 2017
2017
-
[172]
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. Deep learning. Nature, 521(7553): 436–444, 2015
2015
-
[173]
Gradient- based regularization for action smoothness in robotic control with reinforcement learn- ing
I Lee, Hoang-Giang Cao, Cong-Tinh Dao, Yu-Cheng Chen, and I-Chen Wu. Gradient- based regularization for action smoothness in robotic control with reinforcement learn- ing. In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , 2024
2024
-
[174]
G. W. Leibniz. The monadology. In L. E. Loemker (ed.), Philosophical Papers and Letters, pp. 643–653. Springer, Dordrecht, 1714. This edition from: *Philosophical Papers and Letters*, 1989. Originally written in 1714. 274
1989
-
[175]
Nonlinear inverse reinforcement learning with gaussian processes
Sergey Levine, Zoran Popovic, and Vladlen Koltun. Nonlinear inverse reinforcement learning with gaussian processes. InAdvances in Neural Information Processing Systems (NIPS), 2011
2011
-
[176]
Neural image beauty predictor based on bradley- terry model
Shiyu Li, Hao Ma, and Xiangyu Hu. Neural image beauty predictor based on bradley- terry model. CoRR, abs/2111.10127, 2021
2021 arXiv
-
[177]
Infogail: Interpretable imitation learning from visual demonstrations
Yunzhu Li, Jiaming Song, and Stefano Ermon. Infogail: Interpretable imitation learning from visual demonstrations. In Advances in Neural Information Processing Systems (NeurIPS), 2017
2017
-
[178]
Online learning of counter categories and ratings in pvp games
Chiu-Chou Lin and I-Chen Wu. Online learning of counter categories and ratings in pvp games. Proceedings of the Annual Conference of JSAI , JSAI2025:3K5IS2b05– 3K5IS2b05, 2025
2025
-
[179]
An unsupervised video game playstyle metric via state discretization
Chiu-Chou Lin, Wei-Chen Chiu, and I-Chen Wu. An unsupervised video game playstyle metric via state discretization. In Conference on Uncertainty in Artificial Intelligence (UAI), 2021
2021
-
[180]
Method for training ai bot in computer game, February 2022
Chiu-Chou Lin, Ying-Hau Wu, Kuan-Ming Lin, Pei-Wen Huang, I-Chen Wu, and Cheng- Lun Tsai. Method for training ai bot in computer game, February 2022. United States patent
2022
-
[181]
Perceptual similarity for measuring decision-making style and policy diversity in games
Chiu-Chou Lin, Wei-Chen Chiu, and I-Chen Wu. Perceptual similarity for measuring decision-making style and policy diversity in games. Transactions on Machine Learning Research, 2024
2024
-
[182]
Identifying and clustering counter relationships of team com- positions in pvp games for efficient balance analysis
Chiu-Chou Lin, Yu-Wei Shih, Kuei-Ting Kuo, Yu-Cheng Chen, Chien-Hua Chen, Wei- Chen Chiu, and I-Chen Wu. Identifying and clustering counter relationships of team com- positions in pvp games for efficient balance analysis. Transactions on Machine Learning Research, 2024
2024
-
[183]
Method for training ai bot in computer game, October 2024
Chiu-Chou Lin, Ying-Hau Wu, Kuan-Ming Lin, Pei-Wen Huang, I-Chen Wu, and Cheng- Lun Tsai. Method for training ai bot in computer game, October 2024. Taiwan patent
2024
-
[184]
Method for training ai bot in computer game, February 2025
Chiu-Chou Lin, I-Chen Wu, Jung-Chang Kuo, Ying-Hau Wu, An-Lun Teng, and Pei- Wen Huang. Method for training ai bot in computer game, February 2025. United States patent. 275
2025
-
[185]
Breuel, and Jan Kautz
Ming-Yu Liu, Thomas M. Breuel, and Jan Kautz. Unsupervised image-to-image transla- tion networks. In Advances in Neural Information Processing Systems (NeurIPS), 2017
2017
-
[186]
Re-evaluating open-ended evaluation of large language models
Siqi Liu, Ian Gemp, Luke Marris, Georgios Piliouras, Nicolas Heess, and Marc Lanc- tot. Re-evaluating open-ended evaluation of large language models. In International Conference on Learning Representations (ICLR), 2025
2025
-
[187]
Towards unifying behavioral and response diversity for open-ended learning in zero-sum games
Xiangyu Liu, Hangtian Jia, Ying Wen, Yujing Hu, Yingfeng Chen, Changjie Fan, Zhipeng Hu, and Yaodong Yang. Towards unifying behavioral and response diversity for open-ended learning in zero-sum games. In Advances in Neural Information Pro- cessing Systems (NeurIPS), 2021
2021
-
[188]
Imitation from obser- vation: Learning to imitate behaviors from raw video via context translation
Yuxuan Liu, Abhishek Gupta, Pieter Abbeel, and Sergey Levine. Imitation from obser- vation: Learning to imitate behaviors from raw video via context translation. In IEEE International Conference on Robotics and Automation (ICRA), 2018
2018
-
[189]
A unified diversity measure for multiagent reinforcement learning
Zongkai Liu, Chao Yu, Yaodong Yang, Peng Sun, Zifan Wu, and Yuan Li. A unified diversity measure for multiagent reinforcement learning. In Advances in Neural Infor- mation Processing Systems (NeurIPS), 2022
2022
-
[190]
An Essay Concerning Human Understanding
John Locke. An Essay Concerning Human Understanding. Kay & Troutman, Philadel- phia, 1689. This edition: Kay & Troutman, 1847 (Philadelphia). Public domain edition scanned by Google Books. Originally published in 1689
-
[191]
The psychology of curiosity: A review and reinterpretation
George Loewenstein. The psychology of curiosity: A review and reinterpretation. Psy- chological Bulletin, 116(1):75–98, 1994
1994
-
[192]
A study of human-like deep reinforcement learning agents
You-Ren Luo. A study of human-like deep reinforcement learning agents. Master’s thesis, National Chiao Tung University, 2019
2019
-
[193]
Foerster
Andrei Lupu, Brandon Cui, Hengyuan Hu, and Jakob N. Foerster. Trajectory diversity for zero-shot coordination. In International Conference on Machine Learning (ICML) , 2021
2021
-
[194]
Sugarman, and Sarah Hickinbottom
Jack Martin, Jeff H. Sugarman, and Sarah Hickinbottom. Persons: Understanding Psy- chological Selfhood and Agency. Springer, 2009. 276
2009
-
[195]
Abraham H. Maslow. A theory of human motivation. Psychological Review, 50(4): 370–396, 1943
1943
-
[196]
McCulloch and Walter Pitts
Warren S. McCulloch and Walter Pitts. A logical calculus of the ideas immanent in nervous activity. The Bulletin of Mathematical Biophysics, 5(4):115–133, 1943
1943
-
[197]
Kleinberg, and Ashton Ander- son
Reid McIlroy-Young, Yu Wang, Siddhartha Sen, Jon M. Kleinberg, and Ashton Ander- son. Detecting individual decision-making style: Exploring behavioral stylometry in chess. In Advances in Neural Information Processing Systems (NeurIPS), 2021
2021
-
[198]
Charles W. Millard. Fauvism. The Hudson Review, 29(4):576–580, 1976
1976
-
[199]
The Rolling Stone Illustrated History of Rock and Roll
Jim Miller (ed.). The Rolling Stone Illustrated History of Rock and Roll. Random House, New York, 2nd edition, 1981. Originally published in 1976
1981
-
[200]
Perceptrons: An Introduction to Computational Geometry
Marvin Minsky and Seymour Papert. Perceptrons: An Introduction to Computational Geometry. MIT Press, 1969
1969
-
[201]
Conditional generative adversarial nets
Mehdi Mirza and Simon Osindero. Conditional generative adversarial nets. CoRR, 2014
2014
-
[202]
Riedmiller
V olodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin A. Riedmiller. Playing atari with deep reinforcement learning. CoRR, abs/1312.5602, 2013
2013 arXiv
-
[203]
Rusu, Joel Veness, Marc G
V olodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin A. Riedmiller, Andreas Fidjeland, Georg Os- trovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstr...
2015
-
[204]
Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu
V olodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. Asynchronous methods for deep reinforcement learning. In International Conference on Machine Learning (ICML), 2016
2016
-
[205]
Giovanni Molinaro and Andrew G. E. Collins. A goal-centric outlook on learning.Trends in Cognitive Sciences, 27(12):1150–1164, 2023. 277
2023
-
[206]
Position: Levels of agi for operationalizing progress on the path to agi
Meredith Ringel Morris, Jascha Sohl-Dickstein, Noah Fiedel, Tris Warkentin, Allan Dafoe, Aleksandra Faust, Clément Farabet, and Shane Legg. Position: Levels of agi for operationalizing progress on the path to agi. In International Conference on Machine Learning (ICML), 2024
2024
-
[207]
McCaulley, Naomi L
Isabel Briggs Myers, Mary H. McCaulley, Naomi L. Quenk, and Allen L. Hammer.MBTI Manual: A Guide to the Development and Use of the Myers-Briggs Type Indicator. Con- sulting Psychologists Press, 1998
1998
-
[208]
Regularizing action policies for smooth control with reinforcement learning
Siddharth Mysore, Bassel Mabsout, Renato Mancuso, and Kate Saenko. Regularizing action policies for smooth control with reinforcement learning. In IEEE International Conference on Robotics and Automation (ICRA), 2021
2021
-
[209]
The View from Nowhere
Thomas Nagel. The View from Nowhere. Oxford University Press, 1986
1986
-
[210]
Discrete, compositional, and symbolic representations through attractor dynamics
Andrew Nam, Eric Elmoznino, Nikolay Malkin, James McClelland, Yoshua Bengio, and Guillaume Lajoie. Discrete, compositional, and symbolic representations through attractor dynamics. CoRR, abs/2310.01807, 2023
2023 arXiv
-
[211]
Improving image generation with better captions
Charlie Nash, William Chan, Yilun Du, Alexander Kirillov, Ross Girshick, Hanzi Liu, Aditya Ramesh, and Barret Zoph. Improving image generation with better captions. CoRR, 2023
2023
-
[212]
Ng and Stuart Russell
Andrew Y . Ng and Stuart Russell. Algorithms for inverse reinforcement learning. In International Conference on Machine Learning (ICML), 2000
2000
-
[213]
Ng, Daishi Harada, and Stuart Russell
Andrew Y . Ng, Daishi Harada, and Stuart Russell. Policy invariance under reward trans- formations: Theory and application to reward shaping. In International Conference on Machine Learning (ICML), 1999
1999
-
[214]
Information-directed exploration for deep reinforcement learning
Nikolay Nikolov, Johannes Kirschner, Felix Berkenkamp, and Andreas Krause. Information-directed exploration for deep reinforcement learning. In International Con- ference on Learning Representations (ICLR), 2019
2019
-
[215]
Donald A. Norman. The Design of Everyday Things: Revised and Expanded Edition. Ba- sic Books, New York, revised and expanded edition edition, 2013. Originally published in 1988. 278
2013
-
[216]
Game Development Essentials
Jeannie Novak, Meaghan O’Brien, and Jim Gish. Game Development Essentials. Delmar Cengage Learning, 3 edition, 2012
2012
-
[217]
Cuda programming guide
NVIDIA. Cuda programming guide. https://docs.nvidia.com/cuda/, 2008. Ac- cessed: 2025-06-18
2008
-
[218]
Basketball on Paper: Rules and Tools for Performance Analysis
Dean Oliver. Basketball on Paper: Rules and Tools for Performance Analysis. University of Nebraska Press, Lincoln, NE, 2004
2004
-
[219]
Czarnecki, Marc Lanctot, Julien Perolat, and Remi Munos
Shayegan Omidshafiei, Christos Papadimitriou, Georgios Piliouras, Karl Tuyls, Mark Rowland, Jean-Baptiste Lespiau, Wojciech M. Czarnecki, Marc Lanctot, Julien Perolat, and Remi Munos. Alpha-rank: Multi-agent evaluation by evolution. Scientific Reports, 9(1):9937, 2019
2019
-
[220]
OpenAI Five
OpenAI. OpenAI Five. https://blog.openai.com/openai-five/, 2018
2018
-
[221]
Introducing gpt-5
OpenAI. Introducing gpt-5. https://openai.com/zh-Hant/index/ introducing-gpt-5/, 2025. Accessed: 2025-08-15
2025
-
[222]
Deep explo- ration via bootstrapped dqn
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy. Deep explo- ration via bootstrapped dqn. In Advances in Neural Information Processing Systems (NeurIPS), 2016
2016
-
[223]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welin- der, Paul F. Christiano, Jan Le...
2022
-
[224]
Bleu: A method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: A method for automatic evaluation of machine translation. In Annual Meeting of the Association for Computational Linguistics (ACL), 2002
2002
-
[225]
O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S
Joon Sung Park, Joseph C. O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. Generative agents: Interactive simulacra of human behavior. In ACM Symposium on User Interface Software and Technology (UIST), 2023. 279
2023
-
[226]
Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmai- son, Andreas Köpf, Edward Z. Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner...
2019
-
[227]
Efros, and Trevor Darrell
Deepak Pathak, Pulkit Agrawal, Alexei A. Efros, and Trevor Darrell. Curiosity-driven ex- ploration by self-supervised prediction. In International Conference on Machine Learn- ing (ICML), 2017
2017
-
[228]
Ivan P. Pavlov. Conditioned Reflexes: An Investigation of the Physiological Activity of the Cerebral Cortex. Oxford University Press, 1927. Translated by G. V . Anrep
1927
-
[229]
Arithmetices Principia: Nova Methodo Exposita
Giuseppe Peano. Arithmetices Principia: Nova Methodo Exposita. Fratres Bocca, Turin,
-
[230]
Bilevel entropy based mech- anism design for balancing meta in video games
Sumedh Pendurkar, Chris Chow, Luo Jie, and Guni Sharon. Bilevel entropy based mech- anism design for balancing meta in video games. In International Conference on Au- tonomous Agents and Multiagent Systems (AAMAS), 2023
2023
-
[231]
Human- ity’s last exam
Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Chen Bo Calvin Zhang, Mohamed Shaaban, John Ling, Sean Shi, Michael Choi, Anish Agrawal, Arnav Chopra, Adam Khoja, Ryan Kim, Richard Ren, Jason Hausenloy, Oliver Zhang, Mantas Mazeika, Summer Yue, Alexa...
2025 arXiv
-
[232]
Randy C. Ploetz. Panama disease: An old nemesis rears its ugly head: Part 1. the begin- nings of the banana export trades. Plant Health Progress, 6(1):18, 2005
2005
-
[234]
Observe and look further: Achieving consistent performance on atari
Tobias Pohlen, Bilal Piot, Todd Hester, Mohammad Gheshlaghi Azar, Dan Horgan, David Budden, Gabriel Barth-Maron, Hado Van Hasselt, John Quan, Mel Vecerík, Mat- teo Hessel, Rémi Munos, and Olivier Pietquin. Observe and look further: Achieving consistent performance on atari. Co...
2018 arXiv
-
[235]
Pomerleau
Dean A. Pomerleau. Alvinn: An autonomous land vehicle in a neural network. In Advances in Neural Information Processing Systems (NeurIPS), 1988
1988
-
[236]
Matchmaking problems in moba games
Muhammad Farrel Pramono, Kevin Renalda, and Harco Leslie Hendric Spits Warnars. Matchmaking problems in moba games. Indonesian Journal of Electrical Engineering and Computer Science, 2018
2018
-
[237]
Pugh, Lawrence B
Justin K. Pugh, Lawrence B. Soros, and Kenneth O. Stanley. Quality diversity: A new frontier for evolutionary computation. Frontiers in Robotics and AI, 3:40, 2016
2016
-
[238]
Connor, Neil Burch, Thomas W
Julien Pérolat, Bart De Vylder, Daniel Hennes, Eugene Tarassov, Florian Strub, Vin- cent de Boer, Paul Muller, Jerome T. Connor, Neil Burch, Thomas W. Anthony, Stephen McAleer, Romuald Elie, Sarah H. Cen, Zhe Wang, Audrunas Gruslys, Aleksandra Maly- sheva, Mina Khan, Sherjil O...
2022
-
[239]
Ross Quinlan
J. Ross Quinlan. Induction of decision trees. Machine Learning, 1:81–106, 1986
1986
-
[240]
Ross Quinlan
J. Ross Quinlan. C4.5: Programs for Machine Learning. Morgan Kaufmann, 1993
1993
-
[241]
Improving lan- guage understanding by generative pre-training, 2018
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving lan- guage understanding by generative pre-training, 2018. Technical report
2018
-
[242]
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. Technical report, OpenAI, 2019. Technical report
2019
-
[243]
Language models are unsupervised mul- titask learners
Alec Radford, Jeffrey Wu, Rewon Child, et al. Language models are unsupervised mul- titask learners. OpenAI Blog, 1(8), 2019. 281
2019
-
[244]
Learning transferable visual models from natural language supervi- sion
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervi- sion. In International C...
2021
-
[245]
Bayesian inverse reinforcement learning
Deepak Ramachandran and Eyal Amir. Bayesian inverse reinforcement learning. In International Joint Conference on Artificial Intelligence (IJCAI), 2007
2007
-
[246]
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea V oss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In International Conference on Machine Learning (ICML), 2021
2021
-
[247]
Scott E. Reed, Konrad Zolna, Emilio Parisotto, Sergio Gómez Colmenarejo, Alexan- der Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, Tom Eccles, Jake Bruce, Ali Razavi, Ashley Edwards, Nicolas Heess, Yutian Chen, Raia Hadsell, Orio...
2022
-
[248]
Faster r-cnn: Towards real- time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real- time object detection with region proposal networks. In Advances in Neural Information Processing Systems (NeurIPS), 2015
2015
-
[249]
Intrinsic motivation and flow
Falko Rheinberg. Intrinsic motivation and flow. Motivation Science, 2020
2020
-
[250]
The roles of inducer size and distance in the ebbinghaus illusion (titchener circles)
Brian Roberts, Mike G Harris, and Tim A Yates. The roles of inducer size and distance in the ebbinghaus illusion (titchener circles). Perception, 34(7):847–856, 2005
2005
-
[251]
Vision-language models are zero-shot reward models for reinforcement learning
Juan Rocamonde, Victoriano Montesinos, Elvis Nava, Ethan Perez, and David Lindner. Vision-language models are zero-shot reward models for reinforcement learning. In In- ternational Conference on Learning Representations (ICLR), 2024
2024
-
[252]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022
2022
-
[253]
The perceptron: A probabilistic model for information storage and organization in the brain
Frank Rosenblatt. The perceptron: A probabilistic model for information storage and organization in the brain. Psychological Review, 65(6):386–408, 1958. 282
1958
-
[254]
Gordon, and Drew Bagnell
Stéphane Ross, Geoffrey J. Gordon, and Drew Bagnell. A reduction of imitation learning and structured prediction to no-regret online learning. In International Conference on Artificial Intelligence and Statistics (AISTATS), 2011
2011
-
[255]
Rumelhart, Geoffrey E
David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams. Learning representa- tions by back-propagating errors. Nature, 323:533–536, 1986
1986
-
[256]
Human-compatible artificial intelligence
Stuart Russell. Human-compatible artificial intelligence. In Human-Like Machine Intel- ligence, pp. 3–23. Oxford University Press, 2022
2022
-
[257]
Artificial Intelligence: A Modern Approach (4th Edi- tion)
Stuart Russell and Peter Norvig. Artificial Intelligence: A Modern Approach (4th Edi- tion). Pearson, 2020
2020
-
[258]
Denton, Seyed Kamyar Seyed Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L. Denton, Seyed Kamyar Seyed Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J. Fleet, and Mohammad Norouzi. Photorealistic text- to-image diffusion models with d...
2022
-
[259]
Mehdi S. M. Sajjadi, Olivier Bachem, Mario Lucic, Olivier Bousquet, and Sylvain Gelly. Assessing generative models via precision and recall. InAdvances in Neural Information Processing Systems (NeurIPS), 2018
2018
-
[260]
Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen
Tim Salimans, Ian J. Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. Improved techniques for training gans. In Advances in Neural Information Processing Systems (NIPS), 2016
2016
-
[261]
Learning to fly
Claude Sammut, Scott Hurst, Dana Kedzier, and Donald Michie. Learning to fly. In International Workshop on Machine Learning (ML), 1992
1992
-
[262]
The Art of Game Design: A Book of Lenses
Jesse Schell. The Art of Game Design: A Book of Lenses. CRC Press, 2008
2008
-
[263]
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, Timothy Lillicrap, and David Silver. Mastering atari, go, chess and shogi by planning with a learned model. Natur...
2020
-
[264]
Prox- imal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Prox- imal policy optimization algorithms. CoRR, abs/1707.06347, 2017
2017 arXiv
-
[265]
Schumpeter
Joseph A. Schumpeter. The Theory of Economic Development. Harvard University Press, 1934
1934
-
[266]
Time-contrastive net- works: Self-supervised learning from multi-view observation
Pierre Sermanet, Corey Lynch, Jasmine Hsu, and Sergey Levine. Time-contrastive net- works: Self-supervised learning from multi-view observation. In IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPR Workshops), 2017
2017
-
[267]
Neurosymbolic artificial intelligence (why, what, and how)
Amit Sheth, Kaushik Roy, and Manas Gaur. Neurosymbolic artificial intelligence (why, what, and how). IEEE Intelligent Systems, 38(3):56–62, 2023
2023
-
[268]
Exploration into translation-equivariant image quantization
Woncheol Shin, Gyubok Lee, Jiyoung Lee, Eunyi Lyou, Joonseok Lee, and Edward Choi. Exploration into translation-equivariant image quantization. In International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023
2023
-
[269]
Bridging the gap between ethics and practice: Guidelines for reliable, safe, and trustworthy human-centered AI systems
Ben Shneiderman. Bridging the gap between ethics and practice: Guidelines for reliable, safe, and trustworthy human-centered AI systems. ACM Transactions on Interactive Intelligent Systems, 10, 2020
2020
-
[270]
David Silver and Richard S. Sutton. Welcome to the era of experience. https://storage.googleapis.com/deepmind-media/Era-of-Experience% 20/The%20Era%20of%20Experience%20Paper.pdf, 2025
2025
-
[271]
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Vedavyas Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy P. Lillicrap,...
2016
-
[272]
Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy P. Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis. Maste...
2017
-
[273]
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, and Demis Hassabis. A general reinforcement learning algorithm that masters chess, shogi, and go through self-play. Sc...
2018
-
[274]
David Silver, Satinder Singh, Doina Precup, and Richard S. Sutton. Reward is enough. Artificial Intelligence, 299:103535, 2021
2021
-
[275]
Georg Simmel. Fashion. American Journal of Sociology, 62(6):541–558, 1957
1957
-
[276]
Herbert A. Simon. Models of Man: Social and Rational. Wiley, 1957
1957
-
[277]
Very deep convolutional networks for large- scale image recognition
Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large- scale image recognition. International Conference on Learning Representations (ICLR), 2015
2015
-
[278]
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In International Conference on Learning Representations (ICLR), 2021
2021
-
[279]
Baruch Spinoza. Ethics. Anonymous, 1677. English translation included in *A Spinoza Reader: The Ethics and Other Works*, translated and edited by Edwin Curley, Princeton University Press, 1994
1994
-
[280]
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: A simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15:1929–1958, 2014
1929
-
[281]
The Berg Companion to Fashion
Valerie Steele (ed.). The Berg Companion to Fashion. Bloomsbury Publishing, London, 2015
2015
-
[282]
Fusarial Wilt (Panama Disease) of Bananas and Other Musa Species
Robert Harry Stover. Fusarial Wilt (Panama Disease) of Bananas and Other Musa Species. Commonwealth Mycological Institute, 1962
1962
-
[283]
Pilgrim in the Microworld
David Sudnow. Pilgrim in the Microworld. Warner Books, New York, 1983
1983
-
[284]
Richard S. Sutton. The reward hypothesis. http://incompleteideas.net/rlai.cs. ualberta.ca/RLAI/rewardhypothesis.html, 2004. 285
2004
-
[285]
Sutton and Andrew G
Richard S. Sutton and Andrew G. Barto. Reinforcement Learning: An Introduction. MIT Press, second edition, 2018
2018
-
[286]
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015
2015
-
[287]
Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott E. Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015
2015
-
[288]
#exploration: A study of count-based ex- ploration for deep reinforcement learning
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel. #exploration: A study of count-based ex- ploration for deep reinforcement learning. InAdvances in Neural Information Processing Systems (NeurIPS), 2017
2017
-
[289]
Thorndike
Edward L. Thorndike. Animal intelligence: An experimental study of the associative processes in animals. The Psychological Review: Monograph Supplements , 2(4):i–109,
-
[290]
On aims and methods of ethology
Niko Tinbergen. On aims and methods of ethology. Zeitschrift für Tierpsychologie, 20 (4):410–433, 1963
1963
-
[291]
Champandard, Pier Luca Lanzi, Michael Mateas, Ana Paiva, Mike Preuss, and Kenneth O
Julian Togelius, Alex J. Champandard, Pier Luca Lanzi, Michael Mateas, Ana Paiva, Mike Preuss, and Kenneth O. Stanley. Procedural content generation: Goals, challenges and actionable steps. In Artificial and Computational Intelligence in Games . Schloss Dagstuhl - Leibniz-Zent...
2013
-
[292]
Behavioral cloning from observation
Faraz Torabi, Garrett Warnell, and Peter Stone. Behavioral cloning from observation. In International Joint Conference on Artificial Intelligence (IJCAI), 2018
2018
-
[293]
Generative adversarial imitation from observation
Faraz Torabi, Garrett Warnell, and Peter Stone. Generative adversarial imitation from observation. In ICML Workshop on Imitation, Intent, and Interaction (I3), 2019. 286
2019
-
[294]
Klassen, Richard Anthony Valenzano, and Sheila A
Rodrigo Toro Icarte, Toryn Q. Klassen, Richard Anthony Valenzano, and Sheila A. McIl- raith. Reward machines: Exploiting reward function structure in reinforcement learning. Journal of Artificial Intelligence Research, 73:173–208, 2022
2022
-
[295]
Is deep reinforcement learning really superhuman on atari? leveling the playing field
Marin Toromanoff, Emilie Wirbel, and Fabien Moutarde. Is deep reinforcement learning really superhuman on atari? leveling the playing field. CoRR, abs/1908.04683, 2019
1908 arXiv
-
[296]
Pure end-to-end training car racing game ai bot using deep reinforce- ment learning
Cheng-Lun Tsai. Pure end-to-end training car racing game ai bot using deep reinforce- ment learning. Master’s thesis, National Chiao Tung University, 2018
2018
-
[1781]
Translation based on the 1781 (A) and 1787 (B) editions
This edition published in 1998. Translation based on the 1781 (A) and 1787 (B) editions
1998
-
[1889]
Title translation: The Principles of Arithmetic, Presented by a New Method
In Latin. Title translation: The Principles of Arithmetic, Presented by a New Method
-
[1898]
Doctoral dissertation, Columbia University
-
[2025]
URL https://www.blog.udonis.co/mobile-marketing/mobile-games/ gaming-industry. Udonis
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.