Pith. sign in

Paper Citation Record · LEDGER

Large Language Model-Enhanced Multi-Armed Bandits

As of 9 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 4 inbound Pith citation observations for arXiv:2502.01118.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01118 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T16:35:28.447005Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T23:51:50.319798Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:09:45.028347Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7ac073ff-9d07-4bd0-a484-31bec8094c44 · outbound

This paper cites Exploring Large Language Model based Intelligent Agents: Definitions, Methods, and Prospects.

Large Language Model-Enhanced Multi-Armed Bandits Exploring Large Language Model based Intelligent Agents: Definitions, Methods, and Prospects

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.399651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.399651Z digest=sha256:aaf37ccaedfe3d0fae30a0f29d81e337de637b3239562aab9388ad4f83ccee32

Observation 6ae5361e-2716-490a-ac7a-b3e893acc484 · outbound

This paper cites In-context Exploration-Exploitation for Reinforcement Learning.

Large Language Model-Enhanced Multi-Armed Bandits In-context Exploration-Exploitation for Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.402832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.402832Z digest=sha256:96751c51291d36ed64f724c06c3de094cbc00628c63229b59a25b612b35f610e

Observation 9ef088c9-3e97-4d6a-9054-6f8144f7cdf2 · outbound

This paper cites Efficient Exploration for LLMs.

Large Language Model-Enhanced Multi-Armed Bandits Efficient Exploration for LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.405762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.405762Z digest=sha256:f8594bdcdf0b6df83fa4d30dbd851c6d033ed2380c4367c26f5a61c290495509

Observation b3b1b16b-31b2-4072-ac64-9ff18668e5ef · outbound

This paper cites Can large language models explore in-context?.

Large Language Model-Enhanced Multi-Armed Bandits Can large language models explore in-context?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.412559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.412559Z digest=sha256:1e20fb0860d7cc783c268ef3a8b519022577d797c2a4b7adbba9e1c4c23d46ca

Observation 11fce7a2-85a6-4237-a1a5-511c898a5531 · outbound

This paper cites In-context Reinforcement Learning with Algorithm Distillation.

Large Language Model-Enhanced Multi-Armed Bandits In-context Reinforcement Learning with Algorithm Distillation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.414618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.414618Z digest=sha256:ed7fa578c51ad40772bf61c15c5fc8842437aa98626f1d5275439ba2e107486e

Observation 00a4f45e-0efa-4836-b504-e9d32e48dcd5 · outbound

This paper cites Feel-Good Thompson Sampling for Contextual Dueling Bandits.

Large Language Model-Enhanced Multi-Armed Bandits Feel-Good Thompson Sampling for Contextual Dueling Bandits

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.416638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.416638Z digest=sha256:2ce41cc101e9d936cbb8bdd4172c7a05a629fba9d77cdd2bec17753c8385004e

Observation 13f79b30-3afd-46c1-8894-1d1359c4d537 · outbound

This paper cites Prompt Optimization with Human Feedback.

Large Language Model-Enhanced Multi-Armed Bandits Prompt Optimization with Human Feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.419518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.419518Z digest=sha256:5205649f45fdab03eb95e3d1cbe73657442f76bb47a115389e0d99893c473268

Observation e19a04e7-8d86-4d17-8419-f2bd12ee4e85 · outbound

This paper cites DeepSeek-V3 Technical Report.

Large Language Model-Enhanced Multi-Armed Bandits DeepSeek-V3 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.422266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.422266Z digest=sha256:9023330859610e5d06e9a3dbde0f52b79c825187b2795555a95fbc65ad33582b

Observation d9640052-2ea4-4ec3-ac20-e542a866f8f9 · outbound

This paper cites GPT-4 Technical Report.

Large Language Model-Enhanced Multi-Armed Bandits GPT-4 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.427581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.427581Z digest=sha256:8492651f50dc9aba3f28661abb503fd0e62711f9c3eb1790cd701daf61c214f9

Observation ef8e30af-bcb0-4fed-aa18-796d22b2f942 · outbound

This paper cites Wikilinks: A large-scale cross-document coreference cor- pus labeled via links to wikipedia.

Large Language Model-Enhanced Multi-Armed Bandits Wikilinks: A large-scale cross-document coreference cor- pus labeled via links to wikipedia

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:35:29.317407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T16:35:28.430202Z digest=sha256:bd80e98b31f3e27872078fa755dcc402d3ab03d781119d3ed256b65018312cec

Observation 8e777ba9-046d-4891-8604-e21799a6a796 · outbound

This paper cites SmartPlay: A Benchmark for LLMs as Intelligent Agents.

Large Language Model-Enhanced Multi-Armed Bandits SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.435740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.435740Z digest=sha256:97c77924231fad30ea0b4224a52233cdd432eb7555b64a6d3d66161b14132f1a

Observation 81caac2d-8707-4acf-ac4f-aa8752b83d8c · outbound

This paper cites The Rise and Potential of Large Language Model Based Agents: A Survey.

Large Language Model-Enhanced Multi-Armed Bandits The Rise and Potential of Large Language Model Based Agents: A Survey

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.438561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.438561Z digest=sha256:b2d68afdd1ae90c18317cc28e91168b3e42e86d28d0c185c9dde8f778dc0b207

Observation 721aa724-5544-4a75-a6b6-cd19e0770c73 · outbound

This paper cites AgentGym: Evolving Large Language Model-based Agents across Diverse Environments.

Large Language Model-Enhanced Multi-Armed Bandits AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.441317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.441317Z digest=sha256:3418be5482dea772b7dc66ffc1e0536b537752d3d12083cef8e3cd869262365d

Observation da32c694-8115-4e1f-86e5-32a9ef833f86 · outbound

This paper cites Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents.

Large Language Model-Enhanced Multi-Armed Bandits Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.444193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.444193Z digest=sha256:05173c7463380d9dee34a4caaff8a7b73926555d378cb2aa446c107f87569f7b

Observation 4da292a3-76d7-43e0-83ae-7134394e7c3b · outbound

This paper cites Y ., McAleer, S., Fried, D., and Salakhutdinov, R.

Large Language Model-Enhanced Multi-Armed Bandits Y ., McAleer, S., Fried, D., and Salakhutdinov, R

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.410526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.410526Z digest=sha256:0e300b17c667ed18456648d72d9746e7d0da5fdcd7383763ed40d768ff6a484d

Observation 4bfc6eeb-f2ca-4590-b52d-89aa034c0b8f · outbound

This paper cites LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning.

Large Language Model-Enhanced Multi-Armed Bandits LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.447005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.447005Z digest=sha256:26207b7af26e4688218625da2f3548a32bd99149ada37c99abf7c330499439c1

Observation 9f9b5af0-f6c9-4a12-b53c-52a27ba770c1 · outbound

This paper cites Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning.

Large Language Model-Enhanced Multi-Armed Bandits Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.392673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.392673Z digest=sha256:76ae118399f94887acd063a99c254d5ad3a2fa42870492bdd2f36b660a6ed6a1

Observation 6fa7b74c-07aa-4730-814e-0f0e3f15ad62 · outbound

This paper cites Reasoning with Language Model is Planning with World Model.

Large Language Model-Enhanced Multi-Armed Bandits Reasoning with Language Model is Planning with World Model

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.408114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.408114Z digest=sha256:9e529c450485674596b82bf6983c30230542bae0e09a312094128bb49fc1e1f1

Observation 60bec8d0-2692-46cf-91bb-7a64c7da2c42 · outbound

This paper cites Neural Dueling Bandits: Preference-Based Optimization with Human Feedback.

Large Language Model-Enhanced Multi-Armed Bandits Neural Dueling Bandits: Preference-Based Optimization with Human Feedback

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.432870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.432870Z digest=sha256:c7f9a8765356a180e4f00ab635614a3ba9618f4acfeb70c2eb79b002264d4e3b

Observation 14a80937-72c7-4b81-8c85-3dde3647387a · outbound

This paper cites P., Xie, Q., and Nowak, R.

Large Language Model-Enhanced Multi-Armed Bandits P., Xie, Q., and Nowak, R

Reference 2023

Resolution
verified exact
arxiv_id, observed 2026-08-09T16:35:28.672114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T16:35:28.425078Z digest=sha256:d8a1470ba21cb1110dcbbed2678d004f71ac2f53da67209b64798fcb81160a4b

Observation 562233f7-1593-4733-bd40-d1f9d4fb1cd3 · outbound

This paper cites Efficient Sequential Decision Making with Large Language Models.

Large Language Model-Enhanced Multi-Armed Bandits Efficient Sequential Decision Making with Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.396588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.396588Z digest=sha256:1ddcdb6a9dbce6aa113a9a87795862b99ea3ef5e37f45ab473f43c94b7401ec0

Pith citing papers

Observation 7e85714f-417c-4951-a850-a56ad7447dad · inbound

When Do We Need LLMs? A Diagnostic for Language-Driven Bandits cites this paper.

When Do We Need LLMs? A Diagnostic for Language-Driven Bandits Large Language Model-Enhanced Multi-Armed Bandits

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:25:51.129269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T18:32:12.396568Z digest=sha256:5a5a112517b19283bedff23f4c922a100b434f68fcc08b6f7e4729ac653d3445

Observation e24fc470-732f-419d-8a18-c5e79c456709 · inbound

Calibration-Gated LLM Pseudo-Observations for Online Contextual Bandits cites this paper.

Calibration-Gated LLM Pseudo-Observations for Online Contextual Bandits Large Language Model-Enhanced Multi-Armed Bandits

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:25:22.579124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T12:22:28.655926Z digest=sha256:5aa8d2773f6cca0af071bdc98e582e04da2891973c3e6561cdbe702cb7f0e51e

Observation 925a4193-bd9b-4dae-80d2-88e79e406c0b · inbound

GRIMIP: A General Framework for Instance-Specific Configuration of MIP Solvers Using LLMs cites this paper.

GRIMIP: A General Framework for Instance-Specific Configuration of MIP Solvers Using LLMs Large Language Model-Enhanced Multi-Armed Bandits

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:09:45.029930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T09:05:56.687400Z digest=sha256:7095a223eb0d3977adb1d1a3e8bb31652618fee8135caf25a27ecd4b3524bedc

Observation 72a679c9-9451-436b-84b2-a0ae00b3e3ba · inbound

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex cites this paper.

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex Large Language Model-Enhanced Multi-Armed Bandits

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T23:51:50.319798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:51:50.319798Z digest=sha256:84247bd7b1ce970814b39bd7189b7952908a83cbb1326f9ea6285b087796fbf8