Pith. sign in

Paper Citation Record · LEDGER

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison

As of 17 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2506.14448.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14448 v2

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:24:26.960285Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T18:48:30.813878Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T22:37:25.557248Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7f894705-7617-445f-8c1b-2b17ce80d21a · outbound

This paper cites online" 'onlinestring :=.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:23.786452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:23.786452Z digest=sha256:cbaa9dae720d3cc56a13abc8e221ea7e0357b763d1475411fc4c0d634eca2648

Observation 3dd77321-6e4e-4a94-aace-179706d84337 · outbound

This paper cites write newline.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:23.859578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:23.859578Z digest=sha256:a08be6151fbc9800ec411116a5f423409a0ad8af6f01b848c66704960945c60a

Observation bb8ea895-684c-4635-99f3-d9721e5a0e1a · outbound

This paper cites LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:23.916722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:23.916722Z digest=sha256:6f884dba9bd2ae8e77fb55102059b106744ecc626431d8530418526a6afd130a

Observation b4d3df70-0895-4a22-b42d-a4c059332ce1 · outbound

This paper cites GPT-4 Technical Report.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison GPT-4 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:24.013600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:24.013600Z digest=sha256:887ced8d13f18e47ddfc5e126c02d7f47a20df28e00452b108d28cb56c7ec976

Observation c29089f0-f925-4735-9590-1650f110ef0c · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:30.153815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:24:24.080497Z digest=sha256:2f7d3b08e87fcd965bc9e9f57a297a64b24ae8b005a72ad6396dbe763665e3b6

Observation 75ab87cd-d0fa-4021-9bcb-abe44b8215b0 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:24.170968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:24.170968Z digest=sha256:f3f3d3ac2c473c069a3b79106f187fe02f183829419d8534efaa64e0920a7aa8

Observation ac3da65f-7370-49dc-b9df-67f934ddbf01 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Gonzalez, Ion Stoica, and Eric P

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:24.254343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:24.254343Z digest=sha256:b8ee96ee2c3c409083466a2e1ea1817b2c96b237fa04b801e40f3f9c897c314a

Observation b7f39ecf-b494-496b-9b11-0469089520fe · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:24.333492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:24.333492Z digest=sha256:a0e78cc7aa84d43558a579aeb5e84991f39e6960b34cdef60ac63ad1cd28ff9c

Observation 9cfcd0bb-fadb-4e6b-8ccf-fde2171ef875 · outbound

This paper cites RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:24.397362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:24.397362Z digest=sha256:7778cc7e284a328bad79648d9dc8ddd85c4a4eb908b18e2005154cfbd2c7c4d4

Observation ae836f5d-2ec9-4190-ab73-636145995dc8 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:30.005234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:24:24.455525Z digest=sha256:c4d5afa7fb618b55f3ad8cbd904d3e15ae5a89591a4cfe2c558bd711d32e779f

Observation d9002957-21d9-4875-a897-03d251502dee · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:29.877594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:24:24.511202Z digest=sha256:0ab1ac9cc09267aa4403118390b1067d16a19576a6077830f90d5a23ae3447d3

Observation 9e67fbb3-4309-435e-a7bb-998a4539be9f · outbound

This paper cites Amago: Scalable in-context reinforcement learning for adaptive agents.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Amago: Scalable in-context reinforcement learning for adaptive agents

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:24:29.723566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:24:24.568343Z digest=sha256:bbf40c4895cf9298959d6b9d3c608d51ae4796f8ec411c55ecd2cb81b01f6d19

Observation ba936084-338a-4836-8f68-876ad2fc79ea · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:24.623902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:24.623902Z digest=sha256:d6c7fbc1a1a543d0353ce58dec1669c31d6f9fabe9dd84085844c2423b29d4b7

Observation 4e48c53b-f9ff-4c69-9468-c13145611a1f · outbound

This paper cites GPT-4o System Card.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison GPT-4o System Card

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:24.682267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:24.682267Z digest=sha256:55d62ec98f1bb5c4507bbf4043c688e581d51311b2411a74a073e7ed34411ad9

Observation ad5dd136-440b-4ddf-8c69-b71d2f8a32f5 · outbound

This paper cites OpenAI o1 System Card.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison OpenAI o1 System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:24.783283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:24.783283Z digest=sha256:2d11d4a41f2c57c9a4f0de1da32c1c26665e99042327994db60a1f103737c231

Observation 12608b17-a6aa-4d7f-97d5-7625df16776e · outbound

This paper cites SelfEvolve: A Code Evolution Framework via Large Language Models.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison SelfEvolve: A Code Evolution Framework via Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:24.831453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:24.831453Z digest=sha256:bc242e6dd180eef5bddf381f9cb8674c182e4c98a12260ceb59bbd8cf55d446e

Observation 2b5855c9-6d5c-4dd9-8bee-06200e41d7e9 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:29.538934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:24:24.881597Z digest=sha256:2779b459236c9a3d55ca670ca3ea14d6f151153b70aa3b72a976cb1de0073d0a

Observation bba2ab71-24b2-488c-b42d-8d7e8f632fa1 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:29.358637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:24:25.014205Z digest=sha256:37523e5ac79852f2c37cc24fc76cf2e3d01bc6c446d0a2e52c9aaf9f25a8e407

Observation 3b1caedf-ded2-4acc-a86c-4754fe1e3c96 · outbound

This paper cites In-context reinforcement learning with algorithm distillation.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison In-context reinforcement learning with algorithm distillation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:24:29.146080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:24:25.090138Z digest=sha256:e98d73881d959125c01f63e6fa21df185feb9994d58473e0026aac5cedbfb0dc

Observation 4c6af786-b0d3-4876-890e-8458e295b795 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:28.960005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:24:25.170461Z digest=sha256:f62778253ca0d6963f668499fe26ce8b86e573be5849e91c54f150d9ab06df66

Observation 47ec9c2b-cedd-4086-a6e1-e12d4bed5258 · outbound

This paper cites DeepSeek-V3 Technical Report.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison DeepSeek-V3 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:25.256888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:25.256888Z digest=sha256:4c2594ddd6c8ddb654ce2f8e970428f5f272f0c55b4b465446f2a3ecdcc68d75

Observation 573640db-7a0d-4cfd-a009-6ce75db25cc0 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:28.760291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:24:25.350340Z digest=sha256:339d62d86805bc78cf5831c5fbc7ca56cecc150d825fdd18a4d6362fd47114ec

Observation 5071ce91-e83c-46c5-91d4-ad1c1eda275a · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:28.498765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:24:25.479508Z digest=sha256:700205162bc7bba983ecec34b8ee9036c0904608d44c854b4012777092d42ada

Observation 73be322b-c6bf-45bc-9a77-51b52d489d37 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:28.249737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:24:25.552776Z digest=sha256:d48d7378c177494c274e4d959d04c1350ae30bafd0c34ea75d62ca5ff7921445

Observation 16619bf5-2e33-4f78-b84c-67bd1957301d · outbound

This paper cites WizardCoder: Empowering Code Large Language Models with Evol-Instruct.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison WizardCoder: Empowering Code Large Language Models with Evol-Instruct

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:25.657136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:25.657136Z digest=sha256:d54735c0ea72dbb809df980b411b248584bea765ff7429f2338114cf2251655a

Observation 9ec30262-2949-4f68-9c9c-9d176f8d17af · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:28.070041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:24:25.734272Z digest=sha256:4d492fe64cf0b07f26dd206f7d54244b9acd8ddb043aa3980ecf077dfd987317

Observation 50a679f9-b410-483b-b832-9b35f0a99539 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:25.851359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:25.851359Z digest=sha256:672a173b5c1c8b87cd6b68c22b140bc8b902597040e48d4cbcdf930ae588086d

Observation 923dfd12-a7d8-45a1-ab01-8732785c259d · outbound

This paper cites POPGym: Benchmarking Partially Observable Reinforcement Learning.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison POPGym: Benchmarking Partially Observable Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:25.955282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:25.955282Z digest=sha256:2ae84ec6180ea02c1d4c743a1b0c04a35988b96b67cd23b3f2aa55e524e04ebf

Observation 02ef6cc0-cd2b-4919-9ef7-840083d83377 · outbound

This paper cites Investigate-Consolidate-Exploit: A General Strategy for Inter-Task Agent Self-Evolution.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Investigate-Consolidate-Exploit: A General Strategy for Inter-Task Agent Self-Evolution

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:26.043298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:26.043298Z digest=sha256:58a7ccc9535dc805c6d5adba7a40c3e07c7c51ead5cfba60358c5676cce4d0c3

Observation 5b7c9f4d-3da1-4bed-b0ac-24bc2df53085 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:27.925468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:24:26.131603Z digest=sha256:779e77e304f3fa99d70faee0458d14a43a9e3f10043a8478accffcd21f146d54

Observation 8023024d-db22-4845-85ff-8ddb260776a7 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:27.763206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:24:26.237012Z digest=sha256:8187b62b32e2e8299377bc9fec975678e0f4e3e295baf7c1f3cd71f444fefe39

Observation 5f30a8df-16de-4e8e-86ce-a5ff348464a1 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:26.328652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:26.328652Z digest=sha256:f2f5d75b71523542f2ca8c52a15cdecc02703f71c31b28b8fc4387f9e72f0c51

Observation fe21e204-5cd9-4e77-9bd2-4bc281f6dd26 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:27.572140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:24:26.426949Z digest=sha256:50d2ea4f20a77a47aafe85a8f1435e036fe31c21d4c1e14a590597645020dc52

Observation c2ad8fd8-3413-4166-9fde-dfe714fb6a3d · outbound

This paper cites Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:26.491541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:26.491541Z digest=sha256:69cfcc9eacf31aaed6cfd4df82c3619bca52fb630f176c049785067a3ee26274

Observation 6bf691fe-db50-4fee-9b98-96c3a4b4752c · outbound

This paper cites A Survey on Self-Evolution of Large Language Models.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison A Survey on Self-Evolution of Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:26.554434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:26.554434Z digest=sha256:27dd7c386bfc57cdb684d639225736bf16ee268792cd20170c3b18e0bf573b89

Observation 6b4784a3-2c23-40a1-b635-1fd39b2e067a · outbound

This paper cites MAgIC: Investigation of Large Language Model Powered Multi-Agent in Cognition, Adaptability, Rationality and Collaboration.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison MAgIC: Investigation of Large Language Model Powered Multi-Agent in Cognition, Adaptability, Rationality and Collaboration

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:26.690889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:26.690889Z digest=sha256:b2347b38bd0127e22deb4f77f1b12601f1110d9e3405a4ec9090fcb5e49e17ff

Observation f1ec01af-17f8-4dd0-8250-045b5c5082d9 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:26.798146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:26.798146Z digest=sha256:c4e9b70a9192a1461b5a1fabf5980d8cdb52cc174de6ce682b62f525fbd54974

Observation e475ccbc-afae-471d-ad57-36119217c0b6 · outbound

This paper cites PolicyEvol-Agent: Evolving Policy via Environment Perception and Self-Awareness with Theory of Mind.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison PolicyEvol-Agent: Evolving Policy via Environment Perception and Self-Awareness with Theory of Mind

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:24:27.227107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:24:26.878378Z digest=sha256:139efe60392b94c3f1aade702102bd8dd406bb86f5615588d528f04e177f52fe

Observation 0cfb8b71-f14a-44d6-b8f3-c7a381c3cc12 · outbound

This paper cites ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:26.960285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:26.960285Z digest=sha256:0ccfba87b5b7c01ec880a15e4bfe297e8721cce54a71440cc435d45e06160671

Pith citing papers

Observation bada8048-acc7-44a9-9366-d65717504109 · inbound

From Player to Master: Enhancing Test-Time Learning of LLM Agents via Reinforcement Learning over Memory cites this paper.

From Player to Master: Enhancing Test-Time Learning of LLM Agents via Reinforcement Learning over Memory How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:37:25.558920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T18:48:30.813878Z digest=sha256:4dd4fe63b5a351f06d56a42c4fe3e4aa00912e910e73328c8699793f5e72242d