Pith. sign in

Paper Citation Record · LEDGER

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison

As of 12 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2506.14448.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14448 v2

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:24:26.960285Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T18:48:30.813878Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T22:37:25.557248Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7f894705-7617-445f-8c1b-2b17ce80d21a · outbound

This paper cites online" 'onlinestring :=.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:23.786452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:23.786452Z digest=sha256:055960308fba6c2ee838f0d30547714241931ceb3297a4d9da446ad800bb8fc6

Observation 3dd77321-6e4e-4a94-aace-179706d84337 · outbound

This paper cites write newline.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:23.859578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:23.859578Z digest=sha256:117af0597395c1f13388fbfddce903ca8d0c052c18acd59a4fed59091d8d8f75

Observation bb8ea895-684c-4635-99f3-d9721e5a0e1a · outbound

This paper cites LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:23.916722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:23.916722Z digest=sha256:b7d859016b905ef4d4ae73e6009c2be32070f32af8dadf98a642d03a363ac9bb

Observation b4d3df70-0895-4a22-b42d-a4c059332ce1 · outbound

This paper cites GPT-4 Technical Report.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison GPT-4 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:24.013600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:24.013600Z digest=sha256:1c4b3e531e089f7cebb78fe2c479df893e8ca252d2912dfc3049719b7b6bb589

Observation c29089f0-f925-4735-9590-1650f110ef0c · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:30.153815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T00:24:24.080497Z digest=sha256:9c3a73a85777c592bf5d8779d8d27a68b117e85d2361efc7c1c26c5727403ac5

Observation 75ab87cd-d0fa-4021-9bcb-abe44b8215b0 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:24.170968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:24.170968Z digest=sha256:825f933cf17d86d701d292ce69e2f1fe07a6b9c94864c80ca21cc7933a8cc17a

Observation ac3da65f-7370-49dc-b9df-67f934ddbf01 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Gonzalez, Ion Stoica, and Eric P

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:24.254343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:24.254343Z digest=sha256:e37c4da9647e7ec145221e526780524cc16ae2c5e30b72259701376c426b8c2f

Observation b7f39ecf-b494-496b-9b11-0469089520fe · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:24.333492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:24.333492Z digest=sha256:744d943ffc77f7cbbd5b8fc29e286542a328fe8dd6b4ff5c9689fef966680ef8

Observation 9cfcd0bb-fadb-4e6b-8ccf-fde2171ef875 · outbound

This paper cites RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:24.397362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:24.397362Z digest=sha256:470443edaa4f3903bebf8e2f0574f84ad523cc82009a153b8e6467b3849e4596

Observation ae836f5d-2ec9-4190-ab73-636145995dc8 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:30.005234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T00:24:24.455525Z digest=sha256:0dd242f7ce2dd615c7e2151a21c088c13c9a9ec520cda41f3ea207d7a6643c50

Observation d9002957-21d9-4875-a897-03d251502dee · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:29.877594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T00:24:24.511202Z digest=sha256:a459ea70efadd53030a0abbdcd4dcc3504bfdebd9b287bb5e0af6041286c1342

Observation 9e67fbb3-4309-435e-a7bb-998a4539be9f · outbound

This paper cites Amago: Scalable in-context reinforcement learning for adaptive agents.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Amago: Scalable in-context reinforcement learning for adaptive agents

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:24:29.723566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T00:24:24.568343Z digest=sha256:3e6883d46e11cd955e8173416d0fdaf4cce881a8081e30c2e8363beef4706992

Observation ba936084-338a-4836-8f68-876ad2fc79ea · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:24.623902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:24.623902Z digest=sha256:9ca2335057e504ee02462dbf410654765d1a13842aeaef2a37266277e735345d

Observation 4e48c53b-f9ff-4c69-9468-c13145611a1f · outbound

This paper cites GPT-4o System Card.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison GPT-4o System Card

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:24.682267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:24.682267Z digest=sha256:24107a1e6f0a56757e103599bd2e0dd0fe92199fb17a967353b348577fa91cd3

Observation ad5dd136-440b-4ddf-8c69-b71d2f8a32f5 · outbound

This paper cites OpenAI o1 System Card.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison OpenAI o1 System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:24.783283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:24.783283Z digest=sha256:197deb9353f388cf872f86215733f4ca8a4ed9e3518114320260be56ab8ac048

Observation 12608b17-a6aa-4d7f-97d5-7625df16776e · outbound

This paper cites SelfEvolve: A Code Evolution Framework via Large Language Models.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison SelfEvolve: A Code Evolution Framework via Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:24.831453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:24.831453Z digest=sha256:74a848550f45fd7f9827bad4b2570656da009c71c98a84ef12b58d529b580d9e

Observation 2b5855c9-6d5c-4dd9-8bee-06200e41d7e9 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:29.538934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T00:24:24.881597Z digest=sha256:1dbd92c2c534627c0aec0067131b9052bb4af23044825739ba82a5873398da65

Observation bba2ab71-24b2-488c-b42d-8d7e8f632fa1 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:29.358637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T00:24:25.014205Z digest=sha256:5bff13b0ced69b23a1a1fc85831e77b77efc6025108ba4265c3b43d8a6271e6f

Observation 3b1caedf-ded2-4acc-a86c-4754fe1e3c96 · outbound

This paper cites In-context reinforcement learning with algorithm distillation.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison In-context reinforcement learning with algorithm distillation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:24:29.146080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T00:24:25.090138Z digest=sha256:5451168423950299741fd2e874b36e1b7c84f709858f8ae9bcd5a4430c542df2

Observation 4c6af786-b0d3-4876-890e-8458e295b795 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:28.960005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T00:24:25.170461Z digest=sha256:2d5879ac5409de3f33e2e2b3ed9c516827967b0d8f31977caabfd3d5ae175e8f

Observation 47ec9c2b-cedd-4086-a6e1-e12d4bed5258 · outbound

This paper cites DeepSeek-V3 Technical Report.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison DeepSeek-V3 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:25.256888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:25.256888Z digest=sha256:bfc7913b692b70222fb7d47a0c00b1b9ca235951b24ff19641e05be81cfb893a

Observation 573640db-7a0d-4cfd-a009-6ce75db25cc0 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:28.760291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T00:24:25.350340Z digest=sha256:e0ef8390e79c5938ec9c3030dda4208366ca634875b9d02a9bd3c005f223bacd

Observation 5071ce91-e83c-46c5-91d4-ad1c1eda275a · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:28.498765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T00:24:25.479508Z digest=sha256:fbae44b1a4dbb4f12d74b444dbbc755342236a005bfbd5c01cf4175f7e0150d0

Observation 73be322b-c6bf-45bc-9a77-51b52d489d37 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:28.249737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T00:24:25.552776Z digest=sha256:1e81f5a8e073520893c513abe19cccacd7f4ed4e477064060beae3a4c8118ba7

Observation 16619bf5-2e33-4f78-b84c-67bd1957301d · outbound

This paper cites WizardCoder: Empowering Code Large Language Models with Evol-Instruct.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison WizardCoder: Empowering Code Large Language Models with Evol-Instruct

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:25.657136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:25.657136Z digest=sha256:46c9d07f843d1c0a84596248cc593d721badc4d8fba24d12faee822ae8fcfe7c

Observation 9ec30262-2949-4f68-9c9c-9d176f8d17af · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:28.070041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T00:24:25.734272Z digest=sha256:52905944865963f762eaae8b9abb1c4ccd16f39ab652c86db9cc4a15fc378741

Observation 50a679f9-b410-483b-b832-9b35f0a99539 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:25.851359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:25.851359Z digest=sha256:2a3a5f635a72ded970111db561163d8fdab324a29c981175d567c304fb1562ea

Observation 923dfd12-a7d8-45a1-ab01-8732785c259d · outbound

This paper cites POPGym: Benchmarking Partially Observable Reinforcement Learning.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison POPGym: Benchmarking Partially Observable Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:25.955282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:25.955282Z digest=sha256:d2485fbb56f9bf9d79f559789634932ab5d0626eea3a1b8de989337eb004393c

Observation 02ef6cc0-cd2b-4919-9ef7-840083d83377 · outbound

This paper cites Investigate-Consolidate-Exploit: A General Strategy for Inter-Task Agent Self-Evolution.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Investigate-Consolidate-Exploit: A General Strategy for Inter-Task Agent Self-Evolution

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:26.043298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:26.043298Z digest=sha256:4f124f6ea286d9bd28e4177fc49577c4009991dfd712b67abad256336bba3e7d

Observation 5b7c9f4d-3da1-4bed-b0ac-24bc2df53085 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:27.925468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T00:24:26.131603Z digest=sha256:620a5a8eac658f0aee88df65429a40741f451b1550b77caffd4134ff511c6d3b

Observation 8023024d-db22-4845-85ff-8ddb260776a7 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:27.763206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T00:24:26.237012Z digest=sha256:8f77dfd46bc19b26f32bc6f76a52f252a8f9da239e1392b123f1d2997e37a51f

Observation 5f30a8df-16de-4e8e-86ce-a5ff348464a1 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:26.328652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:26.328652Z digest=sha256:2229b77492b06c9ef06188031da5575d408166f123c3067200ebd235068d68f1

Observation fe21e204-5cd9-4e77-9bd2-4bc281f6dd26 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:27.572140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T00:24:26.426949Z digest=sha256:04ba3c4dfc4e8d488715a9678462ec3e781a90c80f6d76d8a5350e81a149cd39

Observation c2ad8fd8-3413-4166-9fde-dfe714fb6a3d · outbound

This paper cites Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:26.491541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:26.491541Z digest=sha256:06e03273c2a8b92d13aeb284e0452632ea29a448f3ca568e2181641d491fe276

Observation 6bf691fe-db50-4fee-9b98-96c3a4b4752c · outbound

This paper cites A Survey on Self-Evolution of Large Language Models.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison A Survey on Self-Evolution of Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:26.554434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:26.554434Z digest=sha256:3c083c8206ec2d151271e1270d527962a42c738bdb63c6f1ca98546d62edc6e5

Observation 6b4784a3-2c23-40a1-b635-1fd39b2e067a · outbound

This paper cites MAgIC: Investigation of Large Language Model Powered Multi-Agent in Cognition, Adaptability, Rationality and Collaboration.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison MAgIC: Investigation of Large Language Model Powered Multi-Agent in Cognition, Adaptability, Rationality and Collaboration

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:26.690889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:26.690889Z digest=sha256:253c704cb453fabe08da190bf692a8e52487112d8e3e6b8cc4dbf2f99fa2be2d

Observation f1ec01af-17f8-4dd0-8250-045b5c5082d9 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:26.798146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:26.798146Z digest=sha256:92dc52e378e67e180fb9f292084013097d6bb896f417a7aac744d334a7067e4f

Observation e475ccbc-afae-471d-ad57-36119217c0b6 · outbound

This paper cites PolicyEvol-Agent: Evolving Policy via Environment Perception and Self-Awareness with Theory of Mind.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison PolicyEvol-Agent: Evolving Policy via Environment Perception and Self-Awareness with Theory of Mind

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:24:27.227107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T00:24:26.878378Z digest=sha256:ca6e43d7aacca462fde5ea5300eb9a2b68de13fd3a94d3858b7dcb17639319e0

Observation 0cfb8b71-f14a-44d6-b8f3-c7a381c3cc12 · outbound

This paper cites ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:26.960285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:26.960285Z digest=sha256:b3d6a529ff0dc6a4cca1257a3491b8d5baff67b086445da751efa44a74593883

Pith citing papers

Observation bada8048-acc7-44a9-9366-d65717504109 · inbound

From Player to Master: Enhancing Test-Time Learning of LLM Agents via Reinforcement Learning over Memory cites this paper.

From Player to Master: Enhancing Test-Time Learning of LLM Agents via Reinforcement Learning over Memory How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:37:25.558920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-27T18:48:30.813878Z digest=sha256:03534cb6a43de437a5590d144d7d91dd016476c1b61a7e948cf2e44179ad782c