Pith. sign in

Paper Citation Record · LEDGER

Generating Symbolic World Models via Test-time Scaling of Large Language Models

As of 9 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 3 inbound Pith citation observations for arXiv:2502.04728.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.04728 v2

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T21:45:45.538720Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:14:21.027074Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T23:42:49.549067Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 03cd33d5-726d-41df-939a-57dd9a284dfe · outbound

This paper cites GPT-4 Technical Report.

Generating Symbolic World Models via Test-time Scaling of Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.239493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.239493Z digest=sha256:84d17f1846aa863a468c1b1bc7486b3e2b6746b12e3c828a1aa85d6c2f54e537

Observation 0333c392-25f1-4a9e-baa1-91417723d13c · outbound

This paper cites Learning discrete world models for heuristic search.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Learning discrete world models for heuristic search

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:45:46.845605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T21:45:45.245571Z digest=sha256:a368bfcebda901998a556d61704fb934c62b2e997c2b06e463b8cc005101b653

Observation 36c5cbae-ff49-461e-acf8-180ea3d9df5d · outbound

This paper cites Program Synthesis with Large Language Models.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Program Synthesis with Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.250220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.250220Z digest=sha256:8249e650680e70f76cd82377fdb879e23a1cc1821ff2a4922748e27b33f3a5ac

Observation 6d31b2ee-9447-45c1-8cab-ecccf8d49965 · outbound

This paper cites Learning warm-start points for ac optimal power flow.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Learning warm-start points for ac optimal power flow

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:45:46.831395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T21:45:45.255223Z digest=sha256:d80f6c5b7aaf33d1ea3d00ec6f53c881c0e1f6c9f1235b468125002fdb113631

Observation 9e4b8858-2e3f-44c4-bf71-3ca19a79f864 · outbound

This paper cites Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path?.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.260053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.260053Z digest=sha256:39caad688a091f0bed2d1ee44f3d2992ee2650397d236afd67e398deae33ca8b

Observation 39f13cc1-2b1d-46b8-a381-21be6515f46a · outbound

This paper cites OpenAI Gym.

Generating Symbolic World Models via Test-time Scaling of Large Language Models OpenAI Gym

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.265018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.265018Z digest=sha256:23a7bb8bad848b33192a63c7e7482713e706b23d0bdcac493e46b40492db540f

Observation 85844672-b9dc-4afe-982a-cc1b173a2e71 · outbound

This paper cites Language models are few-shot learners.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Language models are few-shot learners

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:45:46.816778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T21:45:45.270167Z digest=sha256:e4574cecf67d75474bd22c3222e7ac4f2392ad2264ec77c8f3d39826113feffe

Observation bcbf0079-ae87-45ce-989a-c8b37f8c7cad · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Evaluating Large Language Models Trained on Code

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.275330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.275330Z digest=sha256:914d27593abe1f8b7b26ec1351cab3be31500dbe5ed65259e07a304a38b5f16d

Observation ede1b7a6-2aa2-4a1f-8c27-e9012feacaf2 · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.280344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.280344Z digest=sha256:fbd8e6bbca931910d3d66827af23c9ab0fcf4de643d45db501b712d709d81c42

Observation f07b44fe-b95c-485b-ad97-09d742351135 · outbound

This paper cites Inductive or Deductive? Rethinking the Fundamental Reasoning Abilities of LLMs.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Inductive or Deductive? Rethinking the Fundamental Reasoning Abilities of LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.285649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.285649Z digest=sha256:e11e30277d1d46b4edecfafc1f44778e8cf6dd465e5bff642b2109ec2871d60e

Observation 0ea578a5-37f1-4a1d-a3e8-dc8c91c33b24 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.290407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.290407Z digest=sha256:5827f5ffeaad0dad6dc0dd54338eb57cc4aad11c5a9a7dc7ded56df34bdda234

Observation e46c3c3c-c61a-4dc6-93bf-14d15caae9ad · outbound

This paper cites Generating Code World Models with Large Language Models Guided by Monte Carlo Tree Search.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Generating Code World Models with Large Language Models Guided by Monte Carlo Tree Search

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.295334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.295334Z digest=sha256:0203a1d31da62a969211f251c1db1e79127fe5d5265300301649e006336673f7

Observation 9cae3b3f-2bc0-4940-a254-af6b3f0c1477 · outbound

This paper cites Parameter-efficient fine-tuning of large-scale pre-trained language models.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Parameter-efficient fine-tuning of large-scale pre-trained language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:45:46.802075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T21:45:45.299976Z digest=sha256:82ecd52e558e0347c5852bb5e62755efd30fb5f86a9aceb8fa7e85370c743b7d

Observation bba694be-40d2-4597-b281-d6988b038b87 · outbound

This paper cites A Survey on In-context Learning.

Generating Symbolic World Models via Test-time Scaling of Large Language Models A Survey on In-context Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.304298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.304298Z digest=sha256:30961f60bee5875d62196c4f5f23b0a5efdea972a236d45ee57af091ea69969a

Observation eefa8b78-3ad8-45ac-92c2-2fa09093e5a1 · outbound

This paper cites The Llama 3 Herd of Models.

Generating Symbolic World Models via Test-time Scaling of Large Language Models The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.309015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.309015Z digest=sha256:6a452dbe531be3ef72619a7bdc072a4e27ba7105c0be888050f9899bba25cb8f

Observation 05e9d0c6-f8f0-4c7b-b83d-8b830a7848a5 · outbound

This paper cites Fikes and Nils J.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Fikes and Nils J

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:45:46.787696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T21:45:45.313744Z digest=sha256:b7a4280a355b2e59b45d8646f06fb6ce8d000206578a6b9d418fa6e267f29db3

Observation a8196531-c14f-4dd2-9fa5-80c78856f6b5 · outbound

This paper cites Leveraging pre- trained large language models to construct and utilize world models for model-based task planning.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Leveraging pre- trained large language models to construct and utilize world models for model-based task planning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:45:46.773058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T21:45:45.318349Z digest=sha256:306b3d65a32dd973315767ffc808f3bece13642a035fed3586449e66c69b4a8b

Observation b7d8ce2b-c5dc-43e9-bd3b-a8549d7d7275 · outbound

This paper cites Reasoning with Language Model is Planning with World Model.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Reasoning with Language Model is Planning with World Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.322784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.322784Z digest=sha256:d5b2891e8ed5a64a50cb69f2a1a2ee71327a86d67a06ab1bfe8faf74533aeba0

Observation 7f0a7224-bc1f-4018-8a5d-da4c5e801049 · outbound

This paper cites The fast downward planning system.

Generating Symbolic World Models via Test-time Scaling of Large Language Models The fast downward planning system

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:45:46.758668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T21:45:45.327277Z digest=sha256:27298ddbf5f9f645c8dd0b3a21aed15f2572ee5f932b42da66924d48c74fca9d

Observation ae4c5c4b-61de-4c3b-8da5-288a88e88961 · outbound

This paper cites The Competition: Impact, Organization, Evaluation, Benchmarks.

Generating Symbolic World Models via Test-time Scaling of Large Language Models The Competition: Impact, Organization, Evaluation, Benchmarks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:45:46.744061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T21:45:45.331678Z digest=sha256:9cb1487ff498a181defdb6b5ed3cd6ff6bbfcd15ef1842a224f90920f135605b

Observation 8f457a45-904f-4767-b2c9-6c79ece777c4 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Lora: Low-rank adaptation of large language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:45:46.729161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T21:45:45.335951Z digest=sha256:a830eb43008444c1982e78215925a6b17562c590ff7eba3212bf853307aab291

Observation 6a196f09-0952-4697-b545-3cb315e96399 · outbound

This paper cites OpenAI o1 System Card.

Generating Symbolic World Models via Test-time Scaling of Large Language Models OpenAI o1 System Card

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.340311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.340311Z digest=sha256:91923d9cc41b78b341dbc7f0a2f71845be528ac7f13c3af293c6c78d2059fbbb

Observation 646869af-157b-450d-b043-6057ffe4fe10 · outbound

This paper cites Mistral 7B.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Mistral 7B

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.345137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.345137Z digest=sha256:2cfc04af4a37b7b48a04cc2a90b466dacd5b0d80b263174a87f605866e69fe99

Observation 1a73d1ce-9422-42c9-910b-472c4d627b15 · outbound

This paper cites Can large language models reason and plan?Annals of the New York Academy of Sciences, 1534(1):15–18, 2024.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Can large language models reason and plan?Annals of the New York Academy of Sciences, 1534(1):15–18, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:45:46.714844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T21:45:45.350020Z digest=sha256:2d3a7c74a1e39ae57f27d7fb5d4a65164629aa1c62075946a5f7d53bf2254860

Observation f93ed7bc-4342-4b77-8817-7936a8abda61 · outbound

This paper cites Parameter-efficient orthogonal finetuning via butterfly factorization.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Parameter-efficient orthogonal finetuning via butterfly factorization

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:45:46.699971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T21:45:45.354281Z digest=sha256:a83708164319d7342df0280fc1c5431107422be9d749f73c601851fb79ddb279

Observation 946a65a8-9d97-4a61-a89a-5c4701501fea · outbound

This paper cites Leveraging Environment Interaction for Automated PDDL Translation and Planning with Large Language Models.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Leveraging Environment Interaction for Automated PDDL Translation and Planning with Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.359163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.359163Z digest=sha256:08e977bec217c629e920a7416f0f38533e307c6996bfc40f2ebea0cc0c3e0e08

Observation 54cfd0fa-f274-448c-b515-e2a85c50e5a1 · outbound

This paper cites Howe, Craig A.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Howe, Craig A

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:45:46.685254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T21:45:45.363988Z digest=sha256:79c566b2ad65685c041347de2a1f99f3ebad8662051e917e50e62b0e96e5cd84

Observation 45ab39e7-1ebf-4626-884c-998a24d671d0 · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

Generating Symbolic World Models via Test-time Scaling of Large Language Models GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.368651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.368651Z digest=sha256:8a3bd1855d9dd54b0ec29e358bb0eff70ce2ec70842b80418ad0761839a4c7d4

Observation b772d60b-b75b-43d3-831e-f8a9906aac37 · outbound

This paper cites Fully autonomous ai agents should not be developed.arXiv preprint arXiv:2502.02649, 2025.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Fully autonomous ai agents should not be developed.arXiv preprint arXiv:2502.02649, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.374118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.374118Z digest=sha256:6683477312ebae81f48fa01d421a13bf0fd74fc53316e51c18350aca0c5113c6

Observation 68346e29-ac90-4df6-99d3-75a7b5ab64fc · outbound

This paper cites Large language models as planning domain generators.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Large language models as planning domain generators

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:45:46.671010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T21:45:45.378530Z digest=sha256:78cc1c749ac07cc6a1d0a8879d6aba9c22df6864a1fbb8e326c4e97883a523ee

Observation ed5725ed-d274-40e8-89c1-761ee62b7939 · outbound

This paper cites Automatic Prompt Optimization with "Gradient Descent" and Beam Search.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Automatic Prompt Optimization with "Gradient Descent" and Beam Search

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.383025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.383025Z digest=sha256:4a96d809e6e7850bf208a710c623bcac77f9c16075afe24152465002f3a713bc

Observation 654adc1e-7ed5-4d85-901d-6daa106d887e · outbound

This paper cites O1 Replication Journey: A Strategic Progress Report -- Part 1.

Generating Symbolic World Models via Test-time Scaling of Large Language Models O1 Replication Journey: A Strategic Progress Report -- Part 1

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.387659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.387659Z digest=sha256:479fcba30e950244e296c8586711d64c17b50c004f87d84dec941cea357519a4

Observation d184fcdf-b6bd-47af-9930-beae695c542e · outbound

This paper cites Controlling text-to-image diffusion by orthogonal finetuning.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Controlling text-to-image diffusion by orthogonal finetuning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:45:46.656622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T21:45:45.391868Z digest=sha256:2ed373733ccb4630fe29f196b808b2a468a7ac9ef0566e095903e9653f0bad5c

Observation 0713ec78-5d08-4f1c-bdff-b5e5fc75dd4e · outbound

This paper cites Pearson, 2016.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Pearson, 2016

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:45:46.637822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T21:45:45.396629Z digest=sha256:62a823daa9682a3d2d2a6bc3652bca327f041266b9ae820445d5ff163d3bb6f4

Observation 9fadc9f0-a05a-4a15-8305-74c25f31662e · outbound

This paper cites Learning Multiple Initial Solutions to Optimization Problems.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Learning Multiple Initial Solutions to Optimization Problems

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.401048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.401048Z digest=sha256:5e040e6f98557533aed060aa6c4a7635628bcc597498a8784ac7aea9c3f80ca4

Observation 57bc6fa8-88cf-4830-a177-cab53de0f5b8 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Sys- tems, 36:8634–8652, 2023.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Sys- tems, 36:8634–8652, 2023

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:45:46.622404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T21:45:45.405565Z digest=sha256:ec904d215b1ba13caf57ab3c067fa99c79749607573088e1b6412e5471fee5da

Observation 7a02912c-cd62-4dad-b69b-5270c4090adf · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

Generating Symbolic World Models via Test-time Scaling of Large Language Models ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.409955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.409955Z digest=sha256:8d98e104c343018a139c655200ed2405feaf4bc87cc692cbc2092a2f07184ee8

Observation aa6afb8f-1814-45f2-a363-8dae05abc82a · outbound

This paper cites Generating consistent PDDL domains with Large Language Models.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Generating consistent PDDL domains with Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.414679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.414679Z digest=sha256:babbb1113db876d637925a018b95ccf44ff96d9a762801322e97e4be734d8c56

Observation 8b811352-2a25-4087-9114-a9c5b07014ac · outbound

This paper cites Inference scaling flaws: The limits of llm resampling with imperfect verifiers.arXiv preprint arXiv:2411.17501, 2024.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Inference scaling flaws: The limits of llm resampling with imperfect verifiers.arXiv preprint arXiv:2411.17501, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.419311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.419311Z digest=sha256:9d39aad38da889e24bfbb51d4ee08b482e2d401ac79bb117ae53c4b320c76d30

Observation 138c330b-e603-41c1-9383-3aebb683c63d · outbound

This paper cites WorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the Environment.

Generating Symbolic World Models via Test-time Scaling of Large Language Models WorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the Environment

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.423844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.423844Z digest=sha256:917cacd7cd8aa4855c1e0fbc135f9ab9aba9d11dab40a4d8aca412b71815b87c

Observation 02887bf2-501d-4228-854e-948372ef04ff · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Gemma: Open Models Based on Gemini Research and Technology

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.428515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.428515Z digest=sha256:f2912ca28346358617a2ec0ddd87a4ea95cf444b78b766ec55714e6f2d4651e3

Observation 320494a9-972c-4a9b-aaf3-4c3a54ad206c · outbound

This paper cites Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:45:46.607314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T21:45:45.433859Z digest=sha256:dc47672510dfd5f0503ffc79f131f9a8c4ad67d3372ae23abd9866c71e10c5bc

Observation d1ec938e-c08e-4c54-8447-311a147ecf4c · outbound

This paper cites LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench.

Generating Symbolic World Models via Test-time Scaling of Large Language Models LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.438262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.438262Z digest=sha256:d47e96f8fd366151d15aa03ae95225a3e96be3d7374f2551c31bc9a6f869cd44

Observation 03ae988b-5ada-4c24-970e-d209af5d3491 · outbound

This paper cites On The Planning Abilities of OpenAI's o1 Models: Feasibility, Optimality, and Generalizability.

Generating Symbolic World Models via Test-time Scaling of Large Language Models On The Planning Abilities of OpenAI's o1 Models: Feasibility, Optimality, and Generalizability

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.447954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.447954Z digest=sha256:cd8469c7b594d9216be1120eea1392cd7587754706ce1605af620e9256c950d5

Observation 2b2e7a1c-ce55-4a01-8229-9cd14def1877 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.452286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.452286Z digest=sha256:475e8c9669a984bcae7a1bee414bb0a236eeb0ed71eae14b8c69a5d3974692ec

Observation 26aedf24-c089-4a68-a1c2-24a2e02fbdb8 · outbound

This paper cites Emergent Abilities of Large Language Models.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Emergent Abilities of Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.457229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.457229Z digest=sha256:f5c2740c20bb688b8a49d66b5db5ec31091ea67f81662441b5d3ef36ed3a628e

Observation 7e4dc49f-163a-48a8-ad00-e8ea8df8ae46 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Chain-of-thought prompting elicits reasoning in large language models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:45:46.592486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T21:45:45.461751Z digest=sha256:4dd4391c53c0a1f65440a1754f9916730b3c4da89950157318dbccf471b32c23

Observation b7151f06-4a35-470d-95bf-a263b660de2d · outbound

This paper cites System 2 Attention (is something you might need too).

Generating Symbolic World Models via Test-time Scaling of Large Language Models System 2 Attention (is something you might need too)

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.466137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.466137Z digest=sha256:564c32fc96795ad2673112c27449b2eed453e16e754d96d4770cf7146db16e75

Observation 1045e54b-a72d-44d0-a6ec-35c66fa8dbae · outbound

This paper cites Verbalized Machine Learning: Revisiting Machine Learning with Language Models.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Verbalized Machine Learning: Revisiting Machine Learning with Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.470962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.470962Z digest=sha256:32e798d5aebd8e01c3f1f29285b1d6ddd3d54215635567aab774b59ca43d3d56

Observation 670baf85-dec8-40a9-8b0a-1045fe5539a2 · outbound

This paper cites Qwen2 Technical Report.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Qwen2 Technical Report

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.475667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.475667Z digest=sha256:0e5eca8ba5e3c78835b8c5e38e61c93e3a532c4ae792e63a01458895c5611cb7

Observation 0db9d5d2-2838-4684-adc7-5b4f1e931cde · outbound

This paper cites Le, Denny Zhou, and Xinyun Chen.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Le, Denny Zhou, and Xinyun Chen

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:45:46.576067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T21:45:45.480493Z digest=sha256:6959ee6f735c270a5570618740230d9725351962ae81ae02934b2481dcc8ec93

Observation 5d3b7035-f87b-412a-b727-b261da42ae26 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Yi: Open Foundation Models by 01.AI

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.484908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.484908Z digest=sha256:5d193f5d1028975c808a7fb1aef76c53b6a1b7b73d2bc7b1652923c8e145cf6a

Observation 41055a6c-f42a-4147-ae5f-31763f762d3b · outbound

This paper cites Distilling System 2 into System 1.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Distilling System 2 into System 1

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.489671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.489671Z digest=sha256:a74654d7b3344e85d6b51d1099995d5e0436430ae9333c8589be71717362b2a0

Observation cb0d4312-531b-41b3-a5bd-24554ff9ce11 · outbound

This paper cites TextGrad: Automatic "Differentiation" via Text.

Generating Symbolic World Models via Test-time Scaling of Large Language Models TextGrad: Automatic "Differentiation" via Text

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.494452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.494452Z digest=sha256:7e254881a10b9ac22aab1f47bb362331655f4ea54f836c310b96855d1fa9d34e

Observation e390a31d-ece3-4680-8602-4e9b7f4b24a3 · outbound

This paper cites Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.499263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.499263Z digest=sha256:318f5e02ab1daf83c5ec9fb2da6e384d081890a940902b3958b5f8341a943509

Observation 0135e484-0cab-4eed-8cc7-889552487b0a · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.504109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.504109Z digest=sha256:b4618c413e4b7b3045e04aec6b85e6b00dda3838e1ef0d81f21535fd2f997dc6

Observation e9755324-e6a7-4746-9e07-35f758271294 · outbound

This paper cites MiniF2F: a cross-system benchmark for formal Olympiad-level mathematics.

Generating Symbolic World Models via Test-time Scaling of Large Language Models MiniF2F: a cross-system benchmark for formal Olympiad-level mathematics

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.509398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.509398Z digest=sha256:c4d5c1625b866d80d230abee1ba38653b811f228ec80095b0a4ba0d53ac4d820

Observation f78d30ca-a71c-44cf-bfd8-29d54df6cd81 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:45:46.561071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T21:45:45.515011Z digest=sha256:da0a76afb64513284a6b9552a6e65e95321764b7eb524f9a83a4d10a3224d768

Observation ac2a8a49-3c8e-4897-92cb-edec6482a32d · outbound

This paper cites Large Language Models Are Human-Level Prompt Engineers.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Large Language Models Are Human-Level Prompt Engineers

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.519535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.519535Z digest=sha256:05caf4ac9c136ff6533ad7a30d1d2e62ea71b32dd75bc57822a7b531fd364651

Observation 0017a679-4b3c-4983-9544-091018a6b5b0 · outbound

This paper cites Plane- tarium: A rigorous benchmark for translating text to structured planning languages.arXiv preprint arXiv:2407.03321, 2024.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Plane- tarium: A rigorous benchmark for translating text to structured planning languages.arXiv preprint arXiv:2407.03321, 2024

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T21:45:45.524359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:45:45.524359Z digest=sha256:de144c2dd896e013d734c96261a7f50e1f50776fe651c3eed1e95ce28423a34d

Observation 87adb92d-79da-48f4-b24a-c65f2c4e2186 · outbound

This paper cites object1 is washed and heated.

Generating Symbolic World Models via Test-time Scaling of Large Language Models object1 is washed and heated

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:45:46.545188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T21:45:45.529282Z digest=sha256:ed9f14d90b429c43a0473c3d10fdd47ca7b554febf9aadf5a53031d00d599829

Observation f993925a-28f3-40d8-bf4f-44d289741275 · outbound

This paper cites an unresolved cited work.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-08T21:45:46.528926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T21:45:45.534417Z digest=sha256:a43fa98364be99de10275a6b2fa03d90020c0eb8c6284e8ec8fecfc40b5a351c

Observation 6022571f-e1b0-4246-8129-db0ae523b999 · outbound

This paper cites an unresolved cited work.

Generating Symbolic World Models via Test-time Scaling of Large Language Models Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-08T21:45:46.513526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T21:45:45.538720Z digest=sha256:28960ef759afefa307c22a75e481c73450d12336809541ed47ff70f728fe5e50

Pith citing papers

Observation d1d36373-ba35-402f-b684-c848f9aea5ad · inbound

CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization cites this paper.

CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization Generating Symbolic World Models via Test-time Scaling of Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T19:14:21.027074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:14:21.027074Z digest=sha256:183781880dca425855a29d3833ff0fe824dca03f7db80b05daa6b0d4e1d14aa5

Observation fdba0c82-df4a-4742-a488-ce638e579571 · inbound

Any House Any Task: Scalable Long-Horizon Planning for Abstract Human Tasks cites this paper.

Any House Any Task: Scalable Long-Horizon Planning for Abstract Human Tasks Generating Symbolic World Models via Test-time Scaling of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T06:05:30.314068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:05:30.314068Z digest=sha256:e4f897fc5e99d0c4f9e94d73ca6558fb0269b060d1e36c340b5ff92a295ca1da

Observation b55257f8-e899-462e-a68e-8fb55168961d · inbound

Sample-Efficient Post-Training for LEGO Spatial-Physics Reasoning cites this paper.

Sample-Efficient Post-Training for LEGO Spatial-Physics Reasoning Generating Symbolic World Models via Test-time Scaling of Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:42:49.550411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T23:38:38.127345Z digest=sha256:ae8270edaa7b058fc74ba44cfe5fdfe25694b3f240329ddfe225fa90586e63e7