Pith. sign in

Paper Citation Record · LEDGER

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey

As of 17 August 2026, this Paper Citation Record lists 100 of 122 outbound references and 4 inbound Pith citation observations for arXiv:2504.14520.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.14520 v1

Coverage vector

measured 100 of 122 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:50:28.171568Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:50:30.476760Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T08:58:13.216074Z

Reference resolution

100 of 122 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved98
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1e6c8538-f2b3-46ea-9dbb-bbd4b8ec88a8 · outbound

This paper cites Is creativity without intelligence possible? a necessary condition analysis,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Is creativity without intelligence possible? a necessary condition analysis,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.711553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.711553Z digest=sha256:6ea2f07e0097f850043301db54d98b9c24fa05ffff1693b7494cf491242431ac

Observation d882530a-a2bd-4cb8-9c77-a800c9bc2dc6 · outbound

This paper cites Thinking LLMs: General Instruction Following with Thought Generation.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Thinking LLMs: General Instruction Following with Thought Generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.716555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.716555Z digest=sha256:6ac31f9bd33d1b766ec5b782cab3e4fd0cc78181d0c6be4e48df8faf5625efa3

Observation 096de597-32ce-4b2d-9bc7-38d428171b72 · outbound

This paper cites LLMs for Explainable AI: A Comprehensive Survey.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey LLMs for Explainable AI: A Comprehensive Survey

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.721261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.721261Z digest=sha256:84dae9cf3958dcf3417756c873bb0fbd325f7bb16b531d8472b98b8c381685f5

Observation 1dd879aa-9cc7-4df7-a203-09ba2a5c835b · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Chain-of-thought prompting elicits reasoning in large language models,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.726139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.726139Z digest=sha256:02cd913683db55837f421b633b4da1337663368b36d0c5e943198e788526f917

Observation 42affbab-670e-46d7-b223-df2d468a7cde · outbound

This paper cites To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.730766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.730766Z digest=sha256:df7ced5e7b7bdcf0adfa63271fd76534031c4dc690dfd6be7e3d2d4504ed31c4

Observation 71dc5145-b59d-4802-a989-688a1b3bb8a9 · outbound

This paper cites Retrieval Augmented Generation with Multi-Modal LLM Framework for Wireless Environments.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Retrieval Augmented Generation with Multi-Modal LLM Framework for Wireless Environments

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.735660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.735660Z digest=sha256:afa41e38fe8e709aff83545d9c64d587bb1613911679dfcf2ea4c8075fe06f9c

Observation 154bcae4-8dbc-42c0-a481-b5108176eb91 · outbound

This paper cites A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.741423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.741423Z digest=sha256:561f0e3b1c8105d1669b18f11c04ab5f5259350de968ac638d4a03c6a1151f7d

Observation c9515db6-f872-4c55-b877-87afbf98d45d · outbound

This paper cites Assessing llms for high stakes applications,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Assessing llms for high stakes applications,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.745978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.745978Z digest=sha256:9d4f66165ebd3572ef43e8ad72cc9e6c6920878eb10effc05557c6fc5cbcc071

Observation b65b79f8-6cb2-43d9-945e-cac4f7eaac68 · outbound

This paper cites Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.750368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.750368Z digest=sha256:70a62ab234b737a76d0dbb2a2b8a3d9714407ae4a0de357de0aee8a53f7a8801

Observation 44773207-9de3-461f-a6dc-b6efd21e2ec9 · outbound

This paper cites A review of methods for alleviating hallucination issues in large language models,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey A review of methods for alleviating hallucination issues in large language models,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.755117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.755117Z digest=sha256:81c6424e3400bbe6bf8d1c3bcbfa0a60b96b9d2e8c5c355666b655ac455f03ea

Observation d7c67a13-6121-40a0-80a4-fe585826d8a8 · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey A Long Way to Go: Investigating Length Correlations in RLHF

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.759524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.759524Z digest=sha256:4cbb61e3882ed80657790cc719017d5b5b0d4e3088592315e07ff8ccef1c3feb

Observation 1b9787b9-0b1f-4fe6-ab13-7722ba6d720b · outbound

This paper cites Human-level control through deep reinforcement learning,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Human-level control through deep reinforcement learning,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.764165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.764165Z digest=sha256:72f7a0e22c74aa9677f994cb631b30647addb70c923a406231d5912ef0ddbc55

Observation 6c73aaa1-87b3-4b79-ab48-28cf63316f5a · outbound

This paper cites Continuous control with deep reinforcement learning.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Continuous control with deep reinforcement learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.768476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.768476Z digest=sha256:d59e3912d8b29c667d25da4160f9826b07088d24dddbe91cc48285b88712f390

Observation 6e385f08-7479-43b8-a3d1-aed2a29b7ed7 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.773074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.773074Z digest=sha256:59a1eb9e4a85e118037f609303d18a1ce06669b93e827667f7c61e22009b7fd4

Observation 427aca7f-f001-4028-8712-34c0f4cc7b30 · outbound

This paper cites Mixcl: Mixed contrastive learning for relation extraction,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Mixcl: Mixed contrastive learning for relation extraction,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.777510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.777510Z digest=sha256:997ee2a6abc6d1b1d764e27950503b1356b06fc625492fc413cfc8d487873237

Observation 051a03b3-4b26-40c7-a85c-9253f5c82c3a · outbound

This paper cites Neural Path Hunter: Reducing Hallucination in Dialogue Systems via Path Grounding.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Neural Path Hunter: Reducing Hallucination in Dialogue Systems via Path Grounding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.781720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.781720Z digest=sha256:4cd0843191aa281e7865f90b7565e0c7debb70066ef3e056eef2e0828b5c2275

Observation abe910a9-7111-4836-a2d8-1d669142b0d8 · outbound

This paper cites Mitigating large language model hallucinations via autonomous knowledge graph- based retrofitting,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Mitigating large language model hallucinations via autonomous knowledge graph- based retrofitting,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.786475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.786475Z digest=sha256:44092d526f144262b64bae549241e72b1035a4bcee38be97619807714728e6fb

Observation 6cfc1447-a83b-465b-bcfa-f0b303f4ff94 · outbound

This paper cites A Stitch in Time Saves Nine: Detecting and Mitigating Hallucinations of LLMs by Validating Low-Confidence Generation.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey A Stitch in Time Saves Nine: Detecting and Mitigating Hallucinations of LLMs by Validating Low-Confidence Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.791031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.791031Z digest=sha256:be7bdf0ab3054c50c8d40c280721e48744ba4434969a53ce7727cbe404b5588a

Observation 18d80309-24b0-459c-9331-679058f8a2bd · outbound

This paper cites Reasoning on Graphs: Faithful and Interpretable Large Language Model Reasoning.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Reasoning on Graphs: Faithful and Interpretable Large Language Model Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.795731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.795731Z digest=sha256:6bf908feee19035b74c4412d83e94a9c9faeb5c5e25743dfae5524f5e63e9ab2

Observation f1e05a96-e1a7-42d2-a8b3-725ffe62fe8a · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.800254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.800254Z digest=sha256:82de5b5563497a8737f033f9e195039eff07c89bdc5be34e940604ada6e0e0c0

Observation b177ac6c-635b-4070-9139-0957ba51eeec · outbound

This paper cites SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.804684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.804684Z digest=sha256:f69f7c439648a374972ba76ff6bafa43df6f08f3ee0f8857a559fb8855bd46fb

Observation 03780251-b86b-445b-b0fd-433f624b25a3 · outbound

This paper cites Chain-of-Verification Reduces Hallucination in Large Language Models.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Chain-of-Verification Reduces Hallucination in Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.809408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.809408Z digest=sha256:6a9d58f6779845c663b851a1fd21e7f97437574c702657ed901b74bd61278df1

Observation eb624d59-d1c8-426d-8410-550fd622c07e · outbound

This paper cites Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.814258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.814258Z digest=sha256:a82fc3a3066de2cb89020262f0098df8745365ca5ecd318ac68b3d10611bf537

Observation fa6d9464-ecb5-4ab5-8836-a1589b16fbc5 · outbound

This paper cites Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.818881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.818881Z digest=sha256:ca8a730fcfd4fecd6727124201e77c74eb2d05d1a315643f7141b82bdd669196

Observation 394cbf54-979d-498f-b796-80434bc5d06a · outbound

This paper cites KCTS: Knowledge-Constrained Tree Search Decoding with Token-Level Hallucination Detection.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey KCTS: Knowledge-Constrained Tree Search Decoding with Token-Level Hallucination Detection

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.823504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.823504Z digest=sha256:f1f1cf8905abbf49875d5caa12d09ba6aa6eca89c336561dc1bfd350738f67a0

Observation f8a06022-c61e-49da-b9bb-3b649acac5ac · outbound

This paper cites Inference- time intervention: Eliciting truthful answers from a language model,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Inference- time intervention: Eliciting truthful answers from a language model,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.828277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.828277Z digest=sha256:3f6f2311121e0b927560fd5718a1e38322d12c67652e32c626be52f77f11c215

Observation 66daf7b3-f725-4b40-81f1-135f56b87328 · outbound

This paper cites DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.832665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.832665Z digest=sha256:35825e3bd8fd387f4d990ea1231c907b24048b5bdbccb45bf426ed0d345ef01e

Observation 11b83328-ce46-42b9-bb2b-30733779bead · outbound

This paper cites Self-refine: Iter- ative refinement with self-feedback,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Self-refine: Iter- ative refinement with self-feedback,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.837478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.837478Z digest=sha256:b2b74e20321da42826aea4f173083b4066e62a7e818dee9cf4d8fed10a37a3ef

Observation 5e1110ee-275e-4c93-8f5c-966c12c56193 · outbound

This paper cites Can Large Language Models Really Improve by Self-critiquing Their Own Plans?.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.842026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.842026Z digest=sha256:a572d4adc99575d75e535901f738ae49d5047bccf574d63565a7339215b633b9

Observation 49e02cfd-ec9c-4204-917e-a3e98acae603 · outbound

This paper cites Large language models lack essential metacognition for reliable medical reasoning,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Large language models lack essential metacognition for reliable medical reasoning,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.846740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.846740Z digest=sha256:f043fac174e0a5f0c4b9d0c0327c139db1adbd8356aea8c76cbe0a74d2cc51f2

Observation 42b023a8-cbf4-4719-901a-76c742b31881 · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Large Language Models Cannot Self-Correct Reasoning Yet

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.851212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.851212Z digest=sha256:2877a1c4ad57ff6b275060dbf85491b15e6a35fc6cb26a7e75b71960a5db5c9f

Observation 1be6a8c3-4756-4598-ad56-93b346f8fc36 · outbound

This paper cites LLM-based Multi-Agent Reinforcement Learning: Current and Future Directions.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey LLM-based Multi-Agent Reinforcement Learning: Current and Future Directions

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.855907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.855907Z digest=sha256:7bb69c444181874ff89098262e83d90a29faa0fe3318c371ae4dbc8be6570d98

Observation e3a1c42c-2c14-46da-a5a2-310886184be1 · outbound

This paper cites A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.860924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.860924Z digest=sha256:b128fb6208624d9a70e65afae531625385deb166e708b8d8bce877d197b16b8d

Observation 698330db-ad30-4eb7-8bbd-fab455087f69 · outbound

This paper cites Leveraging large language models for optimised coordination in textual multi-agent reinforcement learning,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Leveraging large language models for optimised coordination in textual multi-agent reinforcement learning,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.865663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.865663Z digest=sha256:6ae45016b09a20ac1aff47c6b88026de1f0fd17d0ed05e2e491a20ca853e1c3a

Observation 364cf4fc-5cf4-4a74-8ac4-69dfc1e583aa · outbound

This paper cites MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.870208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.870208Z digest=sha256:859d54434cdb1990e7c563d6b69175ef33bf4bd606ec4fe0ba1d147da0961ff9

Observation 8777bdbe-28ec-468b-8ee2-64ba41735139 · outbound

This paper cites Building Cooperative Embodied Agents Modularly with Large Language Models.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Building Cooperative Embodied Agents Modularly with Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.874878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.874878Z digest=sha256:2b4d973a2292b88ca189808607c11e9bcf07d2559efdaf6fd16bd4f404911499

Observation 1ee36bf9-6b7b-4bdc-bdfe-5c28d25f82b1 · outbound

This paper cites Smart-llm: Smart multi-agent robot task planning using large language models,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Smart-llm: Smart multi-agent robot task planning using large language models,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.879546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.879546Z digest=sha256:0ee02e3171f0c3d3539d80f83ef18fe736579310379bcd6552d7be0c5a6f223f

Observation 1e6c3d32-4502-4e7e-bd49-c845961de962 · outbound

This paper cites Roco: Dialectic multi-robot col- laboration with large language models,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Roco: Dialectic multi-robot col- laboration with large language models,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.883728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.883728Z digest=sha256:3ed399be96d86215f1e00f4690d28e04afad68b206c4088430d1d07354f54e1d

Observation ffa46914-3744-4eb5-9038-cdd6a3550848 · outbound

This paper cites Co-NavGPT: Multi-Robot Cooperative Visual Semantic Navigation Using Vision Language Models.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Co-NavGPT: Multi-Robot Cooperative Visual Semantic Navigation Using Vision Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.888658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.888658Z digest=sha256:6de65caf4d6aaf6dad6bc8cdab1b33555fcf77358dc8f14a35c36cc9f188f070

Observation 45e3ee58-f44c-4e3f-887b-0270a1e01346 · outbound

This paper cites Embodied LLM Agents Learn to Cooperate in Organized Teams.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Embodied LLM Agents Learn to Cooperate in Organized Teams

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.893387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.893387Z digest=sha256:156e880f643a1768ac58ca4326e4f004e6b0882c8a0c0a70d096afc0264b1d5f

Observation 43151392-5bfa-46e7-8c12-d548975aae4d · outbound

This paper cites Multi-Agent Consensus Seeking via Large Language Models.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Multi-Agent Consensus Seeking via Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.898289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.898289Z digest=sha256:f568e901959d57891b9af9025276c7c818739c4323caaa60545ad60175777bec

Observation ebe41552-9243-42fa-bf8e-cd80ba91705b · outbound

This paper cites Learning Natural Language Constraints for Safe Reinforcement Learning of Language Agents.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Learning Natural Language Constraints for Safe Reinforcement Learning of Language Agents

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:50:29.185967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:50:27.903048Z digest=sha256:1725b18e7a3a697e11c7df8473c4721b79d25330e3088df4dd8c8f42faa87a01

Observation 352db76e-7349-40df-9c3b-bbb078ccf2b1 · outbound

This paper cites ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.907723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.907723Z digest=sha256:6a79f97db11ce0e5cf4f8822f094b435782220ed16a2d41dc53fda25754858b0

Observation e8b33906-5f13-4b0c-acb1-ef29da2f2776 · outbound

This paper cites Theory of Mind for Multi-Agent Collaboration via Large Language Models.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Theory of Mind for Multi-Agent Collaboration via Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.912446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.912446Z digest=sha256:6d24749a1fb2f58bab762aa1ee622b00e3e26ce4de10ded4cf31ea526b24b867

Observation dd710fb3-b6b7-4f08-a328-aa2b344bf477 · outbound

This paper cites Buffer of thoughts: Thought-augmented reasoning with large language models,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Buffer of thoughts: Thought-augmented reasoning with large language models,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.917164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.917164Z digest=sha256:db566fd31a99ec8d34a9082a9b8099fadf59380e1a1b68fe23d37dbb90842b54

Observation f5bf7050-ed37-4906-9f77-cf7cb5e6d5a3 · outbound

This paper cites Meta Learning for Natural Language Processing: A Survey.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Meta Learning for Natural Language Processing: A Survey

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.921574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.921574Z digest=sha256:847be9ea284f305859735e6dc837003e9e361197183913f9e2d45438f1bbc6c4

Observation 7c7d0086-2289-4c4b-ab98-b24412aa91f7 · outbound

This paper cites Continual Learning for Large Language Models: A Survey.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Continual Learning for Large Language Models: A Survey

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.926559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.926559Z digest=sha256:c347aff8b104ed71297d1a4457f8ced8bd6b1570e9ef97c76dc597002e16d90c

Observation 889eb9f5-b774-44aa-be5a-eee1a1c129a3 · outbound

This paper cites Continual Learning of Large Language Models: A Comprehensive Survey.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Continual Learning of Large Language Models: A Comprehensive Survey

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.931335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.931335Z digest=sha256:3fd5a71c1fb3919f40e47942dc25bbac3b6e479ba034ea297f3a53c6f96d8578

Observation 529f355d-e1ec-410b-8770-7601458b3a14 · outbound

This paper cites Towards Incremental Learning in Large Language Models: A Critical Review.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Towards Incremental Learning in Large Language Models: A Critical Review

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.936058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.936058Z digest=sha256:7c5556f434c15ba21ab5da64d0df9cbdb17ed6d67d879f4bd7fba3536c9c6d3a

Observation 97d70b5f-788b-4f6f-af2d-ea48749a2aee · outbound

This paper cites Learning from models beyond fine-tuning.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Learning from models beyond fine-tuning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.940822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.940822Z digest=sha256:e54fcae6c8fa64fd07f36a37b5dac375c190e2862b4abe550ab694df0b7a60dc

Observation 5dc7f790-083b-4d13-a459-46dc0a0c4f84 · outbound

This paper cites Meta Reasoning for Large Language Models.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Meta Reasoning for Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.945469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.945469Z digest=sha256:cc14153ab8c6570e12acac4aaed4733b5157b4ff2df58429b39a60652df2e4fc

Observation 68c5b9f5-1cba-499a-b9c0-edb8a2a9bacf · outbound

This paper cites Reasoning with large language models, a survey,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Reasoning with large language models, a survey,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.950019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.950019Z digest=sha256:94bd2f2bf0b8ea0deb41c10194c87bcd745f64f047cdd6e5c4ae52955e6bb702

Observation 8ef3dec7-8bee-4817-900b-2b2c8f74c115 · outbound

This paper cites Gizaml: A collaborative meta-learning based framework using llm for automated time-series forecasting.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Gizaml: A collaborative meta-learning based framework using llm for automated time-series forecasting

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.954474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.954474Z digest=sha256:3b698b3478b1e340ac40a3f88adbfaa94beada54a873db4779861b1182461e94

Observation 034a4aa5-03f0-460b-9991-e56b5462e6f8 · outbound

This paper cites Towards lifelong learning of large language models: A survey,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Towards lifelong learning of large language models: A survey,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.959408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.959408Z digest=sha256:ef65f4b281fb9916bd0f5c91b7dabeca41acfcab90e75ac19b450ef4d6aa906a

Observation 35511f4b-50f6-4ed2-b0f4-fea4370ae594 · outbound

This paper cites Thinking Machines: A Survey of LLM based Reasoning Strategies.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Thinking Machines: A Survey of LLM based Reasoning Strategies

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.964075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.964075Z digest=sha256:e87973a183d3a754871a0e34e62bd8c05c9340d812b2d08e049472be32a13a8f

Observation 7cb7659b-e411-4e9a-a034-897ef2f3b2a9 · outbound

This paper cites MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.968815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.968815Z digest=sha256:f1f46b107353be4a748c50e4402f66485f1f5fca8b6673430dc9be3630cf2517

Observation 46850007-b9cd-4d20-b553-92a5208124ac · outbound

This paper cites Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.973380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.973380Z digest=sha256:8cda1a0901b94bca8aa9406dc494c99ec31d9cf6ea7168aa8ef66e0e37592d93

Observation ebcbbd79-9a65-47e0-9a1f-9bff521d5042 · outbound

This paper cites Evaluating the Meta- and Object-Level Reasoning of Large Language Models for Question Answering.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Evaluating the Meta- and Object-Level Reasoning of Large Language Models for Question Answering

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:50:28.936737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:50:27.977789Z digest=sha256:e5a9a5415156bbc3596a811ea066e23c4aa5bcd219d9efefc2670d85bae31ff8

Observation 431dabcf-b987-43c7-a8e3-e51b2555374d · outbound

This paper cites Malalgoqa: Pedagogical evaluation of counterfactual reasoning in large language models and implications for ai in education,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Malalgoqa: Pedagogical evaluation of counterfactual reasoning in large language models and implications for ai in education,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.982603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.982603Z digest=sha256:f5a753ecf412bc8c3fde815ad756583a3c31f17755704131cb267344f2f1c39b

Observation 377461d5-1bdc-4e12-9799-a738d93b855c · outbound

This paper cites METAL: Towards Multilingual Meta-Evaluation.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey METAL: Towards Multilingual Meta-Evaluation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.987474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.987474Z digest=sha256:f1bbd294cf52181d6b0073b330a376f631cb08934aff99d18c2990abc7643853

Observation bd47ff9c-9ebf-4db1-a278-a2dc24002790 · outbound

This paper cites MetaICL: Learning to Learn In Context.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey MetaICL: Learning to Learn In Context

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.992183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.992183Z digest=sha256:302e47d1d362bc749fb824077a2601d4f2805c13ed1a65154d4641540a65563e

Observation 4e4b7159-712f-4537-ba66-a7ee4299d250 · outbound

This paper cites Towards metacognitive clinical reasoning: Benchmarking md-pie against state-of-the-art llms in medical decision-making,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Towards metacognitive clinical reasoning: Benchmarking md-pie against state-of-the-art llms in medical decision-making,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.997065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.997065Z digest=sha256:7c905b7cc45600af1bfab37ed6790d06a4935780ae3222738bfe5271fb96ccf5

Observation dc0c558f-ff7e-4608-b522-ec34abe6f280 · outbound

This paper cites Do Large Language Models Know What They Don't Know?.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Do Large Language Models Know What They Don't Know?

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.001877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.001877Z digest=sha256:1471a026ff1cbc5b4cc4223ecb749962b75ddab355b93ccecd5e51f783c52ded

Observation d9ae61b9-778e-49eb-9f06-cb89bc74c7c0 · outbound

This paper cites Language grounded multi- agent reinforcement learning with human-interpretable communica- tion,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Language grounded multi- agent reinforcement learning with human-interpretable communica- tion,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.006674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.006674Z digest=sha256:7dcba764accf54b2307cdfbb003270696bdf3387de121fa0f75edec94b27f7cd

Observation 67fcd80a-e994-4467-af7e-ddc518990c26 · outbound

This paper cites Hypothetical Minds: Scaffolding Theory of Mind for Multi-Agent Tasks with Large Language Models.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Hypothetical Minds: Scaffolding Theory of Mind for Multi-Agent Tasks with Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.011118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.011118Z digest=sha256:54e8f302d4de7d8dfa26c8e250b06f0605677c5ffc4910a91f048a3348a1bbd0

Observation ebbc836d-4e0a-40dc-aacb-1fc3ce636dcd · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Distilling the Knowledge in a Neural Network

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.015604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.015604Z digest=sha256:6d2b000aa0fc2df8f71cbc2c53209d1b65f8f07e2122be4c1188b9de369fd3df

Observation 92a9ccc0-e665-4136-b76a-460f8febf5ee · outbound

This paper cites Knowledge distillation: A survey,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Knowledge distillation: A survey,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.020070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.020070Z digest=sha256:e39a41d20f6cd3a8d877ea447e76e21db5f9fea912ed38df51cd194805a28e6b

Observation e13eeaa1-c6fa-4272-84fa-c432c205caa2 · outbound

This paper cites Large Language Models have Intrinsic Self-Correction Ability.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Large Language Models have Intrinsic Self-Correction Ability

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.024588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.024588Z digest=sha256:2c7fca3dc326d616d716097c4c20a4268f9bd11e7da53dfeb4a88704f216be60

Observation 8e8ead3d-6cf2-4ef4-9157-ea712b0ae964 · outbound

This paper cites Can Rationalization Improve Robustness?.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Can Rationalization Improve Robustness?

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.029078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.029078Z digest=sha256:829ba7e4f53c2e044a7036dbc347b329c696bd638fe75106b6d9fd45dac293d4

Observation 41c59bf3-1b97-4413-8d7a-e4df063ee9e4 · outbound

This paper cites Measuring Compositionality in Representation Learning.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Measuring Compositionality in Representation Learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.033705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.033705Z digest=sha256:fbd4f915f2cdb0225ec6cd161582e1c255e258c235fb393beca85853ff81a1a9

Observation 6affd4e1-55c9-4651-bf20-c75c71cdc5e3 · outbound

This paper cites Multi-agent Reinforcement Learning in Sequential Social Dilemmas.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Multi-agent Reinforcement Learning in Sequential Social Dilemmas

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.038753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.038753Z digest=sha256:ce1efe8c3056e89597a36cbcfed0a2f205f30791bb054856fbac058cf16c5e3c

Observation 4045999f-60d3-410e-8cd0-9c3db45adbaf · outbound

This paper cites Towards Effective GenAI Multi-Agent Collaboration: Design and Evaluation for Enterprise Applications.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Towards Effective GenAI Multi-Agent Collaboration: Design and Evaluation for Enterprise Applications

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.043495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.043495Z digest=sha256:4b3b5fba696430e42fab48d3fcb98a49debe12d09c6a257956210b3425c0c468

Observation 27c89c8d-3764-4f1c-aae9-f40bda59ecfa · outbound

This paper cites AI safety via debate.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey AI safety via debate

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.048060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.048060Z digest=sha256:849b6c1ebef320e99dffdece821e068df7b9f839ecd58a6fda2537869f0ffdf3

Observation 5941701a-ef45-4456-82a8-9b9cbb91d422 · outbound

This paper cites Adversarial training for high-stakes reliability,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Adversarial training for high-stakes reliability,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.052767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.052767Z digest=sha256:28d64f1b9b38b05357f4eb079e9960ff5e8b9a22c05a0b0610b4c56f413df409

Observation 6612d0f1-855f-48ff-8898-5be4bf4fe428 · outbound

This paper cites AD-AutoGPT: An Autonomous GPT for Alzheimer's Disease Infodemiology.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey AD-AutoGPT: An Autonomous GPT for Alzheimer's Disease Infodemiology

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.057392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.057392Z digest=sha256:c5920ccd7642b818101efedd8b9543ca9df73fad7892b99a7d208851175e70ba

Observation d9024881-cdfb-4788-91f7-2dbf9b92107c · outbound

This paper cites Deep reinforcement learning from human preferences,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Deep reinforcement learning from human preferences,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.062051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.062051Z digest=sha256:1b6a8c3c63731a8e34a29efcbe572bb129b89084be9bfc0808a0da49d33408f5

Observation 2b65ccd2-3f01-410c-a3de-2de81a16e07d · outbound

This paper cites Training language models to follow instructions with human feedback,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Training language models to follow instructions with human feedback,

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.066324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.066324Z digest=sha256:ab97bf4f5439667fb77e8e2aed11a4d72a2f917eddf5eb28f7bd881f0a2d2b99

Observation 961f0250-0a70-4bf2-b01b-fdfd2a0a40fa · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Training Language Models to Self-Correct via Reinforcement Learning

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.070616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.070616Z digest=sha256:c6c664a76763b9607e713068f88460a30d9a68e76ee07f184583f66a467dcaef

Observation ce8f0fa8-a1ad-4dbf-8c38-a30ac0bcf42d · outbound

This paper cites A Tutorial on Meta-Reinforcement Learning.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey A Tutorial on Meta-Reinforcement Learning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.075344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.075344Z digest=sha256:b55a03212471ad9e8d06943229c3e38e5473999dbac06c4e895742595b60e92b

Observation 85be3b41-3738-43e8-9fcb-d5c5900d92ad · outbound

This paper cites Eureka: Human-Level Reward Design via Coding Large Language Models.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Eureka: Human-Level Reward Design via Coding Large Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.079809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.079809Z digest=sha256:3dcc2a747c121bd6227bc2addd41dfa2ba770fafb370e08cfa728053086276b5

Observation 7e57c22b-0edb-45fc-b1af-92cf449c920b · outbound

This paper cites Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.084177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.084177Z digest=sha256:1426c2a21e1d9fbe4cdbaf1fb8657b7aaee3bc3e260563f347fd662b8a4c50e6

Observation 93397ad0-a6c0-4119-b5ef-f58d4b6bbc0c · outbound

This paper cites Learning to summarize with human feedback,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Learning to summarize with human feedback,

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.088784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.088784Z digest=sha256:47e7773d190427c54cf9932a878f3469af53cde4bda4009a28146bee71c28888

Observation 0a0131c6-e27b-4d0c-ba80-8fb8cc73294d · outbound

This paper cites Confidence Matters: Revisiting Intrinsic Self-Correction Capabilities of Large Language Models.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Confidence Matters: Revisiting Intrinsic Self-Correction Capabilities of Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.092823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.092823Z digest=sha256:6c8dba27e12cb564cf9389914f93f9a4bd01c2ac366786733618a1d8ba8e1343

Observation c1f33125-130b-4524-aa1c-2a2cadc76e4a · outbound

This paper cites Curiosity-Driven Reinforcement Learning from Human Feedback.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Curiosity-Driven Reinforcement Learning from Human Feedback

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.097413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.097413Z digest=sha256:e4e93855cda9f4eb4521d79e70a945385949ba9519749437ae363a127e118f3f

Observation 4e39e6a5-dffc-4217-b1c4-ae0d3fd4e97a · outbound

This paper cites Online intrinsic rewards for decision making agents from large language model feedback,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Online intrinsic rewards for decision making agents from large language model feedback,

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.101951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.101951Z digest=sha256:a58a285888b92c52b16afde0897cbf48af1b31c83945251268cf04f4e657f3c9

Observation 45c41508-2731-41d1-9af1-c041ccf42002 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.106527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.106527Z digest=sha256:9f37e38c0dc905209a0ba1b8ef7a5494272e6e348ec68249e9814ae95f12b9e0

Observation 4290a4a5-6c3b-4d78-b8e2-93a8058b923c · outbound

This paper cites Grandmaster level in starcraft ii using multi-agent reinforcement learning,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Grandmaster level in starcraft ii using multi-agent reinforcement learning,

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.111357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.111357Z digest=sha256:b3c42c5fcab87fa38db7509d2eeba95a59710c86c1b9f3d58462f49aeb9bac0a

Observation 961f9518-3979-41e8-95cb-bd8151638423 · outbound

This paper cites Human-level play in the game of diplomacy by combining language models with strategic reasoning,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Human-level play in the game of diplomacy by combining language models with strategic reasoning,

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.115812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.115812Z digest=sha256:2ad89f64cca99a683865ad99143894a5db3b1788f206491963cb549b99f2cf49

Observation 77bbbd09-8f67-474a-8f63-6f5ed77d072c · outbound

This paper cites Shall We Team Up: Exploring Spontaneous Cooperation of Competing LLM Agents.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Shall We Team Up: Exploring Spontaneous Cooperation of Competing LLM Agents

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.120233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.120233Z digest=sha256:b977b647613db766530d1de97893bdda1acc47bcebc5edf6bc1211650fa0b5cf

Observation d6b594b8-4cf9-4770-9318-2f3ae99696b8 · outbound

This paper cites Self- playing adversarial language game enhances llm reasoning,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Self- playing adversarial language game enhances llm reasoning,

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.124703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.124703Z digest=sha256:daaed2fe750fa2029188c99ecbea4b78b386c0167d673504dd3dc95e4b59eb89

Observation b50ec277-b147-423a-8e93-544c3680f91e · outbound

This paper cites Logicattack: Adversarial attacks for evaluating logical consistency of natural language inference,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Logicattack: Adversarial attacks for evaluating logical consistency of natural language inference,

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.129150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.129150Z digest=sha256:4e4328f1cff005aedbeb1f2430fb3e86e8934c452a1635958ab60ab409e72d75

Observation 4d0e1425-408a-4eee-8266-419855a42db0 · outbound

This paper cites Model-agnostic meta-learning for fast adaptation of deep networks,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Model-agnostic meta-learning for fast adaptation of deep networks,

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.134129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.134129Z digest=sha256:7484adc7646893a35eeb4f86f70a3a1e5084db7f9af653231136cd146a0b7aa0

Observation 54c4ffe3-a2f7-474c-ba71-1c2918a801bc · outbound

This paper cites Meta-learning for Few-shot Natural Language Processing: A Survey.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Meta-learning for Few-shot Natural Language Processing: A Survey

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.138491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.138491Z digest=sha256:1ccb48d21f49052d8d427deb736aaf9bd3e0f861aea060f61b7fef24897c41f8

Observation b4a7406c-bca5-4b9d-b880-910dfc948c4c · outbound

This paper cites Meta In-Context Learning Makes Large Language Models Better Zero and Few-Shot Relation Extractors.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Meta In-Context Learning Makes Large Language Models Better Zero and Few-Shot Relation Extractors

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.143364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.143364Z digest=sha256:66212f67cf0ada2777631cd72774c895fe979fb9be65e22662ad4755a3fe3e13

Observation df4e08ac-3fce-4d24-877c-f38553618c9b · outbound

This paper cites Improving Consistency in Large Language Models through Chain of Guidance.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Improving Consistency in Large Language Models through Chain of Guidance

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.148408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.148408Z digest=sha256:b54a3482678841183775b8178d10d02fd7a749fb1abdcb7b707329f8bef07b00

Observation 3633bd04-b82f-4a56-8a7a-09f3aeff3b4f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Training Verifiers to Solve Math Word Problems

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.153101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.153101Z digest=sha256:02acd6b8d48457ea3a40433eae0e83ed29daa9147d701a96a71f885d8d192aa9

Observation f1d73e9d-158b-46df-adae-97258db6d2ae · outbound

This paper cites Red teaming language models for contradictory dialogues,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Red teaming language models for contradictory dialogues,

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.158013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.158013Z digest=sha256:f5c798fe8e5c0437a1fb20d34445c30022983fd463b26f83a3d578660a3884a8

Observation 104758c2-e148-4106-b2f1-69ceb8b023b1 · outbound

This paper cites OpenAI o1 System Card.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey OpenAI o1 System Card

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.162336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.162336Z digest=sha256:cdf5481f37c1e6c0b43da68d6137a2a5ae4542b779e1c5000a14028ac2a17e9c

Observation b0de2eb8-aa5b-43e9-b20b-4346b35a212a · outbound

This paper cites Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies,

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.167260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.167260Z digest=sha256:3626ce3097a631afc64fcb1534de61e31f108ef32fc334910e135c294e30e00b

Observation 219f7fc9-5318-4c73-bf61-8388a04b39f8 · outbound

This paper cites Can LLMs Beat Humans in Debating? A Dynamic Multi-agent Framework for Competitive Debate.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Can LLMs Beat Humans in Debating? A Dynamic Multi-agent Framework for Competitive Debate

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.171568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.171568Z digest=sha256:e65c91f76f4d5316bd1a3b075acb301f23a5ed0599a9ba7358d00d5df964b578

Pith citing papers

Observation d795cdaf-31c7-4e31-b172-fad1c89479b0 · inbound

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models cites this paper.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.476760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.476760Z digest=sha256:53e3a6accc0e102bc57d1e2f3be8357dd23e75ea4c0390b4b11445d3c72a3f86

Observation 594d600c-b10e-46ab-b09a-d5a8219ea514 · inbound

Weak-Link Optimization for Multi-Agent Reasoning and Collaboration cites this paper.

Weak-Link Optimization for Multi-Agent Reasoning and Collaboration Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:58:13.218271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T08:53:03.420926Z digest=sha256:b0da3ba80b25a66a2a12eccd6bb88c2a8061b2ec73cf53c62f14d772cc305c46

Observation 0b30f91a-1da6-43b0-b8d6-2a8131d40932 · inbound

Human Cognition in Machines: A Unified Perspective of World Models cites this paper.

Human Cognition in Machines: A Unified Perspective of World Models Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:12:25.995829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T08:12:15.663761Z digest=sha256:b7220a76b2c3325515b2fd47d7c76259f4337da611fdf0f00984f72c8554e038

Observation c3a4fb96-ec94-4311-8ce7-2106133bb173 · inbound

HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs cites this paper.

HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:54:48.570592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T00:51:37.506096Z digest=sha256:36a9ced11fffabfc846af22f4e9bb30d68b7b22cd59b4788dde954c1b39cb0ee