Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:50:28.171568Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 100 of 122 outbound references and 4 inbound Pith citation observations for arXiv:2504.14520.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:50:28.171568Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:50:30.476760Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-10T08:58:13.216074Z
100 of 122 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1e6c8538-f2b3-46ea-9dbb-bbd4b8ec88a8 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Is creativity without intelligence possible? a necessary condition analysis,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d882530a-a2bd-4cb8-9c77-a800c9bc2dc6 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Thinking LLMs: General Instruction Following with Thought Generation
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 096de597-32ce-4b2d-9bc7-38d428171b72 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey LLMs for Explainable AI: A Comprehensive Survey
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dd879aa-9cc7-4df7-a203-09ba2a5c835b · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Chain-of-thought prompting elicits reasoning in large language models,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42affbab-670e-46d7-b223-df2d468a7cde · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71dc5145-b59d-4802-a989-688a1b3bb8a9 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Retrieval Augmented Generation with Multi-Modal LLM Framework for Wireless Environments
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 154bcae4-8dbc-42c0-a481-b5108176eb91 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9515db6-f872-4c55-b877-87afbf98d45d · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Assessing llms for high stakes applications,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b65b79f8-6cb2-43d9-945e-cac4f7eaac68 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44773207-9de3-461f-a6dc-b6efd21e2ec9 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey A review of methods for alleviating hallucination issues in large language models,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7c67a13-6121-40a0-80a4-fe585826d8a8 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey A Long Way to Go: Investigating Length Correlations in RLHF
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b9787b9-0b1f-4fe6-ab13-7722ba6d720b · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Human-level control through deep reinforcement learning,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c73aaa1-87b3-4b79-ab48-28cf63316f5a · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Continuous control with deep reinforcement learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e385f08-7479-43b8-a3d1-aed2a29b7ed7 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 427aca7f-f001-4028-8712-34c0f4cc7b30 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Mixcl: Mixed contrastive learning for relation extraction,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 051a03b3-4b26-40c7-a85c-9253f5c82c3a · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Neural Path Hunter: Reducing Hallucination in Dialogue Systems via Path Grounding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abe910a9-7111-4836-a2d8-1d669142b0d8 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Mitigating large language model hallucinations via autonomous knowledge graph- based retrofitting,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cfc1447-a83b-465b-bcfa-f0b303f4ff94 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey A Stitch in Time Saves Nine: Detecting and Mitigating Hallucinations of LLMs by Validating Low-Confidence Generation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18d80309-24b0-459c-9331-679058f8a2bd · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Reasoning on Graphs: Faithful and Interpretable Large Language Model Reasoning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1e05a96-e1a7-42d2-a8b3-725ffe62fe8a · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b177ac6c-635b-4070-9139-0957ba51eeec · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03780251-b86b-445b-b0fd-433f624b25a3 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Chain-of-Verification Reduces Hallucination in Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb624d59-d1c8-426d-8410-550fd622c07e · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa6d9464-ecb5-4ab5-8836-a1589b16fbc5 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 394cbf54-979d-498f-b796-80434bc5d06a · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey KCTS: Knowledge-Constrained Tree Search Decoding with Token-Level Hallucination Detection
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8a06022-c61e-49da-b9bb-3b649acac5ac · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Inference- time intervention: Eliciting truthful answers from a language model,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66daf7b3-f725-4b40-81f1-135f56b87328 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11b83328-ce46-42b9-bb2b-30733779bead · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Self-refine: Iter- ative refinement with self-feedback,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e1110ee-275e-4c93-8f5c-966c12c56193 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Can Large Language Models Really Improve by Self-critiquing Their Own Plans?
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49e02cfd-ec9c-4204-917e-a3e98acae603 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Large language models lack essential metacognition for reliable medical reasoning,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42b023a8-cbf4-4719-901a-76c742b31881 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Large Language Models Cannot Self-Correct Reasoning Yet
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1be6a8c3-4756-4598-ad56-93b346f8fc36 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey LLM-based Multi-Agent Reinforcement Learning: Current and Future Directions
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3a1c42c-2c14-46da-a5a2-310886184be1 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 698330db-ad30-4eb7-8bbd-fab455087f69 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Leveraging large language models for optimised coordination in textual multi-agent reinforcement learning,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 364cf4fc-5cf4-4a74-8ac4-69dfc1e583aa · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8777bdbe-28ec-468b-8ee2-64ba41735139 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Building Cooperative Embodied Agents Modularly with Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ee36bf9-6b7b-4bdc-bdfe-5c28d25f82b1 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Smart-llm: Smart multi-agent robot task planning using large language models,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e6c3d32-4502-4e7e-bd49-c845961de962 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Roco: Dialectic multi-robot col- laboration with large language models,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffa46914-3744-4eb5-9038-cdd6a3550848 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Co-NavGPT: Multi-Robot Cooperative Visual Semantic Navigation Using Vision Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45e3ee58-f44c-4e3f-887b-0270a1e01346 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Embodied LLM Agents Learn to Cooperate in Organized Teams
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43151392-5bfa-46e7-8c12-d548975aae4d · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Multi-Agent Consensus Seeking via Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebe41552-9243-42fa-bf8e-cd80ba91705b · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Learning Natural Language Constraints for Safe Reinforcement Learning of Language Agents
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 352db76e-7349-40df-9c3b-bbb078ccf2b1 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8b33906-5f13-4b0c-acb1-ef29da2f2776 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Theory of Mind for Multi-Agent Collaboration via Large Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd710fb3-b6b7-4f08-a328-aa2b344bf477 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Buffer of thoughts: Thought-augmented reasoning with large language models,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5bf7050-ed37-4906-9f77-cf7cb5e6d5a3 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Meta Learning for Natural Language Processing: A Survey
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c7d0086-2289-4c4b-ab98-b24412aa91f7 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Continual Learning for Large Language Models: A Survey
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 889eb9f5-b774-44aa-be5a-eee1a1c129a3 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Continual Learning of Large Language Models: A Comprehensive Survey
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 529f355d-e1ec-410b-8770-7601458b3a14 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Towards Incremental Learning in Large Language Models: A Critical Review
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97d70b5f-788b-4f6f-af2d-ea48749a2aee · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Learning from models beyond fine-tuning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dc7f790-083b-4d13-a459-46dc0a0c4f84 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Meta Reasoning for Large Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68c5b9f5-1cba-499a-b9c0-edb8a2a9bacf · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Reasoning with large language models, a survey,
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ef3dec7-8bee-4817-900b-2b2c8f74c115 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Gizaml: A collaborative meta-learning based framework using llm for automated time-series forecasting
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 034a4aa5-03f0-460b-9991-e56b5462e6f8 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Towards lifelong learning of large language models: A survey,
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35511f4b-50f6-4ed2-b0f4-fea4370ae594 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Thinking Machines: A Survey of LLM based Reasoning Strategies
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cb7659b-e411-4e9a-a034-897ef2f3b2a9 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46850007-b9cd-4d20-b553-92a5208124ac · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebcbbd79-9a65-47e0-9a1f-9bff521d5042 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Evaluating the Meta- and Object-Level Reasoning of Large Language Models for Question Answering
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 431dabcf-b987-43c7-a8e3-e51b2555374d · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Malalgoqa: Pedagogical evaluation of counterfactual reasoning in large language models and implications for ai in education,
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 377461d5-1bdc-4e12-9799-a738d93b855c · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey METAL: Towards Multilingual Meta-Evaluation
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd47ff9c-9ebf-4db1-a278-a2dc24002790 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey MetaICL: Learning to Learn In Context
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e4b7159-712f-4537-ba66-a7ee4299d250 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Towards metacognitive clinical reasoning: Benchmarking md-pie against state-of-the-art llms in medical decision-making,
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc0c558f-ff7e-4608-b522-ec34abe6f280 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Do Large Language Models Know What They Don't Know?
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9ae61b9-778e-49eb-9f06-cb89bc74c7c0 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Language grounded multi- agent reinforcement learning with human-interpretable communica- tion,
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67fcd80a-e994-4467-af7e-ddc518990c26 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Hypothetical Minds: Scaffolding Theory of Mind for Multi-Agent Tasks with Large Language Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebbc836d-4e0a-40dc-aacb-1fc3ce636dcd · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Distilling the Knowledge in a Neural Network
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92a9ccc0-e665-4136-b76a-460f8febf5ee · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Knowledge distillation: A survey,
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e13eeaa1-c6fa-4272-84fa-c432c205caa2 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Large Language Models have Intrinsic Self-Correction Ability
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e8ead3d-6cf2-4ef4-9157-ea712b0ae964 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Can Rationalization Improve Robustness?
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41c59bf3-1b97-4413-8d7a-e4df063ee9e4 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Measuring Compositionality in Representation Learning
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6affd4e1-55c9-4651-bf20-c75c71cdc5e3 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Multi-agent Reinforcement Learning in Sequential Social Dilemmas
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4045999f-60d3-410e-8cd0-9c3db45adbaf · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Towards Effective GenAI Multi-Agent Collaboration: Design and Evaluation for Enterprise Applications
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27c89c8d-3764-4f1c-aae9-f40bda59ecfa · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey AI safety via debate
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5941701a-ef45-4456-82a8-9b9cbb91d422 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Adversarial training for high-stakes reliability,
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6612d0f1-855f-48ff-8898-5be4bf4fe428 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey AD-AutoGPT: An Autonomous GPT for Alzheimer's Disease Infodemiology
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9024881-cdfb-4788-91f7-2dbf9b92107c · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Deep reinforcement learning from human preferences,
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b65ccd2-3f01-410c-a3de-2de81a16e07d · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Training language models to follow instructions with human feedback,
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 961f0250-0a70-4bf2-b01b-fdfd2a0a40fa · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Training Language Models to Self-Correct via Reinforcement Learning
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce8f0fa8-a1ad-4dbf-8c38-a30ac0bcf42d · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey A Tutorial on Meta-Reinforcement Learning
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85be3b41-3738-43e8-9fcb-d5c5900d92ad · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Eureka: Human-Level Reward Design via Coding Large Language Models
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e57c22b-0edb-45fc-b1af-92cf449c920b · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93397ad0-a6c0-4119-b5ef-f58d4b6bbc0c · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Learning to summarize with human feedback,
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a0131c6-e27b-4d0c-ba80-8fb8cc73294d · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Confidence Matters: Revisiting Intrinsic Self-Correction Capabilities of Large Language Models
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1f33125-130b-4524-aa1c-2a2cadc76e4a · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Curiosity-Driven Reinforcement Learning from Human Feedback
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e39e6a5-dffc-4217-b1c4-ae0d3fd4e97a · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Online intrinsic rewards for decision making agents from large language model feedback,
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45c41508-2731-41d1-9af1-c041ccf42002 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4290a4a5-6c3b-4d78-b8e2-93a8058b923c · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Grandmaster level in starcraft ii using multi-agent reinforcement learning,
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 961f9518-3979-41e8-95cb-bd8151638423 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Human-level play in the game of diplomacy by combining language models with strategic reasoning,
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77bbbd09-8f67-474a-8f63-6f5ed77d072c · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Shall We Team Up: Exploring Spontaneous Cooperation of Competing LLM Agents
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6b594b8-4cf9-4770-9318-2f3ae99696b8 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Self- playing adversarial language game enhances llm reasoning,
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b50ec277-b147-423a-8e93-544c3680f91e · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Logicattack: Adversarial attacks for evaluating logical consistency of natural language inference,
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d0e1425-408a-4eee-8266-419855a42db0 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Model-agnostic meta-learning for fast adaptation of deep networks,
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54c4ffe3-a2f7-474c-ba71-1c2918a801bc · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Meta-learning for Few-shot Natural Language Processing: A Survey
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4a7406c-bca5-4b9d-b880-910dfc948c4c · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Meta In-Context Learning Makes Large Language Models Better Zero and Few-Shot Relation Extractors
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df4e08ac-3fce-4d24-877c-f38553618c9b · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Improving Consistency in Large Language Models through Chain of Guidance
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3633bd04-b82f-4a56-8a7a-09f3aeff3b4f · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Training Verifiers to Solve Math Word Problems
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1d73e9d-158b-46df-adae-97258db6d2ae · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Red teaming language models for contradictory dialogues,
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 104758c2-e148-4106-b2f1-69ceb8b023b1 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey OpenAI o1 System Card
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0de2eb8-aa5b-43e9-b20b-4346b35a212a · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies,
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 219f7fc9-5318-4c73-bf61-8388a04b39f8 · outbound
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Can LLMs Beat Humans in Debating? A Dynamic Multi-agent Framework for Competitive Debate
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d795cdaf-31c7-4e31-b172-fad1c89479b0 · inbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 594d600c-b10e-46ab-b09a-d5a8219ea514 · inbound
Weak-Link Optimization for Multi-Agent Reasoning and Collaboration Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0b30f91a-1da6-43b0-b8d6-2a8131d40932 · inbound
Human Cognition in Machines: A Unified Perspective of World Models Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c3a4fb96-ec94-4311-8ce7-2106133bb173 · inbound
HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.