Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T04:30:47.578235Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 0 inbound Pith citation observations for arXiv:2608.05139.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T04:30:47.578235Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
87 of 87 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 22f954c5-f7fa-4896-b4de-52cadc159226 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World Scenarios
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85344c15-b040-40d8-8ff8-eda6b9ff7f74 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning STaD: Scaffolded Task Design for Identifying Compositional Skill Gaps in LLMs
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2fc5af15-6fa6-4c42-acba-1ed19670ae6c · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning HeroBench: A Benchmark for Long-Horizon Planning and Structured Reasoning in Virtual Worlds
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d332b2ae-1cc2-4932-99a7-15062ac05017 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Concepts or Skills? Rethinking Instruction Selection for Multi-modal Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 314b5d5d-6905-4b0c-94ec-bb31b4f714f4 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred Skills
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7726b0ef-4a0a-42e4-8438-80c868ed471b · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning SkillCraft: Can LLM agents learn to use tools skillfully?arXiv preprint arXiv:2603.00718, 2026
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79047070-6705-4482-9ea9-ef5261f62b8a · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Self-evolving curriculum for LLM reasoning.arXiv preprint arXiv:2505.14970, 2025
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af86f2c9-eed4-4a18-90d8-f13ccc6943a8 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning WebSRC: A Dataset for Web-Based Structural Reading Comprehension
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e00d57f2-0c17-4e4c-b441-2e87735c546b · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89a2579b-c321-4849-8b09-1032a7c6b001 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Training Verifiers to Solve Math Word Problems
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed4106fb-7792-4233-8261-434ad4b6bfd6 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Metacognitive capabilities of LLMs: An exploration in mathematical problem solving.Advances in Neural Information Processing Systems, 2024
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e249b7d-bd6a-40e0-871e-0056f67c682c · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Thinkless: LLM Learns When to Think
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74652f66-0821-4cb1-a765-b61f5f08c2e9 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e38f4f09-afe9-4b91-9cad-06f220e0180c · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43c2c010-9b67-4309-a4ea-a3cba7e046f5 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning AdaptMI: Adaptive Skill-based In-context Math Instruction for Small Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a21e7b44-59e0-4b5a-b43d-5d0ed2e6e5a6 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning STAT: Skill-targeted adaptive training.arXiv preprint arXiv:2510.10023, 2025
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdfaa97d-e9d8-461e-bef5-8363bdb39ffc · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d40a405c-c813-416e-a2f7-7e2774c6a670 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Measuring Mathematical Problem Solving With the MATH Dataset
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 772fc02a-582d-4e54-b85c-a10830a34981 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15ec8f37-481f-4885-b375-d5cad83cc672 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Open-R1: A fully open reproduction of DeepSeek-R1
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea8571ec-8a18-4a8d-84f0-c290b8239f7e · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cef7bf9b-2075-44b1-95b1-70fdfc3554a4 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning DRP: Distilled Reasoning Pruning with Skill-aware Step Decomposition for Efficient Large Reasoning Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaf7ed6d-2b24-4875-a2cc-0e669720ea6d · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79592456-402e-4e60-bf62-d166c363c77b · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning BIG-Bench Extra Hard
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 634ed69b-c440-426e-8986-7e327cee4b8f · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Benchmark profiling: Mechanistic diagnosis of LLM benchmarks.arXiv preprint arXiv:2510.01232, 2025
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7f1954f-b3c2-46f5-b7cf-9da893fc87d2 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning MSCoRe: A benchmark for multi-stage collaborative reasoning in LLM agents.arXiv preprint arXiv:2509.17628, 2025
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8c72a2b4-2d84-42ef-aa85-1ee0fb024670 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning START: Self-taught Reasoner with Tools
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc749d60-dded-40be-b893-bf3265216b73 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6b5672e-cab7-459a-bf83-17daea23ab7b · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Benchmark test-time scaling of general LLM agents.arXiv preprint arXiv:2602.18998, 2026
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f1b4aa7-6cba-4dd4-972b-d20686689bf4 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1a728ed-df21-49e1-950f-bbadfa80112b · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Let's Verify Step by Step
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cad6590c-cf1c-48fa-a5a2-3e51e0b2dc1d · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a80a4082-64be-48d1-b0ba-890af7af79d9 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning AgentBench: Evaluating LLMs as Agents
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a04cf71-b807-41e6-aed6-8baef6e5fe51 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning GAIA: a benchmark for General AI Assistants
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0bb827e-e8e3-4fb8-9109-87a6f223097e · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Benchmarking and Understanding Compositional Relational Reasoning of LLMs
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c135f8fc-0df8-42c5-849d-509b8887d953 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Reasoning curriculum: Bootstrapping broad LLM reasoning from math.arXiv preprint arXiv:2510.26143, 2025
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 263f343b-08a9-4a51-abf6-b814cf1ab837 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Compositional Semantic Parsing on Semi-Structured Tables
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d070f3e4-b2fb-4427-96fe-6e06ad3690c8 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Learning to reason across parallel samples for LLM reasoning.arXiv preprint arXiv:2506.09014, 2025
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5fb6985-d76d-40f3-a2a0-a51a09c8b4e2 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning EmoAgent: Assessing and safeguarding human-AI interaction for mental health safety
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47ee61b7-1e6d-428b-8a50-87643f3dfebc · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning LogicSkills: A structured benchmark for formal reasoning in large language models.arXiv preprint arXiv:2602.06533, 2026
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5bb94826-154a-4760-abbf-48449447200f · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Reasoning models are test exploiters: Rethinking multiple-choice.arXiv preprint arXiv:2507.15337, 2025
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10a0be47-6ecc-48d0-9f7c-af1da1abf9c6 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Revisiting Compositional Generalization Capability of Large Language Models Considering Instruction Following Ability
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3769a8a2-921a-42ea-a2ef-7aab904d3024 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning AI-Assisted Generation of Difficult Math Questions
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47094969-dc1f-4766-9d03-82be32a2124c · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd41d836-26bf-4ccb-8a8d-7e2b61b4918a · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning DARE- bench: Evaluating modeling and instruction fidelity of LLMs in data science.arXiv preprint arXiv:2602.24288, 2026
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6c3ae39e-671d-4674-a715-775352b9cce4 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 705f7630-3321-4561-bf0d-3c7c08bd223a · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7834b505-9f05-4d7a-8e38-ead4733a71c1 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Reinforcement Learning for Self-Improving Agent with Skill Library
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a28e7c8e-8e10-49e3-937b-da961f2cf94c · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9eac2e05-e348-4e38-86bb-fc182ea509c7 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05187c29-552d-42d3-83c5-fd0d9b42a8ec · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa1f2ede-0d00-40fd-92b6-f54e5b855943 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Reinforcingmulti-turn reasoning in LLM agents via turn-level reward design.arXiv preprint arXiv:2505.11821, 2025
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb388906-b5c2-4b7d-8ce8-4ffefc7cc252 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Towardscompositionalgeneralization of LLMs via skill taxonomy guided data synthesis.arXiv preprint arXiv:2601.03676, 2026
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b0e4634c-9404-4cd6-a681-ab2c31072565 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Hi-ToM: A benchmark for evaluating higher-order theory of mind reasoning in large language models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81e1505c-7efa-49cc-9bca-9615802b1c9b · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning CritICL: Inference-time weak-to-strong generalization from small language model failure modes
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 564db2cd-3f30-4d90-aa42-6f26843d43a8 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning SkillRL: Evolving agents via recursive skill-augmented reinforcement learning, 2026
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e35a80f-b213-4d98-93a5-f462b54b3f1b · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning LaRS: Latent Reasoning Skills for Chain-of-Thought Reasoning
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4be92d90-751e-4fa9-88f4-907765b67579 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning DeepCritic: Deliberate Critique with Large Language Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fdab5d0-8554-4199-93e0-da0123f8527f · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4f045811-840f-41c5-a0ce-eebebd0315f5 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning LongProc: Benchmarking long-context language models on long procedural generation
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8eb82e20-7176-43d1-a1af-6dfee17fb294 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Skill-Mix: a Flexible and Expandable Family of Evaluations for AI models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba4c8a55-fae2-4392-a94e-b9b9559cf248 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning From𝑓(𝑥) and 𝑔(𝑥) to 𝑓(𝑔(𝑥)) : LLMs learn new skills in RL by composing old ones.arXiv preprint arXiv:2509.25123, 2025
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f441f9d-74ea-40fd-baaf-85ae5922310c · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Skill-aware data selection and fine-tuning for data-efficient reasoning distillation
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ecd2ff6-3be4-4195-b060-beed5ce26159 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Skill-awaredataselectionandfine-tuning for data-efficient reasoning distillation, 2026
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9b43e4de-0cd7-482c-9948-122df95a7a07 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Lee, Chenlei Leng, and Fanghui Liu
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0506bd30-63bc-448a-84f7-bf8d05965137 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 985fb27d-d840-456c-bdf2-e1b7307a67ec · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Can Models Learn Skill Composition from Examples?
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 51c4fc5c-8ba0-4182-8d2b-6574abb25afd · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0bce8bd-4215-492e-b0f8-185d3f0d0725 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning NATURAL PLAN: Benchmarking LLMs on Natural Language Planning
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3be68408-0eae-4f51-8cee-15b8fb01407e · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning SkillRouter: Skill Routing for LLM Agents at Scale
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f26bf2fa-4768-4abe-a813-b75e88bafcd9 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63be7e85-8388-469b-a518-0c586be89495 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning It clearly outlines the key factors and their interrelationships
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d34e5cd0-02c8-4f10-990b-62371eb35b0e · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Specific examples from the data are used to support the narrative
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 72849c37-fbd7-4519-87c9-f82842a368fb · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning It captures the reader’s interest and effectively conveys the potential crisis
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 46ee3a08-804c-4b2e-a87f-28fe61b52bd8 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning It uses the data and insights from the previous steps to construct a plausible and coherent narrative
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7d98e6aa-0390-4d6f-b55d-63eab9cd9a22 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning It should reflect the brand’s commitment to sustainability
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 862d3e81-c652-41de-a436-10c783881f85 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning It should not exceed 10 words
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 66e12a0b-a7d2-4dd0-bfe5-a948738e9651 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Unresolved cited work
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8fa281e4-6a35-4557-93b4-bf1458f388eb · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning no-interference
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 505dddaa-bfa2-4486-a3a7-529a6a92c816 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning DON’T CHANGE THE ANSWER, CORE LOGIC, OR THE SKILL REQUIRED
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation aa927f39-65d7-40ad-a2d1-8e1547e1ae1b · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Unresolved cited work
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 97a2f555-a3a9-47c6-954f-b3375f1ef442 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Conference trip on constraints
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6c420ad5-44a8-4954-8aef-9be629177b9d · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning 29 Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning 2.Scenario consistent.All rewritten steps plausibly belong to the samescenario
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d60813c5-b66b-44ed-8dd6-4bcb8aca68b5 · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning The core logic, numerical values, and (for multiple-choice) the option letters and contents must be unchanged
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b72f18a9-bb5b-4b8c-88c0-516cb43f235d · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning FAIL if any step references entities, settings, or framings that contradict the scenario or that read as an unrelated problem pasted in
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a66d5167-c914-4613-9b2f-da05bec9d10b · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning using the value from the previous step
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 37ee22a3-c1f0-4ec2-9688-1ceb5fef7fab · outbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning MAE” is the mean absolute error between the LLM and mean human score in[0, 1]. “Binary agree
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
No inbound Pith citation observations are available.