Pith. sign in

Paper Citation Record · LEDGER

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning

As of 18 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 0 inbound Pith citation observations for arXiv:2608.05139.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05139 v1

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:30:47.578235Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

87 of 87 outbound references displayed

  • verified exact13
  • verified fuzzy13
  • unresolved60
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 22f954c5-f7fa-4896-b4de-52cadc159226 · outbound

This paper cites AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World Scenarios.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World Scenarios

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:38.138268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:38.138268Z digest=sha256:b5c43aa9de4e9ffb6244e8b00f6b9fe33775ee9e12f521e64748aa0cdb84e214

Observation 85344c15-b040-40d8-8ff8-eda6b9ff7f74 · outbound

This paper cites STaD: Scaffolded Task Design for Identifying Compositional Skill Gaps in LLMs.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning STaD: Scaffolded Task Design for Identifying Compositional Skill Gaps in LLMs

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:30:49.765028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:38.264262Z digest=sha256:4bff0dacb72f06f9ce10fae56fbe76ec019f1e916859e3a85434080ff6b4b03e

Observation 2fc5af15-6fa6-4c42-acba-1ed19670ae6c · outbound

This paper cites HeroBench: A Benchmark for Long-Horizon Planning and Structured Reasoning in Virtual Worlds.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning HeroBench: A Benchmark for Long-Horizon Planning and Structured Reasoning in Virtual Worlds

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:38.368705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:38.368705Z digest=sha256:0e6e63006b1ffbc3f211978c77195aba99cc9d95097e4858ec69b4308bb22ae2

Observation d332b2ae-1cc2-4932-99a7-15062ac05017 · outbound

This paper cites Concepts or Skills? Rethinking Instruction Selection for Multi-modal Models.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Concepts or Skills? Rethinking Instruction Selection for Multi-modal Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:30:49.742374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:38.448384Z digest=sha256:4c75105aae3aeece2f6547da348584d32b6e17b34fd4f3df7b2c5212a1f04f21

Observation 314b5d5d-6905-4b0c-94ec-bb31b4f714f4 · outbound

This paper cites Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred Skills.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred Skills

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:38.535214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:38.535214Z digest=sha256:241ad614d8622c8f289b7f29da15de96326cad170f309bf6b3211b4e2f252dc0

Observation 7726b0ef-4a0a-42e4-8438-80c868ed471b · outbound

This paper cites SkillCraft: Can LLM agents learn to use tools skillfully?arXiv preprint arXiv:2603.00718, 2026.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning SkillCraft: Can LLM agents learn to use tools skillfully?arXiv preprint arXiv:2603.00718, 2026

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:38.636140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:38.636140Z digest=sha256:1fa1b0bfbd0f9595f405aae5e1917484c8390944340443ae5ca5c3d184240e17

Observation 79047070-6705-4482-9ea9-ef5261f62b8a · outbound

This paper cites Self-evolving curriculum for LLM reasoning.arXiv preprint arXiv:2505.14970, 2025.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Self-evolving curriculum for LLM reasoning.arXiv preprint arXiv:2505.14970, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:38.777948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:38.777948Z digest=sha256:4d8ec94408b1acdd9ed946909e8824bb645dcf142c566d2cbb4228f41139e3fe

Observation af86f2c9-eed4-4a18-90d8-f13ccc6943a8 · outbound

This paper cites WebSRC: A Dataset for Web-Based Structural Reading Comprehension.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning WebSRC: A Dataset for Web-Based Structural Reading Comprehension

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:38.889955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:38.889955Z digest=sha256:ae849db6bb2bcdf48b2ef651438d116d01a29dd55c379eacc7735b32c4cc6ca0

Observation e00d57f2-0c17-4e4c-b441-2e87735c546b · outbound

This paper cites Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:38.938928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:38.938928Z digest=sha256:7b61dcb59b8c54a61f17f64354c28f377c4f48eb372b1453749e8189168229bc

Observation 89a2579b-c321-4849-8b09-1032a7c6b001 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:39.031956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:39.031956Z digest=sha256:395876e74f52855dac172ef02b5c1e7fb0570837410104c6e20c85443ae25089

Observation ed4106fb-7792-4233-8261-434ad4b6bfd6 · outbound

This paper cites Metacognitive capabilities of LLMs: An exploration in mathematical problem solving.Advances in Neural Information Processing Systems, 2024.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Metacognitive capabilities of LLMs: An exploration in mathematical problem solving.Advances in Neural Information Processing Systems, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:39.141611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:39.141611Z digest=sha256:1749b50e77a6dabfa6591bb99531a9b02c9e06456e00bb58bb1502fecb9ed1d0

Observation 4e249b7d-bd6a-40e0-871e-0056f67c682c · outbound

This paper cites Thinkless: LLM Learns When to Think.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Thinkless: LLM Learns When to Think

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:39.238053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:39.238053Z digest=sha256:7527226e70997417176d2bd5656a23c395f37d8e1524485c97b8ad7540dab337

Observation 74652f66-0821-4cb1-a765-b61f5f08c2e9 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:39.338585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:39.338585Z digest=sha256:f3dff93bc5e7e75e66ee410079e081839daaf893d5fe7032641084b341a52d23

Observation e38f4f09-afe9-4b91-9cad-06f220e0180c · outbound

This paper cites R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:39.422268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:39.422268Z digest=sha256:793eff7a9f0cb1b18c98efd94d3880ee546cf7ffaacdcf3cad89d54604b71705

Observation 43c2c010-9b67-4309-a4ea-a3cba7e046f5 · outbound

This paper cites AdaptMI: Adaptive Skill-based In-context Math Instruction for Small Language Models.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning AdaptMI: Adaptive Skill-based In-context Math Instruction for Small Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:30:49.550875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:39.563825Z digest=sha256:7e2f1b1e49db9ec89f57eaf92b98f53c14069a27670c292a2854dc2275cd1e58

Observation a21e7b44-59e0-4b5a-b43d-5d0ed2e6e5a6 · outbound

This paper cites STAT: Skill-targeted adaptive training.arXiv preprint arXiv:2510.10023, 2025.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning STAT: Skill-targeted adaptive training.arXiv preprint arXiv:2510.10023, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:39.641296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:39.641296Z digest=sha256:db5f7e800ed4b8c109659e896985d4735b8133b9b6607badb82119c686501329

Observation fdfaa97d-e9d8-461e-bef5-8363bdb39ffc · outbound

This paper cites Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:39.705631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:39.705631Z digest=sha256:4a10043eba89a731316649e2fb0c3a47e01043904b70e00fa936efb8cf813e55

Observation d40a405c-c813-416e-a2f7-7e2774c6a670 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:39.862494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:39.862494Z digest=sha256:2a128702f1f373abf3f922ff4b99bdba884f10323f81e69d2da7d0cf7ea42684

Observation 772fc02a-582d-4e54-b85c-a10830a34981 · outbound

This paper cites Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:39.961952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:39.961952Z digest=sha256:b5fd496d015825619a8d7e205399357633b5a36d2f1de94a928e9177bacf8483

Observation 15ec8f37-481f-4885-b375-d5cad83cc672 · outbound

This paper cites Open-R1: A fully open reproduction of DeepSeek-R1.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Open-R1: A fully open reproduction of DeepSeek-R1

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:40.042883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:40.042883Z digest=sha256:6c324a4cb5099583b9fbac970c251bb5686f88bb2e3929873e244509e6619aeb

Observation ea8571ec-8a18-4a8d-84f0-c290b8239f7e · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:40.119348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:40.119348Z digest=sha256:98c9d9e091ef640071928439610a9ebd5018b4855565d580cbf5081a9e424afa

Observation cef7bf9b-2075-44b1-95b1-70fdfc3554a4 · outbound

This paper cites DRP: Distilled Reasoning Pruning with Skill-aware Step Decomposition for Efficient Large Reasoning Models.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning DRP: Distilled Reasoning Pruning with Skill-aware Step Decomposition for Efficient Large Reasoning Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:40.212488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:40.212488Z digest=sha256:2a7192ccc57e42c7b34c5ee951684d793faf693eab1272c82258d73ce7818a2a

Observation eaf7ed6d-2b24-4875-a2cc-0e669720ea6d · outbound

This paper cites Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:40.386775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:40.386775Z digest=sha256:e83e2633943082ed51058b656b6f7b0d305299b22703393ed50ea496dc6edcb4

Observation 79592456-402e-4e60-bf62-d166c363c77b · outbound

This paper cites BIG-Bench Extra Hard.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning BIG-Bench Extra Hard

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:40.572955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:40.572955Z digest=sha256:1562287c2d2cf48b25ca374022a287d8dffe86363b3519e24b6464bd65078a36

Observation 634ed69b-c440-426e-8986-7e327cee4b8f · outbound

This paper cites Benchmark profiling: Mechanistic diagnosis of LLM benchmarks.arXiv preprint arXiv:2510.01232, 2025.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Benchmark profiling: Mechanistic diagnosis of LLM benchmarks.arXiv preprint arXiv:2510.01232, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:40.730539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:40.730539Z digest=sha256:7a4810ebfd1eaf78cdff7cbfea2fd9de4bf10def43b81968ce40ffb8e78dab46

Observation c7f1954f-b3c2-46f5-b7cf-9da893fc87d2 · outbound

This paper cites MSCoRe: A benchmark for multi-stage collaborative reasoning in LLM agents.arXiv preprint arXiv:2509.17628, 2025.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning MSCoRe: A benchmark for multi-stage collaborative reasoning in LLM agents.arXiv preprint arXiv:2509.17628, 2025

Reference 26

Resolution
verified exact
raw_fallback, observed 2026-08-06T04:30:49.323037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:40.868199Z digest=sha256:08a4f4787601aed142c82e7b58655d5eaab3976c94e833ccedb492b4a9962302

Observation 8c72a2b4-2d84-42ef-aa85-1ee0fb024670 · outbound

This paper cites START: Self-taught Reasoner with Tools.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning START: Self-taught Reasoner with Tools

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:41.015987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:41.015987Z digest=sha256:7039068f27f7bc66324f7068fd2ff1f2f509c322eefca2af8c71336402c108a2

Observation bc749d60-dded-40be-b893-bf3265216b73 · outbound

This paper cites SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:41.170677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:41.170677Z digest=sha256:771cddb07006287ac28b11c42f34605b9a26bbac66e3c0d7418cb710ad493a46

Observation b6b5672e-cab7-459a-bf83-17daea23ab7b · outbound

This paper cites Benchmark test-time scaling of general LLM agents.arXiv preprint arXiv:2602.18998, 2026.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Benchmark test-time scaling of general LLM agents.arXiv preprint arXiv:2602.18998, 2026

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:41.324235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:41.324235Z digest=sha256:ee37462b215cec8b28d937ef9d63f41d7a68faa1c30e345f501ff6afe0d860dd

Observation 7f1b4aa7-6cba-4dd4-972b-d20686689bf4 · outbound

This paper cites MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:41.455573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:41.455573Z digest=sha256:c06815641da0d0adfe8d4fac82f42fe247ec03e0bad9975ff0206fb6a5213b0c

Observation d1a728ed-df21-49e1-950f-bbadfa80112b · outbound

This paper cites Let's Verify Step by Step.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Let's Verify Step by Step

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:41.616138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:41.616138Z digest=sha256:2256d3d413aade409f0a79923b1188e25a8a6b6e27faa59fed6cbadc9da68a7b

Observation cad6590c-cf1c-48fa-a5a2-3e51e0b2dc1d · outbound

This paper cites ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:41.776241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:41.776241Z digest=sha256:629df8a50d9bf4a915b5961dd7be74bc302743a87eff4685b1b0b67bae2c3d5f

Observation a80a4082-64be-48d1-b0ba-890af7af79d9 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning AgentBench: Evaluating LLMs as Agents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:41.907092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:41.907092Z digest=sha256:fbb10f5fdaa0ab66c17a99b1ea70903b752caa981b04bfbd2eda3fcd7eddc14b

Observation 5a04cf71-b807-41e6-aed6-8baef6e5fe51 · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning GAIA: a benchmark for General AI Assistants

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:42.065260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:42.065260Z digest=sha256:d683557b2719b6e5357f32b310bfb9e4e4dbc9a0be6bad07052db8ecc30daf12

Observation a0bb827e-e8e3-4fb8-9109-87a6f223097e · outbound

This paper cites Benchmarking and Understanding Compositional Relational Reasoning of LLMs.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Benchmarking and Understanding Compositional Relational Reasoning of LLMs

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:30:49.095104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:42.215749Z digest=sha256:fe0b65021a0647be37cfe346f86f395d2b71985753d49e9ed3c222bbe4cc7bf5

Observation c135f8fc-0df8-42c5-849d-509b8887d953 · outbound

This paper cites Reasoning curriculum: Bootstrapping broad LLM reasoning from math.arXiv preprint arXiv:2510.26143, 2025.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Reasoning curriculum: Bootstrapping broad LLM reasoning from math.arXiv preprint arXiv:2510.26143, 2025

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:42.370880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:42.370880Z digest=sha256:f0ff4976893524e32f90e74603d84bf9c626dbeb10dea45775b1f31e04fc037d

Observation 263f343b-08a9-4a51-abf6-b814cf1ab837 · outbound

This paper cites Compositional Semantic Parsing on Semi-Structured Tables.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Compositional Semantic Parsing on Semi-Structured Tables

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:42.504186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:42.504186Z digest=sha256:1857e4e49fd4f41f998e25b0f72adcc3315d5bebae3ed42aaddc2ef5afe557f1

Observation d070f3e4-b2fb-4427-96fe-6e06ad3690c8 · outbound

This paper cites Learning to reason across parallel samples for LLM reasoning.arXiv preprint arXiv:2506.09014, 2025.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Learning to reason across parallel samples for LLM reasoning.arXiv preprint arXiv:2506.09014, 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:42.598191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:42.598191Z digest=sha256:e85a3e7daba270a2c15cff2d5fd8627735e598af3536e505059613a0cc0359fd

Observation d5fb6985-d76d-40f3-a2a0-a51a09c8b4e2 · outbound

This paper cites EmoAgent: Assessing and safeguarding human-AI interaction for mental health safety.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning EmoAgent: Assessing and safeguarding human-AI interaction for mental health safety

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:42.722899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:42.722899Z digest=sha256:2cee3859693369ca1867d125e3a6be4fe8d3a98b455e4d945897fe6b3edbb317

Observation 47ee61b7-1e6d-428b-8a50-87643f3dfebc · outbound

This paper cites LogicSkills: A structured benchmark for formal reasoning in large language models.arXiv preprint arXiv:2602.06533, 2026.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning LogicSkills: A structured benchmark for formal reasoning in large language models.arXiv preprint arXiv:2602.06533, 2026

Reference 40

Resolution
verified exact
raw_fallback, observed 2026-08-06T04:30:48.916500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:42.911167Z digest=sha256:a13abd8d1b3bfcff19af748e7dc4a2cbe9de335249ced52eea6e8b0c84a67687

Observation 5bb94826-154a-4760-abbf-48449447200f · outbound

This paper cites Reasoning models are test exploiters: Rethinking multiple-choice.arXiv preprint arXiv:2507.15337, 2025.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Reasoning models are test exploiters: Rethinking multiple-choice.arXiv preprint arXiv:2507.15337, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:43.060338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:43.060338Z digest=sha256:ce31ba85aa700dbfcb44c90154e04cda3cad963595769bfbdeae5ef5b7efb548

Observation 10a0be47-6ecc-48d0-9f7c-af1da1abf9c6 · outbound

This paper cites Revisiting Compositional Generalization Capability of Large Language Models Considering Instruction Following Ability.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Revisiting Compositional Generalization Capability of Large Language Models Considering Instruction Following Ability

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:30:48.758172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:43.189580Z digest=sha256:eb0ec57f0fad8d637c70d50fa999b99fe6d8a463fc9d213f7193ce3d66e326bd

Observation 3769a8a2-921a-42ea-a2ef-7aab904d3024 · outbound

This paper cites AI-Assisted Generation of Difficult Math Questions.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning AI-Assisted Generation of Difficult Math Questions

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:43.301574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:43.301574Z digest=sha256:0e2f6403831167b599ed31b5bcc40aad0bed5a62b83389c3e05692296e23ae84

Observation 47094969-dc1f-4766-9d03-82be32a2124c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:43.408023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:43.408023Z digest=sha256:39d8c6781a886bff54962234323a5324f67b613c889fa1d965eab4786d0a3317

Observation fd41d836-26bf-4ccb-8a8d-7e2b61b4918a · outbound

This paper cites DARE- bench: Evaluating modeling and instruction fidelity of LLMs in data science.arXiv preprint arXiv:2602.24288, 2026.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning DARE- bench: Evaluating modeling and instruction fidelity of LLMs in data science.arXiv preprint arXiv:2602.24288, 2026

Reference 45

Resolution
verified exact
raw_fallback, observed 2026-08-06T04:30:48.717785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:43.543973Z digest=sha256:27b79cc856c6a08fcc4b70f82ac9ca269abddbc2bc75323f42b6fe9f38dde17d

Observation 6c3ae39e-671d-4674-a715-775352b9cce4 · outbound

This paper cites PaperBench: Evaluating AI's Ability to Replicate AI Research.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:43.695729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:43.695729Z digest=sha256:6c3eee962c16efb31eedfa7d2e04b6ecf4784d5fbfac404d3442839ab7d5571c

Observation 705f7630-3321-4561-bf0d-3c7c08bd223a · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:43.796711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:43.796711Z digest=sha256:ebe95667e1609f25e6df8ccfd56a200ab6fa012d77acdd45630a4d02f6a036f1

Observation 7834b505-9f05-4d7a-8e38-ead4733a71c1 · outbound

This paper cites Reinforcement Learning for Self-Improving Agent with Skill Library.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Reinforcement Learning for Self-Improving Agent with Skill Library

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:43.949208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:43.949208Z digest=sha256:00dcf6b16cc09f33c34a0db8f745108ec435ef813b62d9fff67c4fe2b8885b02

Observation a28e7c8e-8e10-49e3-937b-da961f2cf94c · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:44.073590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:44.073590Z digest=sha256:47174304d11d1651c4d24468902a778ac8f766072bd854c6f1da9ed5fee4487f

Observation 9eac2e05-e348-4e38-86bb-fc182ea509c7 · outbound

This paper cites OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:44.200213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:44.200213Z digest=sha256:080a770f836699c6c48e465691c67144770923525359b9f758a79c96f42291a8

Observation 05187c29-552d-42d3-83c5-fd0d9b42a8ec · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:44.345592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:44.345592Z digest=sha256:2e60ee867d18eb90928f60ed92b216f821a904b9505aeeabf4343e1467b4cab0

Observation aa1f2ede-0d00-40fd-92b6-f54e5b855943 · outbound

This paper cites Reinforcingmulti-turn reasoning in LLM agents via turn-level reward design.arXiv preprint arXiv:2505.11821, 2025.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Reinforcingmulti-turn reasoning in LLM agents via turn-level reward design.arXiv preprint arXiv:2505.11821, 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:44.416172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:44.416172Z digest=sha256:68c22194a02d80d2caaafafe454dc640290f615e657b00ab3e18e35ab14e2d53

Observation fb388906-b5c2-4b7d-8ce8-4ffefc7cc252 · outbound

This paper cites Towardscompositionalgeneralization of LLMs via skill taxonomy guided data synthesis.arXiv preprint arXiv:2601.03676, 2026.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Towardscompositionalgeneralization of LLMs via skill taxonomy guided data synthesis.arXiv preprint arXiv:2601.03676, 2026

Reference 53

Resolution
verified exact
raw_fallback, observed 2026-08-06T04:30:48.497373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:44.481066Z digest=sha256:9a038cca22fddaf955df74fff7a3b9f67ef422885ad54c371b1ba2bf4fa2dd05

Observation b0e4634c-9404-4cd6-a681-ab2c31072565 · outbound

This paper cites Hi-ToM: A benchmark for evaluating higher-order theory of mind reasoning in large language models.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Hi-ToM: A benchmark for evaluating higher-order theory of mind reasoning in large language models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:44.546621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:44.546621Z digest=sha256:8c686e1b5284e17d508d8db3d44220167b29d5fdfa5e52ef75c5a51e8ce84420

Observation 81e1505c-7efa-49cc-9bca-9615802b1c9b · outbound

This paper cites CritICL: Inference-time weak-to-strong generalization from small language model failure modes.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning CritICL: Inference-time weak-to-strong generalization from small language model failure modes

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:44.623109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:44.623109Z digest=sha256:9c6ae23cf8811c2e2c3c38f97ab7fbc3820a0079caf133d509bbd12b5f4e368d

Observation 564db2cd-3f30-4d90-aa42-6f26843d43a8 · outbound

This paper cites SkillRL: Evolving agents via recursive skill-augmented reinforcement learning, 2026.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning SkillRL: Evolving agents via recursive skill-augmented reinforcement learning, 2026

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:44.689567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:44.689567Z digest=sha256:82ef561b800691ca656eaa5dd269f3c559e96c0b41bca10850c163227cb60cb9

Observation 4e35a80f-b213-4d98-93a5-f462b54b3f1b · outbound

This paper cites LaRS: Latent Reasoning Skills for Chain-of-Thought Reasoning.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning LaRS: Latent Reasoning Skills for Chain-of-Thought Reasoning

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:30:48.427051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:44.748285Z digest=sha256:2367c71e1946891908add5eef0d004d4dcfa6bf1b9db7bddd0281054caeca50a

Observation 4be92d90-751e-4fa9-88f4-907765b67579 · outbound

This paper cites DeepCritic: Deliberate Critique with Large Language Models.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning DeepCritic: Deliberate Critique with Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:44.844166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:44.844166Z digest=sha256:5ca8d83a9043634fd0a7aaede3e6a97a8acd7ed0dbf638438a2cbccc87d94407

Observation 5fdab5d0-8554-4199-93e0-da0123f8527f · outbound

This paper cites Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:30:48.399971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:44.909089Z digest=sha256:b51c808acfd8251d3dd85ce5ba6f74cf5303ebe871747441045de1664cc2aabb

Observation 4f045811-840f-41c5-a0ce-eebebd0315f5 · outbound

This paper cites LongProc: Benchmarking long-context language models on long procedural generation.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning LongProc: Benchmarking long-context language models on long procedural generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:44.984844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:44.984844Z digest=sha256:becc2182768123019d6a3bf223d86ba6a0e2a00908050cd7a3a702f4332bdf0b

Observation 8eb82e20-7176-43d1-a1af-6dfee17fb294 · outbound

This paper cites Skill-Mix: a Flexible and Expandable Family of Evaluations for AI models.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Skill-Mix: a Flexible and Expandable Family of Evaluations for AI models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:45.049452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:45.049452Z digest=sha256:0a319b27c6f8372ae08f19f6fffd4143c9c4859b2a8216c29446d64c8905252d

Observation ba4c8a55-fae2-4392-a94e-b9b9559cf248 · outbound

This paper cites From𝑓(𝑥) and 𝑔(𝑥) to 𝑓(𝑔(𝑥)) : LLMs learn new skills in RL by composing old ones.arXiv preprint arXiv:2509.25123, 2025.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning From𝑓(𝑥) and 𝑔(𝑥) to 𝑓(𝑔(𝑥)) : LLMs learn new skills in RL by composing old ones.arXiv preprint arXiv:2509.25123, 2025

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:45.118337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:45.118337Z digest=sha256:606e26b070469aa1dc1ec9bb9b63cf553bc036c9028897c81a73b8837e13b4df

Observation 4f441f9d-74ea-40fd-baaf-85ae5922310c · outbound

This paper cites Skill-aware data selection and fine-tuning for data-efficient reasoning distillation.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Skill-aware data selection and fine-tuning for data-efficient reasoning distillation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:45.193661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:45.193661Z digest=sha256:7ff9f410a65994739e2e21abfb8b3500eff3bf483b419c0ba3a03427f4dd9563

Observation 7ecd2ff6-3be4-4195-b060-beed5ce26159 · outbound

This paper cites Skill-awaredataselectionandfine-tuning for data-efficient reasoning distillation, 2026.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Skill-awaredataselectionandfine-tuning for data-efficient reasoning distillation, 2026

Reference 64

Resolution
verified exact
raw_fallback, observed 2026-08-06T04:30:48.157980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:45.286495Z digest=sha256:a2d9a01f0c567822d587c62a412b8c1ccdb93819d87d96901421ffd4c5c17ce8

Observation 9b43e4de-0cd7-482c-9948-122df95a7a07 · outbound

This paper cites Lee, Chenlei Leng, and Fanghui Liu.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Lee, Chenlei Leng, and Fanghui Liu

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:45.384534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:45.384534Z digest=sha256:52c9ffa3d30628ce1a54b72253c93aca9781592fbea35b8239ab90d64220fda0

Observation 0506bd30-63bc-448a-84f7-bf8d05965137 · outbound

This paper cites RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:45.506154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:45.506154Z digest=sha256:6cda878d63dc159cdf782f108eb41e456eec00c85a9b354307e7cf9980b24881

Observation 985fb27d-d840-456c-bdf2-e1b7307a67ec · outbound

This paper cites Can Models Learn Skill Composition from Examples?.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Can Models Learn Skill Composition from Examples?

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:30:47.894990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:45.607770Z digest=sha256:bcc4a33be394098ffa8e0b147bcc10f2422450d7c5c94d8bd17c6af5ce9c5a7d

Observation 51c4fc5c-8ba0-4182-8d2b-6574abb25afd · outbound

This paper cites A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:45.669787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:45.669787Z digest=sha256:cc290509aff2132c1f4c5ee7f7ed61c48b3d03803ed3839d9a6d24693869e781

Observation c0bce8bd-4215-492e-b0f8-185d3f0d0725 · outbound

This paper cites NATURAL PLAN: Benchmarking LLMs on Natural Language Planning.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning NATURAL PLAN: Benchmarking LLMs on Natural Language Planning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:45.725769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:45.725769Z digest=sha256:d299638817efe9ac17e082e7e880001b2ccdc2a5205b57a97f501d413fd6a4bb

Observation 3be68408-0eae-4f51-8cee-15b8fb01407e · outbound

This paper cites SkillRouter: Skill Routing for LLM Agents at Scale.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning SkillRouter: Skill Routing for LLM Agents at Scale

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:45.785107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:45.785107Z digest=sha256:5377973eb8dfe0b9f6f532d9fa57aba649009a1419a0aeeb9e0ff315459c27b5

Observation f26bf2fa-4768-4abe-a813-b75e88bafcd9 · outbound

This paper cites SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:45.842422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:45.842422Z digest=sha256:f40dcd51f0829acbe52e2cc7aa914cc589c8f4925e37e887102f9e76c2ba36de

Observation 63be7e85-8388-469b-a518-0c586be89495 · outbound

This paper cites It clearly outlines the key factors and their interrelationships.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning It clearly outlines the key factors and their interrelationships

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:49.927126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:45.896367Z digest=sha256:fee1ca369cf644c29a2fbee7075891c658bcf73d83c2410d34b9e071ffd7acda

Observation d34e5cd0-02c8-4f10-990b-62371eb35b0e · outbound

This paper cites Specific examples from the data are used to support the narrative.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Specific examples from the data are used to support the narrative

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:49.918035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:45.957925Z digest=sha256:264a060fec0f0c33039b5771e84aeef5aca15bf153663299303b1c12d71d1b25

Observation 72849c37-fbd7-4519-87c9-f82842a368fb · outbound

This paper cites It captures the reader’s interest and effectively conveys the potential crisis.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning It captures the reader’s interest and effectively conveys the potential crisis

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:49.909058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:46.040304Z digest=sha256:7200570970a939a74a6a82dbe4e9fda5a05f1d82f9939c218c4b4d3ac0f32959

Observation 46ee3a08-804c-4b2e-a87f-28fe61b52bd8 · outbound

This paper cites It uses the data and insights from the previous steps to construct a plausible and coherent narrative.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning It uses the data and insights from the previous steps to construct a plausible and coherent narrative

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:49.900061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:46.110513Z digest=sha256:63c2f28e93b2a156f8daacb5dd07bb19d64990be0c03242134bc802955c9896c

Observation 7d98e6aa-0390-4d6f-b55d-63eab9cd9a22 · outbound

This paper cites It should reflect the brand’s commitment to sustainability.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning It should reflect the brand’s commitment to sustainability

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:49.890925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:46.212338Z digest=sha256:48c6b47e4f9f7200f04c898eaf9c1a91b9e3440b5b94ed5e68c5e191243be901

Observation 862d3e81-c652-41de-a436-10c783881f85 · outbound

This paper cites It should not exceed 10 words.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning It should not exceed 10 words

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:49.881322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:46.277290Z digest=sha256:366c0e6b4109f223decae61870ddbe32d2b2c18fa2058d6e915c3abb8a62b313

Observation 66e12a0b-a7d2-4dd0-bfe5-a948738e9651 · outbound

This paper cites an unresolved cited work.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-06T04:30:49.872486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:46.369607Z digest=sha256:3da3dc2a6dacc17db35924420895fd416ca2e84d8e6dfdf6538c1e1e1435c180

Observation 8fa281e4-6a35-4557-93b4-bf1458f388eb · outbound

This paper cites no-interference.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning no-interference

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:49.863445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:46.432911Z digest=sha256:24e171028a419a03b2bd37766d7545865e0fa6900fb146c092dda3bb747ead47

Observation 505dddaa-bfa2-4486-a3a7-529a6a92c816 · outbound

This paper cites DON’T CHANGE THE ANSWER, CORE LOGIC, OR THE SKILL REQUIRED.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning DON’T CHANGE THE ANSWER, CORE LOGIC, OR THE SKILL REQUIRED

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:49.854613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:46.473121Z digest=sha256:ccddeef0fbf25b4403fcb7802ab71c315080fdc3a1449c1c3572091b677d2d73

Observation aa927f39-65d7-40ad-a2d1-8e1547e1ae1b · outbound

This paper cites an unresolved cited work.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-06T04:30:49.846320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:46.578880Z digest=sha256:ff0757a41f8149fe533683d4ae510b747dc206f09d4957ec1c6f2e27bf9cec39

Observation 97a2f555-a3a9-47c6-954f-b3375f1ef442 · outbound

This paper cites Conference trip on constraints.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Conference trip on constraints

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:49.837044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:46.749688Z digest=sha256:580bc493462df5158113c88d87b397fe1579fd948fcc60df123f964ea797f2a3

Observation 6c420ad5-44a8-4954-8aef-9be629177b9d · outbound

This paper cites 29 Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning 2.Scenario consistent.All rewritten steps plausibly belong to the samescenario.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning 29 Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning 2.Scenario consistent.All rewritten steps plausibly belong to the samescenario

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:49.826925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:46.966544Z digest=sha256:4408b871ba5240db3271c455488c6e94b2ae88ed41e6f7882e397027bb5340fe

Observation d60813c5-b66b-44ed-8dd6-4bcb8aca68b5 · outbound

This paper cites The core logic, numerical values, and (for multiple-choice) the option letters and contents must be unchanged.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning The core logic, numerical values, and (for multiple-choice) the option letters and contents must be unchanged

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:49.817051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:47.080943Z digest=sha256:ada550392a42810c4470a5db98895e8034900b65f62d221d3dd0fa9f2f07f4d6

Observation b72f18a9-bb5b-4b8c-88c0-516cb43f235d · outbound

This paper cites FAIL if any step references entities, settings, or framings that contradict the scenario or that read as an unrelated problem pasted in.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning FAIL if any step references entities, settings, or framings that contradict the scenario or that read as an unrelated problem pasted in

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:49.806753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:47.257228Z digest=sha256:3cee6b8c408a6c16e8b46ac7a0fc06a594237793e4e535632e61a24cbbf2d005

Observation a66d5167-c914-4613-9b2f-da05bec9d10b · outbound

This paper cites using the value from the previous step.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning using the value from the previous step

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:49.796334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:47.390521Z digest=sha256:ae3c8a591c9cf54aef587f7202392be0f36e61755299654707a629b4ae51dbdf

Observation 37ee22a3-c1f0-4ec2-9688-1ceb5fef7fab · outbound

This paper cites MAE” is the mean absolute error between the LLM and mean human score in[0, 1]. “Binary agree.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning MAE” is the mean absolute error between the LLM and mean human score in[0, 1]. “Binary agree

Reference 87

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T04:30:49.786033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T04:30:47.578235Z digest=sha256:9c45abc6a900c9aff745281e994cfabf16fa15ccd7a4635ab1f5f78c7f606ecc

Pith citing papers

No inbound Pith citation observations are available.