Pith. sign in

Paper Citation Record · LEDGER

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy

As of 14 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2507.01327.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.01327 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:02:38.886258Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved60
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 31a208ec-737b-4732-823d-902ac8dc57a8 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:36.942746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:36.942746Z digest=sha256:918e8ad65f4fdc484dee2be26da5cd62f86efdf14a34391d41d52dc5c68609df

Observation 5f9e17e1-3cc5-483a-b922-d22bd53228f7 · outbound

This paper cites o3-mini vs DeepSeek-R1: Which One is Safer?.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy o3-mini vs DeepSeek-R1: Which One is Safer?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:37.032293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:37.032293Z digest=sha256:6a8b9157d701985d690c67ad24049233e80df84e615089a655754c34286c2f8c

Observation 14b9f2ce-14ee-4cf7-a291-fb1d8309115d · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:37.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:37.183211Z digest=sha256:ff8d7a533f66e17b2e78a6c9e94b9ae8c575f6ec6b1d975fca8a642cb877c5a1

Observation 0d5c8270-2ff8-49f1-b33a-2dbbc9d82af1 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.234899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T21:02:37.319463Z digest=sha256:7f20d4f955df70f0eb54af2d7715384b6ddbf653b1eb365b9ee7c00157352fc3

Observation 023fce87-47f5-49f7-8100-39b5698eaf4c · outbound

This paper cites M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:37.502465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:37.502465Z digest=sha256:256ece4bba631c2ae41d82ae2943a6159126bd091e125ea967f71f73087257cf

Observation 0aa882ab-7b75-4db8-8aef-e46eea1ddc3a · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:37.699455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:37.699455Z digest=sha256:6f036284110b4521d295b9fa4d1320f5fda451b8ff137f9c6269d0e13166992f

Observation 86e7978e-5dd8-4787-9961-5830e10eec17 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:37.820378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:37.820378Z digest=sha256:8f8eed9f2c22e0def7eb2754eb9644806bd8ee116014f2692fba329722efa3a3

Observation ce74fb56-f1ee-4738-a1c1-2e2a673cc6c3 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.220572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T21:02:37.955406Z digest=sha256:3465ddf676f3831fb8c85c543368371db04e6d56f45d700e494eed82eafb5833

Observation c657feba-84a2-4a22-8038-9a9c4d28439b · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.205589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.044785Z digest=sha256:0154ead0c21b7fd956454f9f535838b67fde95c8ead57d156c818dbdb5851046

Observation b4835359-7fc4-40df-b9fd-6224fa5dd951 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.151830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.151830Z digest=sha256:a47dce5222aa115743b732c83bdec157b23992b15d69d7bd3512fce80dc41e8a

Observation 2c2ad01c-1d6e-4ccf-abb3-9329f1992d0e · outbound

This paper cites Token-Hungry, Yet Precise: DeepSeek R1 Highlights the Need for Multi-Step Reasoning Over Speed in MATH.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Token-Hungry, Yet Precise: DeepSeek R1 Highlights the Need for Multi-Step Reasoning Over Speed in MATH

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.200430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.200430Z digest=sha256:80735652cba5198742ccfa7b484be6b6a01783e3cc0ca9eff850a032e0159959

Observation a1327e2c-c0e8-4beb-85d4-dfe89e595765 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.180524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.263726Z digest=sha256:3e4ba82cf6c4fc6af6495fc88ffd649e68d6682c8d707d5353bc4462879891a1

Observation e57f1145-b8fc-42c4-bed9-231ca0008c03 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.448293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.448293Z digest=sha256:14c6c2a8d1fac46deb8252e0f2c89968e9477f066da405b4ea62997a23f67b8f

Observation 8878cbaa-47b4-4a76-80a2-7262b7f49da0 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.165410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.528388Z digest=sha256:bd869a9e06f6845bec7ccbc02c998d125c074f29c6891bc55eee3ae7a011f547

Observation dd28b2ec-b78f-4368-bd84-ba7e374f707d · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.646820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.646820Z digest=sha256:c3493aac91607f2f439c136945913a1c998685d2652feeb111e8fa380ab9cac8

Observation 9c259246-689b-4490-928f-5ad5de36cd71 · outbound

This paper cites OpenAI o1 System Card.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy OpenAI o1 System Card

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.664449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.664449Z digest=sha256:80d108d41970ab97dc6c2f3e74c54705bec4d98c4e68b0c39e5f5454d8063db4

Observation b7e29413-02af-4dc8-8d42-a5e0a05c3671 · outbound

This paper cites Scaling Laws for Neural Language Models.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Scaling Laws for Neural Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.671165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.671165Z digest=sha256:ba1271c25498f0d4a660ac4af6393e559caaf807b6ce24f2a2dcaf93f3e3368c

Observation f43a6cd5-3164-47a8-a521-720e6188e685 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Gonzalez, Hao Zhang, and Ion Stoica

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.676052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.676052Z digest=sha256:61ec8c53ada6ca5b5f6e96d360adf2aa5b3e0b469bfa88a7a69417810506d3a1

Observation 12f06d24-1b9c-4673-bf13-d373f5c1b33a · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.681209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.681209Z digest=sha256:c61e942f0314aad7d0e164d5fb14a4bffb7e0303102ea2b4626cb97e5034fc3b

Observation 3b8e6642-fd7f-4db2-babd-ce61d14f1753 · outbound

This paper cites Towards General Text Embeddings with Multi-stage Contrastive Learning.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Towards General Text Embeddings with Multi-stage Contrastive Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.685458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.685458Z digest=sha256:2c0d54de6f589835ba549fe97f28b39bba984699bf88e0cd7a58a7ca9060eb89

Observation 6cb29d71-2aa2-4135-8315-01a3eedbb761 · outbound

This paper cites ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.690014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.690014Z digest=sha256:f9f6c505ea6a39a4714b6c6cf026526be83e87337a7eb92099802837ba476d6b

Observation a3d2dc45-2930-4ced-9b96-f57ef3b601f6 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.694881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.694881Z digest=sha256:fc584a215eae35af51dc8a013ccd26fc35010acfc0161759d0eb639d43a2ad53

Observation 806618c4-0cc6-4a35-9b9a-34319bb0c842 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Understanding R1-Zero-Like Training: A Critical Perspective

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.699466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.699466Z digest=sha256:a57f9bb84f1303db4be6093e8f6bdd9edbd511bfe95fa9c9d504c5da26b0b943

Observation 7586d958-0dc2-4586-9fd6-44af33d8ac51 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.704508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.704508Z digest=sha256:654b388a5238e4ad49b7a90c6032bb4b4331baece81f7d14d7bf033afffbcf64

Observation 07392453-e916-4d8d-8ef8-bc9b199f4f63 · outbound

This paper cites Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.709251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.709251Z digest=sha256:aedf1810d9e71eccbfcdec20fcd0de27f328742d4db3a551521528cf93d109d2

Observation 9aeb3252-57f0-4b23-9cf4-20a64e488caa · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.129593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.713899Z digest=sha256:15f4c9c077003a85f31ad94fc4ce56cbbdb40b45e7fae7cc63aa695be42d1099

Observation b96b7442-c81b-4c59-80d4-041fdf07e4dd · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.115473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.718779Z digest=sha256:e8cc258c9d56fb5787813bc16d87974f6bba24357d608298a1a7d3e75b3bf165

Observation 8d6d9956-b807-470b-ad2e-2f02d026acf9 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.100897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.723190Z digest=sha256:8d0de24ced662f3710a660f79133348b0d69b208d4859cb467db31cf5c33245a

Observation ceacf498-c68d-439d-9547-2842c0e27f5d · outbound

This paper cites GPT-4 Technical Report.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy GPT-4 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.727798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.727798Z digest=sha256:8e7da7251755760d1ed8bb504ab2cc382fcf6fdcab347263b21407f8fd0d009c

Observation f1d61937-ff6d-4540-bd52-7cfa4ed879ee · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.085501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.732475Z digest=sha256:e59fd8b731be18230b2a8f8d30c4773f857b5e962f0be368cc4b120db8fb4895

Observation 04ea2fdc-426e-4521-8c9c-92d6778a9918 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.736901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.736901Z digest=sha256:16d58113e3421cdfc01ce4b42eb4f2225560c5ce164f260011891bc786d7685e

Observation a00037ce-f4da-4fb2-ba28-2a7ffbd0fbe1 · outbound

This paper cites Multi-class Text Classification using BERT-based Active Learning.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Multi-class Text Classification using BERT-based Active Learning

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:02:39.432882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.741316Z digest=sha256:e383bcf5cc6c65fee490340bc616e93c40a4c971b9558ea0657c96eaaac08362

Observation 32ffb3d0-ec8e-4c7b-b520-c4910cf7dc8b · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.059976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.746100Z digest=sha256:5a431f3f429e6165a9187b5ed388c22f54ce6a177b3ccdb1c23c7d3c0f65f35a

Observation 0591341e-520c-4e6c-af89-90d0f5dffecf · outbound

This paper cites Qwen2.5 Technical Report.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Qwen2.5 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.750400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.750400Z digest=sha256:467c01411dc49594f9797f6ddab068906229406a7c09b23d37b9405f770cc0d6

Observation c7ace076-2e60-4f14-8da7-2890c9818394 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.754928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.754928Z digest=sha256:a6b7c7a3c8e6345e72ca1a402102106ccf335f77c659c2419add8812b04e01cf

Observation d9babff3-870f-4bcd-b9bc-6ed2395b8610 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.035515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.758816Z digest=sha256:a406003190c51626d4c01ba5f1086c7bf276b5cc7d373bb98f20faee9cdead9a

Observation 2715de2e-6c1d-45f2-ac73-649341fc1bea · outbound

This paper cites Ray Interference: a Source of Plateaus in Deep Reinforcement Learning.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Ray Interference: a Source of Plateaus in Deep Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.763176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.763176Z digest=sha256:13eefbd640db94398692c18453638459e8dee85123029dcc156a8f9a56ac8c93

Observation 7b522326-3a75-4857-915a-6ccf1c508f45 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Proximal Policy Optimization Algorithms

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.767702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.767702Z digest=sha256:618d578fc4e5155596af757f9e0fbbdac04071208627862c2e1f98cf00c272f2

Observation fdb864e8-2e6b-4ad6-92e9-26720e25ec49 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.776906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.776906Z digest=sha256:f8dc776118176bb76053b33134d68d680477b144686eba1cc7b0046e13b852cd

Observation 4571d06d-73c8-4ff6-83e7-161d0e3e36d0 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy HybridFlow: A Flexible and Efficient RLHF Framework

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.781944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.781944Z digest=sha256:7fb8f7894cababdd23f1a7c6ad8a758485dd0defa9ac6dba1ea6e4a8c201427f

Observation 042cd7d2-6082-4890-bce1-bb2480a4d2e9 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.786360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.786360Z digest=sha256:72dc40fa445a7faeb73ec0277dd38c29d0df6f0df74f72ee770d3312b0765117

Observation e05d6af9-440d-4d16-bb6d-5ce7cc0b3662 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.021093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.790628Z digest=sha256:a12fdc74be038a2c933a548934672b167c2165d56a23c3152db014a8dc063923

Observation 11905553-2ea7-4744-ad24-d71fc44b1561 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.794746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.794746Z digest=sha256:2f217d7637b0c065da73b4a2b0fe6bf564ef40b4126d368a0078215f8a527e0b

Observation 9e72aea4-7d5f-474f-be6f-67a07e0a119e · outbound

This paper cites Fine-tuned vs. Prompt-tuned Supervised Representations: Which Better Account for Brain Language Representations?.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Fine-tuned vs. Prompt-tuned Supervised Representations: Which Better Account for Brain Language Representations?

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:02:39.295404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.799491Z digest=sha256:69c8d46d1bc13f61f497b1970a935e421adeaffb77eb9f0cf3005d0ff8723d71

Observation 8a02cf4b-ebfd-4335-aac8-44d35fa87c26 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.804100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.804100Z digest=sha256:2cecf984ba0d368755c4bfeff975c9befab4bda87826e8cb1ffaf2c873f08119

Observation df9f09b6-8183-466e-98d8-68bd965f2c22 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy LLaMA: Open and Efficient Foundation Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.808953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.808953Z digest=sha256:b18de3273c77a550716b8cc5281f7b7652d8cd5d19de8c265f2db6a993dce9e0

Observation 33e7ef3a-5699-49e0-8229-fa15229b2143 · outbound

This paper cites ChatGPT Empowered Long-Step Robot Control in Various Environments: A Case Application.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy ChatGPT Empowered Long-Step Robot Control in Various Environments: A Case Application

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:02:39.239638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.813932Z digest=sha256:8ebf493926facce5944fbe5f4f7efc7ad352d6fa1097f9bf78994bd274c5e557

Observation 5765637a-d77e-4077-9058-a0f4022bc6d1 · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Finetuned Language Models Are Zero-Shot Learners

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.818681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.818681Z digest=sha256:b11943fb403daa2907015e6152e67d382a9f5c289c184632160dd5f079a61ab2

Observation 8e8d32ef-5428-4420-81a6-7d0ec150fb46 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:40.006437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.823905Z digest=sha256:408cb276deba8a37aa395f2adc5fb010a05a3a58c212847ddff15997353085ad

Observation dd18e72e-1c9a-4bbb-92b7-fee9f47c2141 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.828256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.828256Z digest=sha256:1f7a3eb3e6f1c74d3443864d49b4c39e69a79a8acf1b0e49dbda4387d441b58f

Observation 20842445-d568-402d-b50f-703292704157 · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.833031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.833031Z digest=sha256:b9d6eb8a1b46123f39cf667eefd5c0af6eecec4f11aab92d96def26c3f796479

Observation 4dc20609-ed91-4ef5-a216-9aa1bdba487c · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:39.981747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.837355Z digest=sha256:0e2cab0476e0b27905107fbd6fc1489900c8ef4715adf31afe6f4464251b9b5c

Observation 777fff92-f2f1-415f-9946-188530af05c8 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.841898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.841898Z digest=sha256:b3af514cffa44fc9baa8ce548abd1b2bdc4a1ef703248728677367fff309553a

Observation c1c42d31-9ed7-48d7-8c42-8cefef7bf7ad · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.846079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.846079Z digest=sha256:f65f8c81f3883a91a3f74850483792c6322c3d2f885a12083237792e3ae1a0d8

Observation 1bb0e20d-945b-41d7-a4a4-30fc039d6810 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:39.957748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.851156Z digest=sha256:b9f7331ba606d76182e1020b0fc2cfb95e472fbe61cdf7a5291e6ce994bbe802

Observation 48d27688-5e9e-45c2-a2d9-51bd55bfa345 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.855807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.855807Z digest=sha256:58f080478f96929041b922501dd32e130edad491d1721c0a6b46f4e67fa74d5d

Observation 71613651-96da-44e0-bf60-01d6a1a7ebf9 · outbound

This paper cites 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.860058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.860058Z digest=sha256:a9bcd9cbec62bd666c0b7bef042a4bb979cca515f61ecdca8f09d61bcb0fa3b8

Observation c0df93c6-3af7-4ce3-ad93-beb0d070b678 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.864617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.864617Z digest=sha256:9c8c6b6d06214f4b75c77443aec2bb68c4ca97618c3f0189edfa732eb6b80870

Observation 6c168912-83aa-46d6-94cc-af00a997d0dd · outbound

This paper cites Least-to-Most Prompting Enables Complex Reasoning in Large Language Models.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.868488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.868488Z digest=sha256:a6b7be2846500f11adadbee2c27019ec041eda7c4c00cfb674be34c5496f75ec

Observation c0586836-fa80-412a-920f-13e07fa9f098 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.873273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.873273Z digest=sha256:8703cfd8dcc868dd96c4c3d144d57ba72b3ab14864fac27a4840d7b3e6d51197

Observation 656c26ef-ca14-4248-8233-0421db02e172 · outbound

This paper cites an unresolved cited work.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:39.942268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T21:02:38.877329Z digest=sha256:231c8b3bf8d9aa06db0e1db59bab4da6b07637f99871923ac8fefdf0c8ad8ae1

Observation 1d90440a-2dd7-43ce-8ece-082d594f8f6d · outbound

This paper cites online" 'onlinestring :=.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy online" 'onlinestring :=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.881524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.881524Z digest=sha256:e587920895560b7a1e30b499ab33656953b1b1a5374331d3ef7f8d5fc94d1649

Observation 16f3e57d-a978-4b96-9d35-2180688f1449 · outbound

This paper cites write newline.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy write newline

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.886258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.886258Z digest=sha256:0d50b07a6b04da054ad1d46002cb04cb6f747bb275500170b62b46abba0c1f39

Pith citing papers

No inbound Pith citation observations are available.