Pith. sign in

Paper Citation Record · LEDGER

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models

As of 21 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 2 inbound Pith citation observations for arXiv:2506.22950.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22950 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:01:24.023367Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:26:49.821867Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation ae54f573-f745-4daf-ba65-108b384d61d9 · outbound

This paper cites Language Models are Few-Shot Learners.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Language Models are Few-Shot Learners

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.593717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.593717Z digest=sha256:810f68d7b7a076f0afbb3391c9e8f6e5871a3c3cb93a90708c1fb399676a1a61

Observation a0ffcfce-48fd-4696-8714-fe408679f181 · outbound

This paper cites an unresolved cited work.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:01:25.031292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:01:23.600719Z digest=sha256:4cf25feccfdde2a458e10ab797763f2ea340f53c5975ac63d1613875154ecedc

Observation 35687e8f-703a-4f9d-a023-0057372a4b2d · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Training Verifiers to Solve Math Word Problems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.612403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.612403Z digest=sha256:ef11ee8e020b3955b2c43549133b2b5efd23f716801ed5a871618b918d23f06d

Observation 1fd0fafb-220a-499d-8d60-0cd0fefc6f90 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.617039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.617039Z digest=sha256:8a705ced23fcab2f2a20756ea016635719a16a36cc798182954043c7b4727dc9

Observation 9ce02750-4d96-4171-96db-47d2b1064493 · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.623374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.623374Z digest=sha256:8d663fea6d711ccee0c289f3433f9e2f929f7e5e0720c41b5dd7d179bb8f3038

Observation 02ca16c4-bb02-4f1a-9a5d-807e42616aab · outbound

This paper cites DeepSeek-V3 Technical Report.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models DeepSeek-V3 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.629095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.629095Z digest=sha256:2601c8c4f8a3fbf4a4fd09a11e6402164d1f3dff4563ff7e0a8925be98c5a44b

Observation 7873e5f0-ff6d-4453-90e0-d934f9058916 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.637326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.637326Z digest=sha256:7ed510110bdda8ebb2dea847467277d59add3c738adaeebeaf7484f0cbca9cb0

Observation 7ab8bd83-868e-4506-8a6c-85bd7c9ace69 · outbound

This paper cites an unresolved cited work.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:01:25.004161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:01:23.643295Z digest=sha256:90f29cd2d0d2351b43e515caa677fce7db131b18a6ce495a4525e48f0a583f20

Observation 21cdab9b-a7df-42a8-b479-38dd412f7c39 · outbound

This paper cites The Llama 3 Herd of Models.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.653086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.653086Z digest=sha256:ea4c68540353daf0430640da214c363a6782d7a53701f21e0ea6652105a0d7dc

Observation 91e4a796-2636-4899-b357-0aaf145a915b · outbound

This paper cites AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.648438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.648438Z digest=sha256:d7491e04486763882837a48ad04c70224301801015fae1b32f62e1ff6e7d5e73

Observation e2a12bd6-bce4-45fc-a22b-36429e1786d8 · outbound

This paper cites an unresolved cited work.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.668619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.668619Z digest=sha256:d705d4f5af02c1239d56bfa8879a491f2c33df776103cf742dec0adafe268162

Observation c0a218a9-1eb2-46de-ad02-6bbad8992223 · outbound

This paper cites an unresolved cited work.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:01:24.980796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:01:23.657733Z digest=sha256:a3a110c68841294a69796d111a349f3e9cdb888f09450ba35b4fe3787ed21786

Observation 3a7bf3c4-4390-4b18-83b9-b8bd7623e577 · outbound

This paper cites Let's Verify Step by Step.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Let's Verify Step by Step

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.678791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.678791Z digest=sha256:00192711427ee3929d4cf3ddc4f00473c35ba0e67fc636426b5517afd85d08cd

Observation e5c73d4f-ed82-425e-bdd1-a28487e13bd2 · outbound

This paper cites an unresolved cited work.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:01:24.951744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:01:23.683935Z digest=sha256:f8486d1fb8fd6302e8f1f5967c76c10b5a69d1c8d67c437253965f884b944a7f

Observation 0324c5d3-8bfb-43a9-b25e-f78c9f0978cd · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.672951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.672951Z digest=sha256:59cde82973e24655b54fa059e23de91e5b8c726ef6f74179988507607c808684

Observation 83df7646-07b1-49a2-8b27-64def2f8506a · outbound

This paper cites Training language models to follow instructions with human feedback.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Training language models to follow instructions with human feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.696728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.696728Z digest=sha256:c2abdfad170a9b72eb36a45f8856bf28d9d665939c271d2ca233ccbbde953eb8

Observation 556c831d-9d4d-4ebb-a79a-5b862a0c33b8 · outbound

This paper cites Efficiently Scaling Transformer Inference.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Efficiently Scaling Transformer Inference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.701288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.701288Z digest=sha256:f305c2baad84a579df94f8567fb9704449bfe7c634abb7ceff56106605d6afc8

Observation 1b26baf4-b9f6-4004-89de-59e4080fa7a5 · outbound

This paper cites an unresolved cited work.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.691994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.691994Z digest=sha256:2889c1cf92a85a1258dc70233f0efae0fbe87a7cc0b6932e5636cbaf26a9f2c0

Observation 075eb96c-3cca-4954-8d59-bee904c6f77b · outbound

This paper cites an unresolved cited work.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.716259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.716259Z digest=sha256:6be9a86f1d8780b6721ba81d9c5765a30d9dd59342c720ce92a6d9dbfd046d92

Observation e24bbe5b-553d-4f20-8d0e-ab7256611dc3 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.722022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.722022Z digest=sha256:9242d02992548eb5b9e54abc554f79d115354cb2b695e113dc12238294fae8bb

Observation e501a27e-f3b9-4dd4-a11f-b5cb318a822c · outbound

This paper cites Kalbarczyk, Tamer Başar, and Ravishankar K.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Kalbarczyk, Tamer Başar, and Ravishankar K

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:01:24.923125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:01:23.706120Z digest=sha256:7e29a24550f36e8b586f3db0a4ec6184f1b5ca478c92d9fa7564d6ad99f96063

Observation 7831e371-0d79-4ee5-8c34-0a809bbc6b12 · outbound

This paper cites Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.711196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.711196Z digest=sha256:bdd8f11dece6bd9e9a2da3e2d6d550ccf5129c8d29f9196e93eaeb9469fc30e2

Observation 6342af0b-bbc0-4141-996f-1ae6c8fbad3a · outbound

This paper cites an unresolved cited work.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.740249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.740249Z digest=sha256:793dd25d414ca8a3af8634315ac6fb566c99beb2553d85f91b58ee7a9ba42ab3

Observation 752630ae-3137-4788-a1d3-27aaf5239c58 · outbound

This paper cites an unresolved cited work.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.749311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.749311Z digest=sha256:cdd6877ec400668d8c4a3f25d47b4fc7aea1e91aee3b85cd8b24538510df390f

Observation 526b10c1-46cf-4202-9774-2f0bc85bace8 · outbound

This paper cites ZeRO: Memory Optimizations Toward Training Trillion Parameter Models.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.728685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.728685Z digest=sha256:cfcb474044c2720dd5be7e3e328390c5731f88eca57f6107d52641bed8e328eb

Observation e2b3638e-38a1-4326-946c-6c6f81a290f7 · outbound

This paper cites ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.734112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.734112Z digest=sha256:61a61f689a17fd8acf0dd727d01742c8fa50d2ee80166182497f0f75f251a46e

Observation 461a5b00-b6bf-454a-900d-d65b113bee91 · outbound

This paper cites Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.768568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.768568Z digest=sha256:8d40fe1335d8b1acd434eeba655e4b75e97948f8705580717cbd26f68b05860b

Observation e8c2a1ec-fe9e-4a25-8d62-26aaef5ee500 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.773742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.773742Z digest=sha256:dfd4d2ab40ea68cb3c7627d1cf0feed7c466061cfc29d5222d997e1099f1174a

Observation 43b9d60f-dec3-466c-932b-8145eee0f448 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.778031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.778031Z digest=sha256:dd10e505b81fde74f399a644f3b7a8a46ebd34bc4f94b1ee892be6297c37303d

Observation 5058069b-c939-46af-9846-4f54c05fc867 · outbound

This paper cites On Memorization of Large Language Models in Logical Reasoning.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models On Memorization of Large Language Models in Logical Reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.782702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.782702Z digest=sha256:c5fbb88378bc00609ed8f3601e5f04f168e96f5bafaddf505b26619f4f6d2d4a

Observation 1ca1adcf-f569-4157-91bf-3b52a2fd4607 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.760043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.760043Z digest=sha256:19619241f5d395667d6187eeeda019d5fc3912ceab32f4e51b0fb2a7019caafb

Observation f3b65236-2515-4339-9181-4264bc3c47bf · outbound

This paper cites an unresolved cited work.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.764379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.764379Z digest=sha256:3bd1b35a0b8c20a4b132aee5ec096bcb19b98a45e2189d2435e6ff95ef3ee6e5

Observation fe41b33c-e4ed-4d4c-b436-b80dc72d7493 · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models SGLang: Efficient Execution of Structured Language Model Programs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.812966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.812966Z digest=sha256:2c12c25b2a30f08c4f63939c9e2c3ac2872c04a12cb51f73bc9db77d4fd2c149

Observation dfb51034-43cd-4758-80a4-22ac06a00aa4 · outbound

This paper cites an unresolved cited work.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.837341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.837341Z digest=sha256:7f72e2db68482890ecf6c6c432f9e09e99174e76f8e56e5a90a193ecb826fcdd

Observation 09a90acc-3fd0-4efd-8eda-2e08a99ff9bb · outbound

This paper cites StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:24.023367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:24.023367Z digest=sha256:c7ab4e0525513458a824ab5aaaa84d35b240f1c7b35f0f89658012cbe5db6335

Observation 3a569497-5cc8-4f5e-89d9-40f93c10517d · outbound

This paper cites Qwen3 Technical Report.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Qwen3 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.788123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.788123Z digest=sha256:d89e46614fd084d10cbe108c233197572755877bdcbd8d57b5f6e0d8b582660d

Observation 0abfcd2f-2d7b-478b-bbc5-bb11e48aa632 · outbound

This paper cites LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.797749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.797749Z digest=sha256:6819df5841b0bebf2a44118524faa756978e95e34eb49ad173541a78faa507ae

Observation 584b6f7e-317a-418d-96e6-b6d1030f27e3 · outbound

This paper cites BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.849216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.849216Z digest=sha256:84adadba43aab7dcd4dfbfe4c927ad88c362c3478de32b818f84bf72e72c6b01

Observation f353a0ea-3b89-42f4-bdd7-4bd10b908e8e · outbound

This paper cites Proximal Policy Optimization Algorithms.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Proximal Policy Optimization Algorithms

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.755199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.755199Z digest=sha256:8b3e845ae19900cf1a6ea0448cb4508e4ed9037957d59ca7302d5b64c40fddaa

Observation 5e13f927-1c64-4418-a0bd-da4134782729 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.744845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.744845Z digest=sha256:bf560b94fae1a4384956e128a2f5522c02482fba5f98e5e3c10dbff48acaaa69

Observation 0226e5e9-c2e2-4ecc-b316-0636eeba918c · outbound

This paper cites Multi-Bin Batching for Increasing LLM Inference Throughput.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Multi-Bin Batching for Increasing LLM Inference Throughput

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.662435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.662435Z digest=sha256:14130e52fa16041259f14869554be6e40c2c2aef77dcf3a33dfb2c7ed56e563d

Observation 3b90de7b-6217-4c90-9730-19029bdad7de · outbound

This paper cites ELIS: Efficient LLM Iterative Scheduling System with Response Length Predictor.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models ELIS: Efficient LLM Iterative Scheduling System with Response Length Predictor

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.606728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.606728Z digest=sha256:605da51e95c166f16bdcaa3726f6f3ec4ae9ed72b32384bb8eae11500423b07e

Pith citing papers

Observation 9cd97a3e-0019-4e55-a89f-21905f666d43 · inbound

F-TIS: Harnessing Diverse Models in Collaborative GRPO cites this paper.

F-TIS: Harnessing Diverse Models in Collaborative GRPO Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:41:14.382554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T07:40:35.113858Z digest=sha256:4661b6c2d46ff6accccee6d0fb4de433ac7a5dc35df7e011fb601d50afa403a6

Observation 5e194027-7fbd-4cf7-816b-ace09f840c68 · inbound

TideRL: Boosting Agentic RL Goodput with Readiness-Aware Scheduling cites this paper.

TideRL: Boosting Agentic RL Goodput with Readiness-Aware Scheduling Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T14:26:49.821867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:26:49.821867Z digest=sha256:855a7aaa4ea7aeea0232faf7310fdf5e959184e1edc35bf512180d3212663fe0