Pith. sign in

Paper Citation Record · LEDGER

Tina: Tiny Reasoning Models via LoRA

As of 23 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 15 inbound Pith citation observations for arXiv:2504.15777.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.15777 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:23:45.610470Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:44:35.167445Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:08:21.920833Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 93e319e8-24f5-4ae5-b0a7-ae5b9802b4a1 · outbound

This paper cites an unresolved cited work.

Tina: Tiny Reasoning Models via LoRA Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:23:46.373412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T11:23:45.605382Z digest=sha256:e9f65fa325870da3a7448d0f06fa14282ee69553e9e0d33a84eb7b703d65e9d2

Observation ee612c02-d691-44f1-b249-b1eb26906b3f · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Tina: Tiny Reasoning Models via LoRA Process Reinforcement through Implicit Rewards

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:23:45.308358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:23:45.308358Z digest=sha256:0bd44e1374da434902f16fae00723adb229fb52acefa34691d73a0f949c481b9

Observation 8df13247-50ae-41a8-b0e4-cff5b629206a · outbound

This paper cites DeepSeek-AI.

Tina: Tiny Reasoning Models via LoRA DeepSeek-AI

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:23:45.353089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:23:45.353089Z digest=sha256:eb33ac17b372d55fa807c69ef06c533469717ad0ff4935350c42f315473259be

Observation 0333a5c3-edd7-4876-8a57-d7f74b48414d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Tina: Tiny Reasoning Models via LoRA DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:23:45.358170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:23:45.358170Z digest=sha256:d560a4e1c2e1895da053cc61d8677d1764d96fba92467e7ea38a0df86c83642e

Observation 672c5063-0b93-46b8-9135-ef6f5327bf90 · outbound

This paper cites Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey.

Tina: Tiny Reasoning Models via LoRA Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:23:45.368112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:23:45.368112Z digest=sha256:9b37ca99f41c68f9700cf07a584a4ca535681ac7d2cc87c8b2a5215d3185de9e

Observation c5d92ddf-84b4-4c1f-b3a6-4b2eb3652684 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Tina: Tiny Reasoning Models via LoRA LoRA: Low-Rank Adaptation of Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T11:23:45.378486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:23:45.378486Z digest=sha256:2eae2811cdb755ef04b7de10adc7d9bfdc044dfd3bfec9ba50ba21b0fb8c66ea

Observation d1d64196-4f45-4af1-9615-00bb001e7f69 · outbound

This paper cites O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?.

Tina: Tiny Reasoning Models via LoRA O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:23:45.383186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:23:45.383186Z digest=sha256:8bdeaae9f0073132facf62c322e55840759d2f1faf0c363dbfd7d4c630d8545c

Observation cf495bd9-6d10-4434-a121-7c1523b3eb8e · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Tina: Tiny Reasoning Models via LoRA Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:23:45.387459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:23:45.387459Z digest=sha256:00a7d8b1b595e01938524f4e7653a2516677458a07f075a19383133f4a6348ab

Observation 2ba93b16-607f-4523-bbb9-445bfa16c279 · outbound

This paper cites ReFT: Reasoning with Reinforced Fine-Tuning.

Tina: Tiny Reasoning Models via LoRA ReFT: Reasoning with Reinforced Fine-Tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T11:23:45.393231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:23:45.393231Z digest=sha256:7b80c245e24a91f96b0a899e7eb5ee92f1f91f96d526bf0daae48667236df4cf

Observation d3fa8dc7-43d5-4184-8561-086666a81455 · outbound

This paper cites Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning.

Tina: Tiny Reasoning Models via LoRA Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:23:45.398469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:23:45.398469Z digest=sha256:8c469451a604a0f82bdd27694c225added756db5cafe249c44f68883a6f5c1ef

Observation 3b4ef344-0a96-484c-8b2f-82c8d69b565e · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

Tina: Tiny Reasoning Models via LoRA Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:23:45.403389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:23:45.403389Z digest=sha256:5ebd1fd435123b030023bd5915c1b9d68534a8af6e326abcfab120bdd31bdd26

Observation bf225b99-10bf-42f2-ac2f-ec0162c2cc2b · outbound

This paper cites s1: Simple test-time scaling.

Tina: Tiny Reasoning Models via LoRA s1: Simple test-time scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:23:45.409312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:23:45.409312Z digest=sha256:224744f789e17d47c594b55d72756210922de03bf05adc6aceaeb4d410bf15d3

Observation 1475e24b-a7fb-4bcd-b20d-e46c73213376 · outbound

This paper cites OpenAI o1 System Card.

Tina: Tiny Reasoning Models via LoRA OpenAI o1 System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:23:45.471487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:23:45.471487Z digest=sha256:b785bee470ea6d5f5ac7dbbe42de0c4f164a12e4d8a261ef45e66c8489f693f1

Observation ead0f956-3be5-4904-bbca-b5672bb68b71 · outbound

This paper cites Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers.

Tina: Tiny Reasoning Models via LoRA Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:23:45.548721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:23:45.548721Z digest=sha256:4a05c7b2440f66c4da7fc90d5c884a6cbea410f2446ea7b29ef7389ab18d1a7c

Observation f5c56728-56cd-4861-bd6e-e87a258b960c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Tina: Tiny Reasoning Models via LoRA DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:23:45.584679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:23:45.584679Z digest=sha256:428dabd0cdecdae8616955ce157be040ba472550a09c63ff0d3481d132e0b13c

Observation 8c500d0c-fc56-4b20-9602-d778023207bf · outbound

This paper cites URL http://dx.doi.org/10.1145/3689031.3696075.

Tina: Tiny Reasoning Models via LoRA URL http://dx.doi.org/10.1145/3689031.3696075

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T11:23:45.590102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:23:45.590102Z digest=sha256:dff297a567809273972f57ebebd2ae2e77fa977457415215bf2509918c464812

Observation 29bed85c-3af7-48f9-a4a8-359ddbb6dfa9 · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

Tina: Tiny Reasoning Models via LoRA Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:23:45.595331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:23:45.595331Z digest=sha256:3d2498310e2bf66f61ad4d8ad2db24037f30513fef49c17c7dfce203f0649233

Observation a66c1970-8e8c-4ff3-95f5-23e5d83bd78b · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Tina: Tiny Reasoning Models via LoRA SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:23:45.600284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:23:45.600284Z digest=sha256:e2291090b07943ae38cb43f623829ff740616011f72ea20bbdb8a41792162778

Observation 906306fb-3cd2-4f06-bece-673d7a4a8c9f · outbound

This paper cites lighteval vllm $MODEL_ARGS.

Tina: Tiny Reasoning Models via LoRA lighteval vllm $MODEL_ARGS

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:23:46.328979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T11:23:45.610470Z digest=sha256:ab64be995a0851b0125251cc9d52079d2f330c5e7be179b1a87d481d5748b1f7

Observation 976172a2-c570-4479-aff5-4a1ba8979b61 · outbound

This paper cites ZeRO: Memory Optimizations Toward Training Trillion Parameter Models.

Tina: Tiny Reasoning Models via LoRA ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-16T11:23:45.579251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:23:45.579251Z digest=sha256:e4100ae6e99ace08242050b2868f0192b11069273505e9b1a62adaebc8ac9f4a

Observation 5cff9742-1706-4ac7-aea8-1cf28b566d0a · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Tina: Tiny Reasoning Models via LoRA Measuring Mathematical Problem Solving With the MATH Dataset

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-16T11:23:45.373588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:23:45.373588Z digest=sha256:492b89126a864fe9764a38fedb94808b48f7d9e036834059fb98e0c6ecc96d33

Observation e189e2a0-512a-4af4-9621-e6db0cf5cc4b · outbound

This paper cites an unresolved cited work.

Tina: Tiny Reasoning Models via LoRA Unresolved cited work

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T11:23:45.363347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:23:45.363347Z digest=sha256:5d7d32f538efec0a07c4c76eeb48b71022169dab8169c64cea88ad8315bc144f

Observation 51193c9b-d6ea-45a3-a06b-207ad1df89b9 · outbound

This paper cites Cudo Compute.

Tina: Tiny Reasoning Models via LoRA Cudo Compute

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T11:23:45.257899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:23:45.257899Z digest=sha256:2bd226f33f78a581cb20d4e88341e4d263a22363ba85b55df3671b38516879e9

Observation 047a7e4d-dcb1-4d2d-bd78-32f1e2b4e7b4 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

Tina: Tiny Reasoning Models via LoRA L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-16T11:23:45.237203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:23:45.237203Z digest=sha256:c23b93a550fa785437846d09f468fa3f974287e4a00689f836e20e940d3ae72b

Pith citing papers

Observation 9e7e68fb-7fad-46bd-a097-8001bb4f84c5 · inbound

The Hallucination Tax of Reinforcement Finetuning cites this paper.

The Hallucination Tax of Reinforcement Finetuning Tina: Tiny Reasoning Models via LoRA

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:35.167445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:35.167445Z digest=sha256:f38e03f3195a392def215983c95bca96f3eae28c9853fd788c8d681028dc9956

Observation 26c69b02-4de2-4ea3-a0f2-f3937dc3a5f9 · inbound

Get Experience from Practice: LLM Agents with Record & Replay cites this paper.

Get Experience from Practice: LLM Agents with Record & Replay Tina: Tiny Reasoning Models via LoRA

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:29.928994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:29.928994Z digest=sha256:efce87cf2a2e32d445b5a23bd7ad554edbe1af30fd2bb47a4768d9b086ecf47f

Observation 60dacb50-4306-4873-9a21-8dd28b6d64cd · inbound

RECIPE-TKG: From Sparse History to Structured Reasoning for LLM-based Temporal Knowledge Graph Completion cites this paper.

RECIPE-TKG: From Sparse History to Structured Reasoning for LLM-based Temporal Knowledge Graph Completion Tina: Tiny Reasoning Models via LoRA

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:00.411568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:46:00.411568Z digest=sha256:0935c53558476209812c3fc0f301f563b98a4730a129d993f4eb6be161292e55

Observation c81686f5-7c51-41d6-8b98-efe0d11c8567 · inbound

Resa: Transparent Reasoning Models via SAEs cites this paper.

Resa: Transparent Reasoning Models via SAEs Tina: Tiny Reasoning Models via LoRA

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:42.738769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:42.738769Z digest=sha256:04115f2e7e1f0adc1719749252b1bf8f055a8d444a788606fed8c1e8ef6f14cc

Observation 69f5e61a-5f4d-4ef7-a578-3c7f63f99790 · inbound

Sharp Generalization Bounds for Foundation Models with Asymmetric Randomized Low-Rank Adapters cites this paper.

Sharp Generalization Bounds for Foundation Models with Asymmetric Randomized Low-Rank Adapters Tina: Tiny Reasoning Models via LoRA

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T00:28:43.564233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:28:43.564233Z digest=sha256:1c916567715251b115d3a7ec80b3cb5f3266355fb49b3db0e96870211ced1193

Observation f803da0e-4eef-40ce-bd83-f7d3c4e7a549 · inbound

Efficient Reasoning on the Edge cites this paper.

Efficient Reasoning on the Edge Tina: Tiny Reasoning Models via LoRA

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T23:28:12.790404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:28:12.790404Z digest=sha256:dadcf8d67881e122532faa0893187c3fe503515cde94e4bec5e953012d697143

Observation aa6226c3-a3b8-4163-89af-c413d067b0cf · inbound

ALTO: Adaptive LoRA Tuning and Orchestration for Heterogeneous LoRA Training Workloads cites this paper.

ALTO: Adaptive LoRA Tuning and Orchestration for Heterogeneous LoRA Training Workloads Tina: Tiny Reasoning Models via LoRA

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:48.481454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T20:21:27.428360Z digest=sha256:5a1b53788f9a2469e304ff1fe2d9791e088d5dae3c625a2d2f8a4d73465a3d39

Observation 778680e8-6a2d-4598-a560-5dee44fcee29 · inbound

PubSwap: Public-Data Off-Policy Coordination for Federated RLVR cites this paper.

PubSwap: Public-Data Off-Policy Coordination for Federated RLVR Tina: Tiny Reasoning Models via LoRA

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:21:00.199168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T16:05:27.319466Z digest=sha256:9504408a2003f3a06d15742d91ac07bf7e389da398a4515f4f19635027e35192

Observation 5aa77fee-cc7c-4e87-b549-895aaab2cd76 · inbound

Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes cites this paper.

Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes Tina: Tiny Reasoning Models via LoRA

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:10:09.643586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T04:55:18.468593Z digest=sha256:ca57ad9a5b63d91087ab55ddd119f0724be62ac6e7bda60d11800aca4ac30d8b

Observation bca76dd8-d124-488b-86e6-3bdd0ca0e7d0 · inbound

Fine-Tuning Small Reasoning Models for Quantum Field Theory cites this paper.

Fine-Tuning Small Reasoning Models for Quantum Field Theory Tina: Tiny Reasoning Models via LoRA

Reference 217

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:31:05.510711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T03:23:18.770963Z digest=sha256:766650a361a7115f26d3a51b527aa1c418fc133081e58054966cf838d1a200e7

Observation be1cb815-e181-4e9f-a394-4b51bd1878a6 · inbound

Low-Rank Adaptation Redux for Large Models cites this paper.

Low-Rank Adaptation Redux for Large Models Tina: Tiny Reasoning Models via LoRA

Reference 199

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:26:03.892303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-09T21:48:48.992712Z digest=sha256:859d99a0da8fbfb16aa1ef4519b5e894952ebade921f6d184afd2042c173fd60

Observation 2d50b0c6-e537-4a7e-9773-4074901fff47 · inbound

Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning cites this paper.

Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning Tina: Tiny Reasoning Models via LoRA

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:01:09.753151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T10:33:02.741857Z digest=sha256:023bbc7d3c1eda01be2173929cda145a3a57b4c2a21748a520b6856c2188a73d

Observation 40df2ba8-c027-4b3b-bd91-bf2d4ff37b25 · inbound

Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning cites this paper.

Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning Tina: Tiny Reasoning Models via LoRA

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:26:24.155287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T01:10:54.313745Z digest=sha256:aac29b732aefd996cef40e8e3ee5c75ca26994ddba73fe1699d8ac93d61e1b88

Observation 71d24292-4caf-4320-a2d2-ed5383256db0 · inbound

The Hidden Power of Scaling Factor in LoRA Optimization cites this paper.

The Hidden Power of Scaling Factor in LoRA Optimization Tina: Tiny Reasoning Models via LoRA

Reference 123

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:08:21.922153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T07:14:08.479610Z digest=sha256:4a3a3cf9003599b6e7f50e0c81057c85209c5ad0212bc4a48769dc18fb0531db

Observation 0f0a37e7-5dd7-44d0-9cc5-833ec769b190 · inbound

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts cites this paper.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Tina: Tiny Reasoning Models via LoRA

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:55.596043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:55.596043Z digest=sha256:e1ae499bcde74c919cb88f8d7974925dca256a78e555cf3495626e8f2620fce6