Pith. sign in

Paper Citation Record · LEDGER

BIG-Bench Extra Hard

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2502.19187.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.19187 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:07.541003Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1e22fd9f-ce59-403f-9f87-77c97a994e24 · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review BIG-Bench Extra Hard

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:57:38.353786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:9319440c1c1f3f4e418549f55b83c5c7ecd22c5686b3d5fbc59fce70cc1589d1

Observation 7abd6682-5d79-4dbe-8f1b-764533b579b5 · inbound

General-Reasoner: Advancing LLM Reasoning Across All Domains cites this paper.

General-Reasoner: Advancing LLM Reasoning Across All Domains BIG-Bench Extra Hard

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:07.541003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:07.541003Z digest=sha256:bb7eb55017d706800beb01db94dbf07fe9fc6d41f4856e0c44d4df8e476ecc87

Observation eb7e8847-5ffd-44c6-9aa3-af7f4a129e46 · inbound

Do Large Language Models Excel in Complex Logical Reasoning with Formal Language? cites this paper.

Do Large Language Models Excel in Complex Logical Reasoning with Formal Language? BIG-Bench Extra Hard

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:16.192399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:16.192399Z digest=sha256:571577ef819e75318a24f3a01429419598efa0361b8dea804e65ad0ddc075bfb

Observation 4f86518a-073f-4ddd-bdd5-8a3e01d618af · inbound

CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting cites this paper.

CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting BIG-Bench Extra Hard

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:34.218264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:34.218264Z digest=sha256:030ea0208294d75399fdcc7514915c8b01847b6a522b8e58f7f6e0d5672c687a

Observation 28e58b02-1e23-4a62-afba-442c3ec3094a · inbound

SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond cites this paper.

SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond BIG-Bench Extra Hard

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:32.717640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:32.717640Z digest=sha256:89294ebaab22ceaa6fca6bb5e1c43229d8ade7be137b34853ee91238b991c4d5

Observation dab77a8c-5779-48ce-9821-740a9524f26c · inbound

Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem cites this paper.

Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem BIG-Bench Extra Hard

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:13:53.511955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:13:53.511955Z digest=sha256:cadfee7ead55e31c4f7aaab62a719bdfde44a90d9e0a1e111e9495ea7ded2422

Observation 757417b5-63df-4558-ae92-acedc68f49aa · inbound

Real-Time Progress Prediction in Reasoning Language Models cites this paper.

Real-Time Progress Prediction in Reasoning Language Models BIG-Bench Extra Hard

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:08.585573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:08.585573Z digest=sha256:46389aed3cd25411a46e55947a6f5025c3bf2d0aded7fd22750a606b87bfa281

Observation f8e50d8c-0241-45b1-ab33-318d1ddbc40f · inbound

A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search cites this paper.

A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search BIG-Bench Extra Hard

Reference 204

Resolution
unresolved
no resolver link, observed 2026-08-07T05:07:40.349660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:07:40.349660Z digest=sha256:895c395b541ff4c815a5067e051f3e23280cef52ab925e1beece5c6cacdb32e2

Observation f347257f-16a5-479d-89cd-7c123902c146 · inbound

From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation cites this paper.

From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation BIG-Bench Extra Hard

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:16:42.770379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:16:42.770379Z digest=sha256:fc66d7ee648681ef58e859cc4e052e5d054a0567d43af4da749f95e2a4c6cbea

Observation b58082f3-69fe-46e9-8724-3a756e89cf61 · inbound

Bridging Brains and Machines: A Unified Frontier in Neuroscience, Artificial Intelligence, and Neuromorphic Systems cites this paper.

Bridging Brains and Machines: A Unified Frontier in Neuroscience, Artificial Intelligence, and Neuromorphic Systems BIG-Bench Extra Hard

Reference 163

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.974566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T04:37:33.928616Z digest=sha256:083ef4d22e7e11740f689d9ff447abc9ed9c2f3d9d873fe8ac51874b44046214

Observation 525b72be-22f4-478c-837d-5394d13c17d0 · inbound

The Serial Scaling Hypothesis cites this paper.

The Serial Scaling Hypothesis BIG-Bench Extra Hard

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:12:01.955842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T04:08:11.344622Z digest=sha256:5f4d047a8454143090632930708983c8016fdd581c994361811a866cded831c5

Observation 62bbde9c-ac16-4bc1-92d8-1c737c8671eb · inbound

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench cites this paper.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench BIG-Bench Extra Hard

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:07.971609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:07.971609Z digest=sha256:3fa24fe3a7548ead16f3d5148599b8014f42f2ed215335633c40e998d2d767c8

Observation 68e54fea-97ac-4c93-bd86-39cd9a0f4fb6 · inbound

Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead cites this paper.

Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead BIG-Bench Extra Hard

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:12:55.785825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T02:12:48.586913Z digest=sha256:20645a89b8e5c20b3950785a08c3cfacf45deb46b4f4dac288fbc24f352f5209

Observation 72ebc64b-5e92-417b-804f-dfd571453004 · inbound

InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling cites this paper.

InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling BIG-Bench Extra Hard

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:34:24.026350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T22:33:09.674822Z digest=sha256:f69ecfaebe528bb20ae171472eb6c3e2d172d58ac304224c6e67375bde4d7638

Observation 8b2e1b14-a71e-47b6-804f-b2a0955f44f9 · inbound

PuzzleClone: A DSL-Powered Framework for Synthesizing Verifiable Data cites this paper.

PuzzleClone: A DSL-Powered Framework for Synthesizing Verifiable Data BIG-Bench Extra Hard

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T18:08:40.834563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:08:40.834563Z digest=sha256:f84f3e28c46ac0af2e9f6f8ac968555443faebbd663d758d7367a798728aa445

Observation 920baffd-8102-4ed6-a9a0-9c5665ef8fb3 · inbound

Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions cites this paper.

Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions BIG-Bench Extra Hard

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T14:35:50.806216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:35:50.806216Z digest=sha256:7451d3bd40f212243411ffce598863a710325c97fb83f4e56a52c7415f0a66a6

Observation fba4441d-5e24-41df-b266-a4e9756fcbd5 · inbound

DP-FedLoRA: Privacy-Enhanced Federated Fine-Tuning for On-Device Large Language Models cites this paper.

DP-FedLoRA: Privacy-Enhanced Federated Fine-Tuning for On-Device Large Language Models BIG-Bench Extra Hard

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T19:48:48.964322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:48:48.964322Z digest=sha256:6a967434c5895e5d3705910ba9c00c80bfa423c3aadb8d1087c557f009257188

Observation 5ff89a9f-4a92-4457-9983-5683a446d207 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle BIG-Bench Extra Hard

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:31.880939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:31.880939Z digest=sha256:2c491eb5404d8d5fb6783967f2aaf9ba6e7c9aa56c685b18611e0a1797730130

Observation ff9ca1ad-82bb-41f4-8021-9346ea726981 · inbound

Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following cites this paper.

Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following BIG-Bench Extra Hard

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T06:56:01.585103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T06:54:21.052326Z digest=sha256:00376ba76104852d51583e85c17bd18f4d786a3d2b29a814989f1e77e39aa078

Observation 3414d351-c086-43d3-a653-39a06647375d · inbound

Learning to Pose Problems: Reasoning-Driven and Solver-Adaptive Data Synthesis cites this paper.

Learning to Pose Problems: Reasoning-Driven and Solver-Adaptive Data Synthesis BIG-Bench Extra Hard

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:00:25.507895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T22:59:48.260121Z digest=sha256:6e1ab92e43ff7502ec1ed48cf377d21d281e48a018ecac44145768c554ad0ec3

Observation 262e58f0-d614-47dd-a1d8-0cc99518eb4c · inbound

Too long; didn't solve cites this paper.

Too long; didn't solve BIG-Bench Extra Hard

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:42.966189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:29:25.830962Z digest=sha256:1ccff356c375ce8e0b541800e7c87ebfce0f56e0a6271dc246c3dcca919bf8b5

Observation af3a7a04-df78-4956-a6e3-4b5c252b217f · inbound

Too long; didn't solve cites this paper.

Too long; didn't solve BIG-Bench Extra Hard

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T08:26:49.097626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T08:26:49.097626Z digest=sha256:b6e6c5f174f16785b93daed3a7dbf762e014d0a48963b5755cffc7ece9871d21

Observation 55b65bfc-9684-4b7a-a584-ced511ec03d4 · inbound

Agentic Frameworks for Reasoning Tasks: An Empirical Study cites this paper.

Agentic Frameworks for Reasoning Tasks: An Empirical Study BIG-Bench Extra Hard

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:08:26.177810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T08:24:51.573913Z digest=sha256:ee4c1394e7d9b985bd52009b201a6a2c19fcd091dd1b41e268444bdba1dee50a

Observation c8c91847-9a8f-442d-8c80-6c4a83d67baa · inbound

Hypothesis generation and updating in large language models cites this paper.

Hypothesis generation and updating in large language models BIG-Bench Extra Hard

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-08T21:14:12.648060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T14:43:56.125546Z digest=sha256:55d9f167bc330150363661bc6c9a59fb8d0bdc359c82fcc101cb32422654761f

Observation 9453ea26-0d61-4ffa-855b-c66cd20a4537 · inbound

Trading Human Curation for Synthetic Augmentation in RLVR cites this paper.

Trading Human Curation for Synthetic Augmentation in RLVR BIG-Bench Extra Hard

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.172298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T11:10:06.685591Z digest=sha256:88005ced85ec6af27bb6a1b8502c86b6fabd5fb556d3b476b3a3abfd2c090022

Observation 1f92c4b6-1373-4289-801e-80a71a97f175 · inbound

Trading Human Curation for Synthetic Augmentation in RLVR cites this paper.

Trading Human Curation for Synthetic Augmentation in RLVR BIG-Bench Extra Hard

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T15:15:09.675778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:15:09.675778Z digest=sha256:51651d509a2be47e231b65ea1cbacd78e58eb7b2555c5ef0d9a4795f80b5354d

Observation 5edf16bd-706a-439c-8d55-eec16a58dea3 · inbound

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale cites this paper.

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale BIG-Bench Extra Hard

Reference 196

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:17:25.662736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-02T22:10:59.568675Z digest=sha256:06866986a13d6175c01e5757e5915607aaad287aed02a8b85766ea869b395d2d

Observation 6d765560-0150-4a8f-8aa0-663e6a371c68 · inbound

What Drives Interactive Improvement from Feedback? cites this paper.

What Drives Interactive Improvement from Feedback? BIG-Bench Extra Hard

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T12:05:43.484641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T02:36:20.649687Z digest=sha256:95245a0710eac5cb8ba78b023ef25463136311e9b842d58a638961b145697387

Observation f61e585a-de91-4fa7-a9a9-671c044eb43c · inbound

Gemma 4 Technical Report cites this paper.

Gemma 4 Technical Report BIG-Bench Extra Hard

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T07:10:06.281832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:10:06.281832Z digest=sha256:a975d644b01f498298e6d4b5036d587351cca04bcc4d6739444893690a316bc7

Observation 72a0b73d-67eb-43be-86bf-886bb74949aa · inbound

Think Through a Bottleneck: Hourglass Reasoning for Rigorous Induction cites this paper.

Think Through a Bottleneck: Hourglass Reasoning for Rigorous Induction BIG-Bench Extra Hard

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T03:48:31.314623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T03:48:31.314623Z digest=sha256:bc45ae2f2c98b734a331e7f220a09c194bf134409e13679cf908e52d7d8d67e8

Observation 79592456-402e-4e60-bf62-d166c363c77b · inbound

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning cites this paper.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning BIG-Bench Extra Hard

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:40.572955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:40.572955Z digest=sha256:3f967c4f530cf03321e0d41a68cb7fccfd0d782f1f801b395d3a27f57135c8e9