Pith. sign in

Paper Citation Record · LEDGER

BIG-Bench Extra Hard

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2502.19187.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.19187 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:07.541003Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1e22fd9f-ce59-403f-9f87-77c97a994e24 · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review BIG-Bench Extra Hard

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:57:38.353786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:be64c59b3122542dc7d5b7da0da54bf314832d2aed3f5238b1300eca200f1e66

Observation 7abd6682-5d79-4dbe-8f1b-764533b579b5 · inbound

General-Reasoner: Advancing LLM Reasoning Across All Domains cites this paper.

General-Reasoner: Advancing LLM Reasoning Across All Domains BIG-Bench Extra Hard

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:07.541003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:07.541003Z digest=sha256:bb7eb55017d706800beb01db94dbf07fe9fc6d41f4856e0c44d4df8e476ecc87

Observation eb7e8847-5ffd-44c6-9aa3-af7f4a129e46 · inbound

Do Large Language Models Excel in Complex Logical Reasoning with Formal Language? cites this paper.

Do Large Language Models Excel in Complex Logical Reasoning with Formal Language? BIG-Bench Extra Hard

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:16.192399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:16.192399Z digest=sha256:571577ef819e75318a24f3a01429419598efa0361b8dea804e65ad0ddc075bfb

Observation 4f86518a-073f-4ddd-bdd5-8a3e01d618af · inbound

CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting cites this paper.

CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting BIG-Bench Extra Hard

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:34.218264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:34.218264Z digest=sha256:030ea0208294d75399fdcc7514915c8b01847b6a522b8e58f7f6e0d5672c687a

Observation 28e58b02-1e23-4a62-afba-442c3ec3094a · inbound

SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond cites this paper.

SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond BIG-Bench Extra Hard

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:32.717640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:32.717640Z digest=sha256:89294ebaab22ceaa6fca6bb5e1c43229d8ade7be137b34853ee91238b991c4d5

Observation dab77a8c-5779-48ce-9821-740a9524f26c · inbound

Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem cites this paper.

Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem BIG-Bench Extra Hard

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:13:53.511955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:13:53.511955Z digest=sha256:cadfee7ead55e31c4f7aaab62a719bdfde44a90d9e0a1e111e9495ea7ded2422

Observation 757417b5-63df-4558-ae92-acedc68f49aa · inbound

Real-Time Progress Prediction in Reasoning Language Models cites this paper.

Real-Time Progress Prediction in Reasoning Language Models BIG-Bench Extra Hard

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:08.585573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:08.585573Z digest=sha256:46389aed3cd25411a46e55947a6f5025c3bf2d0aded7fd22750a606b87bfa281

Observation f8e50d8c-0241-45b1-ab33-318d1ddbc40f · inbound

A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search cites this paper.

A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search BIG-Bench Extra Hard

Reference 204

Resolution
unresolved
no resolver link, observed 2026-08-07T05:07:40.349660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:07:40.349660Z digest=sha256:895c395b541ff4c815a5067e051f3e23280cef52ab925e1beece5c6cacdb32e2

Observation f347257f-16a5-479d-89cd-7c123902c146 · inbound

From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation cites this paper.

From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation BIG-Bench Extra Hard

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:16:42.770379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:16:42.770379Z digest=sha256:4a679aa0a35876d665d0cd65c02520f93d72b52f811341cbc0c4fe10b69da2c9

Observation b58082f3-69fe-46e9-8724-3a756e89cf61 · inbound

Bridging Brains and Machines: A Unified Frontier in Neuroscience, Artificial Intelligence, and Neuromorphic Systems cites this paper.

Bridging Brains and Machines: A Unified Frontier in Neuroscience, Artificial Intelligence, and Neuromorphic Systems BIG-Bench Extra Hard

Reference 163

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.974566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T04:37:33.928616Z digest=sha256:02870e96de876c83b5db8096a66f1729c9d3a550e61985185964702f6e039b9d

Observation 525b72be-22f4-478c-837d-5394d13c17d0 · inbound

The Serial Scaling Hypothesis cites this paper.

The Serial Scaling Hypothesis BIG-Bench Extra Hard

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:12:01.955842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-19T04:08:11.344622Z digest=sha256:edf48a1eb85d6493532e86df0035b7e753260ec879ba13e4400337eaccd15df5

Observation 62bbde9c-ac16-4bc1-92d8-1c737c8671eb · inbound

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench cites this paper.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench BIG-Bench Extra Hard

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:07.971609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:07.971609Z digest=sha256:99a577178bc8bb7d4e67f1f54c17472ea44bf1a3bf71f9a348f179c2389f672e

Observation 68e54fea-97ac-4c93-bd86-39cd9a0f4fb6 · inbound

Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead cites this paper.

Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead BIG-Bench Extra Hard

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:12:55.785825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T02:12:48.586913Z digest=sha256:fa040d0296d342877ead6e5a9ceb21a92f7d43b679a79e6a759f35b47acc1d25

Observation 72ebc64b-5e92-417b-804f-dfd571453004 · inbound

InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling cites this paper.

InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling BIG-Bench Extra Hard

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:34:24.026350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T22:33:09.674822Z digest=sha256:d844263d563c5baba52607ec9861fd9ba5ee3842915dd97d1a6a2a86d1dbbe40

Observation 8b2e1b14-a71e-47b6-804f-b2a0955f44f9 · inbound

PuzzleClone: A DSL-Powered Framework for Synthesizing Verifiable Data cites this paper.

PuzzleClone: A DSL-Powered Framework for Synthesizing Verifiable Data BIG-Bench Extra Hard

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T18:08:40.834563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:08:40.834563Z digest=sha256:f84f3e28c46ac0af2e9f6f8ac968555443faebbd663d758d7367a798728aa445

Observation 920baffd-8102-4ed6-a9a0-9c5665ef8fb3 · inbound

Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions cites this paper.

Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions BIG-Bench Extra Hard

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T14:35:50.806216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:35:50.806216Z digest=sha256:0a61a30c778e29c228ecf11b3b9aa71ffae7a04b3ac300981d957323e21e006e

Observation fba4441d-5e24-41df-b266-a4e9756fcbd5 · inbound

DP-FedLoRA: Privacy-Enhanced Federated Fine-Tuning for On-Device Large Language Models cites this paper.

DP-FedLoRA: Privacy-Enhanced Federated Fine-Tuning for On-Device Large Language Models BIG-Bench Extra Hard

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T19:48:48.964322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:48:48.964322Z digest=sha256:85b26a9d805281101802c64dda399dfa5940e6017c1f81a71516c51efd8eacc9

Observation 5ff89a9f-4a92-4457-9983-5683a446d207 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle BIG-Bench Extra Hard

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:31.880939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:31.880939Z digest=sha256:2c491eb5404d8d5fb6783967f2aaf9ba6e7c9aa56c685b18611e0a1797730130

Observation ff9ca1ad-82bb-41f4-8021-9346ea726981 · inbound

Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following cites this paper.

Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following BIG-Bench Extra Hard

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T06:56:01.585103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T06:54:21.052326Z digest=sha256:e733ae99763e034a23abe9599f68bca8b2e2dc8b7ba8a92fd68ebac7ccff245d

Observation 3414d351-c086-43d3-a653-39a06647375d · inbound

Learning to Pose Problems: Reasoning-Driven and Solver-Adaptive Data Synthesis cites this paper.

Learning to Pose Problems: Reasoning-Driven and Solver-Adaptive Data Synthesis BIG-Bench Extra Hard

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:00:25.507895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T22:59:48.260121Z digest=sha256:a4968c329a6b87986dc06f1ec95016c1ea2f3b524b4aadfe71f9e9b5952d3426

Observation 262e58f0-d614-47dd-a1d8-0cc99518eb4c · inbound

Too long; didn't solve cites this paper.

Too long; didn't solve BIG-Bench Extra Hard

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:42.966189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:29:25.830962Z digest=sha256:ccbff24f02f63ea5d6d70b5b948ece269bf55b1d2160af35354bd2080ffe697e

Observation af3a7a04-df78-4956-a6e3-4b5c252b217f · inbound

Too long; didn't solve cites this paper.

Too long; didn't solve BIG-Bench Extra Hard

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T08:26:49.097626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T08:26:49.097626Z digest=sha256:b6e6c5f174f16785b93daed3a7dbf762e014d0a48963b5755cffc7ece9871d21

Observation 55b65bfc-9684-4b7a-a584-ced511ec03d4 · inbound

Agentic Frameworks for Reasoning Tasks: An Empirical Study cites this paper.

Agentic Frameworks for Reasoning Tasks: An Empirical Study BIG-Bench Extra Hard

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:08:26.177810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:24:51.573913Z digest=sha256:8d221e3bc3b7a4acf089068cafe7f35c3cc2d7d5973fc1846f7271a92da8e363

Observation c8c91847-9a8f-442d-8c80-6c4a83d67baa · inbound

Hypothesis generation and updating in large language models cites this paper.

Hypothesis generation and updating in large language models BIG-Bench Extra Hard

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-08T21:14:12.648060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T14:43:56.125546Z digest=sha256:23256ebf894336aef01c4ad7736e12143788a95b5eef255edf321b0dd3181a86

Observation 9453ea26-0d61-4ffa-855b-c66cd20a4537 · inbound

Trading Human Curation for Synthetic Augmentation in RLVR cites this paper.

Trading Human Curation for Synthetic Augmentation in RLVR BIG-Bench Extra Hard

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.172298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T11:10:06.685591Z digest=sha256:d719eee5955943cb71a292d53915a3b2f0b88505b6cc748275fbb2afa4483a3c

Observation 1f92c4b6-1373-4289-801e-80a71a97f175 · inbound

Trading Human Curation for Synthetic Augmentation in RLVR cites this paper.

Trading Human Curation for Synthetic Augmentation in RLVR BIG-Bench Extra Hard

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T15:15:09.675778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:15:09.675778Z digest=sha256:51651d509a2be47e231b65ea1cbacd78e58eb7b2555c5ef0d9a4795f80b5354d

Observation 5edf16bd-706a-439c-8d55-eec16a58dea3 · inbound

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale cites this paper.

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale BIG-Bench Extra Hard

Reference 196

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:17:25.662736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-02T22:10:59.568675Z digest=sha256:bd879d710740508b8c638d48d33aa7a9c6722591386c975a786961576b94ec04

Observation 6d765560-0150-4a8f-8aa0-663e6a371c68 · inbound

What Drives Interactive Improvement from Feedback? cites this paper.

What Drives Interactive Improvement from Feedback? BIG-Bench Extra Hard

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T12:05:43.484641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T02:36:20.649687Z digest=sha256:4c10216d7f127a89f2b3f22926a404bcadf2c38000a87e6362b26b03f7447d38

Observation f61e585a-de91-4fa7-a9a9-671c044eb43c · inbound

Gemma 4 Technical Report cites this paper.

Gemma 4 Technical Report BIG-Bench Extra Hard

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T07:10:06.281832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:10:06.281832Z digest=sha256:e35ab85c00109d0466116acd6cfacc6d0ee5d58a834e1e0e589d7da162c9f4dd

Observation 72a0b73d-67eb-43be-86bf-886bb74949aa · inbound

Think Through a Bottleneck: Hourglass Reasoning for Rigorous Induction cites this paper.

Think Through a Bottleneck: Hourglass Reasoning for Rigorous Induction BIG-Bench Extra Hard

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T03:48:31.314623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T03:48:31.314623Z digest=sha256:bc45ae2f2c98b734a331e7f220a09c194bf134409e13679cf908e52d7d8d67e8

Observation 79592456-402e-4e60-bf62-d166c363c77b · inbound

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning cites this paper.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning BIG-Bench Extra Hard

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:40.572955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:40.572955Z digest=sha256:eceb2f2fd3fb9cfadb9ad71a45f88bb040ff011c5b9b9c032f723b1a135b4792