Pith. sign in

Paper Citation Record · LEDGER

CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2503.13517.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.13517 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:22:25.206982Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:56:14.503168Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9993ea64-1444-49b2-b6e0-7d4512665de9 · inbound

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models cites this paper.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:25.206982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:25.206982Z digest=sha256:540d8c007eebdd1dc5b0b0151608ed164c01628064b00ca1d6fc069611f71467

Observation a309628d-d656-4ba9-b38f-d1c6fecd7713 · inbound

MuSciClaims: Multimodal Scientific Claim Verification cites this paper.

MuSciClaims: Multimodal Scientific Claim Verification CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:30.727270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:30.727270Z digest=sha256:f564940aec3c935df77f9ba0553548f2260fa2c4c3a3d908da54c14c41a241dd

Observation 5a236325-4353-425f-887a-5ef92784f916 · inbound

CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics cites this paper.

CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:06:32.145533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T15:05:37.519850Z digest=sha256:99d3bf2d8b0aa6534fb11c5ef214ffb66f8bc538a66b383e303098dd3d2cad0f

Observation f8767d78-1f19-4143-b60d-ccc37d77edd3 · inbound

AInstein: Can LLMs Solve Research Problems From Parametric Memory Alone? cites this paper.

AInstein: Can LLMs Solve Research Problems From Parametric Memory Alone? CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T09:31:11.201411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T09:30:16.609270Z digest=sha256:41ffd9383ea934fc90099672e7403112dc945357ddf940db9440bd4f94b672d8

Observation 8671f932-8a1a-4756-bf91-baf9a3ad94aa · inbound

Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression cites this paper.

Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:00:27.007010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-21T17:59:23.826110Z digest=sha256:0a03adb9f361607a37fb73942dc9dd7617b2715a12f17e147f6869096485b27a

Observation 23e35262-1098-4fca-9608-4990d58a9f3c · inbound

FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs cites this paper.

FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T14:21:26.946925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:21:26.946925Z digest=sha256:7ec43b49a73d596b2f738eb33718226b88de2acda8f821f8a0cf037ca2e0e50a

Observation 26b0604f-ee09-48bc-9c59-2712c8963ba0 · inbound

PaperScope: A Multi-Modal Multi-Document Benchmark for Agentic Deep Research Across Massive Scientific Papers cites this paper.

PaperScope: A Multi-Modal Multi-Document Benchmark for Agentic Deep Research Across Massive Scientific Papers CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:46:00.185498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T15:51:39.424754Z digest=sha256:b71db19922c55abf756697e77c2b5717283f32b207ed6ced076d233bf50d9e65

Observation 0c95b5e2-6721-4d11-ba3c-0b15876c6eef · inbound

The Agentification of Scientific Research: A Physicist's Perspective cites this paper.

The Agentification of Scientific Research: A Physicist's Perspective CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:15:11.111291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T11:12:32.165086Z digest=sha256:64a97fbf33848b5f89b83b740dc2510a81803490bc781986a05a05fff829501d

Observation 43ef07e9-4470-420d-aa5f-39b6e8ae1c70 · inbound

AI scientists produce results without reasoning scientifically cites this paper.

AI scientists produce results without reasoning scientifically CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:16:06.962359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T03:56:34.581133Z digest=sha256:a57257978dcebd82964bae5c5350a77f5c1a9570380b2d550a097e789f823922

Observation d8570ee1-ef87-4da3-a1d2-2c8b25054755 · inbound

Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction cites this paper.

Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:05:06.650089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-15T06:04:03.605898Z digest=sha256:ea4c65443d538eb9bc0dd238d9a72582c84f064cb29fb5710fd647e2588e4112

Observation 515e8379-feb9-452c-ab0b-6a011dd340e5 · inbound

Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches cites this paper.

Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:14.504795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T17:35:01.285534Z digest=sha256:b7123f7f67e6ca9b42d1f3a59613615ee23f4aabe1b99c778f636b415d271dd7

Observation 0f8b3c20-782e-4678-9443-99cb829e2592 · inbound

Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches cites this paper.

Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:35:34.124483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-01T07:22:25.349398Z digest=sha256:032b017ead385d924a2e15f3252e9fe9103b4f5f1ea7e1944c9d6855f1f213d9