Pith. sign in

Paper Citation Record · LEDGER

Multi-stage Large Language Model Pipelines Can Outperform GPT-4o in Relevance Assessment

As of 10 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 2 inbound Pith citation observations for arXiv:2501.14296.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.14296 v1

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T15:19:01.309104Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:31:00.001334Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T14:43:58.335104Z

Reference resolution

11 of 11 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d2dd6df6-9cac-4078-b18b-4d7f2b5391d4 · outbound

This paper cites an unresolved cited work.

Multi-stage Large Language Model Pipelines Can Outperform GPT-4o in Relevance Assessment Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:19:01.508069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:19:01.259114Z digest=sha256:29e7a7deee62dc384facd3c39d4014e03eea19c475639460d9155340ec1b48cc

Observation bea5a6c1-dcde-4fb9-a69c-95fcad2426e3 · outbound

This paper cites Voorhees, and Ian Soboroff.

Multi-stage Large Language Model Pipelines Can Outperform GPT-4o in Relevance Assessment Voorhees, and Ian Soboroff

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:19:01.492445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:19:01.264326Z digest=sha256:9b7119b1ccb36d7d3ea8cd71d0c389041607287bc5f6af4667a687527ebe16a6

Observation fad2b0a8-4ccc-4e25-9597-09e6f0c09ad9 · outbound

This paper cites an unresolved cited work.

Multi-stage Large Language Model Pipelines Can Outperform GPT-4o in Relevance Assessment Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:19:01.476177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:19:01.269198Z digest=sha256:c5067dcb97afb9dbfc8cfa8b9fe4d08bef5b1fd569cb1b9e30a4cbd3427b1e0c

Observation 619723c5-648e-4c5f-99e6-f73ae0c4c7eb · outbound

This paper cites an unresolved cited work.

Multi-stage Large Language Model Pipelines Can Outperform GPT-4o in Relevance Assessment Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:19:01.448899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:19:01.273985Z digest=sha256:d2734d9855321c3c90b3fc11356b688abc6f3ca40795e05ad0e266898f192ecc

Observation fc1da27e-3d1a-4abf-a072-d8ed35f8fe4a · outbound

This paper cites an unresolved cited work.

Multi-stage Large Language Model Pipelines Can Outperform GPT-4o in Relevance Assessment Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:19:01.434673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:19:01.278667Z digest=sha256:9683753232b90da87a2e7b28fa1ebbf781f24464f7efd95833a47a838d0e2d13

Observation bf399a4e-3747-462e-ac6e-ac93afe38f9d · outbound

This paper cites an unresolved cited work.

Multi-stage Large Language Model Pipelines Can Outperform GPT-4o in Relevance Assessment Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:19:01.406121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:19:01.288620Z digest=sha256:b17000747ceaf10abcfc6be792fdb6091181f493c7ef1801882ad74f7b456c22

Observation d7d6b61f-4c44-47e7-b4e6-b96a6a891013 · outbound

This paper cites an unresolved cited work.

Multi-stage Large Language Model Pipelines Can Outperform GPT-4o in Relevance Assessment Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:19:01.391508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:19:01.293412Z digest=sha256:eec0675e4bf5c7f67f87fe946251a4d8d52e5220010ca413c54a29860dd44e30

Observation fa85362c-fafb-47bd-a761-5c992578f605 · outbound

This paper cites UMBRELA: UMbrela is the (Open-Source Reproduction of the) Bing RELevance Assessor.

Multi-stage Large Language Model Pipelines Can Outperform GPT-4o in Relevance Assessment UMBRELA: UMbrela is the (Open-Source Reproduction of the) Bing RELevance Assessor

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T15:19:01.298408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:19:01.298408Z digest=sha256:a5b72ed0389f428784d87a0e0ae0baa0ed67d818dec7566cb1636dc28596f877

Observation 052c5493-1504-4ca9-b7ab-61b3037d0773 · outbound

This paper cites an unresolved cited work.

Multi-stage Large Language Model Pipelines Can Outperform GPT-4o in Relevance Assessment Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T15:19:01.304353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:19:01.304353Z digest=sha256:237f1bdc1a2a987d29e43e94b213f8036a7bf07d2d6d3a0871111830a4201e3b

Observation 8c977d79-12de-4c87-a3ab-368fb4686a14 · outbound

This paper cites Shane Culpepper, Falk Scholer, and Paul Thomas.

Multi-stage Large Language Model Pipelines Can Outperform GPT-4o in Relevance Assessment Shane Culpepper, Falk Scholer, and Paul Thomas

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:19:01.365684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:19:01.309104Z digest=sha256:87cb5fe647e4c63aad357773ed34eae2dfb82f1b6705a3d5b58037d6ff3864c9

Observation 16bde970-20a3-4ad1-8f6f-418d7b102371 · outbound

This paper cites an unresolved cited work.

Multi-stage Large Language Model Pipelines Can Outperform GPT-4o in Relevance Assessment Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:19:01.420325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:19:01.283905Z digest=sha256:809d66fd434be600d8c54fbc7a3a33d6153810e37939ad478e9acda36025c652

Pith citing papers

Observation 5a07de35-388b-415a-a0f9-79f19c4ac1ec · inbound

Large Language Models in the Task of Automatic Validation of Text Classifier Predictions cites this paper.

Large Language Models in the Task of Automatic Validation of Text Classifier Predictions Multi-stage Large Language Model Pipelines Can Outperform GPT-4o in Relevance Assessment

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:00.001334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:00.001334Z digest=sha256:40f1e9640cbde5781dec35a722f6c72e0622ec396120431b53e7f4ee04511198

Observation 230f0533-fcec-4b05-ba69-87b8accca1b6 · inbound

Synthetic Data Generation for Phrase Break Prediction with Large Language Model cites this paper.

Synthetic Data Generation for Phrase Break Prediction with Large Language Model Multi-stage Large Language Model Pipelines Can Outperform GPT-4o in Relevance Assessment

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:43:58.338469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:43:58.263199Z digest=sha256:b6d5edb14d3077fcc0e556684c1ff4cfdf126ebe9c9563ff2e7ed173581b4456