Pith. sign in

Paper Citation Record · LEDGER

A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation

As of 19 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 4 inbound Pith citation observations for arXiv:2412.15298.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15298 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:01:22.032652Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:43:26.773582Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T22:35:49.803600Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 33505e2b-e0f1-476b-9dd0-b5567c2d35b2 · outbound

This paper cites A Comprehensive Overview of Large Language Models.

A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation A Comprehensive Overview of Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T12:01:21.914803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:01:21.914803Z digest=sha256:d04e4fc0c2cee3adc53f9dcab700e8a45242f655495bf22ba62690e2283493be

Observation 5df2753b-4781-433b-bffc-69ab70202a16 · outbound

This paper cites A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications.

A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T12:01:21.922532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:01:21.922532Z digest=sha256:fd24172a881b02eb8281c70b22101144a17a8220f125f009eabb2d204ba3fef4

Observation 099164e1-a8bd-457d-8a8a-807d7b4b7b82 · outbound

This paper cites Can an unsupervised clustering algorithm re produce a catego- rization system? In Proceedings of the 5th ACM International Conference on AI in Finance, pages 213–221, 2024.

A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation Can an unsupervised clustering algorithm re produce a catego- rization system? In Proceedings of the 5th ACM International Conference on AI in Finance, pages 213–221, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:01:22.428379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:01:21.933095Z digest=sha256:1eff9ade47ceb8a8fb312158f0100a3f96cd9efd33c40005c77dbc12197d5d45

Observation 1878ef86-c5fb-42e5-88c1-8fdeeba3d7e1 · outbound

This paper cites Human-calibrate d automated testing and validation of generative language models: An overview.

A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation Human-calibrate d automated testing and validation of generative language models: An overview

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:01:22.407632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:01:21.943107Z digest=sha256:d108e770c2450e64706dd12ea19851220318b7a760fbbd4a9d6c518a70aa7a69

Observation 3ef7d701-b159-4d06-8586-750fabea96ea · outbound

This paper cites How to choose a t hreshold for an evaluation metric for large language models, 2024.

A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation How to choose a t hreshold for an evaluation metric for large language models, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:01:22.384672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:01:21.952896Z digest=sha256:2e85db26cb08282f18948a6a27789d6c50d82bbb0cabb2d37c0d3fbd988de5f4

Observation 462d8068-cb13-4302-a522-efa972e1be84 · outbound

This paper cites DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines.

A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T12:01:21.960419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:01:21.960419Z digest=sha256:7bad08e25e8ea5772867cde094b1f77aaba3c0951b74aad0a9dfbe8badedbbfb

Observation 704fc6ff-9df9-479d-a19c-df556e83c010 · outbound

This paper cites Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey.

A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T12:01:21.971357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:01:21.971357Z digest=sha256:f1a0a9a51c93fcd2d2d5262330afc6fe0f709d530bd21a18584d0ecf27a08e75

Observation 5105474f-e8a6-483c-aee5-f50ed0c5cbf7 · outbound

This paper cites Domain Specialization as the Key to Make Large Language Models Disruptive: A Comprehensive Survey.

A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation Domain Specialization as the Key to Make Large Language Models Disruptive: A Comprehensive Survey

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T12:01:21.978252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:01:21.978252Z digest=sha256:26f66fcc8c57731843820e813c04498b991ada1eb59f601aacbae2f209df15eb

Observation 89de5ff8-6d67-42be-9eef-db2eb56a8c67 · outbound

This paper cites Lynx: An Open Source Hallucination Evaluation Model.

A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation Lynx: An Open Source Hallucination Evaluation Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T12:01:21.990131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:01:21.990131Z digest=sha256:7a07b6e5800c875dfbf6168d21f2d0a5bcbbcb1b0c52a554ea49c3d789a5c769

Observation d3807932-4338-43f7-a734-d1abdf327ab2 · outbound

This paper cites Teach Better or Show Smarter? On Instructions and Exemplars in Automatic Prompt Optimization.

A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation Teach Better or Show Smarter? On Instructions and Exemplars in Automatic Prompt Optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T12:01:21.998399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:01:21.998399Z digest=sha256:9d86d8051dfbb65a18aff832553c2df7ea6d0ebeafa6dc230ac44faeb199da99

Observation 56942437-dec7-4fe3-9b22-49bf6b255150 · outbound

This paper cites Dspy guardrails: Building safe ll m applications via self- refining language model pipelines.

A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation Dspy guardrails: Building safe ll m applications via self- refining language model pipelines

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:01:22.364078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:01:22.004134Z digest=sha256:ba19f4f29e4c09347649041d2c2860f48210b0285c2f11a11a6f04ae7b442b39

Observation a28ef56c-fa48-4288-bacb-cd30988bde59 · outbound

This paper cites Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs.

A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T12:01:22.009758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:01:22.009758Z digest=sha256:71f9f81bdf09520d02ef3cdf284aeda92649a4d924c241771f2101e97d68523d

Observation 861fef18-8df1-45c1-bf8d-008d652bf422 · outbound

This paper cites Optuna: A next-generation hyperparameter optimiz ation framework.

A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation Optuna: A next-generation hyperparameter optimiz ation framework

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:01:22.344131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:01:22.019255Z digest=sha256:b944e854fdf6ee2117ea9bf86beb69c53adc425b6edc3f0d6e560adc40d856cb

Observation c8302609-f110-4f57-8e79-b037138b71c1 · outbound

This paper cites Ragas: Automated Evaluation of Retrieval Augmented Generation.

A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation Ragas: Automated Evaluation of Retrieval Augmented Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T12:01:22.025506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:01:22.025506Z digest=sha256:da15f8a0c850e9defd743f3e19b9f277bd4823d08d997a2ca93aea241185d1c6

Observation ae1b5330-cc1b-4245-88db-5d7c768385a0 · outbound

This paper cites Fin e-tuning and prompt op- timization: Two great steps that work better together.

A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation Fin e-tuning and prompt op- timization: Two great steps that work better together

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:01:22.318380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:01:22.032652Z digest=sha256:c023847686eb826742802f09bdb4f37d1e8feaf4ca475737667196a068e9267a

Pith citing papers

Observation 9fda36b6-16e5-4f65-a47e-93d2b28a15bd · inbound

Data Diversification Methods In Alignment Enhance Math Performance In LLMs cites this paper.

Data Diversification Methods In Alignment Enhance Math Performance In LLMs A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:26.773582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:43:26.773582Z digest=sha256:35b4c3176b684978bd9e852f110eb9f924ae932c0ee0735887984171db276a31

Observation 6dc5b067-9875-4d2b-bd40-fa61ba0030e2 · inbound

Toxicity-Aware Few-Shot Prompting for Low-Resource Singlish Translation cites this paper.

Toxicity-Aware Few-Shot Prompting for Low-Resource Singlish Translation A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:03:24.474243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:03:24.474243Z digest=sha256:500c395c68695eb5330d100d59709f1737fe63210351f12bc523a9cd26a1ff55

Observation 783859c4-2f48-4371-9601-93430e7dd6e7 · inbound

FMI@SU ToxHabits: Evaluating LLMs Performance on Toxic Habit Extraction in Spanish Clinical Texts cites this paper.

FMI@SU ToxHabits: Evaluating LLMs Performance on Toxic Habit Extraction in Spanish Clinical Texts A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:35:49.809037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T19:43:18.106356Z digest=sha256:8bfb731429f304249a06a3599c3a35158cac3287aa8434fde86c45c9ab4d2efd

Observation 6a27738a-a448-4059-bf05-c18fee98723b · inbound

From Errors to Rules: Iterative Prompt Optimization for Text Classification cites this paper.

From Errors to Rules: Iterative Prompt Optimization for Text Classification A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T11:09:22.275369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:09:22.275369Z digest=sha256:23fa353e8922e9a17215a45d7c968767d0bcc42b49a6098759294208cc4132cd