Pith. sign in

Paper Citation Record · LEDGER

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny

As of 10 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2506.22370.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22370 v4

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:10:20.375361Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d4c4811e-c837-481f-9b17-92d39f489c0f · outbound

This paper cites an unresolved cited work.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:10:24.100340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:19.122363Z digest=sha256:8512cdfff1b94143fb00afee7fc055c11e721878d86f39afacff0ece414c6e29

Observation 61c8ef18-3838-4605-b2eb-af93dbae9874 · outbound

This paper cites Do AI assistants help students write formal specifications? A study with ChatGPT and the B-Method.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Do AI assistants help students write formal specifications? A study with ChatGPT and the B-Method

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T22:10:20.557101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:19.146590Z digest=sha256:03f431039e3675c7d17b7f8f3e5702939267ab28cab525aca19f91bc7b69c3f4

Observation 2b587b74-2640-4695-97a5-38c078019567 · outbound

This paper cites In: Proceedings of the 23rd ACM International Workshop on Formal Techniques for Java-like Programs.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: Proceedings of the 23rd ACM International Workshop on Formal Techniques for Java-like Programs

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:23.900773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:19.210768Z digest=sha256:fb453e17f8f7503094440835b66008043e76c8d51f8212deba8f6c1249b72230

Observation 54a65751-e22c-4207-afe0-22fc08e0a5ed · outbound

This paper cites Computers and Education: Artificial Intelligence7, 100290 (2024).

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Computers and Education: Artificial Intelligence7, 100290 (2024)

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:23.768504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:19.243840Z digest=sha256:fd68db45968b76e29f8f3e1c72582808ccda396fcf648047f3e031e38986e7e3

Observation b76a109a-0889-4aa4-9422-14de82073c15 · outbound

This paper cites an unresolved cited work.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:10:23.588382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:19.290210Z digest=sha256:de263289544ca828265432fc653460af298f573f8a4a67898e635dda516c0b06

Observation 4370b81c-38ba-4fd1-9bf5-a84598731bb6 · outbound

This paper cites Applied Sciences14(10), 4115 (2024).

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Applied Sciences14(10), 4115 (2024)

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:23.410110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:19.378703Z digest=sha256:8a1b81f329dae344c8492ac54378f0e3a3a458b1c2d654e3ef7fa9716f5bfa26

Observation c0889314-3feb-4250-8d20-554ca277e91f · outbound

This paper cites In: International conference on logic for programming artificial intelligence and reasoning.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: International conference on logic for programming artificial intelligence and reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:19.451424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:19.451424Z digest=sha256:5ee9d5500c8f78bb7281646ef0e9b56f1e3bac670b862e8cac1f462986b7ee6a

Observation 56e87d08-1a2d-41db-ba05-793a0798aa78 · outbound

This paper cites DafnyBench: A Benchmark for Formal Software Verification.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny DafnyBench: A Benchmark for Formal Software Verification

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:19.488681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:19.488681Z digest=sha256:4bc2c41278b021f8f22158eb560e11dd6eff86a75a2cc00174cb832030a9a1dd

Observation 75f1cf6d-cdbf-4b22-aade-a14bb249f638 · outbound

This paper cites In: Proceedings of the 56th ACM Technical Symposium on Computer Science Education V.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: Proceedings of the 56th ACM Technical Symposium on Computer Science Education V

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:23.156920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:19.493600Z digest=sha256:905d5dc42da3a1318fc83941505ced8d15b7643a14e12986ed4044288b2f5aa8

Observation 6073b738-31f2-4d1c-9042-a9549d87fde2 · outbound

This paper cites Assured Automatic Programming via Large Language Models.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Assured Automatic Programming via Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:19.502790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:19.502790Z digest=sha256:5921e1f3ed1302d97d308ecb5d4429f124d46d99168b69f210c3ba5e03beeb25

Observation b3f59be7-b167-40a1-a408-850c86a0c474 · outbound

This paper cites Proceedings of the ACM on Software Engineering1(FSE), 812–835 (2024).

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Proceedings of the ACM on Software Engineering1(FSE), 812–835 (2024)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:22.970377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:19.547781Z digest=sha256:968394b033eb9ad1ffb9410034f0c7a1bcfde08fd2450c91843733d500cab797

Observation 46ed1b5e-4f0f-4645-aa9d-219e279763e4 · outbound

This paper cites Proceedings of the ACM on Programming Languages9(OOPSLA1), 1519–1545 (2025).

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Proceedings of the ACM on Programming Languages9(OOPSLA1), 1519–1545 (2025)

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:22.803786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:19.614436Z digest=sha256:44955b6e4f4fc211d7605387f8a0d5679031cd47ab19e88bc8f09951fc969417

Observation 97463543-45bf-4e0f-bb42-b083cbb0bec9 · outbound

This paper cites In: NASA Formal Methods Symposium.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: NASA Formal Methods Symposium

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:22.635095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:19.712977Z digest=sha256:15d2b2e3ba689fdea5c59ea20edd9bb46edc5d327623601c34019131938d69cd

Observation 68cd8d5a-b1e2-4995-9a80-49f1b69cee0c · outbound

This paper cites In: International Conference on Fundamentals of Software Engineering.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: International Conference on Fundamentals of Software Engineering

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:22.475039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:19.840113Z digest=sha256:29ded19dfd7933d73e03ed9cdd9adda05639162ef4d029847b2deae42355cc3f

Observation 77098d97-8aa7-4a23-9456-0ae2eec6f4d3 · outbound

This paper cites dafny-annotator: AI-Assisted Verification of Dafny Programs.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny dafny-annotator: AI-Assisted Verification of Dafny Programs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:19.945974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:19.945974Z digest=sha256:cdc626bc5ca9d42b8e06a642eb2126b8584130c549504872245798b3a593fc0b

Observation 6e4b1156-728a-455f-a3d4-d2cd697e98b6 · outbound

This paper cites In: Proceedings of the ACM Con- ference on Global Computing Education Vol 1.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: Proceedings of the ACM Con- ference on Global Computing Education Vol 1

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:22.307291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:20.024034Z digest=sha256:8d404bcc49b0948e506d379f4c99c08e2c9ba50e51d1b3d335cf488995017dc6

Observation 810191bd-857c-4856-9448-11c0964c1342 · outbound

This paper cites It’s weird that it knows what I want.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny It’s weird that it knows what I want

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:22.151705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:20.103786Z digest=sha256:e1a01f348800fc5d6e6a0d1aa813e49cb55be250b6e6f68cd3d59150b84e1564

Observation 6ba07279-6064-41a6-9c1e-752edd9fbfa4 · outbound

This paper cites In: Proceedings of the 2024 ACM Conference on International Computing Education Research-Volume 1.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: Proceedings of the 2024 ACM Conference on International Computing Education Research-Volume 1

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:22.026977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:20.159230Z digest=sha256:7e09ecde03589e1d9ded66e8849496a33844ed82b5e175daafc1293bfe281de2

Observation e952621b-b581-44f9-b11e-b93d7decd9b3 · outbound

This paper cites Exploring the Use of ChatGPT as a Tool for Learning and Assessment in Undergraduate Computer Science Curriculum: Opportunities and Challenges.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Exploring the Use of ChatGPT as a Tool for Learning and Assessment in Undergraduate Computer Science Curriculum: Opportunities and Challenges

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:20.208512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:20.208512Z digest=sha256:7f514479ce0f9ce6a7269799e1c1292c29ee1babdac41f8e68ff1736eb3ef7c7

Observation 1c3fc999-e5e6-4e00-81fe-42c42ecf9885 · outbound

This paper cites an unresolved cited work.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:10:21.812585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:20.225656Z digest=sha256:0223c30ab1c1e44085d2624512b757d3bd4c6f2636f798fed0b627cc57aa71be

Observation ec95da06-b675-4a8f-a4f9-072de1460dcd · outbound

This paper cites In: Proceedings of the 2024 IEEE/ACM 12th International Conference on Formal Methods in Software Engineering (FormaliSE).

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: Proceedings of the 2024 IEEE/ACM 12th International Conference on Formal Methods in Software Engineering (FormaliSE)

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:21.643639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:20.255691Z digest=sha256:652c95a214266fe65817ac916422fe045a0f47e926bfb916a4aa5b10abc2c10f

Observation bf8e2275-0295-4037-b569-d1d6f8c44caa · outbound

This paper cites In: International Symposium on AI Verification.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: International Symposium on AI Verification

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:21.420615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:20.273290Z digest=sha256:6c0179c4f2a4d40c10ec0078195beaa512c959aafa4843eb272aa8d459ced39d

Observation 0b6970fd-49df-42d5-98c6-229f88c64605 · outbound

This paper cites International Journal of Educational Technology in Higher Education21(1), 14 (2024).

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny International Journal of Educational Technology in Higher Education21(1), 14 (2024)

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:21.182525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:20.292400Z digest=sha256:5cdc647dfd56812f8234f0361a681de20e718e5a61e49ea803fd25f5f00d2616

Observation e6684581-ff6d-4104-8e4d-7a10a94d377d · outbound

This paper cites In: 23rd International Conference on Software Engineering and Formal Methods (SEFM) (2025).

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: 23rd International Conference on Software Engineering and Formal Methods (SEFM) (2025)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:20.989008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:20.317739Z digest=sha256:8f463d672c2767a4f324165de7a8ef7be3299cb574190e6689042061f907fe59

Observation f0e85018-93e9-4385-a92c-0cef8b8d3ce6 · outbound

This paper cites In: Proceedings of the 46th International Conference on Software Engineering: Software Engineering Education and Training.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: Proceedings of the 46th International Conference on Software Engineering: Software Engineering Education and Training

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:20.839689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:20.375361Z digest=sha256:7941e68868ca10e65411ec58c268720386bc30b46bbefa24899074260e9722d5

Pith citing papers

No inbound Pith citation observations are available.