Pith. sign in

Paper Citation Record · LEDGER

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny

As of 16 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2506.22370.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22370 v4

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:10:20.375361Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d4c4811e-c837-481f-9b17-92d39f489c0f · outbound

This paper cites an unresolved cited work.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:10:24.100340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:10:19.122363Z digest=sha256:0c4da69fe7a76fa5f1959d3904055e803f53f6bbcc7d415e8db2b39eeff193e8

Observation 61c8ef18-3838-4605-b2eb-af93dbae9874 · outbound

This paper cites Do AI assistants help students write formal specifications? A study with ChatGPT and the B-Method.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Do AI assistants help students write formal specifications? A study with ChatGPT and the B-Method

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T22:10:20.557101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:10:19.146590Z digest=sha256:a578482866061ecc362a3845d6ad22e7eff3ff99c0a70a1beae04c1076175696

Observation 2b587b74-2640-4695-97a5-38c078019567 · outbound

This paper cites In: Proceedings of the 23rd ACM International Workshop on Formal Techniques for Java-like Programs.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: Proceedings of the 23rd ACM International Workshop on Formal Techniques for Java-like Programs

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:23.900773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:10:19.210768Z digest=sha256:3dd749cfd762c54b7ac631c4deddbe76570a637fda819414a90d581311afd4c0

Observation 54a65751-e22c-4207-afe0-22fc08e0a5ed · outbound

This paper cites Computers and Education: Artificial Intelligence7, 100290 (2024).

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Computers and Education: Artificial Intelligence7, 100290 (2024)

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:23.768504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:10:19.243840Z digest=sha256:e2ce4bab6b9ef0497dc4f82822f74629176b7af98363671acbc22e9465bdbea5

Observation b76a109a-0889-4aa4-9422-14de82073c15 · outbound

This paper cites an unresolved cited work.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:10:23.588382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:10:19.290210Z digest=sha256:65eade51fd27016563e8c8f7b02074dea7450ed80473a78d4e6080d51932cbb5

Observation 4370b81c-38ba-4fd1-9bf5-a84598731bb6 · outbound

This paper cites Applied Sciences14(10), 4115 (2024).

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Applied Sciences14(10), 4115 (2024)

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:23.410110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:10:19.378703Z digest=sha256:753e12a16da8776550eea453d08ff08578bb7361cb3427da728a5d91f89a32b5

Observation c0889314-3feb-4250-8d20-554ca277e91f · outbound

This paper cites In: International conference on logic for programming artificial intelligence and reasoning.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: International conference on logic for programming artificial intelligence and reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:19.451424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:19.451424Z digest=sha256:dd65b0980daf97d7d6ae49db9ba32d52c1d153acf57f9ca063d0a59e38db3d4f

Observation 56e87d08-1a2d-41db-ba05-793a0798aa78 · outbound

This paper cites DafnyBench: A Benchmark for Formal Software Verification.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny DafnyBench: A Benchmark for Formal Software Verification

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:19.488681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:19.488681Z digest=sha256:e2fc14650b1dfc30810e40de70d80e366df55216a471d0d4bb46a8c13bdadf54

Observation 75f1cf6d-cdbf-4b22-aade-a14bb249f638 · outbound

This paper cites In: Proceedings of the 56th ACM Technical Symposium on Computer Science Education V.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: Proceedings of the 56th ACM Technical Symposium on Computer Science Education V

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:23.156920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:10:19.493600Z digest=sha256:7deb358c40484f947cf76ce593efe7a8ddf8d1f15255a966cac2b2b4106ced27

Observation 6073b738-31f2-4d1c-9042-a9549d87fde2 · outbound

This paper cites Assured Automatic Programming via Large Language Models.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Assured Automatic Programming via Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:19.502790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:19.502790Z digest=sha256:8069bf17bd75cbf11a3ce513189fb4e7be4ac9b06aa36edea751872d4fade823

Observation b3f59be7-b167-40a1-a408-850c86a0c474 · outbound

This paper cites Proceedings of the ACM on Software Engineering1(FSE), 812–835 (2024).

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Proceedings of the ACM on Software Engineering1(FSE), 812–835 (2024)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:22.970377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:10:19.547781Z digest=sha256:fa675cfda7bdd2802ce00b835f933d8f30a8969adbf23c5efce9e5b175405e60

Observation 46ed1b5e-4f0f-4645-aa9d-219e279763e4 · outbound

This paper cites Proceedings of the ACM on Programming Languages9(OOPSLA1), 1519–1545 (2025).

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Proceedings of the ACM on Programming Languages9(OOPSLA1), 1519–1545 (2025)

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:22.803786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:10:19.614436Z digest=sha256:538fbfaf695c0511ebbdc706e7db89639a60c5a8039b1c4ed4898d69d4922c11

Observation 97463543-45bf-4e0f-bb42-b083cbb0bec9 · outbound

This paper cites In: NASA Formal Methods Symposium.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: NASA Formal Methods Symposium

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:22.635095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:10:19.712977Z digest=sha256:7fa201c76e597ddb5839a8997faddd1183901f9d1970b04f6c61de78443c6f53

Observation 68cd8d5a-b1e2-4995-9a80-49f1b69cee0c · outbound

This paper cites In: International Conference on Fundamentals of Software Engineering.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: International Conference on Fundamentals of Software Engineering

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:22.475039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:10:19.840113Z digest=sha256:63e631a125be9443fbf5f8684f546f5f5a311ea7c4000b30a00926d794c5ed79

Observation 77098d97-8aa7-4a23-9456-0ae2eec6f4d3 · outbound

This paper cites dafny-annotator: AI-Assisted Verification of Dafny Programs.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny dafny-annotator: AI-Assisted Verification of Dafny Programs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:19.945974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:19.945974Z digest=sha256:079518e070d779bdca5d3dee2fe9743a1e81cb8451768b0b881eeb5256f5e765

Observation 6e4b1156-728a-455f-a3d4-d2cd697e98b6 · outbound

This paper cites In: Proceedings of the ACM Con- ference on Global Computing Education Vol 1.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: Proceedings of the ACM Con- ference on Global Computing Education Vol 1

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:22.307291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:10:20.024034Z digest=sha256:a6ad199e504b32e488e7e6bf112073f54aa6589e64a1c14c37d47d2d932a5fb1

Observation 810191bd-857c-4856-9448-11c0964c1342 · outbound

This paper cites It’s weird that it knows what I want.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny It’s weird that it knows what I want

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:22.151705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:10:20.103786Z digest=sha256:ea17ec34748c8dfa967090655ebe16efc97a1fa6ab4f133fa056d56e734475ff

Observation 6ba07279-6064-41a6-9c1e-752edd9fbfa4 · outbound

This paper cites In: Proceedings of the 2024 ACM Conference on International Computing Education Research-Volume 1.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: Proceedings of the 2024 ACM Conference on International Computing Education Research-Volume 1

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:22.026977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:10:20.159230Z digest=sha256:40d294455183a95abf44d11118f2b7e069cc9d12366446509cfdb72d01c9ce96

Observation e952621b-b581-44f9-b11e-b93d7decd9b3 · outbound

This paper cites Exploring the Use of ChatGPT as a Tool for Learning and Assessment in Undergraduate Computer Science Curriculum: Opportunities and Challenges.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Exploring the Use of ChatGPT as a Tool for Learning and Assessment in Undergraduate Computer Science Curriculum: Opportunities and Challenges

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:20.208512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:20.208512Z digest=sha256:716cb6ec82587a322ad5f9a17de81ef65ee1e75ac6691a2b3ddc371a25702d10

Observation 1c3fc999-e5e6-4e00-81fe-42c42ecf9885 · outbound

This paper cites an unresolved cited work.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:10:21.812585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:10:20.225656Z digest=sha256:838ff0a4d05a8252ae1b143c2db40faedfaaec2b07d7cf69129107eef3cf20b2

Observation ec95da06-b675-4a8f-a4f9-072de1460dcd · outbound

This paper cites In: Proceedings of the 2024 IEEE/ACM 12th International Conference on Formal Methods in Software Engineering (FormaliSE).

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: Proceedings of the 2024 IEEE/ACM 12th International Conference on Formal Methods in Software Engineering (FormaliSE)

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:21.643639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:10:20.255691Z digest=sha256:141bf5a90230b61c3b7b102061dc896b0dcfe402ce42c88a0a33ada81d939dd2

Observation bf8e2275-0295-4037-b569-d1d6f8c44caa · outbound

This paper cites In: International Symposium on AI Verification.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: International Symposium on AI Verification

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:21.420615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:10:20.273290Z digest=sha256:3a00e860c42a67247e446d7d0573646d727006f5117b79aa7900f5d0c161a01f

Observation 0b6970fd-49df-42d5-98c6-229f88c64605 · outbound

This paper cites International Journal of Educational Technology in Higher Education21(1), 14 (2024).

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny International Journal of Educational Technology in Higher Education21(1), 14 (2024)

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:21.182525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:10:20.292400Z digest=sha256:90c578bad5034aed182312330460fa280e1bf046a3829a1141b3be5fd956013e

Observation e6684581-ff6d-4104-8e4d-7a10a94d377d · outbound

This paper cites In: 23rd International Conference on Software Engineering and Formal Methods (SEFM) (2025).

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: 23rd International Conference on Software Engineering and Formal Methods (SEFM) (2025)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:20.989008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:10:20.317739Z digest=sha256:286b4a929ee88fc26f18ec028dededdbaecf090d954051c166ada74bf6df5aa5

Observation f0e85018-93e9-4385-a92c-0cef8b8d3ce6 · outbound

This paper cites In: Proceedings of the 46th International Conference on Software Engineering: Software Engineering Education and Training.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: Proceedings of the 46th International Conference on Software Engineering: Software Engineering Education and Training

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:20.839689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:10:20.375361Z digest=sha256:2300fc16d1f57cd3914f7ba7cda1716df76fddec313a3b4c59b05f28e6797ea2

Pith citing papers

No inbound Pith citation observations are available.