Pith. sign in

Paper Citation Record · LEDGER

Evalita-LLM: Benchmarking Large Language Models on Italian

As of 12 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 3 inbound Pith citation observations for arXiv:2502.02289.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02289 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T12:43:45.354815Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:21:45.017840Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T08:19:52.876928Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved7
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation abf912cf-0ee2-4d2c-9b4b-ec3b2f5fe5ff · outbound

This paper cites an unresolved cited work.

Evalita-LLM: Benchmarking Large Language Models on Italian Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-09T12:43:45.599975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T12:43:45.280100Z digest=sha256:23fea4bdd9a3b483feaf79315332bf7b9ca51aabdc7c2c2d742ee45a692f91e5

Observation 06f21b7e-8ceb-4120-bece-b841f10558fd · outbound

This paper cites In: Proceedings of the Tenth Italian Conference on Computational Linguistics (CLiC-it 2024) (2024).

Evalita-LLM: Benchmarking Large Language Models on Italian In: Proceedings of the Tenth Italian Conference on Computational Linguistics (CLiC-it 2024) (2024)

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:43:45.590296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T12:43:45.284235Z digest=sha256:739722cca76527b36b2ab14459c8b772aed1b2df438cec5623be0671c9b49c8e

Observation bbebef4f-9824-4e69-b8ba-1678644bac8d · outbound

This paper cites Proceedings of the International Conference on Learning Representations (ICLR) (2021).

Evalita-LLM: Benchmarking Large Language Models on Italian Proceedings of the International Conference on Learning Representations (ICLR) (2021)

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:43:45.580088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T12:43:45.287638Z digest=sha256:fed695e2bed1cb7f014633eaa39b21bf5ca5104149c3998f0acffc823ca0e3cb

Observation c57a7210-cbe3-48b4-9703-e4e04704c413 · outbound

This paper cites : Calamita: Challenge the abilities of language models in italian.

Evalita-LLM: Benchmarking Large Language Models on Italian : Calamita: Challenge the abilities of language models in italian

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:43:45.566832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T12:43:45.291268Z digest=sha256:dc36f43ffa6c0610b5fc4f8827d2d05204d08ce31b620964910c08aa795937d3

Observation cb3fe706-8351-42ce-bf0a-9f3710f5efa7 · outbound

This paper cites Lessons from the Trenches on Reproducible Evaluation of Language Models.

Evalita-LLM: Benchmarking Large Language Models on Italian Lessons from the Trenches on Reproducible Evaluation of Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T12:43:45.294991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:43:45.294991Z digest=sha256:0e3b2fe27c89751437d79e37fd49138773296035f4e47bdb29810f8da8b542f7

Observation 6920c427-e1d6-4133-b7d9-4f9e8be038d1 · outbound

This paper cites EV ALITA (2023).

Evalita-LLM: Benchmarking Large Language Models on Italian EV ALITA (2023)

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:43:45.553386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T12:43:45.298802Z digest=sha256:51ac29138388cfb0a48367b00f3bca37d6ec0000338ccf1cc3cab934b1cb2b41

Observation ead6d0df-8031-4bf2-9e27-aaaff0679a76 · outbound

This paper cites In: Proceedings of EV ALITA 2009 (2009).

Evalita-LLM: Benchmarking Large Language Models on Italian In: Proceedings of EV ALITA 2009 (2009)

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:43:45.543034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T12:43:45.302542Z digest=sha256:5207f2c9447ddf857c7bbcd24d46a8a91fd6e86c15c6cebea976cd43b527d446

Observation 1347b221-df26-4b53-aa1d-d7a74955ef98 · outbound

This paper cites : Overview of the evalita 2016 sentiment polarity classification task.

Evalita-LLM: Benchmarking Large Language Models on Italian : Overview of the evalita 2016 sentiment polarity classification task

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:43:45.533865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T12:43:45.305966Z digest=sha256:65003f51396d759a6910bb606ad269f5f0053b250948a50e5791934d9981bd13

Observation 7a7357ae-53f6-4b29-8afd-3a1e8077e144 · outbound

This paper cites In: of the Final Workshop 7 December 2016, Naples, p.

Evalita-LLM: Benchmarking Large Language Models on Italian In: of the Final Workshop 7 December 2016, Naples, p

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:43:45.523365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T12:43:45.309114Z digest=sha256:65bf877c0171844b302886cdaf48d01e0ce5bcc6fcc9d63791183b7a974395c7

Observation 0ffd585c-4df3-4aa6-a722-bf6d7c10a5df · outbound

This paper cites In: CLiC-it (2023).

Evalita-LLM: Benchmarking Large Language Models on Italian In: CLiC-it (2023)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:43:45.511680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T12:43:45.312306Z digest=sha256:f00713f24bd1bfeb5a37327162e243cdcedec924f281d2d887a1bdd6ad74874f

Observation fb4d7656-c2c6-4ab2-87e3-db8dfc592642 · outbound

This paper cites In: Proceedings of 40 EV ALITA Workshop, 11th Congress of Italian Association for Artificial Intelli- gence, Reggio Emilia, Italy (2009).

Evalita-LLM: Benchmarking Large Language Models on Italian In: Proceedings of 40 EV ALITA Workshop, 11th Congress of Italian Association for Artificial Intelli- gence, Reggio Emilia, Italy (2009)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:43:45.501695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T12:43:45.315607Z digest=sha256:aba3b8d38bc745a5c02ca6b0616730c6b4c8b40c1aa9dc6c22b25b3dc0d97896

Observation 8adcf9e7-a95f-4127-8b70-cda3a06b2933 · outbound

This paper cites : Nermud at evalita 2023: overview of the named-entities recognition on multi-domain documents task.

Evalita-LLM: Benchmarking Large Language Models on Italian : Nermud at evalita 2023: overview of the named-entities recognition on multi-domain documents task

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:43:45.492061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T12:43:45.318948Z digest=sha256:b9cea152cc7dd03ba28a28db9268f267c0636925ae943f1a52c633c5216e9c85

Observation a7881a81-c4b7-485b-8189-222ef4cc21aa · outbound

This paper cites EV ALITA (2023).

Evalita-LLM: Benchmarking Large Language Models on Italian EV ALITA (2023)

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:43:45.481568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T12:43:45.322246Z digest=sha256:5b0d0c57e42db182add9d6f4f46ebfcad47576e4b4ebb03a182956cd1957c840

Observation 8e0ef92c-e636-49a7-b7b6-75a5407318dc · outbound

This paper cites Proceedings of the Seventh Italian Conference on Computational Linguistics CLiC-it 2020 (2020).

Evalita-LLM: Benchmarking Large Language Models on Italian Proceedings of the Seventh Italian Conference on Computational Linguistics CLiC-it 2020 (2020)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:43:45.471342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T12:43:45.326021Z digest=sha256:ab3de6e0bc251739972a0cf3498922f242ee938170f239445dc9dd7e664a86f0

Observation 1ddf87fa-6bb8-4879-b433-fb169fd6f789 · outbound

This paper cites Information 13(5) (2022) https://doi.

Evalita-LLM: Benchmarking Large Language Models on Italian Information 13(5) (2022) https://doi

Reference 15

Resolution
verified exact
doi, observed 2026-08-09T12:43:45.391920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T12:43:45.329801Z digest=sha256:9b69e6269cc650d250c538324b9e97502c86eec39d21e924defd2ea5ff3766a8

Observation 0666da02-62d6-47eb-a534-844480b1a6ca · outbound

This paper cites Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing.

Evalita-LLM: Benchmarking Large Language Models on Italian Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T12:43:45.333606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:43:45.333606Z digest=sha256:e13288bf2a1d5b875deaed07c9a8ceabbdcc8e98d427e9c830253f550b6d5652

Observation ff26a153-3347-4ae6-9a21-ae095f436840 · outbound

This paper cites In: Toutanova, K., Rumshisky, A., Zettlemoyer, L., Hakkani-Tur, D., Beltagy, I., Bethard, S., Cotterell, R., Chakraborty, T., Zhou, Y.

Evalita-LLM: Benchmarking Large Language Models on Italian In: Toutanova, K., Rumshisky, A., Zettlemoyer, L., Hakkani-Tur, D., Beltagy, I., Bethard, S., Cotterell, R., Chakraborty, T., Zhou, Y

Reference 17

Resolution
malformed identifier
raw_fallback, observed 2026-08-09T12:43:45.460943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T12:43:45.337075Z digest=sha256:74d1ceb3e210799d382d7286e76b75141267eafde70e6c9b746ec6f4ad12b901

Observation 66da9b07-76a6-4a1d-97b5-e7922a4032a2 · outbound

This paper cites A NotSo Simple Way to Beat Simple Bench.

Evalita-LLM: Benchmarking Large Language Models on Italian A NotSo Simple Way to Beat Simple Bench

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-08-09T12:43:45.432046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T12:43:45.340196Z digest=sha256:51395a37903bc39d69127e853a82903a7098c6a3f1723a971f2b15658f52b615

Observation ebf17e8a-803a-46f9-9a86-1f73a7e839db · outbound

This paper cites State of What Art? A Call for Multi-Prompt LLM Evaluation.

Evalita-LLM: Benchmarking Large Language Models on Italian State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T12:43:45.343766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:43:45.343766Z digest=sha256:6823dcdc90985902846337b83fd4565d8e57492c1f1003c1099d6d3afd8b31e8

Observation 348d50de-ceee-43ed-b366-50d91bf6ddea · outbound

This paper cites Efficient multi-prompt evaluation of LLMs.

Evalita-LLM: Benchmarking Large Language Models on Italian Efficient multi-prompt evaluation of LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T12:43:45.347746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:43:45.347746Z digest=sha256:f1410dcfce72954f4aaf83f0c214dc722275a34fc39ebe2f4e49f0ee79d9b488

Observation e476783c-281b-4f14-9f18-dc83110409c6 · outbound

This paper cites In: Al-Onaizan, Y., Bansal, M., Chen, Y.-N.

Evalita-LLM: Benchmarking Large Language Models on Italian In: Al-Onaizan, Y., Bansal, M., Chen, Y.-N

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T12:43:45.351620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:43:45.351620Z digest=sha256:c1cbb9fee229925ab602d4e2cce0bf5f8668a910dc0fa946a5fddd55d270efc0

Observation 3d433e1b-b5f9-4e19-85ee-a92506f5868e · outbound

This paper cites A Survey of Large Language Models.

Evalita-LLM: Benchmarking Large Language Models on Italian A Survey of Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T12:43:45.354815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:43:45.354815Z digest=sha256:60e098d7e7e91c6f2c33cdf81c95181384eb44638d4ceb03a726ba7cabc762db

Pith citing papers

Observation a7adaa20-c0dc-4b4e-aa30-f04429f2e38f · inbound

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs cites this paper.

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs Evalita-LLM: Benchmarking Large Language Models on Italian

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:53.515641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T03:40:04.692279Z digest=sha256:1770e64200199d9298ad6e74965606c6a223bdaebfa526dfde3585707eeac65a

Observation c98e34a3-206b-42d1-962d-74163b7c14d3 · inbound

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs cites this paper.

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs Evalita-LLM: Benchmarking Large Language Models on Italian

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:19:52.878692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T08:14:55.858466Z digest=sha256:4b03a3ff6abcac6ee616cb5741cd801f26dbcf75af223917dfe6694c312029e8

Observation 50242c54-3c37-4d0a-8311-5e6ccd0e23b5 · inbound

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark cites this paper.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Evalita-LLM: Benchmarking Large Language Models on Italian

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.017840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.017840Z digest=sha256:ca8f6fa40d74cdee543c03b5b6a9899487c5725cd890f7847cd6dca0f16cac56