Pith. sign in

Paper Citation Record · LEDGER

Evalita-LLM: Benchmarking Large Language Models on Italian

As of 12 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 3 inbound Pith citation observations for arXiv:2502.02289.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02289 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T12:43:45.354815Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:21:45.017840Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T08:19:52.876928Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved7
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation abf912cf-0ee2-4d2c-9b4b-ec3b2f5fe5ff · outbound

This paper cites an unresolved cited work.

Evalita-LLM: Benchmarking Large Language Models on Italian Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-09T12:43:45.599975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T12:43:45.280100Z digest=sha256:adb4a837ce132d4744f51329b26e74ae8d945b1ce6a7356d0b0588ab54f0c1b1

Observation 06f21b7e-8ceb-4120-bece-b841f10558fd · outbound

This paper cites In: Proceedings of the Tenth Italian Conference on Computational Linguistics (CLiC-it 2024) (2024).

Evalita-LLM: Benchmarking Large Language Models on Italian In: Proceedings of the Tenth Italian Conference on Computational Linguistics (CLiC-it 2024) (2024)

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:43:45.590296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T12:43:45.284235Z digest=sha256:541cbef21981df66cf6caeb7b4eaabcc2c7da49aafd21656c24c8146098eeb52

Observation bbebef4f-9824-4e69-b8ba-1678644bac8d · outbound

This paper cites Proceedings of the International Conference on Learning Representations (ICLR) (2021).

Evalita-LLM: Benchmarking Large Language Models on Italian Proceedings of the International Conference on Learning Representations (ICLR) (2021)

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:43:45.580088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T12:43:45.287638Z digest=sha256:f623bb58379a5dea103c783f059f6dac8ff407ef55e6ad2b10f21c9215b1e84a

Observation c57a7210-cbe3-48b4-9703-e4e04704c413 · outbound

This paper cites : Calamita: Challenge the abilities of language models in italian.

Evalita-LLM: Benchmarking Large Language Models on Italian : Calamita: Challenge the abilities of language models in italian

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:43:45.566832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T12:43:45.291268Z digest=sha256:d3c3a333e2c7b12f5b10c419f055afd68a195218f7c4d405d2ade5a5ab7de785

Observation cb3fe706-8351-42ce-bf0a-9f3710f5efa7 · outbound

This paper cites Lessons from the Trenches on Reproducible Evaluation of Language Models.

Evalita-LLM: Benchmarking Large Language Models on Italian Lessons from the Trenches on Reproducible Evaluation of Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T12:43:45.294991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:43:45.294991Z digest=sha256:0e3b2fe27c89751437d79e37fd49138773296035f4e47bdb29810f8da8b542f7

Observation 6920c427-e1d6-4133-b7d9-4f9e8be038d1 · outbound

This paper cites EV ALITA (2023).

Evalita-LLM: Benchmarking Large Language Models on Italian EV ALITA (2023)

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:43:45.553386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T12:43:45.298802Z digest=sha256:016cc0d0b8720cfcef7755fe109592e1ffbb2b66f62fd6c95333b79e2babfc9a

Observation ead6d0df-8031-4bf2-9e27-aaaff0679a76 · outbound

This paper cites In: Proceedings of EV ALITA 2009 (2009).

Evalita-LLM: Benchmarking Large Language Models on Italian In: Proceedings of EV ALITA 2009 (2009)

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:43:45.543034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T12:43:45.302542Z digest=sha256:35387f881e0d89696d37cfc30e69d8c27dccb3a6079689e7c258c633a4fe0b6a

Observation 1347b221-df26-4b53-aa1d-d7a74955ef98 · outbound

This paper cites : Overview of the evalita 2016 sentiment polarity classification task.

Evalita-LLM: Benchmarking Large Language Models on Italian : Overview of the evalita 2016 sentiment polarity classification task

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:43:45.533865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T12:43:45.305966Z digest=sha256:0f882aa7dbc289c59d99eaa709d94e6e535707e53cf80b3cbe520b2df7130288

Observation 7a7357ae-53f6-4b29-8afd-3a1e8077e144 · outbound

This paper cites In: of the Final Workshop 7 December 2016, Naples, p.

Evalita-LLM: Benchmarking Large Language Models on Italian In: of the Final Workshop 7 December 2016, Naples, p

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:43:45.523365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T12:43:45.309114Z digest=sha256:cfd0346cabf175c23bdb1f9a578dcd2b4750d315a0a6a796d430a25d093b9e7f

Observation 0ffd585c-4df3-4aa6-a722-bf6d7c10a5df · outbound

This paper cites In: CLiC-it (2023).

Evalita-LLM: Benchmarking Large Language Models on Italian In: CLiC-it (2023)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:43:45.511680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T12:43:45.312306Z digest=sha256:43ae1e65b81499e56a4681e3f4d443cb972d01e90ead2f06f82a76429031b780

Observation fb4d7656-c2c6-4ab2-87e3-db8dfc592642 · outbound

This paper cites In: Proceedings of 40 EV ALITA Workshop, 11th Congress of Italian Association for Artificial Intelli- gence, Reggio Emilia, Italy (2009).

Evalita-LLM: Benchmarking Large Language Models on Italian In: Proceedings of 40 EV ALITA Workshop, 11th Congress of Italian Association for Artificial Intelli- gence, Reggio Emilia, Italy (2009)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:43:45.501695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T12:43:45.315607Z digest=sha256:59b691234bc4e714ba26526bff375e3c9d078a2baea6e8c595a99a8247d19e82

Observation 8adcf9e7-a95f-4127-8b70-cda3a06b2933 · outbound

This paper cites : Nermud at evalita 2023: overview of the named-entities recognition on multi-domain documents task.

Evalita-LLM: Benchmarking Large Language Models on Italian : Nermud at evalita 2023: overview of the named-entities recognition on multi-domain documents task

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:43:45.492061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T12:43:45.318948Z digest=sha256:f29e36b763a1b9a8e19897336454c10c6f7328daff177c05abcac23ef1cc4522

Observation a7881a81-c4b7-485b-8189-222ef4cc21aa · outbound

This paper cites EV ALITA (2023).

Evalita-LLM: Benchmarking Large Language Models on Italian EV ALITA (2023)

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:43:45.481568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T12:43:45.322246Z digest=sha256:ac1b3014e72d3ec9c636b1fa912e4c0a8326dad99108500ce188c14c290bbd81

Observation 8e0ef92c-e636-49a7-b7b6-75a5407318dc · outbound

This paper cites Proceedings of the Seventh Italian Conference on Computational Linguistics CLiC-it 2020 (2020).

Evalita-LLM: Benchmarking Large Language Models on Italian Proceedings of the Seventh Italian Conference on Computational Linguistics CLiC-it 2020 (2020)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:43:45.471342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T12:43:45.326021Z digest=sha256:abef20c64ee52458d24e4e49008d6adce10887ca933038b3cae33019dfac8fc7

Observation 1ddf87fa-6bb8-4879-b433-fb169fd6f789 · outbound

This paper cites Information 13(5) (2022) https://doi.

Evalita-LLM: Benchmarking Large Language Models on Italian Information 13(5) (2022) https://doi

Reference 15

Resolution
verified exact
doi, observed 2026-08-09T12:43:45.391920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T12:43:45.329801Z digest=sha256:b8469756e28f8802911619911262150395accd29c2e51995eece20bdf2c7c5eb

Observation 0666da02-62d6-47eb-a534-844480b1a6ca · outbound

This paper cites Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing.

Evalita-LLM: Benchmarking Large Language Models on Italian Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T12:43:45.333606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:43:45.333606Z digest=sha256:e13288bf2a1d5b875deaed07c9a8ceabbdcc8e98d427e9c830253f550b6d5652

Observation ff26a153-3347-4ae6-9a21-ae095f436840 · outbound

This paper cites In: Toutanova, K., Rumshisky, A., Zettlemoyer, L., Hakkani-Tur, D., Beltagy, I., Bethard, S., Cotterell, R., Chakraborty, T., Zhou, Y.

Evalita-LLM: Benchmarking Large Language Models on Italian In: Toutanova, K., Rumshisky, A., Zettlemoyer, L., Hakkani-Tur, D., Beltagy, I., Bethard, S., Cotterell, R., Chakraborty, T., Zhou, Y

Reference 17

Resolution
malformed identifier
raw_fallback, observed 2026-08-09T12:43:45.460943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T12:43:45.337075Z digest=sha256:9e4f418eeda900ecbe5c62df5a233c62b9c99365a9bccff38b422eeb7de4ab18

Observation 66da9b07-76a6-4a1d-97b5-e7922a4032a2 · outbound

This paper cites A NotSo Simple Way to Beat Simple Bench.

Evalita-LLM: Benchmarking Large Language Models on Italian A NotSo Simple Way to Beat Simple Bench

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-08-09T12:43:45.432046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T12:43:45.340196Z digest=sha256:d8846cd540276b766dfd90342d985394ee3a085ff6cffbbcc7064c16ed320897

Observation ebf17e8a-803a-46f9-9a86-1f73a7e839db · outbound

This paper cites State of What Art? A Call for Multi-Prompt LLM Evaluation.

Evalita-LLM: Benchmarking Large Language Models on Italian State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T12:43:45.343766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:43:45.343766Z digest=sha256:6823dcdc90985902846337b83fd4565d8e57492c1f1003c1099d6d3afd8b31e8

Observation 348d50de-ceee-43ed-b366-50d91bf6ddea · outbound

This paper cites Efficient multi-prompt evaluation of LLMs.

Evalita-LLM: Benchmarking Large Language Models on Italian Efficient multi-prompt evaluation of LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T12:43:45.347746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:43:45.347746Z digest=sha256:f1410dcfce72954f4aaf83f0c214dc722275a34fc39ebe2f4e49f0ee79d9b488

Observation e476783c-281b-4f14-9f18-dc83110409c6 · outbound

This paper cites In: Al-Onaizan, Y., Bansal, M., Chen, Y.-N.

Evalita-LLM: Benchmarking Large Language Models on Italian In: Al-Onaizan, Y., Bansal, M., Chen, Y.-N

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T12:43:45.351620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:43:45.351620Z digest=sha256:c1cbb9fee229925ab602d4e2cce0bf5f8668a910dc0fa946a5fddd55d270efc0

Observation 3d433e1b-b5f9-4e19-85ee-a92506f5868e · outbound

This paper cites A Survey of Large Language Models.

Evalita-LLM: Benchmarking Large Language Models on Italian A Survey of Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T12:43:45.354815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:43:45.354815Z digest=sha256:60e098d7e7e91c6f2c33cdf81c95181384eb44638d4ceb03a726ba7cabc762db

Pith citing papers

Observation a7adaa20-c0dc-4b4e-aa30-f04429f2e38f · inbound

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs cites this paper.

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs Evalita-LLM: Benchmarking Large Language Models on Italian

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:53.515641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T03:40:04.692279Z digest=sha256:eee792b1a8fe2891254e57f170c31e733b3ecdbcfb4ade72425a85778d3dffa9

Observation c98e34a3-206b-42d1-962d-74163b7c14d3 · inbound

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs cites this paper.

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs Evalita-LLM: Benchmarking Large Language Models on Italian

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:19:52.878692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T08:14:55.858466Z digest=sha256:7457b6f59d1b73556c8f8d1036618ecbd7a81b2235cbb3f122684d011fba0c60

Observation 50242c54-3c37-4d0a-8311-5e6ccd0e23b5 · inbound

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark cites this paper.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Evalita-LLM: Benchmarking Large Language Models on Italian

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.017840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.017840Z digest=sha256:ca8f6fa40d74cdee543c03b5b6a9899487c5725cd890f7847cd6dca0f16cac56