Pith. sign in

Paper Citation Record · LEDGER

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation

As of 21 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2506.01920.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01920 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:34:51.004940Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact3
  • verified fuzzy2
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8f7d8dbf-4588-4510-aae3-9d8b16ce6974 · outbound

This paper cites ArabicaQA: A Comprehensive Dataset for Arabic Question Answering.

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation ArabicaQA: A Comprehensive Dataset for Arabic Question Answering

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:34:51.384322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T11:34:49.699170Z digest=sha256:f4435cad57e326ba5099d85676a354e9127953e245fe0e7b64b9083dff890f84

Observation d17ebd33-56b1-4c43-bffe-2c2ea598c0ed · outbound

This paper cites AraTrust: An Evaluation of Trustworthiness for LLMs in Arabic.

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation AraTrust: An Evaluation of Trustworthiness for LLMs in Arabic

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:34:51.295733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T11:34:49.726917Z digest=sha256:266c71f004fe2e17f57e6236c80ef44fa34b9a804edc18f26adb6401559696c8

Observation 7a85c6bf-8121-49e8-8b79-fc5720f1f99a · outbound

This paper cites an unresolved cited work.

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:34:49.802082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:34:49.802082Z digest=sha256:4037b4fac08197c3a5e8917fde92a61ee8f1a25d159270b8fcc2a69982efc1a2

Observation a5bd6df8-e191-4d1c-9331-48aeba50ff24 · outbound

This paper cites ALLaM: Large Language Models for Arabic and English.

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation ALLaM: Large Language Models for Arabic and English

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:34:49.890709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:34:49.890709Z digest=sha256:424873be68b4149339e3701d24cc1fccc20c8a2a4f4a0e53d52f768dccb1b520

Observation cb319d80-6cc7-4ece-818f-8f616ef391c8 · outbound

This paper cites Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier.

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:34:49.985966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:34:49.985966Z digest=sha256:2614f8fec709295ac258c33683ca0300ccedd26cfe796c8fe70c44669907d89b

Observation 50f7f1dd-fe43-4733-8a89-65cec3f12bb1 · outbound

This paper cites an unresolved cited work.

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation Unresolved cited work

Reference 6

Resolution
verified exact
doi, observed 2026-08-07T11:34:51.098268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T11:34:50.069963Z digest=sha256:4a2bdf54b5be7464d9a96ccbc2fca665c74a6f3cb1d347aa2fda63c7a322fc9e

Observation 0820906a-b442-4c1b-96fa-127cc3db3aa2 · outbound

This paper cites Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation.

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:34:50.154288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:34:50.154288Z digest=sha256:7e63c89b22dd9b56bc20108d43e096b29fa6a1e59a96c0ac97b9842db1c5aca8

Observation 855f6657-0bce-4b7c-9f61-b864dec77795 · outbound

This paper cites an unresolved cited work.

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:34:51.834139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T11:34:50.239113Z digest=sha256:177557c9a0499f5492360014b8577ed1180076f86728f6e38d3fd79f7da34a0c

Observation 2fde341b-6ecf-44b0-93d8-2c963b5141d1 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation Measuring Massive Multitask Language Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:34:50.332940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:34:50.332940Z digest=sha256:3e7d3340b5b8c0a1484bb73a70f3d223e4cdcbd5354dff0ed434ec642fbf4c1c

Observation fd8e2869-78cc-4063-874c-cdfe1402124e · outbound

This paper cites an unresolved cited work.

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:34:50.403505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:34:50.403505Z digest=sha256:2c5c04ecd2bb637410303dbb4fd4844bc9d60f065194ffef497a5e413cda7fe6

Observation ace0a10a-b685-44a5-8d4d-b2259fc7dce9 · outbound

This paper cites an unresolved cited work.

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:34:50.470442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:34:50.470442Z digest=sha256:29119c050e0d6d9030ba2fd66730dd672d3095a2413f54d4820a4fa44abd73b0

Observation 4bc2f5ad-db21-46d8-ab38-cedbfcf43adf · outbound

This paper cites an unresolved cited work.

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:34:51.734852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T11:34:50.511383Z digest=sha256:4b9748970392e5bfcf819b95d62b7bb2d29a99e6835635875a1c1b7774c33818

Observation 11bc0d2f-8310-4f50-906f-2f9c468e9e19 · outbound

This paper cites DeepSeek-V3 Technical Report.

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation DeepSeek-V3 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:34:50.559074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:34:50.559074Z digest=sha256:f0f19142e1a8cdb3120e044fd2586860238b413164e42769921400268b69f2fc

Observation 95ec516b-d1f6-4e6a-8f19-ed54f3a30943 · outbound

This paper cites Arid Hasan, Maram Hasanain, Tameem Kabbani, Fahim Dalvi, Shammur Absar Chowdhury, and Firoj Alam.

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation Arid Hasan, Maram Hasanain, Tameem Kabbani, Fahim Dalvi, Shammur Absar Chowdhury, and Firoj Alam

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:34:51.634114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T11:34:50.596991Z digest=sha256:eca79c7b4519016d0c12d39cd32bbd06d10bf836fcfbb0faa719d8f1a0e86276

Observation 511f2e6c-ad49-4129-b0e6-ca653a0abae8 · outbound

This paper cites AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects.

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:34:50.648503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:34:50.648503Z digest=sha256:d6bc10848a22b1131c6cc6da6e44040ab4f57977f409f6327b54bce746e1b997

Observation c2f1f7bb-01c7-42a2-a458-7e5cc6e8f851 · outbound

This paper cites Al-Batati, Arwa Alsehibani, Nour Qandos, Omar Elshehy, Mohamed Abdelkader, and Anis Koubaa.

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation Al-Batati, Arwa Alsehibani, Nour Qandos, Omar Elshehy, Mohamed Abdelkader, and Anis Koubaa

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:34:51.545735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T11:34:50.693248Z digest=sha256:b315f2209ea51eb9d3f327c7947cc5e4a1b6f807a63bc2687ea076fa4cc66469

Observation 09d78c35-2ec7-4c0c-aa8f-3d143facbab3 · outbound

This paper cites an unresolved cited work.

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:34:51.471881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T11:34:50.752337Z digest=sha256:aab9e970c7a26d6a1e25bca503e93f871219125ddc9f6fd22591f657af61dfe6

Observation f0f6db6a-3284-4eb0-a59d-da0585b39560 · outbound

This paper cites INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge.

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:34:50.802917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:34:50.802917Z digest=sha256:1ee173f03feb5d7f2f4c3945a32fccbb7d8710ad7f1218440c4f718f52abc462

Observation 037462d2-db88-4024-9a5b-01b34de86ba4 · outbound

This paper cites Jais and Jais-chat: Arabic-Centric Foundation and Instruction-Tuned Open Generative Large Language Models.

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation Jais and Jais-chat: Arabic-Centric Foundation and Instruction-Tuned Open Generative Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:34:50.844752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:34:50.844752Z digest=sha256:bb5be4f04f9d985d3993ab5b0423648909f6efd27c2a57bf1d8d3d1075a06724

Observation 5a4dc060-92f1-4ac6-b295-377b85b3c08d · outbound

This paper cites Fanar: An Arabic-Centric Multimodal Generative AI Platform.

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation Fanar: An Arabic-Centric Multimodal Generative AI Platform

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:34:50.894175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:34:50.894175Z digest=sha256:f2f9d0f93d72948b9410111fdda3a82a8a5bbb9903436d1b6e04cb3cc8bd4113

Observation 15872eeb-4ca4-4493-98b8-91bfb477cc0c · outbound

This paper cites The Bitter Lesson Learned from 2,000+ Multilingual Benchmarks.

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation The Bitter Lesson Learned from 2,000+ Multilingual Benchmarks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:34:50.933392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:34:50.933392Z digest=sha256:9d0efe95e376aeed5f5b9e7d07d446b6c8a975d8f15e1dd13bdb1926ad528a18

Observation d0ff7317-27c3-4750-a1aa-6784c20fc1b7 · outbound

This paper cites online" 'onlinestring :=.

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation online" 'onlinestring :=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:34:50.960111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:34:50.960111Z digest=sha256:8c08361063982e2144083a5f6cc8718583ec36e700cd2a6158dae56a831898a6

Observation 0aeb4db4-44d9-4a75-a9af-5b03c0ab2649 · outbound

This paper cites write newline.

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation write newline

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:34:51.004940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:34:51.004940Z digest=sha256:c5a087a7fc31b22e6341897239b1e90d8b2302d7bfd362793ede51c091dc48cb

Pith citing papers

No inbound Pith citation observations are available.