Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:50:25.827128Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 4 inbound Pith citation observations for arXiv:2501.00559.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:50:25.827128Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:55:11.855633Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T15:25:49.002359Z
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8f7fa140-edc8-4c4f-b6e8-3a8e43a81bb3 · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Crosslingual generalization through multitask finetuning,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 25c9595b-51ce-4554-a64a-e500223d17df · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Aya dataset: An open-access collection for multilingual in- struction tuning,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4cad6f03-d6fc-4890-ba00-bedd0ce602f6 · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Jais and jais-chat: Arabic-centric foundation and instruction-tuned open generative large language models,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 251e5570-fbc0-489e-82cd-bbc1941ff4b7 · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Mea- suring massive multitask language understand- ing,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 20ebc0df-5dda-4e6d-8263-aad8bdb623b0 · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Hellaswag: Can a machine really finish your sentence?,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9a893492-2b8a-4cf8-ad6b-b0c78a24e1fe · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Winogrande: An adversarial winograd schema challenge at scale,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e90c6536-dfc4-48b7-a21f-c3536e41cdcb · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Judging llm-as-a-judge with mt-bench and chatbot arena,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1e07d139-2196-48e8-bbed-7e26c84602b7 · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Measuring mathematical problem solving with the math dataset,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fa3d8a6c-9d9f-4f0c-b4a1-c51cd64e3d30 · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Challenging big-bench tasks and whether chain-of-thought can solve them,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c2dcbc2f-d836-467a-862a-4f9c72a737dd · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Drop: A read- ing comprehension benchmark requiring discrete reasoning over paragraphs,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0967165d-5df9-44d6-96ba-657c9798e496 · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Mmlu-pro: A more ro- bust and challenging multi-task language under- standing benchmark,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 99d13ad5-ceb7-4bce-85ef-d22f90b7f08c · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Gpqa: A graduate-level google-proof q&a benchmark,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 33fc12c3-00e1-4c7b-b7ad-d491992fd3b9 · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Instruction- following evaluation for large language models,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2ee9c884-b2e6-4fca-b519-c606146024c4 · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Multilingual massive multitask language under- standing (mmmlu)
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 65a3444b-fc16-423f-8363-f85fd0b53da9 · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Arabicmmlu: As- sessing massive multitask language understand- ing in arabic,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1559eae9-5e70-4f95-8b64-ce0214183032 · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects A deep neural network optimized by a genetic al- gorithm to improve arabic sentiment classifica- tion,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b2eea073-3cd3-4ce3-a0e4-868d6ffcad6a · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Weighted entropy corti- cal algorithms for isolated arabic speech recogni- tion,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8f758b26-e894-4052-9d0e-51b2a14a9fba · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Non-diacritized arabic speech recog- nition based on cnn-lstm and attention-based models,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 66fbe260-a5a0-4d33-bb2f-3fbe3cf36c84 · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Qalam : A multimodal llm for arabic optical character and handwriting recognition,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9df2ab79-21ad-46fa-be96-342a6bb9f6c9 · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Towards a deep learning question-answering specialized chatbot for objective structured clin- ical examinations,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b88ee1b3-b643-43ea-b098-bd4569a4bbf1 · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Question dif- ficulty prediction for multiple choice problems in medical exams,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 02facbb7-9309-4129-8b9e-21abc0636d30 · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects The data provenance initiative: A large scale audit of dataset licensing & attribution in ai
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0cd14c0d-e33a-46cd-af10-50ea4cb1a807 · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Multilingual E5 Text Embeddings: A Technical Report
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9cc5540-5084-4d64-b510-b9430ae73578 · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Umap: Uniform manifold approximation and projection for dimension reduction,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a01ead31-dbd3-44f0-a1a2-b4b8a3921d49 · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Chain-of-thought prompting elicits reasoning in large language models,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a36ca7a-f179-4abb-9c38-9028a4cb7a61 · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects AceGPT, localizing large language models in Arabic,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 450d3577-6d69-4c0c-ac5b-5623c7825041 · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects The llama 3 herd of models,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 30543f7c-8846-4bd2-a516-af7f421575fc · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Training compute-optimal large language models,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f688259f-5bb3-4d9e-b146-0d1f81faa891 · outbound
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects On the explain- ability of natural language processing deep mod- els,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ed6f077c-0ebe-4b4b-927f-6ba65d56b2f5 · inbound
MedArabiQ: Benchmarking Large Language Models on Arabic Medical Tasks AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 262647ab-2064-4ced-94fe-dab5a64ec751 · inbound
ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 511f2e6c-ad49-4129-b0e6-ca653a0abae8 · inbound
From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1ff8792-e50f-4d4f-b47d-4e555fa3cdb5 · inbound
3LM: Bridging Arabic, STEM, and Code through Benchmarking AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.