Pith. sign in

Paper Citation Record · LEDGER

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects

As of 21 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 4 inbound Pith citation observations for arXiv:2501.00559.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00559 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:50:25.827128Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:55:11.855633Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T15:25:49.002359Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy26
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8f7fa140-edc8-4c4f-b6e8-3a8e43a81bb3 · outbound

This paper cites Crosslingual generalization through multitask finetuning,.

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Crosslingual generalization through multitask finetuning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:26.394761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:50:25.673674Z digest=sha256:719a6a341573253b00cf1f125e437038fb37e5ba07dc4e5883f193632db1ca25

Observation 25c9595b-51ce-4554-a64a-e500223d17df · outbound

This paper cites Aya dataset: An open-access collection for multilingual in- struction tuning,.

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Aya dataset: An open-access collection for multilingual in- struction tuning,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:26.381517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:50:25.678874Z digest=sha256:f715c533159d895c4f6be6bbeff28024bb32913517dca14b7258db30a3b82807

Observation 4cad6f03-d6fc-4890-ba00-bedd0ce602f6 · outbound

This paper cites Jais and jais-chat: Arabic-centric foundation and instruction-tuned open generative large language models,.

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Jais and jais-chat: Arabic-centric foundation and instruction-tuned open generative large language models,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:26.368550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:50:25.684961Z digest=sha256:03ae0cd0c777ff1a9cd5e2dd3dd0fc16525b200ef4b6bd38c60efab3f324ac29

Observation 251e5570-fbc0-489e-82cd-bbc1941ff4b7 · outbound

This paper cites Mea- suring massive multitask language understand- ing,.

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Mea- suring massive multitask language understand- ing,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:26.351593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:50:25.694090Z digest=sha256:54e94264f7acdd7aa95a258406009224b875528f6da1897dcc8c6a8c0c50b131

Observation 20ebc0df-5dda-4e6d-8263-aad8bdb623b0 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence?,.

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Hellaswag: Can a machine really finish your sentence?,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:26.336484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:50:25.701107Z digest=sha256:2172d3f375fe3eabe6966480d563893e813c54ab1d55092061fcdf5a9911cba5

Observation 9a893492-2b8a-4cf8-ad6b-b0c78a24e1fe · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale,.

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Winogrande: An adversarial winograd schema challenge at scale,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:26.320322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:50:25.708748Z digest=sha256:e1376fd92b69f3853cc0d8ba790441e4e1c1d51aeb90075934df10d2e2a0d1fc

Observation e90c6536-dfc4-48b7-a21f-c3536e41cdcb · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena,.

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Judging llm-as-a-judge with mt-bench and chatbot arena,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:26.297264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:50:25.715834Z digest=sha256:6e169fb1669db3e9f33c41ad5b9daa333d2cf79c679b9ef49a89ad454ef0c4fd

Observation 1e07d139-2196-48e8-bbed-7e26c84602b7 · outbound

This paper cites Measuring mathematical problem solving with the math dataset,.

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Measuring mathematical problem solving with the math dataset,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:26.281207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:50:25.722543Z digest=sha256:3b51688f7490f75e495305991120c19654432a6fae6f9e7ffb0478e81ee3d975

Observation fa3d8a6c-9d9f-4f0c-b4a1-c51cd64e3d30 · outbound

This paper cites Challenging big-bench tasks and whether chain-of-thought can solve them,.

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Challenging big-bench tasks and whether chain-of-thought can solve them,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:26.262895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:50:25.727060Z digest=sha256:70de6a331cc4e1430e1ddacb67c0298027f8c2d53bb4150dd0e370ae540a67d7

Observation c2dcbc2f-d836-467a-862a-4f9c72a737dd · outbound

This paper cites Drop: A read- ing comprehension benchmark requiring discrete reasoning over paragraphs,.

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Drop: A read- ing comprehension benchmark requiring discrete reasoning over paragraphs,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:26.242711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:50:25.732307Z digest=sha256:691777fc32f356789a7c824e8387013889ce32e1607cc026d683f175e8cbc9b3

Observation 0967165d-5df9-44d6-96ba-657c9798e496 · outbound

This paper cites Mmlu-pro: A more ro- bust and challenging multi-task language under- standing benchmark,.

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Mmlu-pro: A more ro- bust and challenging multi-task language under- standing benchmark,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:26.211781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:50:25.737991Z digest=sha256:207833d5f722d74cf12ea149f7bb514f6399bb523307addfec19b605ed36c95b

Observation 99d13ad5-ceb7-4bce-85ef-d22f90b7f08c · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark,.

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Gpqa: A graduate-level google-proof q&a benchmark,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:26.178193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:50:25.743440Z digest=sha256:4bb0ae10b8279759771e1aeab02a6737f68fc49efdd8e965ea830d5f43cc4a3e

Observation 33fc12c3-00e1-4c7b-b7ad-d491992fd3b9 · outbound

This paper cites Instruction- following evaluation for large language models,.

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Instruction- following evaluation for large language models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:26.160518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:50:25.748566Z digest=sha256:ef4325a8228a492f16f773d6f4ceb7bc3f46b78d841ac077d9be97e63304812b

Observation 2ee9c884-b2e6-4fca-b519-c606146024c4 · outbound

This paper cites Multilingual massive multitask language under- standing (mmmlu).

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Multilingual massive multitask language under- standing (mmmlu)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:26.144233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:50:25.752785Z digest=sha256:eed99e98a719fbc427dfc7a5a8c4fb32a86d65c507e3351bd9fcee64ab34ceef

Observation 65a3444b-fc16-423f-8363-f85fd0b53da9 · outbound

This paper cites Arabicmmlu: As- sessing massive multitask language understand- ing in arabic,.

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Arabicmmlu: As- sessing massive multitask language understand- ing in arabic,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:26.129803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:50:25.763735Z digest=sha256:4fb1f3d311d5683bf08bec8e12d30810b8c8260e6c93cd585315a3aa30a42f0c

Observation 1559eae9-5e70-4f95-8b64-ce0214183032 · outbound

This paper cites A deep neural network optimized by a genetic al- gorithm to improve arabic sentiment classifica- tion,.

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects A deep neural network optimized by a genetic al- gorithm to improve arabic sentiment classifica- tion,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:26.113352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:50:25.768548Z digest=sha256:0f2a652633055bb12c7aee8c47a671c474768cf4a77b79d0860dddd5e3402b64

Observation b2eea073-3cd3-4ce3-a0e4-868d6ffcad6a · outbound

This paper cites Weighted entropy corti- cal algorithms for isolated arabic speech recogni- tion,.

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Weighted entropy corti- cal algorithms for isolated arabic speech recogni- tion,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:26.092491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:50:25.775409Z digest=sha256:0921af6d4d6b03bbf73771bf42671e5995fb6792def955aca95b57eb922552d2

Observation 8f758b26-e894-4052-9d0e-51b2a14a9fba · outbound

This paper cites Non-diacritized arabic speech recog- nition based on cnn-lstm and attention-based models,.

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Non-diacritized arabic speech recog- nition based on cnn-lstm and attention-based models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:26.074893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:50:25.780443Z digest=sha256:8765cb8fa1a2cac57b3f66bc1c881c87c069a544a76c7cf7157b613b501eb31b

Observation 66fbe260-a5a0-4d33-bb2f-3fbe3cf36c84 · outbound

This paper cites Qalam : A multimodal llm for arabic optical character and handwriting recognition,.

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Qalam : A multimodal llm for arabic optical character and handwriting recognition,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:26.050761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:50:25.784802Z digest=sha256:4c1d3684ebbabf9bcbd34fcfc5972e902c4f351b5e5f2607884bf7241a58b68a

Observation 9df2ab79-21ad-46fa-be96-342a6bb9f6c9 · outbound

This paper cites Towards a deep learning question-answering specialized chatbot for objective structured clin- ical examinations,.

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Towards a deep learning question-answering specialized chatbot for objective structured clin- ical examinations,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:26.031709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:50:25.789605Z digest=sha256:c5002ace1913987df3bd2dae7c5d6ba2dbf8fbdc1cd2faf1b503fc22ec4ba70a

Observation b88ee1b3-b643-43ea-b098-bd4569a4bbf1 · outbound

This paper cites Question dif- ficulty prediction for multiple choice problems in medical exams,.

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Question dif- ficulty prediction for multiple choice problems in medical exams,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:26.010151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:50:25.793902Z digest=sha256:7956730bae11d209d1b43b5859c45ff9dbff1dc7a221f8a238d56db6c42123b6

Observation 02facbb7-9309-4129-8b9e-21abc0636d30 · outbound

This paper cites The data provenance initiative: A large scale audit of dataset licensing & attribution in ai.

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects The data provenance initiative: A large scale audit of dataset licensing & attribution in ai

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:25.989565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:50:25.797889Z digest=sha256:f4efd37682a094febd31ab0eefd470dce837bd9483976c89fe2761d2379266cc

Observation 0cd14c0d-e33a-46cd-af10-50ea4cb1a807 · outbound

This paper cites Multilingual E5 Text Embeddings: A Technical Report.

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Multilingual E5 Text Embeddings: A Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:25.802278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:25.802278Z digest=sha256:b69d863d6938df88180666af52f16f2bf84adb56660edcaa4f83f7200cc7b741

Observation c9cc5540-5084-4d64-b510-b9430ae73578 · outbound

This paper cites Umap: Uniform manifold approximation and projection for dimension reduction,.

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Umap: Uniform manifold approximation and projection for dimension reduction,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:25.806971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:25.806971Z digest=sha256:acfdd1ef439d8ac4a39ba76f8cc2d78a09f9a5a2783b2a0231b2159d670f377a

Observation a01ead31-dbd3-44f0-a1a2-b4b8a3921d49 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Chain-of-thought prompting elicits reasoning in large language models,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:25.810599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:25.810599Z digest=sha256:4db8d4cefd20c252f4bfbf7ca4eb6b1dfa647e7ceb005fb044615a3b703c4e70

Observation 3a36ca7a-f179-4abb-9c38-9028a4cb7a61 · outbound

This paper cites AceGPT, localizing large language models in Arabic,.

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects AceGPT, localizing large language models in Arabic,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:25.941658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:50:25.814847Z digest=sha256:3945292dc001200d66be4760533b798c9d41fca06356c488044855654f6bdb7d

Observation 450d3577-6d69-4c0c-ac5b-5623c7825041 · outbound

This paper cites The llama 3 herd of models,.

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects The llama 3 herd of models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:25.924738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:50:25.818608Z digest=sha256:161e5530515f4dba9fa968455ee9ffc45354da5c7bdcda9124a7ba9608ad6bbc

Observation 30543f7c-8846-4bd2-a516-af7f421575fc · outbound

This paper cites Training compute-optimal large language models,.

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects Training compute-optimal large language models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:25.904875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:50:25.822710Z digest=sha256:c7834a7d2eba4a7172e3bf4c6818fc86e2a080382956f1ca4f737d24f5579188

Observation f688259f-5bb3-4d9e-b146-0d1f81faa891 · outbound

This paper cites On the explain- ability of natural language processing deep mod- els,.

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects On the explain- ability of natural language processing deep mod- els,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:25.885704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:50:25.827128Z digest=sha256:eb97a22a9cd2e86e30129ceb762148dcdb54c8ab59b97dee0b63dd5fecbd21f3

Pith citing papers

Observation ed6f077c-0ebe-4b4b-927f-6ba65d56b2f5 · inbound

MedArabiQ: Benchmarking Large Language Models on Arabic Medical Tasks cites this paper.

MedArabiQ: Benchmarking Large Language Models on Arabic Medical Tasks AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T23:55:11.855633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:55:11.855633Z digest=sha256:010e8ee39ebf565d8f6f88a9f9d1f60e7c0033d808a741ed41b031c2ec40a7c2

Observation 262647ab-2064-4ced-94fe-dab5a64ec751 · inbound

ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark cites this paper.

ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:25.382481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:25.382481Z digest=sha256:54ef44c822c1ad63f7696fd9ea147addbf2ed25b760b64653c24002409ac2ac7

Observation 511f2e6c-ad49-4129-b0e6-ca653a0abae8 · inbound

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation cites this paper.

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:34:50.648503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:34:50.648503Z digest=sha256:d6bc10848a22b1131c6cc6da6e44040ab4f57977f409f6327b54bce746e1b997

Observation b1ff8792-e50f-4d4f-b47d-4e555fa3cdb5 · inbound

3LM: Bridging Arabic, STEM, and Code through Benchmarking cites this paper.

3LM: Bridging Arabic, STEM, and Code through Benchmarking AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects

Reference 2025

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T15:25:49.007226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:25:48.963935Z digest=sha256:62ee285ed2cc2a284cfdd0ab55cf670a516315cf31dc969af37307e03f8e8876