Pith. sign in

Paper Citation Record · LEDGER

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu

As of 10 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2508.01006.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.01006 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T06:00:10.327109Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 548a52aa-2253-4475-81d1-0927df9c063e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T06:00:10.281920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:00:10.281920Z digest=sha256:e75a43c88b42ca4158c1b8dd247c7f54eaf574720692ae54cac54dd9920ea386

Observation 47537837-aa1d-4b6d-8a6b-7c73a663c655 · outbound

This paper cites MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T06:00:10.296880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:00:10.296880Z digest=sha256:9af1cd1db8f4e0868cc4d410031168ff490b125fea7ef23935ba2242593dfc64

Observation f6ac2c71-d32d-4fac-8f5a-8aa11ec5954a · outbound

This paper cites Controlled Evaluation of Syntactic Knowledge in Multilingual Language Models.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu Controlled Evaluation of Syntactic Knowledge in Multilingual Language Models

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T06:00:10.374786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T06:00:10.300890Z digest=sha256:d7c91848ea06af84e1aa9bcc9bdf07628f7f389128a115e50ea5e6f489846000

Observation 9b21b02a-1d0d-4cfb-b3bb-a103a75aaca7 · outbound

This paper cites In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 175–184, Online and Punta Cana, Dominican Republic.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 175–184, Online and Punta Cana, Dominican Republic

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:00:17.008156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T06:00:10.304908Z digest=sha256:2983900e768f00d3c56ea029ba6afec3bb8936c8ac7fd242022f3e7b47d2769f

Observation 325c05ce-bb73-4f46-91eb-d7f90fe50223 · outbound

This paper cites In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 4606–4634, Abu Dhabi, United Arab Emirates.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 4606–4634, Abu Dhabi, United Arab Emirates

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:00:16.984655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T06:00:10.316058Z digest=sha256:c1248f736608258614bf32cf500a6f189f9690eabe5002147793adfe876336b0

Observation 658c8df3-eb7a-4f0c-90ec-0c39dc780deb · outbound

This paper cites https: //huggingface.co/large-traversaal/ Alif-1.0-8B-Instruct.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu https: //huggingface.co/large-traversaal/ Alif-1.0-8B-Instruct

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:00:16.973213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T06:00:10.319635Z digest=sha256:17f85799b5611c74db90661800ffb20c0ccdf5878807a4aad2d53aae5f207320

Observation d6c2eb2c-9940-4d45-9194-ce4a9ebbffa2 · outbound

This paper cites In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:00:16.962267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T06:00:10.323430Z digest=sha256:59e4afb30f1eb867092532079e0242240d98e0baa11238a3a680342d2e608985

Observation 1a95b647-5a3a-400a-b0b8-921d650a2b73 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T06:00:10.285644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:00:10.285644Z digest=sha256:e38554b40ec9bffff68490805402deada155032aba8a1144cad5412f76f6384a

Observation 833c7619-d4b4-4f6c-8b1c-ef0c6e9dde59 · outbound

This paper cites an unresolved cited work.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu Unresolved cited work

Reference 2020

Resolution
unresolved
raw_fallback, observed 2026-08-06T06:00:16.950764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T06:00:10.327109Z digest=sha256:001106c3525e25b4cf6625f20448d13e54f0d3579242c43161a5eaae20950c81

Observation 903cf07f-df61-470c-a178-bc27745dfb5a · outbound

This paper cites In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7038–7051, Online and Punta Cana, Dominican Republic.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7038–7051, Online and Punta Cana, Dominican Republic

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:00:17.020156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T06:00:10.293383Z digest=sha256:40e265f6ff6895fd6c569eda0a5254af6f7d6ca0262b42c24dc5b0c5751bd322

Observation 62d44409-ffa4-4a70-acb5-b3d8dcc8b068 · outbound

This paper cites Crosslingual Generalization through Multitask Finetuning.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu Crosslingual Generalization through Multitask Finetuning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T06:00:10.308640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:00:10.308640Z digest=sha256:b0b0fba4076b7e5b4f8647afc19169c7ad08edb43685d1730c38c7ee6068d991

Observation 12c83f7b-ce1b-4839-b665-57a31b8abe16 · outbound

This paper cites In Findings of the Association for Computational Linguistics: EACL 2023 , pages 1581–1594, Dubrovnik, Croatia.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu In Findings of the Association for Computational Linguistics: EACL 2023 , pages 1581–1594, Dubrovnik, Croatia

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:00:16.995724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T06:00:10.312630Z digest=sha256:45a327837eb0a2032aa91d9946061d99d6fe16af808ed1f58a535a1ae796ae05

Observation c3f58f95-83ac-4dde-b071-2c9f97a73365 · outbound

This paper cites The Llama 3 Herd of Models.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu The Llama 3 Herd of Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T06:00:10.289493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:00:10.289493Z digest=sha256:2c96f1914c6560570983ca3fa5d26b333f8b16f4860719971405475b3ce82bde

Observation b6b58527-6c82-42f4-836d-554e91e84c54 · outbound

This paper cites Preprint, arXiv:2506.13487.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu Preprint, arXiv:2506.13487

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T06:00:10.277918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:00:10.277918Z digest=sha256:0d874ae0d19502761d44eb0a68b81babff1729ab63924e776dd43522b21b0345

Pith citing papers

No inbound Pith citation observations are available.