Pith. sign in

Paper Citation Record · LEDGER

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks?

As of 18 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 0 inbound Pith citation observations for arXiv:2507.20419.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20419 v1

Coverage vector

measured 94 of 94 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:38:01.264257Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

94 of 94 outbound references displayed

  • verified exact20
  • verified fuzzy25
  • unresolved47
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 689c45d2-87cc-4baa-a256-e26be0f53ed6 · outbound

This paper cites Principles of Evaluation in Natural Language Processing,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Principles of Evaluation in Natural Language Processing,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:51.994781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:51.994781Z digest=sha256:7e4f0bb1b13786a65145f8aec148f4319ba0b7ea1e34f30a5263a716a83a9738

Observation 7680c4c5-ab05-40de-bfb9-5b3bdd50f0df · outbound

This paper cites Natural Language Inference ,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Natural Language Inference ,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:52.365652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:52.365652Z digest=sha256:9f6f824398e4667b7763866667697a82364a278ef83211d996fe0348f4f46805

Observation 1259cf48-029a-46e1-a7d5-840d469b0b86 · outbound

This paper cites PROBABILISTIC TEXTUAL ENTAILMENT: GENERIC APPLIED MODELING OF LANGUAGE VARIABILITY,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? PROBABILISTIC TEXTUAL ENTAILMENT: GENERIC APPLIED MODELING OF LANGUAGE VARIABILITY,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:52.487079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:52.487079Z digest=sha256:23f62ece3ce6c4ee100905cdde2f0665596fecf330604e38704086728c1af3d9

Observation 927f39e9-f3cc-454e-bf7d-d4a33ded723b · outbound

This paper cites Probabilistic textual entailment: Generic applied modeling of language variability,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Probabilistic textual entailment: Generic applied modeling of language variability,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:52.725915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:52.725915Z digest=sha256:efee579a7d0b2c87a43917165436e4fe4acc70b656679c7d30ecaf0cc6e4ea4e

Observation f96c17a0-920b-4d4d-8f46-579190916b18 · outbound

This paper cites The Seventh PASCAL Recognizing Textual Entailment Challenge,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? The Seventh PASCAL Recognizing Textual Entailment Challenge,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:52.872893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:52.872893Z digest=sha256:7e038731a5908c308e8ad2aebbd3ded0487c40d4bb2fb4715eb8086623adf4b5

Observation 55544bb1-66cf-4be4-b593-6d34d9c0c05d · outbound

This paper cites The Sixth PASCAL Recognizing Textual Entailment Challenge,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? The Sixth PASCAL Recognizing Textual Entailment Challenge,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:53.153624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:53.153624Z digest=sha256:12597ec1f39c05c3d8c1a2bb5a1259ca87b12f2511aed4b84b492487e6820cf8

Observation 7fe02df0-38d3-44dc-8c54-1e59a0b1d7b2 · outbound

This paper cites The Fifth PASCAL Recognizing Textual Entailment Challenge,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? The Fifth PASCAL Recognizing Textual Entailment Challenge,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:53.279800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:53.279800Z digest=sha256:b5833bb97ceebec3a387968c252459287c95e6a7c52f234a2f988b3eac9f4dd9

Observation 9ece044d-c086-4df3-bf6b-0e44cb3cd691 · outbound

This paper cites The Third PASCAL Recognizing Textual Entailment Challenge,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? The Third PASCAL Recognizing Textual Entailment Challenge,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:53.425491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:53.425491Z digest=sha256:5cca4c599d343aa3cfb2de5890f734214d102a5fd76516e15d11256e85dc207d

Observation 92ba844e-0949-47ea-addb-9ab860166dde · outbound

This paper cites The Second PASCAL Recognising Textual Entailment Challenge,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? The Second PASCAL Recognising Textual Entailment Challenge,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:53.601192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:53.601192Z digest=sha256:b2c9a3f8c827ce59b1c6975fdabd7f498733dedaad81677f482cce1833aebcd9

Observation 1b9991c9-4ae9-4ed4-88b2-123b54dcb7e0 · outbound

This paper cites The PASCAL Recognising Textual Entailment Challenge,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? The PASCAL Recognising Textual Entailment Challenge,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:53.769882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:53.769882Z digest=sha256:e21c76fd4f687afe4e343119972306886fe8f5777953071e25d086a402975666

Observation cae6e79b-fc41-44fd-9fbe-c77460cc6ebf · outbound

This paper cites The Fourth PASCAL Recognizing Textual Entailment Challenge,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? The Fourth PASCAL Recognizing Textual Entailment Challenge,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:53.910758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:53.910758Z digest=sha256:c6095c576c1a0f80600009e0e70d15f595fd87cc8fa18d08fcabbb3183dbd0b8

Observation f78a0d00-f5f3-4b3f-b020-b63fb381a23b · outbound

This paper cites Recognizing Textual Entailment: Models and Applications,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Recognizing Textual Entailment: Models and Applications,

Reference 14

Resolution
verified exact
doi, observed 2026-08-06T13:38:04.882356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:54.002707Z digest=sha256:23ca13e0c8ac65af0e2176bd78379e54f980c54e3ee646bc15e4ae51634cb4a9

Observation ebe93d99-08dc-4bbf-b71a-f2a226fd7535 · outbound

This paper cites The Winograd Schema Challenge,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? The Winograd Schema Challenge,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:54.120712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:54.120712Z digest=sha256:341b214e5362fff13dc79534fdb7ad0e39cef4d6bd3da18ab5e4a9ff55545828

Observation 08857552-76f4-423c-8d58-10b41f0d7ba5 · outbound

This paper cites A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:54.187851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:54.187851Z digest=sha256:1803d543c3a4fb5b948d7b5c82b0a2102daae19b82e91d90e07ff08a27021e58

Observation 66995a61-e64d-4e3c-9ff5-05e584fbf5a5 · outbound

This paper cites A large annotated corpus for learning natural language inference,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? A large annotated corpus for learning natural language inference,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:54.254334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:54.254334Z digest=sha256:24fa59c5b95be25ac9c84a1d1f50da8cd6cc14973857ec4000e6042ee3952fc1

Observation 3597f193-4c39-413e-ad2a-facdc16169b5 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:54.336819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:54.336819Z digest=sha256:a4418fad0bc98e5a13c3acd8248926291662ca0ed87cfc16ae0ce6d4745b1ab5

Observation f1d0f05a-eadd-4e8d-b8ea-4d6e4d976f47 · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? PaLM: Scaling Language Modeling with Pathways

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:54.375277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:54.375277Z digest=sha256:37f86f1f498cd7ae6b30a681b28645957bfdaf25b6b0d468df61394d24fd9a0e

Observation 77c1d34e-f5b0-49df-a804-4d570799d029 · outbound

This paper cites Toward Efficient Language Model Pretraining and Downstream Adaptation via Self-Evolution: A Case Study on SuperGLUE.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Toward Efficient Language Model Pretraining and Downstream Adaptation via Self-Evolution: A Case Study on SuperGLUE

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:38:06.437083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:54.473408Z digest=sha256:936defcc15904a51b3e715857e8ce8f294651b46783056536655444ba1712ae0

Observation 3368c16d-2b2d-4928-be27-0ec4dcc1c3e8 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? RoBERTa: A Robustly Optimized BERT Pretraining Approach,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.816504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:54.552526Z digest=sha256:16b246c060c97df91a396644fbc3d60e9b4f692595fa8902fef4014be15fc937

Observation bb834d31-58b1-488c-a015-781ebdc0d131 · outbound

This paper cites Semantics-aware BERT for Language Understanding,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Semantics-aware BERT for Language Understanding,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.803775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:54.618700Z digest=sha256:9882768a7b5c4ecc2f2ec73a7255279764fa94ea0a3893a20144d1b858f35043

Observation 6d1c5b9b-b0dd-4cf2-8f26-47ae3d0b1789 · outbound

This paper cites XLNet: Generalized Autoregressive Pretraining for Language Understanding,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? XLNet: Generalized Autoregressive Pretraining for Language Understanding,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.791260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:54.677424Z digest=sha256:9b38ee3e490f52c8dfd8c0d960a415265d2d4629a23b60387ee62d8f33a57d95

Observation 92347395-5567-476b-b319-f3dd66da57aa · outbound

This paper cites AlexaTM 20B: Few-Shot Learning Using a Large-Scale Multilingual Seq2Seq Model.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? AlexaTM 20B: Few-Shot Learning Using a Large-Scale Multilingual Seq2Seq Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:54.767217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:54.767217Z digest=sha256:97a15cf84cff986a050e9dffdcac960887b9510d8e2dc57716e9d377dea94188

Observation c865e660-91f3-44ca-8312-1150c6491cd4 · outbound

This paper cites BloombergGPT: A Large Language Model for Finance.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? BloombergGPT: A Large Language Model for Finance

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:54.862401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:54.862401Z digest=sha256:2be9563b05c2dd1ea52bee06d44a0b58e6ef0c4f16e721e1ad601b67edaeb214

Observation 3ff88a52-1719-405a-80ac-5075153207c8 · outbound

This paper cites First Train to Generate, then Generate to Train: UnitedSynT5 for Few-Shot NLI.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? First Train to Generate, then Generate to Train: UnitedSynT5 for Few-Shot NLI

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:38:04.472952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:54.934359Z digest=sha256:2d7bbc24de1bfca8767ffa56af1eb14e674ee1dbed0e75489006ffd668fd96fd

Observation a6ec2638-a7ae-49e6-968c-5deaf8a1b8fe · outbound

This paper cites ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:54.981376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:54.981376Z digest=sha256:4cf2f38b0c53ac31e3d7d711bbcc83a09d10eace83a7ace33f9b3e46a327bb0a

Observation 3fa32c45-f2d4-4d46-9df6-b83f036f226a · outbound

This paper cites Rethinking embedding coupling in pre-trained language models.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Rethinking embedding coupling in pre-trained language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:55.076184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:55.076184Z digest=sha256:95e88bcf2dbe6f7070641ef7002d0940affc16c84993cb07eaabb6971081c775

Observation 5b3cc215-56f3-4f27-9ab0-4497e504c3bd · outbound

This paper cites mGPT: Few-Shot Learners Go Multilingual,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? mGPT: Few-Shot Learners Go Multilingual,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:55.158177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:55.158177Z digest=sha256:2c9e80fd7f552515eda490edaa54248c07d90473092c038f78bec062d8569652

Observation c363be4e-caae-41c6-acfb-d19496db7f62 · outbound

This paper cites DeBERTa: Decoding-enhanced BERT with Disentangled Attention,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? DeBERTa: Decoding-enhanced BERT with Disentangled Attention,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.778137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:55.202142Z digest=sha256:e96df5ac4bafb435524aa29d9198fc605ea5ea9fa1de0bf7575fb0e4d9ff9336

Observation c9666e37-80af-4a35-9a5a-f5a4ba0d4a8e · outbound

This paper cites Available: https://aclanthology.org/2007.tal-1.1.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Available: https://aclanthology.org/2007.tal-1.1

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:52.167222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:52.167222Z digest=sha256:0e0cfecfa5f9f4ecaac9a3f5a603bd9e45c7c1b7e305a609125c12ddf2e58be9

Observation 5f610143-a1b9-4151-adf0-e3e4041827f1 · outbound

This paper cites SpanBERT: Improving Pre-training by Representing and Predicting Spans,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? SpanBERT: Improving Pre-training by Representing and Predicting Spans,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:55.271734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:55.271734Z digest=sha256:9c745788a40306ff76f2526b5393477c7efcac782af960bdce88e25a6e72513b

Observation 282848ce-90e3-4530-9594-25e0b0ed6caa · outbound

This paper cites SqueezeBERT: What can computer vision teach NLP about efficient neural networks?,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? SqueezeBERT: What can computer vision teach NLP about efficient neural networks?,

Reference 33

Resolution
verified exact
doi, observed 2026-08-06T13:38:04.069737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:55.343223Z digest=sha256:043985517ce964076902608b5000beb49800ad281435635cfeabc611fb638f1b

Observation 9a9672f9-0a89-4189-b7b4-577db920ad26 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:55.387682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:55.387682Z digest=sha256:886d9e5f92fe42961883fc38463526bde7e194157b7ce8bc0895e216ccbec38e

Observation ee6d3f64-5603-453d-9f90-1a7a75c00cf4 · outbound

This paper cites A Dataset for Arabic Textual Entailment,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? A Dataset for Arabic Textual Entailment,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.765397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:55.479023Z digest=sha256:abc7865a591c18039dc752ff9c9b696a01265c88a8a842e781ce154881e0871c

Observation 3278f0e6-242a-41c4-b4e6-b98a3facfe33 · outbound

This paper cites ARNLI: ARABIC NATURAL LANGUAGE INFERENCE ENTAILMENT AND CONTRADICTION DETECTION,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? ARNLI: ARABIC NATURAL LANGUAGE INFERENCE ENTAILMENT AND CONTRADICTION DETECTION,

Reference 36

Resolution
verified exact
doi, observed 2026-08-06T13:38:03.913885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:55.613730Z digest=sha256:61327d2c24d790303d4318eb55007c01e07725d98e16b090295a87600af89ef6

Observation bd9dd812-c1de-45af-9eb6-6662e3eb03f9 · outbound

This paper cites ArEntail: manually-curated Arabic natural language inference dataset from news headlines,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? ArEntail: manually-curated Arabic natural language inference dataset from news headlines,

Reference 37

Resolution
verified exact
doi, observed 2026-08-06T13:38:03.723052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:55.710306Z digest=sha256:b6d0cdd10c5fb7c7b1bf69cf774eddcf904716a350b95cd969401e87f597b834

Observation 1a673ff0-d2b4-49a6-98f6-a8169e2a117e · outbound

This paper cites Baselines and Test Data for Cross-Lingual Inference,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Baselines and Test Data for Cross-Lingual Inference,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.752969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:55.827796Z digest=sha256:87c6abe02b98671e76fa9314c3b83e419d003387feae58047469c079ebe2be61

Observation 72509850-399e-4755-a1a4-301479c834bc · outbound

This paper cites XNLI: Evaluating Cross-lingual Sentence Representations,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? XNLI: Evaluating Cross-lingual Sentence Representations,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:55.900030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:55.900030Z digest=sha256:e4f09b5e52ef7058234d70c373ea06f991ccd87a4859782599e98b71e2b37a49

Observation 5eef9a04-89dc-4f4f-915a-623b072993da · outbound

This paper cites Benchmarking Zero-shot Text Classification: Datasets, Evaluation and Entailment Approach,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Benchmarking Zero-shot Text Classification: Datasets, Evaluation and Entailment Approach,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:55.976655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:55.976655Z digest=sha256:1afe48b5c457f167f1139f673dace46b472437178d74efa78c76c6a41c66df82

Observation fed28c6d-80b1-4c58-a42d-1a3e195df645 · outbound

This paper cites BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:56.072296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:56.072296Z digest=sha256:e6344bef8eb8a12c4057d1b5263e748c147c2cb7d56ad092dc2d78b234b39e2f

Observation 134ae3ce-000c-4867-b024-63f1140fcd0f · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:56.192293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:56.192293Z digest=sha256:7b26feca860a43ad20e94c45ef89903eed88b6e407b524b88f8c53a680d900b0

Observation 4a4fdf83-36d2-47a1-9a3a-bf5c34d8e745 · outbound

This paper cites MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:56.242266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:56.242266Z digest=sha256:f780668285fbb844c7c6e973c41f3f602c96f2771decf92d891f7f27bd50443c

Observation 6e8521c3-0535-431a-adb0-cf7b87e61d38 · outbound

This paper cites DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:56.318397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:56.318397Z digest=sha256:d2e835f777328c8b419b81e05c22ae7471a48b897d1bdf19d896367c1be7c777

Observation 16c9bc1a-6cad-42d4-b394-b7fbd6ae360d · outbound

This paper cites Less annotating, more classifying: Addressing the data scarcity issue of supervised machine learning with deep transfer learning and BERT-NLI,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Less annotating, more classifying: Addressing the data scarcity issue of supervised machine learning with deep transfer learning and BERT-NLI,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:56.411095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:56.411095Z digest=sha256:3c993cfbd6005ffd423f3000a77b286daf299d511e863eaaa20860c5922f7e23

Observation 667d3618-e2f9-4aad-aed6-1437fbae2362 · outbound

This paper cites GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLP,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLP,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:56.506650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:56.506650Z digest=sha256:eeee1507e6349410bb2d06157df5ba9c24aa1a668fb7298d310bd9e62c8893f3

Observation 40349f99-cd76-426c-8308-65af190c6935 · outbound

This paper cites CLSE: Corpus of Linguistically Significant Entities,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? CLSE: Corpus of Linguistically Significant Entities,

Reference 47

Resolution
verified exact
doi, observed 2026-08-06T13:38:03.524186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:56.576929Z digest=sha256:3e1b71c769b096881e9162357e93d751a7851d259d05bc3eee954114e2efd728

Observation 6780fe5d-533e-4a91-b726-fad1dc3b9773 · outbound

This paper cites The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:56.721456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:56.721456Z digest=sha256:ae0e472da5b96aa5ecb25766aab8bf52978534a217feba0a6b93f2618e364ed7

Observation d9bdb60a-d4bf-4124-a838-90cf873f7543 · outbound

This paper cites GEMv2: Multilingual NLG Benchmarking in a Single Line of Code,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? GEMv2: Multilingual NLG Benchmarking in a Single Line of Code,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:56.833242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:56.833242Z digest=sha256:0f6d747c1992b276ee1c91e680dc655414594b5fe1a4d47d56b593199d19b59f

Observation 8d70fd07-bf64-4b58-9148-ee1014d64153 · outbound

This paper cites IndicNLG Benchmark: Multilingual Datasets for Diverse NLG Tasks in Indic Languages,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? IndicNLG Benchmark: Multilingual Datasets for Diverse NLG Tasks in Indic Languages,

Reference 50

Resolution
verified exact
doi, observed 2026-08-06T13:38:03.343289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:56.920313Z digest=sha256:072a79cc10a6e8b4159f20bd17f6c8b93b8fb020e65cd1d484186089fd1ac099

Observation cefb55ee-4493-4e6d-a886-cfd1123f63fa · outbound

This paper cites MTG: A Benchmark Suite for Multilingual Text Generation,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? MTG: A Benchmark Suite for Multilingual Text Generation,

Reference 51

Resolution
malformed identifier
no resolver link, observed 2026-08-06T13:37:57.001448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:57.001448Z digest=sha256:381dbf6d0502175651c69455f3226dd8a64f4498724aa62083d4384a9567d1fd

Observation d9de5286-75d9-4675-8b07-727b8602c4a2 · outbound

This paper cites IndoNLG: Benchmark and Resources for Evaluating Indonesian Natural Language Generation,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? IndoNLG: Benchmark and Resources for Evaluating Indonesian Natural Language Generation,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:57.121190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:57.121190Z digest=sha256:08be519930cd33c8c38ce4789735bbcebf40cd09ebbcf5f8cfcedecf61214c11

Observation 5bcbe44a-2669-4ef9-a89d-6eb813094b39 · outbound

This paper cites Dolphin: A Challenging and Diverse Benchmark for Arabic NLG,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Dolphin: A Challenging and Diverse Benchmark for Arabic NLG,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.740001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:57.234294Z digest=sha256:ac15fdfd2a3800fd6869d61837ce6b29244e22e2d34650250d4074f106990e95

Observation 8720e990-6952-456b-85f3-0f4c21cbe7cb · outbound

This paper cites TURJUMAN: A Public Toolkit for Neural Arabic Machine Translation,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? TURJUMAN: A Public Toolkit for Neural Arabic Machine Translation,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.727123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:57.449397Z digest=sha256:4de84a80ce9c5fcfb692e52feb6470bbd21e2537a868f450c2c2b41689668456

Observation ddab1a97-4f81-4959-a22c-a35ea5dffcae · outbound

This paper cites AraBench: Benchmarking Dialectal Arabic-English Machine Translation,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? AraBench: Benchmarking Dialectal Arabic-English Machine Translation,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:57.376864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:57.376864Z digest=sha256:8dc87c550f48c94139bb207e95fbcd909ee446022c672e76e88bdb2133eb4865

Observation 71743e22-ff20-481c-9453-e97fd137ca06 · outbound

This paper cites BanglaNLG and BanglaT5: Benchmarks and Resources for Evaluating Low-Resource Natural Language Generation in Bangla,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? BanglaNLG and BanglaT5: Benchmarks and Resources for Evaluating Low-Resource Natural Language Generation in Bangla,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:57.717646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:57.717646Z digest=sha256:a06f9fa67c192a93f90d4ddf1875cb7d24fc694bfe1833e97f28c9c84169948d

Observation e9f11951-f937-49e1-9aee-a16f37f9f3d7 · outbound

This paper cites AraT5: Text-to-Text Transformers for Arabic Language Generation,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? AraT5: Text-to-Text Transformers for Arabic Language Generation,

Reference 57

Resolution
verified exact
doi, observed 2026-08-06T13:38:03.156193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:57.586733Z digest=sha256:6e387d7a24601f273086c464c1bbb4c219e1c2543af37c6c689103932eafd18b

Observation dfb6a474-923b-401a-b713-bd14cc433a19 · outbound

This paper cites CUGE: A Chinese Language Understanding and Generation Evaluation Benchmark.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? CUGE: A Chinese Language Understanding and Generation Evaluation Benchmark

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:38:05.964795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:57.924811Z digest=sha256:adbd72212670c2f99468ad584f1d515d7f7118aa4f3e32e1454813c30d642302

Observation 697a2908-5568-4ec5-aa20-637db9fab701 · outbound

This paper cites PhoMT: A High-Quality and Large-Scale Benchmark Dataset for Vietnamese-English Machine Translation,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? PhoMT: A High-Quality and Large-Scale Benchmark Dataset for Vietnamese-English Machine Translation,

Reference 59

Resolution
verified exact
doi, observed 2026-08-06T13:38:02.957624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:57.828392Z digest=sha256:82df5c04ed1953308230c204ce5d2f0b821ac5087cf7faabdd114c41426ee26b

Observation 6026dce6-d303-4a06-95d5-120a540ddcfb · outbound

This paper cites Benchmarking Multidomain English-Indonesian Machine Translation,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Benchmarking Multidomain English-Indonesian Machine Translation,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.714116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:58.142558Z digest=sha256:32a886ac5562b10623a760252294fa6bce20d9fcaa249bcaa70bbaddf549c1d3

Observation 733fc5c9-80fc-4805-99d5-8ee6d3a4d02c · outbound

This paper cites LOT: A Story-Centric Benchmark for Evaluating Chinese Long Text Understanding and Generation,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? LOT: A Story-Centric Benchmark for Evaluating Chinese Long Text Understanding and Generation,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:57.998337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:57.998337Z digest=sha256:bfe7d91b4f52609768a4f7bfb21d5f3aa15e3f1febbe1f9d4f5d0a5fd9dad018

Observation ec315073-34a0-4046-908d-7d5c555da901 · outbound

This paper cites XGLUE: A New Benchmark Dataset for Cross-lingual Pre-training, Understanding and Generation,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? XGLUE: A New Benchmark Dataset for Cross-lingual Pre-training, Understanding and Generation,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.701733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:58.317146Z digest=sha256:ec4c8a9b7c4e79432b745ff3cfe05cf97dff8b7e7f4c038b79dd53fdc8475102

Observation 834cf8cc-1dd7-40e3-a06a-18e24224514f · outbound

This paper cites GLGE: A New General Language Generation Evaluation Benchmark,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? GLGE: A New General Language Generation Evaluation Benchmark,

Reference 63

Resolution
verified exact
doi, observed 2026-08-06T13:38:02.737908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:58.240465Z digest=sha256:cffa1bf32fce8cd4c721a6f1096534e2e477879228324486b3027fe9f44c659f

Observation cd68d2a5-e9de-4fd4-84b3-337170cf595c · outbound

This paper cites XTREME-R: Towards More Challenging and Nuanced Multilingual Evaluation,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? XTREME-R: Towards More Challenging and Nuanced Multilingual Evaluation,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:58.638320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:58.638320Z digest=sha256:3f1d3eca885131910b1b1b2e32fdbafb41b87a7cb332fdf37afac228e8ff00c7

Observation 5d447d3c-7d5e-4330-9a75-7b9517c6d21f · outbound

This paper cites SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.676377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:58.735056Z digest=sha256:f78d0d1266dc030bff69dacbbc18d397967dce6390c5bb004ca2c669b5a44636

Observation b22bb799-41d6-4867-8b87-9c6c7755d77d · outbound

This paper cites XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.689029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:58.530760Z digest=sha256:e9dcb27afa84252c72bf8a4563a13424952c7ea4632a501969af7716c4e5804a

Observation 91ba2886-2bae-4cb1-b74f-6be8ae3ec06f · outbound

This paper cites ARBERT & MARBERT: Deep Bidirectional Transformers for Arabic,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? ARBERT & MARBERT: Deep Bidirectional Transformers for Arabic,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:58.891569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:58.891569Z digest=sha256:c891be8c939052e788e153bb4c2fe5a408d01f456069160216c9974bc9d15386

Observation 6c24f0c1-3e5d-44a9-bc2a-ea910c5188b6 · outbound

This paper cites ORCA: A Challenging Benchmark for Arabic Language Understanding,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? ORCA: A Challenging Benchmark for Arabic Language Understanding,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.651628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:58.994414Z digest=sha256:826d14c656e3e8c9257d191ea262672b8ca818f2e36a506d94cf27760985b679

Observation b32b011a-7ce8-4541-b6bd-ae0eff8c7655 · outbound

This paper cites ALUE: Arabic Language Understanding Evaluation,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? ALUE: Arabic Language Understanding Evaluation,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.663782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:58.815279Z digest=sha256:7fb9dc60c04b14c293856b9c0644c7a5c8b645b81e27c6b9f563d953d29ddf38

Observation aa3195d2-a51d-4dc0-b2af-9a6328055803 · outbound

This paper cites LAraBench: Benchmarking Arabic AI with Large Language Models,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? LAraBench: Benchmarking Arabic AI with Large Language Models,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.638619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:59.258663Z digest=sha256:fba619110413238bd52207496850f883a9d6e714fd7055dd4796a55910ec53b7

Observation d53b659a-8b8c-4b69-b0d3-297b7cee87ca · outbound

This paper cites Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:59.448235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:59.448235Z digest=sha256:9560858256ee8a8c2f45cf393cc22519a1efd5cc46f9d4353214cd11654aa985

Observation a748ce81-d50b-4170-9c70-a9b4a5157d34 · outbound

This paper cites ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic,

Reference 72

Resolution
malformed identifier
no resolver link, observed 2026-08-06T13:37:59.160290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:59.160290Z digest=sha256:853c7c31d6c08aa10ecd0b44e5772b340bec1daa6c5c498a3dfb18d4fa5f3c70

Observation 9877f715-1110-4b79-8c3f-227868f75d03 · outbound

This paper cites KorNLI and KorSTS: New Benchmark Datasets for Korean Natural Language Understanding,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? KorNLI and KorSTS: New Benchmark Datasets for Korean Natural Language Understanding,

Reference 73

Resolution
verified exact
doi, observed 2026-08-06T13:38:02.509856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:59.675564Z digest=sha256:3676754f24e3256188d91c93d849086fd6b5fab6e6f9ceb3e9c421c8d0091d5c

Observation 9a5d9f2e-eccc-4638-a16a-58b6e0d13866 · outbound

This paper cites CLUE: A Chinese Language Understanding Evaluation Benchmark,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? CLUE: A Chinese Language Understanding Evaluation Benchmark,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:59.782498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:59.782498Z digest=sha256:7971b8cf07f509bd5e205f9ee5db33dfc72452f74ba055918f970a2d37f66daf

Observation 3c0eb517-fdff-4864-8015-0a2649c2db02 · outbound

This paper cites JGLUE: Japanese General Language Understanding Evaluation,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? JGLUE: Japanese General Language Understanding Evaluation,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.597294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:59.854991Z digest=sha256:44281ceddaf49ce0ae26ae83fead67587b20f967a75363d27d4fbf935d018a0c

Observation 1a814e55-d8bb-433c-aca1-1e236472a633 · outbound

This paper cites FlauBERT: Unsupervised Language Model Pre-training for French,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? FlauBERT: Unsupervised Language Model Pre-training for French,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.611359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:59.569273Z digest=sha256:941dac3b357b98c1f37a07e8e3cf92ecc8d0ad6d84583ef67bae657f82107ced

Observation 41de1071-0512-47c3-a8fe-401ca88f2706 · outbound

This paper cites KLUE: Korean Language Understanding Evaluation,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? KLUE: Korean Language Understanding Evaluation,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.583677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:00.038452Z digest=sha256:035bda61da78101017c12269e167a289ee94f56e1bb07ecd0d1091d357e2a7da

Observation 1f389aa9-9204-45a9-b2f4-5567479ee156 · outbound

This paper cites UINAUIL: A Unified Benchmark for Italian Natural Language Understanding,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? UINAUIL: A Unified Benchmark for Italian Natural Language Understanding,

Reference 78

Resolution
verified exact
doi, observed 2026-08-06T13:38:02.321727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:00.166860Z digest=sha256:c7f6aef5ed3596082916d37573393546d0f5b7b95affbada45b128cce0273dc9

Observation f7a1f880-dcd6-4806-b4b3-97735c014b95 · outbound

This paper cites SuperGLEBer: German Language Understanding Evaluation Benchmark,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? SuperGLEBer: German Language Understanding Evaluation Benchmark,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:00.244011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:00.244011Z digest=sha256:458efd29971b82dea8f4973f9798190cc6c689a251228d136b47cfdc1d0a4a52

Observation 8d1de291-c078-4bfd-a723-c1c8c9e3acb6 · outbound

This paper cites IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding,

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:59.916339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:59.916339Z digest=sha256:b67b61e04548190a9f0fd229fe639e98659245ccd184aa6eab992d446449cdab

Observation 76900af7-102a-4306-8a0b-67ccdec9f9c5 · outbound

This paper cites IndicNLPSuite: Monolingual Corpora, Evaluation Benchmarks and Pre-trained Multilingual Language Models for Indian Languages,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? IndicNLPSuite: Monolingual Corpora, Evaluation Benchmarks and Pre-trained Multilingual Language Models for Indian Languages,

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.569610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:00.370586Z digest=sha256:2bb00f83586bebd5778a810f2117872163e9f09eff97cf280fb9166cc6e36b33

Observation 1e878b29-8a63-4574-a305-5822fc162d67 · outbound

This paper cites Using the Framework,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Using the Framework,

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.555603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:00.526646Z digest=sha256:22b905341df90aa47582972fa34446129e059773567bd95d89decabbb3035874

Observation ba872268-3e9f-40d9-9560-5bbaf7689c15 · outbound

This paper cites An extended model of natural logic.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? An extended model of natural logic

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.542099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:00.643216Z digest=sha256:55d6ad2a3b713ddda36520c627fb9c560de71f9fc418cdbfd569dbde9f842f90

Observation 200d0496-5d0a-4469-b32f-05bc9bee392d · outbound

This paper cites VLUE: A New Benchmark and Multi-task Knowledge Transfer Learning for Vietnamese Natural Language Understanding,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? VLUE: A New Benchmark and Multi-task Knowledge Transfer Learning for Vietnamese Natural Language Understanding,

Reference 84

Resolution
verified exact
doi, observed 2026-08-06T13:38:02.179538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:00.314390Z digest=sha256:e520e51a3e5fa59a55ab60a15a2e83346acff28bf1966d1225ca8290a5bb11e3

Observation b17cb167-b484-401e-a616-35348e3c564f · outbound

This paper cites Analysis of identifying linguistic phenomena for recognizing inference in text,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Analysis of identifying linguistic phenomena for recognizing inference in text,

Reference 85

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-06T13:38:05.500337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:00.819109Z digest=sha256:29f40f9100c7a6e70f9737d32aa9cb9d52787a194547dbcb9fa1d2f1ca51201b

Observation 2351178e-b1e2-4ac3-aae2-2b68cae1df7e · outbound

This paper cites EQUATE: A Benchmark Evaluation Framework for Quantitative Reasoning in Natural Language Inference,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? EQUATE: A Benchmark Evaluation Framework for Quantitative Reasoning in Natural Language Inference,

Reference 86

Resolution
verified exact
doi, observed 2026-08-06T13:38:01.970328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:00.875892Z digest=sha256:0c869bff2049f2929b93f570537c5b18d4600aec1b1e5a097ee274815c855ffa

Observation 40b6093e-fe1b-4e39-b590-8aa0b747a202 · outbound

This paper cites Evaluation Metrics for Machine Reading Comprehension: Prerequisite Skills and Readability,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Evaluation Metrics for Machine Reading Comprehension: Prerequisite Skills and Readability,

Reference 87

Resolution
verified exact
doi, observed 2026-08-06T13:38:01.777765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:00.932270Z digest=sha256:4fc8c19270e98968ade24fd594fe1b31b3153b6b79f588ee8902699aa4029012

Observation acffebb4-f9ff-4e31-b663-cfa23187988a · outbound

This paper cites NaturalLI: Natural Logic Inference for Common Sense Reasoning,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? NaturalLI: Natural Logic Inference for Common Sense Reasoning,

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.125392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:00.991401Z digest=sha256:1a1735674b034089b4200b0bbb6f8a42aa956794ae4be590f227ecd68b03d8d5

Observation 5dec987b-ab00-4b50-ba09-6af7a885574b · outbound

This paper cites Building Textual Entailment Specialized Data Sets: a Methodology for Isolating Linguistic Phenomena Relevant to Inference.,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Building Textual Entailment Specialized Data Sets: a Methodology for Isolating Linguistic Phenomena Relevant to Inference.,

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.450314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:00.733334Z digest=sha256:20632bdcf535a2fda85b3f1ec293aed8e3e2f5360213506a1003de78593d3397

Observation 1106c391-1b24-4171-b2f6-dfb5fc5e9610 · outbound

This paper cites Not another Negation Benchmark: The NaN-NLI Test Suite for Sub-clausal Negation.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Not another Negation Benchmark: The NaN-NLI Test Suite for Sub-clausal Negation

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:38:01.441104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:01.153777Z digest=sha256:90fbcc76387e3f6390696bf17db649aca9520873e21eb04e141d75433673b75c

Observation c4943d82-5393-44fc-b3cd-d9f8ea6645d3 · outbound

This paper cites Neural Networks and Textual Inference: How did we get here and where do we go now?,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Neural Networks and Textual Inference: How did we get here and where do we go now?,

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:06.760253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:01.264257Z digest=sha256:3874c0ea1fca4230ca4ab62ffe49f42b5571900a6cabc0bc6ccfee27590ef36e

Observation 8fdf345c-55e3-41bd-8aec-bfd2f400b9e4 · outbound

This paper cites On the Evaluation of Semantic Phenomena in Neural Machine Translation Using Natural Language Inference,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? On the Evaluation of Semantic Phenomena in Neural Machine Translation Using Natural Language Inference,

Reference 94

Resolution
verified exact
doi, observed 2026-08-06T13:38:01.605861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:38:01.049741Z digest=sha256:258ce0919a65e39487dfcade5f73346b936e68518bbece2fd88b32ff49a5935f

Observation 40049495-c160-45a7-b31a-d5732925ded0 · outbound

This paper cites Available: https://aclanthology.org/2024.eacl-long.30/.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Available: https://aclanthology.org/2024.eacl-long.30/

Reference 520

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.625044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:37:59.366997Z digest=sha256:8e6d5f7bc95b5a5cf4d5b0899716971836c7c8240e97b78201c7cf7f66a9f932

Observation 066998a0-309a-4784-bc5e-efbdddc4b793 · outbound

This paper cites an unresolved cited work.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Unresolved cited work

Reference 1422

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:57.304442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:57.304442Z digest=sha256:c10fa5eb8a7a2fdda763c3801ff9ce1a5aaf8c167b9d8f4fe47faabe9ed435d1

Observation f7cab572-fdec-491a-aa9c-c7017d1b647a · outbound

This paper cites an unresolved cited work.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Unresolved cited work

Reference 4961

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:00.432949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:00.432949Z digest=sha256:dcf7f6143589fafea6f21abcf88bf408ec13178c4ca14e1e053dffa379d30219

Observation b815d477-6ea3-4d30-bf6b-0f182e0c0899 · outbound

This paper cites an unresolved cited work.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Unresolved cited work

Reference 6018

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:58.401646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:58.401646Z digest=sha256:81041cbb0601bd7c96673952afec1617cdbbba3bf3907b871db58052eddb1a4b

Pith citing papers

No inbound Pith citation observations are available.