Pith. sign in

Paper Citation Record · LEDGER

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams

As of 15 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2606.22723.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.22723 v2

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-01T07:04:21.555971Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact3
  • verified fuzzy5
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d6f6dc6d-7fe9-4f01-87ba-2a4a0304bae2 · outbound

This paper cites In: Naldi, M.C., Bianchi, R.A.C.

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams In: Naldi, M.C., Bianchi, R.A.C

Reference 1

Resolution
metadata mismatch
doi, observed 2026-07-01T07:05:27.451600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:835a67662a2b9239f2deaae5a4a6e2e6739a380bff4eb70188a9cad890fd0cfd

Observation 8aaa95cc-5364-4878-9a16-49ab4f223af2 · outbound

This paper cites Educational and Psy- chological Measurement20, 37 – 46 (1960).

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams Educational and Psy- chological Measurement20, 37 – 46 (1960)

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:32:39.189627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:55aeae513e7f4beaf4b47482dbdfb2e750d24393864f6c477f670da5a86f3b3d

Observation 6768d605-eaa4-4f0a-bd77-68dfca62459c · outbound

This paper cites an unresolved cited work.

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-07-06T16:32:39.184125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:67c5bd0b8f1d39dfc9994d963e3e16797ede4afb23871b119fc8fd4101658c9c

Observation cf096ed5-4d03-4dff-84e5-49364cdfeaab · outbound

This paper cites In: International Con- ference on Learning Representations (ICLR) (2021).

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams In: International Con- ference on Learning Representations (ICLR) (2021)

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:32:39.180142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:8e91175c34c2c6a733474c1471b39c1e13fe9a9a2508e8cccbee58a4a0ad2fa9

Observation c50c3c1a-dc09-4438-b367-781a115730b7 · outbound

This paper cites CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation.

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:05:28.220496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:55d8cd562f1a674572e590b446977613095296d4af42f263490cef2d1b2f411b

Observation ab4cfba3-2939-4087-b9f4-adc36828036e · outbound

This paper cites Sabi\’a-4 technical report.

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams Sabi\’a-4 technical report

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T07:05:28.233471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:98f8eb49a0ef08e3ad432f3a27cca36faef5c975d75b1b0c7de7a16cc23887e0

Observation 4633b72e-c1f7-4624-a852-2eaee7e97732 · outbound

This paper cites Biometrics pp.

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams Biometrics pp

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:32:39.182134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:60afe3164d48350ff075e9aa814cc340a6b2c1d30cfbea9e8d9feee3a7993ad8

Observation 903c2315-5fbb-4c01-ac25-a02cb6a5e750 · outbound

This paper cites an unresolved cited work.

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-07-06T16:32:39.187877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:d1627432c1f99665216fb43531dcd8888ffdf4cb9ebcaf5ad3af403c91a74dbb

Observation 8cc3dfe2-3c08-4624-9138-c49696b8fe6d · outbound

This paper cites Evaluating GPT-3.5 and GPT-4 Models on Brazilian University Admission Exams.

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams Evaluating GPT-3.5 and GPT-4 Models on Brazilian University Admission Exams

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:05:28.224466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:17bbc2874eb7b70df989ba0cc60af9144a1c4dd5a9e7f77f1c4e8b03ec0688de

Observation 86f13892-1268-4f79-8134-52414b1ea590 · outbound

This paper cites Automatic Legal Writing Evaluation of LLMs.

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams Automatic Legal Writing Evaluation of LLMs

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:05:28.238188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:5d0b854958b0c0b6556c4c989a7d63ebe7fe68bf4f9a52f46b403efd83a0a6d7

Observation ae62caed-c618-49ef-8860-2ab6a44fa524 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T07:05:28.228708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:31eb539889d520c836b0993b89654f3ce006cfc4774064da6b31433f7ca08afb

Observation 61be3dab-fabc-48f8-8dba-bd09c9f92fbd · outbound

This paper cites In: Proceedings of ENIAC (2025).

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams In: Proceedings of ENIAC (2025)

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:32:39.186002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:8638472718b266a1602ce636e5f7687e4ff1db7982b48d48467e23aaa1a18b9d

Observation 12eb00bc-923a-46cb-9586-1717bdcca0e4 · outbound

This paper cites In: Advances in Neural Information Processing Systems (NeurIPS) (2023).

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams In: Advances in Neural Information Processing Systems (NeurIPS) (2023)

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:32:39.178157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:a29947b686cc732e1e37e19318a7d256788dbc3e2008c559edb08fc481035520

Observation bfa5fb53-4e67-4b36-8283-784db1cf8bab · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T07:05:28.242446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:775807cf939e1ce002f3c6eac20dde327a9adc12bc25f7734a10e8d3bad6896e

Pith citing papers

No inbound Pith citation observations are available.