Pith. sign in

Paper Citation Record · LEDGER

Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2508.05464.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.05464 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:16:23.976746Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T09:57:00.715643Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 86f057f0-2625-46ef-9d51-c8308039786c · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:08.084480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T05:50:28.114140Z digest=sha256:22de1fa966853336a2d821b6696faf9db686deb83665c234e3b22d11bfa9d690

Observation bfd147ea-3651-43f1-9d4e-4fae56c0fcd8 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:06:42.972093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-25T06:05:27.736494Z digest=sha256:584f323c65fefde1a793b14cc657dd3c704342c924693d5d3e33c94bd895a272

Observation b74665d4-a5f9-4569-b87d-52766c75e6b5 · inbound

Reverse Engineering Compliance: A Dual-Graph Verification Framework for Auditing Legacy IT Security Concepts cites this paper.

Reverse Engineering Compliance: A Dual-Graph Verification Framework for Auditing Legacy IT Security Concepts Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:57:00.717055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T09:55:40.421587Z digest=sha256:1264c5ed4ce7214afadb3e3f2a85616ce0bc25021bdb9e9be8302983c37313ac

Observation 31d06651-36fd-46e0-9cfc-da8a565f789e · inbound

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation cites this paper.

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T04:18:49.064486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:18:49.064486Z digest=sha256:1582b45053e6835b40437747e9c187aaba3f7a54a39483242e55c9638e59f332

Observation 026fb2b0-78e5-4b3a-913f-e87df15f1e1c · inbound

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation cites this paper.

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T04:16:23.976746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:16:23.976746Z digest=sha256:4150bb748917fee4dcdd8909d563c483e61f92a7c89dc6774a690256820f9419