Pith. sign in

Paper Citation Record · LEDGER

ORCA-bench: How Ready Are Language Model Agents for Oncall?

As of 9 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2607.28545.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.28545 v2

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:40:33.122717Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved6
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a67d0831-c5a3-4c43-8f87-13ac6660e72f · outbound

This paper cites Long Code Arena: a Set of Benchmarks for Long-Context Code Models.

ORCA-bench: How Ready Are Language Model Agents for Oncall? Long Code Arena: a Set of Benchmarks for Long-Context Code Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T04:40:31.765264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:40:31.765264Z digest=sha256:c5fecca2a362e4d0063f47aca956e2853a3c832308a2306c7a6cf8b34a8a887e

Observation b6c465db-c1d9-4df1-92c7-714a303b96f3 · outbound

This paper cites SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios.

ORCA-bench: How Ready Are Language Model Agents for Oncall? SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T04:40:31.853540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:40:31.853540Z digest=sha256:6c2d83cf7440ee20bd3b7fc98bb135ec58bc6e2ccdd08d10f9598f6ab1dc386f

Observation 7c631941-4708-4774-8bf1-f51af23ab502 · outbound

This paper cites Itbench: Evaluating ai agents across diverse real-world it automation tasks.

ORCA-bench: How Ready Are Language Model Agents for Oncall? Itbench: Evaluating ai agents across diverse real-world it automation tasks

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:40:34.585210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T04:40:31.932782Z digest=sha256:b98df94c2a86e62c131345a64d266c23e50020ec8853fb189bab8dcfda68bed3

Observation d4f1571b-17c5-4f18-9ce8-1fb6f9658f37 · outbound

This paper cites Swe-bench: Can language models resolve real-world github issues? In International Conference on Learning Representations, volume 2024, pages 54107--54157, 2024.

ORCA-bench: How Ready Are Language Model Agents for Oncall? Swe-bench: Can language models resolve real-world github issues? In International Conference on Learning Representations, volume 2024, pages 54107--54157, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:40:34.571700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T04:40:32.007547Z digest=sha256:e4612e2866f12c4b1c687d8a467bb5d13376df77b38e805c854e383033e4331e

Observation c1f0203f-d55f-46db-a1e5-0f68d07d964b · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

ORCA-bench: How Ready Are Language Model Agents for Oncall? Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T04:40:32.178531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:40:32.178531Z digest=sha256:baf8dd89a48409701cfa3164213cc05b17225f62799303b17dc36be5f463b0c7

Observation 3caebaa7-a09c-4176-83a1-95f96065a7ce · outbound

This paper cites Rcaeval: A benchmark for root cause analysis of microservice systems with telemetry data.

ORCA-bench: How Ready Are Language Model Agents for Oncall? Rcaeval: A benchmark for root cause analysis of microservice systems with telemetry data

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:40:34.556265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T04:40:32.274533Z digest=sha256:dacb02944839cefe3c7531cc453bfe2864273a5678597f70ba64bbae31d27db9

Observation 703cdbae-f913-497b-85eb-0dc67a3bfe5e · outbound

This paper cites Building ai agents for autonomous clouds: Challenges and design principles.

ORCA-bench: How Ready Are Language Model Agents for Oncall? Building ai agents for autonomous clouds: Challenges and design principles

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:40:34.412446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T04:40:32.340183Z digest=sha256:a6971966ddec4b09e6aa802417b04c5d475c836e47259f35fc5272cc455f9e96

Observation 5b613cb5-f74a-422f-abcf-d769fe4da5d5 · outbound

This paper cites Openrca: Can large language models locate the root cause of software failures? In The Thirteenth International Conference on Learning Representations, 2025.

ORCA-bench: How Ready Are Language Model Agents for Oncall? Openrca: Can large language models locate the root cause of software failures? In The Thirteenth International Conference on Learning Representations, 2025

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:40:34.158418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T04:40:32.442926Z digest=sha256:a32bec10934672c6376551399be83f9e701c1145ef0b2aff7ee86f85c70ff77e

Observation 29c43866-cf73-42dc-badd-965ef8c563a2 · outbound

This paper cites Swe-bench multimodal: Do ai systems generalize to visual software domains? In The Thirteenth International Conference on Learning Representations, 2025.

ORCA-bench: How Ready Are Language Model Agents for Oncall? Swe-bench multimodal: Do ai systems generalize to visual software domains? In The Thirteenth International Conference on Learning Representations, 2025

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:40:33.982615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T04:40:32.524970Z digest=sha256:0f5af0b0f1a3053bebe20aec6c6fe602c74a6ac3f05d13ad712611c6fcd851fb

Observation 47f386de-8d47-4459-919b-1bd9a2799903 · outbound

This paper cites Swe-smith: Scaling data for software engineering agents.

ORCA-bench: How Ready Are Language Model Agents for Oncall? Swe-smith: Scaling data for software engineering agents

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:40:33.716878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T04:40:32.658578Z digest=sha256:4bdd90a4f8dcb7ba6cdad27acb4fd8d5946720bd7d37697626c80a16c5fc5247

Observation 0bdee2f2-9d83-4c48-81fd-52effa5b812e · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

ORCA-bench: How Ready Are Language Model Agents for Oncall? Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T04:40:32.768418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:40:32.768418Z digest=sha256:1b192d44f3d587c4b5fd9a2e089fdb9b90f661d60a33f78eec8227e8ffa42b7a

Observation b1b018bd-c6e6-4698-8063-b7a62fff3da5 · outbound

This paper cites Graders should cheat: privileged information enables expert-level automated evaluations.

ORCA-bench: How Ready Are Language Model Agents for Oncall? Graders should cheat: privileged information enables expert-level automated evaluations

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:40:33.507625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T04:40:32.836167Z digest=sha256:efd5ae6c4d8e52e04c6b6275ed09d7fc4a6029cf53c2679c31c453b654fa3949

Observation f595175f-81f7-4df1-b643-eb90207a3f55 · outbound

This paper cites @esa (Ref.

ORCA-bench: How Ready Are Language Model Agents for Oncall? @esa (Ref

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T04:40:32.940042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:40:32.940042Z digest=sha256:d83fcf0f02ab6f1dd58f4907dc997611c26abb34636593dcdba2addca51fe3bd

Observation 136b50f1-8939-4830-8fd4-305350631b9a · outbound

This paper cites an unresolved cited work.

ORCA-bench: How Ready Are Language Model Agents for Oncall? Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T04:40:33.014494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:40:33.014494Z digest=sha256:4a11cbc33c52fe56b829f4d45e23fd13b86cf66248dd7c0818f56148110b9c36

Observation 2397f58b-10e4-4d29-b909-fa110d9a7746 · outbound

This paper cites an unresolved cited work.

ORCA-bench: How Ready Are Language Model Agents for Oncall? Unresolved cited work

Reference 15

Resolution
malformed identifier
no resolver link, observed 2026-08-06T04:40:33.122717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:40:33.122717Z digest=sha256:2fd94057e46d9ec1cc9c52f206ef78024f03440d5be5aed59e21ab9687a2ee3d

Pith citing papers

No inbound Pith citation observations are available.