Pith. sign in

Paper Citation Record · LEDGER

OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2508.09124.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.09124 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T05:13:41.995563Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T01:37:42.801023Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 65d346f4-7218-450b-a8d3-1038b655014f · inbound

Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows cites this paper.

Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:48:38.315116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T22:43:48.618334Z digest=sha256:f6dfedf8431f0c3b7f0a94ee61d662a6757179b0935ef94bfa44346dc7f5b9e0

Observation 5072d6c0-9b4a-4c27-8feb-cd934da8508e · inbound

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions cites this paper.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:12.299526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:12.299526Z digest=sha256:80c3b912516dcdfb63e5aefba23f80adcb1a07ee76dff836fe53b5ec071f94ae

Observation c62324e5-15e4-4783-944e-7ea13ed7d6c7 · inbound

RetailBench: Evaluating Long-Horizon Autonomous Decision-Making and Strategy Stability of LLM Agents in Realistic Retail Environments cites this paper.

RetailBench: Evaluating Long-Horizon Autonomous Decision-Making and Strategy Stability of LLM Agents in Realistic Retail Environments OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T23:45:24.809791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:45:24.809791Z digest=sha256:d1eb29b246d31452aa013defc2ee25f6eb100f00e361bff5022331ec228c445a

Observation 24c38b31-9cb7-4baa-9c58-338a028a99f0 · inbound

Opal: Private Memory for Personal AI cites this paper.

Opal: Private Memory for Personal AI OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 239

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:13:16.664821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T21:09:06.320543Z digest=sha256:54a62c3fff32971d4d0868ec2265693dcf13a8ee1e612ec082bb2985387063d4

Observation 6b0e752a-c0bd-4e4b-ae6e-4ebeb9b0e139 · inbound

GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows cites this paper.

GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:43:01.095541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:42:57.035226Z digest=sha256:2b59f03af5a503e2f5764185b59402e7a275ec5fa4179b9538e7945a2f50b8b3

Observation 1bea0ef8-cb73-4380-9fa3-c3f46052a504 · inbound

CUJBench: Benchmarking LLM-Agent on Cross-Modal Failure Diagnosis from Browser to Backend cites this paper.

CUJBench: Benchmarking LLM-Agent on Cross-Modal Failure Diagnosis from Browser to Backend OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:51:13.490188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T07:47:51.195882Z digest=sha256:26a8d3051a972855d344552b8a929ff80f4884e2ee0028846e1c0d67044bafbb

Observation 4fa2cec6-8f8b-4a61-b35c-6e1a2595bf5d · inbound

Tools as Continuous Flow for Evolving Agentic Reasoning cites this paper.

Tools as Continuous Flow for Evolving Agentic Reasoning OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:45:51.700933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:30:19.859374Z digest=sha256:084b9043ae799d91b3a4dfd3b634c8bb7c035fc285b79760fb4427cec7d7d4ca

Observation 4d5e23ff-a394-4099-9135-f654742b73de · inbound

WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation cites this paper.

WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:23.957989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:40:00.725327Z digest=sha256:556fd2ddb5b8a3378898668ce3d2a7ef709336ba2909f67124967439a87961e1

Observation f5a64f86-eeb9-4c68-ba41-61067e440367 · inbound

EnergyAgentBench: Benchmarking LLM Agents on Live Energy Infrastructure Data cites this paper.

EnergyAgentBench: Benchmarking LLM Agents on Live Energy Infrastructure Data OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T17:47:42.076382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T17:43:11.038874Z digest=sha256:57395c1cdff4a6600fc0d7b9fc942090c7b47878533979706068d3e4cdf19f70

Observation 6e0420ee-eb3c-4c33-8f04-924fbe58b065 · inbound

EnergyAgentBench: Benchmarking LLM Agents on Live Energy Infrastructure Data cites this paper.

EnergyAgentBench: Benchmarking LLM Agents on Live Energy Infrastructure Data OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T05:13:41.995563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:13:41.995563Z digest=sha256:84df6fb1f0e5ae92684aeb686598af398c7ffaf12c12cff8bea85cf16257baee

Observation 4d295bdd-2f3e-46ed-8708-8acc3d463185 · inbound

SentinelBench: A Benchmark for Long-Running Monitoring Agents cites this paper.

SentinelBench: A Benchmark for Long-Running Monitoring Agents OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T08:16:48.301325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T06:09:38.698353Z digest=sha256:004dccbec720ed65f9d6669d1b4416f27db1ef445c5af440a0f386c6919cbf47

Observation 2fb866c7-3830-4bb4-b3fb-a9ecb1dd1d5a · inbound

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents cites this paper.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:57:26.489813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:6edf4b7f85354e4a1895ba29d41518abe2bea7db46c83f2ae76fd200587c019d

Observation a4eb0f97-5ceb-46e3-8957-5d90ae170888 · inbound

Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments cites this paper.

Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:08:43.394326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T04:35:38.527140Z digest=sha256:4b270235ff31e2c0830f988a47cfdfccd450aa23ad0c3c1d5dc4454e457921f6

Observation 3ab79eeb-82cd-4ef1-80d5-1a74c4fe99e4 · inbound

CEO-Bench: Can Agents Play the Long Game? cites this paper.

CEO-Bench: Can Agents Play the Long Game? OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 91

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T21:38:58.671790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T00:19:47.396643Z digest=sha256:0191981c89fea9a2541091d78ebced8429c9410b76accb87afcedfa16a7d6ce0

Observation f8e0eeee-4f78-4a5a-94ab-21a3e66e5da5 · inbound

CEO-Bench: Can Agents Play the Long Game? cites this paper.

CEO-Bench: Can Agents Play the Long Game? OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T11:01:20.504958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:01:20.504958Z digest=sha256:a51f76b5da740a8a2e4fc3a43418a205567233b25cb51a040844069f2a70e281

Observation ad13d3d4-f532-4450-96d8-0ce87595a72a · inbound

ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld Tasks cites this paper.

ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld Tasks OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:49:38.855626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T14:05:28.423719Z digest=sha256:1aceaccedf72814062c98ef04cbc8d56d1a9775e99b0e602b37f65de755323a7

Observation 561a7d5e-68da-4aaa-9250-13b69716b9d0 · inbound

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? cites this paper.

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:29:56.687129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T01:34:10.103638Z digest=sha256:b65392e2e57fee4527027ab9b0929e7d2be462626fa715dec897f1cbcaa5d603

Observation 0adcedff-cf5d-4e65-ac83-237afa52fb7e · inbound

Office Comprehension Benchmark cites this paper.

Office Comprehension Benchmark OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:29:15.017566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-04T00:20:42.974208Z digest=sha256:a4ba2f314db52c43f9880574b5d5f635b06cc7f1f7c595f14b060d04c6f02b3f

Observation 662aa882-8fc7-40db-979c-6a923384bff0 · inbound

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents cites this paper.

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-08T20:05:34.092723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-08T20:03:47.212663Z digest=sha256:f3d103d7efd7016c2dab0f15373529067826f5393406ec0db9ba75987e0808b8

Observation 3d6e37f4-3f48-4ee5-9232-60d5934d89f5 · inbound

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents cites this paper.

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:37:42.831140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-11T01:37:23.646537Z digest=sha256:79048ea20d90225f8268360b9f9875bcc5b91c1a121f16f8563363bfb06e6a86

Observation 35539d15-0a15-41c9-9662-f7ecf4bb764c · inbound

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding cites this paper.

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-30T10:50:09.643083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T10:50:09.643083Z digest=sha256:066f791caa00c19b4487b596f18b8fe61a9a77b5786051c0485f7ff3d64bf8d2