Pith. sign in

Paper Citation Record · LEDGER

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions

As of 9 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 5 inbound Pith citation observations for arXiv:2602.05843.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.05843 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:08:14.068241Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:38:40.363598Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T13:29:51.596096Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation abcf9bb2-67e3-4b98-a511-9cdc52c13b92 · outbound

This paper cites write newline.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:08.370940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:08.370940Z digest=sha256:c502b62584a6a88844235f0089e77fcb83f0a7519fd39b4b35d283ffb7c8a812

Observation 07be5ed1-70db-4ffe-b7a9-17b09f240e50 · outbound

This paper cites J., Bethge, M., and Schulz, E.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions J., Bethge, M., and Schulz, E

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:08.453393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:08.453393Z digest=sha256:d317948121e58619be730a30bfed464bcce08c9bf7a2b118b01a879ac12516df

Observation 94800fc7-0706-4bf2-8fb4-90643d4fc806 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions The claude 3 model family: Opus, sonnet, haiku

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:08.572461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:08.572461Z digest=sha256:6c590fc2d4549fcae475e907b00499142ad29c2f51f979e6cafb23c0b5e9ee24

Observation 57ffcf1c-e6a2-4347-b8e3-5696591186d4 · outbound

This paper cites H., and Bengio, Y.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions H., and Bengio, Y

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:08.670628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:08.670628Z digest=sha256:61a1892dcf692bc35e53c591374544b74306efcae3e8be8994e1db32a37217ca

Observation 16f03f94-b2e8-4fb0-b521-7d6d0098e075 · outbound

This paper cites On the Measure of Intelligence.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions On the Measure of Intelligence

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:08.735295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:08.735295Z digest=sha256:9b708fd2e6df4d2ecb496cc2dd00c92d64ccdd8d7b701b25d7e7640dbb9fa1cc

Observation 3fe0ebb9-7175-4866-a10c-225c9e067939 · outbound

This paper cites Evaluating long-context reasoning in llm-based webagents.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Evaluating long-context reasoning in llm-based webagents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:08.828452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:08.828452Z digest=sha256:b9607183cb5500116d23f37daac333e1238aaa161458974d6e64315de17eef98

Observation be7fd5c4-2ec1-406e-a03b-57bd71bd983c · outbound

This paper cites The dynamical challenge.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions The dynamical challenge

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:09.003347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:09.003347Z digest=sha256:08ec947486ea76ec0b4d296c48e6585326ac1acebb0c02d857c9514e36a1f856

Observation 62382c70-fc00-4411-b000-1b900451b63a · outbound

This paper cites Mind2web: Towards a generalist agent for the web.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Mind2web: Towards a generalist agent for the web

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:09.106776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:09.106776Z digest=sha256:80ec525933a25ab685b191b39a9ca827871b1a184ac6ffa50780dc344602f622

Observation 44797934-6322-4048-b2ff-eb7641e983e9 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:09.262712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:09.262712Z digest=sha256:6a72cde0d4349b6367e0f658e866f3ac75acebe8faaed3bbceb54299a3a06340

Observation af6b907a-38c2-4a0f-be18-bc1e33218f2f · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:09.404493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:09.404493Z digest=sha256:b28404812c4a89ff8d592fb846b5beffa5d13f8bab98775ac44e982f4b11398f

Observation 1e10462f-7d50-4565-9820-223e71e219ff · outbound

This paper cites The Llama 3 Herd of Models.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions The Llama 3 Herd of Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:09.547578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:09.547578Z digest=sha256:818530d5f05ff761ce387b288dc1d72efe3b88536877292d5563cdfa8c12d5e9

Observation ed61fcd1-4de9-4f12-b221-08dfed5f1429 · outbound

This paper cites and Schmidhuber, J.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions and Schmidhuber, J

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:09.716497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:09.716497Z digest=sha256:2702e29ebb95280274141a930a13a82e59f17e6d97ad50bed77e1cc498530d02

Observation 9cec0ab3-7747-4553-af99-c627e6a1a99b · outbound

This paper cites H., Gonzalez, J.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions H., Gonzalez, J

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:09.905299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:09.905299Z digest=sha256:f02703be441575a6247742dfb6636fe398fb959cdc26b84fc98616ec73f47ea4

Observation 8ff668df-7ab4-408c-8124-43c317ff134d · outbound

This paper cites M., Ullman, T.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions M., Ullman, T

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:10.115144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:10.115144Z digest=sha256:6d980f3bdf7b87eaf5712317a020b27a3c800c18f884e525b1e9820834c7c0d0

Observation 5ac8cd97-6d31-4f57-9d06-5cf1245b2007 · outbound

This paper cites State space models on temporal graphs: A first-principles study.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions State space models on temporal graphs: A first-principles study

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:10.290501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:10.290501Z digest=sha256:e13090a59a6632ce0174a69c245dd650d5187a28710098555b5fc85241fbd3cb

Observation 2cd6a435-f4ba-42ab-b273-ca7db33800c4 · outbound

This paper cites Y., Le Bras, R., Richardson, K., Sabharwal, A., Poovendran, R., Clark, P., and Choi, Y.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Y., Le Bras, R., Richardson, K., Sabharwal, A., Poovendran, R., Clark, P., and Choi, Y

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:10.419920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:10.419920Z digest=sha256:9d222b23dec070eedd99daf0be759f1e4a4c6651083227b47737cbd7a103d639

Observation 1852803b-9dba-413b-b13f-e8e029d63b47 · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:10.546967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:10.546967Z digest=sha256:d61e48483e13426194d97d96ba04bce1c9d678ceedc21e61140a02862be3a3bf

Observation c604330b-a965-4290-9c52-a8cef996ec79 · outbound

This paper cites Agentbench: Evaluating llms as agents.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Agentbench: Evaluating llms as agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:10.676369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:10.676369Z digest=sha256:f2faf73a53d5fc974e0ef7e7fe52e7f91ce6ff6b83d763e9e2add77ad56d3bf2

Observation dfa35dbd-6872-4218-ac85-b64b4971485e · outbound

This paper cites Gaia: a benchmark for general ai assistants.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Gaia: a benchmark for general ai assistants

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:10.801903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:10.801903Z digest=sha256:420d7cbfc5f8f5ea60e2aa020a42ad5e5ab80919edd566773b6f0efd246d897d

Observation 85a10ad0-493f-4700-91ed-7e7dd255d1c3 · outbound

This paper cites gpt-oss-120b & gpt-oss-20b Model Card.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions gpt-oss-120b & gpt-oss-20b Model Card

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:10.909307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:10.909307Z digest=sha256:04fdc109759c9091419ad6c39a88d58c3105bab5d15e64d7ca890769da14739e

Observation 23bf426b-6ed8-4bef-ac19-2290650ab49b · outbound

This paper cites G., Mao, H., Yan, F., Ji, C.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions G., Mao, H., Yan, F., Ji, C

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:11.042144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:11.042144Z digest=sha256:b365df9c13306e5e361b39a07304158401e9d5460d6e3e5da7614d0570a72a4f

Observation b09a2020-4b4a-47e9-954c-030b1dd6c2b7 · outbound

This paper cites E., Li, W., Campbell-Ajala, F., Toyama, D.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions E., Li, W., Campbell-Ajala, F., Toyama, D

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:11.173177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:11.173177Z digest=sha256:633409d457b040f654d99d1893c0af70ccb15cf3a66d78baa032a33e0c0b0053

Observation 0a45872a-b04e-43aa-ac18-15242df16764 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Reflexion: Language agents with verbal reinforcement learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:11.342942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:11.342942Z digest=sha256:316ba339824e6f7da7a165fbdbfd286da60d3e4c4b3abe54f431b3265c0c146d

Observation a6e9b9cd-6bbb-496e-b748-f03955026ff5 · outbound

This paper cites Alfworld: Aligning text and embodied environments for interactive learning.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Alfworld: Aligning text and embodied environments for interactive learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:11.473340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:11.473340Z digest=sha256:90a22eaf188398a398911bfc99b380085c4804447e18a6009c9634d25f496669

Observation 631d89fc-a2a0-4684-9021-44f972d26f30 · outbound

This paper cites Corex: Pushing the Boundaries of Complex Reasoning through Multi-Model Collaboration.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Corex: Pushing the Boundaries of Complex Reasoning through Multi-Model Collaboration

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:11.637834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:11.637834Z digest=sha256:cc50366079253faf5c1b7a658c7d2ffdd3cf897921179660f333bea01b03c1cd

Observation 0ba0fedd-aab4-41ee-9055-510c069421ca · outbound

This paper cites Os-genesis: Automating gui agent trajectory construction via reverse task synthesis.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Os-genesis: Automating gui agent trajectory construction via reverse task synthesis

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:11.767750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:11.767750Z digest=sha256:8d5c84130d0db7133ba14c1e3da1b2ac60d3361404dcdf97be0faf88b94a485a

Observation b568cec5-36ea-4275-8471-e1aaa957c2bd · outbound

This paper cites ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:11.873882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:11.873882Z digest=sha256:0107bab337cf7252bd6d83ae0fe6466fdd26b4b26e09508cf5410a1f8077adfc

Observation 5457a67c-5fef-457e-a935-c1c5ea5a7185 · outbound

This paper cites Mars: Situated inductive reasoning in an open-world environment.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Mars: Situated inductive reasoning in an open-world environment

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:12.006468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:12.006468Z digest=sha256:778f62f6addef601fbe27ad88fd2137a75ad31d4a8ed538e87387a3b42a9ea86

Observation 5db826d9-a19c-4df5-8b05-6c831e77eb1d · outbound

This paper cites Michelangelo: Long context evaluations beyond haystacks via latent structure queries.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Michelangelo: Long context evaluations beyond haystacks via latent structure queries

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:12.083678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:12.083678Z digest=sha256:351d9d354875b42c21291a28fcfee63b834f0bb4204094e8a0e9faff6759c509

Observation b861cdd6-77b0-4f70-816b-14f279c5dd1e · outbound

This paper cites Large language models for robotics: Opportunities, challenges, and perspectives.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Large language models for robotics: Opportunities, challenges, and perspectives

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:12.161041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:12.161041Z digest=sha256:a7c5beb9e7ba7ff0f32081c0d9e288166c552304189ae51bba685d3b3385fa63

Observation 5072d6c0-9b4a-4c27-8feb-cd934da8508e · outbound

This paper cites OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:12.299526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:12.299526Z digest=sha256:d032a38fc6ef4684d00f740353dc496e3964c31ba694f53910f9ea9666833f26

Observation 7dc572a8-0df3-451e-b9ac-a00441bee7b8 · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:12.409116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:12.409116Z digest=sha256:773abdb095babdb81e0109b8f037d65e4d65579cd745925ba79816d717bd633a

Observation 3d64ee9c-600e-4045-91ef-3e82ee3fccb2 · outbound

This paper cites J., Cheng, Z., Shin, D., Lei, F., et al.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions J., Cheng, Z., Shin, D., Lei, F., et al

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:12.589507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:12.589507Z digest=sha256:59c2ab1675d5c7303e2514e59ff4c1aff547585bf103260b411ede762be2de55

Observation 44a7aa4d-0634-4392-a1d1-38bc028e9b73 · outbound

This paper cites -decoding: Adaptive foresight sampling for balanced inference-time exploration and exploitation.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions -decoding: Adaptive foresight sampling for balanced inference-time exploration and exploitation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:12.822243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:12.822243Z digest=sha256:ec75f5796487fdece61ffa0b3917537657480458ed2a8170336a1de106d05049

Observation 8845d5c2-7533-412e-9d91-966758b90c40 · outbound

This paper cites Genius: A generalizable and purely unsupervised self-training framework for advanced reasoning.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Genius: A generalizable and purely unsupervised self-training framework for advanced reasoning

Reference 35

Resolution
verified exact
doi, observed 2026-08-03T04:08:22.679861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-03T04:08:13.072792Z digest=sha256:a1a65fe5e387f1ffc60cd371d1e195d83acb248714f85f6946a248011a902c25

Observation 75d7f364-33cf-4678-b519-71aa044782ed · outbound

This paper cites F., Song, Y., Li, B., Tang, Y., Jain, K., Bao, M., Wang, Z.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions F., Song, Y., Li, B., Tang, Y., Jain, K., Bao, M., Wang, Z

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:13.225646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:13.225646Z digest=sha256:26318d845422f3ef1cf4d0d6beda337b25758955a0c919c3c4d2441d6be00e9a

Observation dd8d17e4-1205-4db8-a761-618e5d003638 · outbound

This paper cites Tide: Trajectory-based diagnostic evaluation of test-time improvement in llm agents.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Tide: Trajectory-based diagnostic evaluation of test-time improvement in llm agents

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:13.333287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:13.333287Z digest=sha256:06d2c573bb8328496fb1ff27ff54dae4354695fd273352350fb95247a3842f26

Observation 2e1afbbe-60b3-4e4a-b98c-8b4a7d61c8ff · outbound

This paper cites Qwen3 Technical Report.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Qwen3 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:13.491585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:13.491585Z digest=sha256:c6928cf7698d7cbb5e89d0bfaf0c93359d8b9b755f3aa5709be0a3a3389acfbd

Observation a70c3329-8764-4e79-8ab6-3bf806c448e3 · outbound

This paper cites R., and Cao, Y.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions R., and Cao, Y

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:13.702944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:13.702944Z digest=sha256:376a072b2ada73ab578f7bb1aa9ea04fd0abde099c859fd40ff7fbaa32cbbcbb

Observation 9fcf34c6-6b1a-4498-9210-432e801688f0 · outbound

This paper cites an unresolved cited work.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:13.920742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:13.920742Z digest=sha256:a0631c0ff0b2a3eecfedeeef6a7292fa75b5d94fcb91425c3819ce1df45cc061

Observation 610b432b-f0ae-41c2-9860-b129f41b0f91 · outbound

This paper cites F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Ou, T., Bisk, Y., Fried, D., et al.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Ou, T., Bisk, Y., Fried, D., et al

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:14.068241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:14.068241Z digest=sha256:325dfc4583f90bb2b3bec0bcd1a963f5e970476d82103c61e8dbee77afd8b1b5

Pith citing papers

Observation 1605ec29-59e7-4226-a88a-3b8ec39d37e8 · inbound

Data-Driven Boundary Control of Distributed Port-Hamiltonian Systems cites this paper.

Data-Driven Boundary Control of Distributed Port-Hamiltonian Systems OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-13T10:21:59.840828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T10:21:59.840828Z digest=sha256:c29593dbe1305ce5b665982c2dc6b9b4a9678ababc47ed9b99d6338916ddbab0

Observation 21c0901f-91b2-4d8e-9d21-c68fd5fcdc63 · inbound

OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning cites this paper.

OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T13:29:51.597617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T05:07:16.926615Z digest=sha256:3f7389094599c692286f64e043d970033ef149897c768a69f23e82a248f7b13a

Observation af3179df-4eb4-4276-ac8f-2bcd261284e6 · inbound

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning cites this paper.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.243699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.243699Z digest=sha256:5b3c22b104a083c1a62edfda4c41528426f9fc0d3ddb53a256913bbf33b753ac

Observation e53c8bbd-d6b2-4c87-a4c7-57212f0f7113 · inbound

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models cites this paper.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.888233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.888233Z digest=sha256:d2682c1a72b182bd06da742cbb576750f6c9592ff9195267de28bbebcd1cdfdd

Observation 7ec780ab-da57-4e1c-ae97-bc8fa791f9dd · inbound

OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents cites this paper.

OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:40.363598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:38:40.363598Z digest=sha256:095fa38e1c1b599e55ab16f128f90619070e1bc0b7912fa0b4903d77087a254d