Pith. sign in

Paper Citation Record · LEDGER

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions

As of 7 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 4 inbound Pith citation observations for arXiv:2602.05843.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.05843 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:08:14.068241Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T01:10:21.243699Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T13:29:51.596096Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation abcf9bb2-67e3-4b98-a511-9cdc52c13b92 · outbound

This paper cites write newline.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:08.370940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:08.370940Z digest=sha256:1abd5f739f98d3b19d26e1c6517b660b3a7943a16110fe6923b280b11577e077

Observation 07be5ed1-70db-4ffe-b7a9-17b09f240e50 · outbound

This paper cites J., Bethge, M., and Schulz, E.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions J., Bethge, M., and Schulz, E

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:08.453393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:08.453393Z digest=sha256:4aff112e96897eddc19a536dda44e5d375dabf2e04041e43ba71694cdd142785

Observation 94800fc7-0706-4bf2-8fb4-90643d4fc806 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions The claude 3 model family: Opus, sonnet, haiku

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:08.572461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:08.572461Z digest=sha256:ef282b3255acab88f2df942ad4ae42e4f4b75962a3d5dd9b01b1b0b78457500d

Observation 57ffcf1c-e6a2-4347-b8e3-5696591186d4 · outbound

This paper cites H., and Bengio, Y.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions H., and Bengio, Y

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:08.670628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:08.670628Z digest=sha256:65fcf01f5170c57dc4c85a73e3c9aebf70d9ab01d25c5df5894a56541b75c10c

Observation 16f03f94-b2e8-4fb0-b521-7d6d0098e075 · outbound

This paper cites On the Measure of Intelligence.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions On the Measure of Intelligence

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:08.735295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:08.735295Z digest=sha256:5fb4ef16bce6687e5128fafc182b9eeb2e68df0bffb56f2d57be29a360930b3c

Observation 3fe0ebb9-7175-4866-a10c-225c9e067939 · outbound

This paper cites Evaluating long-context reasoning in llm-based webagents.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Evaluating long-context reasoning in llm-based webagents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:08.828452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:08.828452Z digest=sha256:0d6e544ce1858149cda038a1ed609d66c75b7798fa4a1d9d6008555d75f3ae74

Observation be7fd5c4-2ec1-406e-a03b-57bd71bd983c · outbound

This paper cites The dynamical challenge.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions The dynamical challenge

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:09.003347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:09.003347Z digest=sha256:e0cc8cbf183acb534e9c616709f11b3eb8b4d354e36e67afed05200c88ef6e6c

Observation 62382c70-fc00-4411-b000-1b900451b63a · outbound

This paper cites Mind2web: Towards a generalist agent for the web.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Mind2web: Towards a generalist agent for the web

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:09.106776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:09.106776Z digest=sha256:06c470ece85732ab96b9ef9260717190ae092211939131cb21309388d206618f

Observation 44797934-6322-4048-b2ff-eb7641e983e9 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:09.262712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:09.262712Z digest=sha256:0afb343968184976a946de4b5bd53c98e944c5ece4648f3affd178d9f9057953

Observation af6b907a-38c2-4a0f-be18-bc1e33218f2f · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:09.404493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:09.404493Z digest=sha256:e87a3e929bad63c258e767befa19d1108df8408154e2ddec1cfb93aa2fdd79d2

Observation 1e10462f-7d50-4565-9820-223e71e219ff · outbound

This paper cites The Llama 3 Herd of Models.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions The Llama 3 Herd of Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:09.547578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:09.547578Z digest=sha256:c5a33c92d4c3664605c093d2a7c505b224c5383f722fa759562d20e19eda0c80

Observation ed61fcd1-4de9-4f12-b221-08dfed5f1429 · outbound

This paper cites and Schmidhuber, J.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions and Schmidhuber, J

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:09.716497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:09.716497Z digest=sha256:c6943de7f050039ebce8e240a50a6b5689fb03ececc187cecc363975b3770cf2

Observation 9cec0ab3-7747-4553-af99-c627e6a1a99b · outbound

This paper cites H., Gonzalez, J.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions H., Gonzalez, J

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:09.905299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:09.905299Z digest=sha256:053f00274616bc790ad80363befaf58ce70811174a31c0ce0746e4b3536da450

Observation 8ff668df-7ab4-408c-8124-43c317ff134d · outbound

This paper cites M., Ullman, T.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions M., Ullman, T

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:10.115144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:10.115144Z digest=sha256:0588c877dea31a82c6511a17dd2c17005fe2573aa956cb5bf95cd3f990cc5e67

Observation 5ac8cd97-6d31-4f57-9d06-5cf1245b2007 · outbound

This paper cites State space models on temporal graphs: A first-principles study.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions State space models on temporal graphs: A first-principles study

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:10.290501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:10.290501Z digest=sha256:0ef948d70c99d55daa4be83ae9b3e4d590a0d463f3747955ccb60a7ef2d18d6b

Observation 2cd6a435-f4ba-42ab-b273-ca7db33800c4 · outbound

This paper cites Y., Le Bras, R., Richardson, K., Sabharwal, A., Poovendran, R., Clark, P., and Choi, Y.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Y., Le Bras, R., Richardson, K., Sabharwal, A., Poovendran, R., Clark, P., and Choi, Y

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:10.419920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:10.419920Z digest=sha256:8ba78e0f02a4bbcd633c302b25f3727eff564893c6b08a42669b35dd00eb7c88

Observation 1852803b-9dba-413b-b13f-e8e029d63b47 · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:10.546967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:10.546967Z digest=sha256:2a677045c1da8ca59a1ca0b7704a92f77dc9ec172149e0793552c71862720688

Observation c604330b-a965-4290-9c52-a8cef996ec79 · outbound

This paper cites Agentbench: Evaluating llms as agents.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Agentbench: Evaluating llms as agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:10.676369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:10.676369Z digest=sha256:3540e29dd2a2848e8126373c80aaf03887a19ce7437d0b95d3ce327b4079446b

Observation dfa35dbd-6872-4218-ac85-b64b4971485e · outbound

This paper cites Gaia: a benchmark for general ai assistants.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Gaia: a benchmark for general ai assistants

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:10.801903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:10.801903Z digest=sha256:6c83a3caa04db0e29c7c5f310a913eda6d16304b5fbf55a4ddc66d90e9c77eb4

Observation 85a10ad0-493f-4700-91ed-7e7dd255d1c3 · outbound

This paper cites gpt-oss-120b & gpt-oss-20b Model Card.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions gpt-oss-120b & gpt-oss-20b Model Card

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:10.909307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:10.909307Z digest=sha256:3812780f13a4517c1b086df820c22a5d61fd2f42cdab73527885845d7d4b3ccd

Observation 23bf426b-6ed8-4bef-ac19-2290650ab49b · outbound

This paper cites G., Mao, H., Yan, F., Ji, C.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions G., Mao, H., Yan, F., Ji, C

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:11.042144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:11.042144Z digest=sha256:1341a2812bfac1a571bda54d739ec1c544a285991f9513b47fbb46647cca92f6

Observation b09a2020-4b4a-47e9-954c-030b1dd6c2b7 · outbound

This paper cites E., Li, W., Campbell-Ajala, F., Toyama, D.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions E., Li, W., Campbell-Ajala, F., Toyama, D

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:11.173177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:11.173177Z digest=sha256:404170f937ed829e6030783a37112f0215e6edf5a26f3c7653d5e23e59445f52

Observation 0a45872a-b04e-43aa-ac18-15242df16764 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Reflexion: Language agents with verbal reinforcement learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:11.342942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:11.342942Z digest=sha256:cc9cfae6177e35cef6156ac8ac544284fd6b0a1eaa04cec90a9bba6c0d41d53f

Observation a6e9b9cd-6bbb-496e-b748-f03955026ff5 · outbound

This paper cites Alfworld: Aligning text and embodied environments for interactive learning.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Alfworld: Aligning text and embodied environments for interactive learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:11.473340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:11.473340Z digest=sha256:0d58dc9e40da60015b1e17fa394da093258c78ca6979eb793e438275e7b6107b

Observation 631d89fc-a2a0-4684-9021-44f972d26f30 · outbound

This paper cites Corex: Pushing the Boundaries of Complex Reasoning through Multi-Model Collaboration.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Corex: Pushing the Boundaries of Complex Reasoning through Multi-Model Collaboration

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:11.637834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:11.637834Z digest=sha256:eb9f76bec436551d0bbc2cedb73ee05782cba068f6986497ec71e23c4d60f45e

Observation 0ba0fedd-aab4-41ee-9055-510c069421ca · outbound

This paper cites Os-genesis: Automating gui agent trajectory construction via reverse task synthesis.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Os-genesis: Automating gui agent trajectory construction via reverse task synthesis

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:11.767750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:11.767750Z digest=sha256:dbf21b04f0d6cd879305e805436a25080537934770fed928e79e83b6b3ffab87

Observation b568cec5-36ea-4275-8471-e1aaa957c2bd · outbound

This paper cites ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:11.873882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:11.873882Z digest=sha256:4a1208d6d7df817d048698d55f8e4523366a7b14ae1fa60354d56540e209a1c4

Observation 5457a67c-5fef-457e-a935-c1c5ea5a7185 · outbound

This paper cites Mars: Situated inductive reasoning in an open-world environment.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Mars: Situated inductive reasoning in an open-world environment

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:12.006468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:12.006468Z digest=sha256:c5fd9ccc5d646b2757c88b3761b89c9842898a31b34ca04c944093c7f74e7103

Observation 5db826d9-a19c-4df5-8b05-6c831e77eb1d · outbound

This paper cites Michelangelo: Long context evaluations beyond haystacks via latent structure queries.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Michelangelo: Long context evaluations beyond haystacks via latent structure queries

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:12.083678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:12.083678Z digest=sha256:4e759c83ea633ddbaf6857315df0c40540367352c2996585f7623689175e6c61

Observation b861cdd6-77b0-4f70-816b-14f279c5dd1e · outbound

This paper cites Large language models for robotics: Opportunities, challenges, and perspectives.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Large language models for robotics: Opportunities, challenges, and perspectives

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:12.161041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:12.161041Z digest=sha256:d4b5cb4c0a5cb81d5d0ad038a2669cc62874f0e415bd392fc3d8a87cb8ed6c3e

Observation 5072d6c0-9b4a-4c27-8feb-cd934da8508e · outbound

This paper cites OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:12.299526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:12.299526Z digest=sha256:fa570fe91f7b751ab533427aef8af812515e2261aaebb3f25099cd90f54d9a8e

Observation 7dc572a8-0df3-451e-b9ac-a00441bee7b8 · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:12.409116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:12.409116Z digest=sha256:bc6656586e6ecc00f6835ef37c187e74738cebb59b56df2fdd559066b33fe430

Observation 3d64ee9c-600e-4045-91ef-3e82ee3fccb2 · outbound

This paper cites J., Cheng, Z., Shin, D., Lei, F., et al.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions J., Cheng, Z., Shin, D., Lei, F., et al

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:12.589507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:12.589507Z digest=sha256:2a0246b8405c82bc4f7fdd6a46b0be739d8125f4cd21ffa96ccb3254679a9983

Observation 44a7aa4d-0634-4392-a1d1-38bc028e9b73 · outbound

This paper cites -decoding: Adaptive foresight sampling for balanced inference-time exploration and exploitation.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions -decoding: Adaptive foresight sampling for balanced inference-time exploration and exploitation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:12.822243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:12.822243Z digest=sha256:3e245c9264c23244020b8a93d8290b5fb487aaa718d26e63b391dda8e0f3f006

Observation 8845d5c2-7533-412e-9d91-966758b90c40 · outbound

This paper cites Genius: A generalizable and purely unsupervised self-training framework for advanced reasoning.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Genius: A generalizable and purely unsupervised self-training framework for advanced reasoning

Reference 35

Resolution
verified exact
doi, observed 2026-08-03T04:08:22.679861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-03T04:08:13.072792Z digest=sha256:109d81215b7f33de01b10123fa28eb05e183c6a658df70dcc60dd802d115c363

Observation 75d7f364-33cf-4678-b519-71aa044782ed · outbound

This paper cites F., Song, Y., Li, B., Tang, Y., Jain, K., Bao, M., Wang, Z.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions F., Song, Y., Li, B., Tang, Y., Jain, K., Bao, M., Wang, Z

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:13.225646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:13.225646Z digest=sha256:b9134f54fbc7edc992b362f8656bb2d21c1c878bb3fb7b8550243746394b439c

Observation dd8d17e4-1205-4db8-a761-618e5d003638 · outbound

This paper cites Tide: Trajectory-based diagnostic evaluation of test-time improvement in llm agents.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Tide: Trajectory-based diagnostic evaluation of test-time improvement in llm agents

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:13.333287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:13.333287Z digest=sha256:ee1a1389f27e625f88898f0473b6d06e237668c7176b340c4c91acf857553b3f

Observation 2e1afbbe-60b3-4e4a-b98c-8b4a7d61c8ff · outbound

This paper cites Qwen3 Technical Report.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Qwen3 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:13.491585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:13.491585Z digest=sha256:7b6d88daf5c092f61a6d0b5146c80e929ad70e79ffea7de4fb09381f4d706173

Observation a70c3329-8764-4e79-8ab6-3bf806c448e3 · outbound

This paper cites R., and Cao, Y.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions R., and Cao, Y

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:13.702944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:13.702944Z digest=sha256:9f571b0b6c9371e9371851a0cf21bf39c6d7e618b3949e0fdab55b196b31338e

Observation 9fcf34c6-6b1a-4498-9210-432e801688f0 · outbound

This paper cites an unresolved cited work.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:13.920742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:13.920742Z digest=sha256:f3d7f16ca1bce93626269b71ecfec0c54527d49c75c29bb5fdbff137decfa75a

Observation 610b432b-f0ae-41c2-9860-b129f41b0f91 · outbound

This paper cites F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Ou, T., Bisk, Y., Fried, D., et al.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Ou, T., Bisk, Y., Fried, D., et al

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:14.068241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:14.068241Z digest=sha256:42772366517f8df96fbfd1eda0e7436a1c31167766992a4ed3ec4575cfbde0a1

Pith citing papers

Observation 1605ec29-59e7-4226-a88a-3b8ec39d37e8 · inbound

Data-Driven Boundary Control of Distributed Port-Hamiltonian Systems cites this paper.

Data-Driven Boundary Control of Distributed Port-Hamiltonian Systems OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-13T10:21:59.840828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T10:21:59.840828Z digest=sha256:ccf97c36a5908e1e6715d09fe9efb378480b74fa9f5e79d889dbc07ff66ae642

Observation 21c0901f-91b2-4d8e-9d21-c68fd5fcdc63 · inbound

OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning cites this paper.

OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T13:29:51.597617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T05:07:16.926615Z digest=sha256:021762d62ba04baf0663dd4bbc71c72748fee46fb7bee2ad08085d68e0f5c5f3

Observation af3179df-4eb4-4276-ac8f-2bcd261284e6 · inbound

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning cites this paper.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.243699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.243699Z digest=sha256:4a63d45f43ff3ede73c9a7a7ec43a027d240f592640e7f13f3aa8a235e0a031e

Observation e53c8bbd-d6b2-4c87-a4c7-57212f0f7113 · inbound

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models cites this paper.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.888233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.888233Z digest=sha256:5728999efeecf9e5e3677688a80847cfdd35eb0d589e1371a5fd7fc0068b88a1