Pith. sign in

Paper Citation Record · LEDGER

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios

As of 10 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2607.14989.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.14989 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T00:36:15.896432Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T03:38:16.681296Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f92f1d4b-46d7-4c3f-81e9-094a084376a9 · outbound

This paper cites Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.686510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.686510Z digest=sha256:1842b897c32c76489e7b4a8e00fee7403beb2695fc2248af17a35f1750f51528

Observation ee70646b-4454-4e07-8009-a1e1c2d75d7b · outbound

This paper cites MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.692389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.692389Z digest=sha256:edb22c41e0d1713e942bd25cfa6444da5f2397e589d5644f67de22565a2d7498

Observation 98ecf2b2-2b0c-4c1a-8cae-d008ec31a237 · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.697891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.697891Z digest=sha256:c86a8073df5bdac2dcd07731d06d47c015aebf4fac5884447a0a4302a14ad7dd

Observation 69d89396-9c19-42f5-9f52-506cad14fab0 · outbound

This paper cites Workarena++: Towards compositional planning and reasoning-based common knowledge work tasks.Advances in Neural Information Processing Systems, 37:5996–6051, 2024.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Workarena++: Towards compositional planning and reasoning-based common knowledge work tasks.Advances in Neural Information Processing Systems, 37:5996–6051, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.703133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.703133Z digest=sha256:e6b949492afa23003ef105894f2267ddca95c7dd9889ce974775efa047b56b8e

Observation 92e1de46-1928-4ce4-ba49-12da0a77ca13 · outbound

This paper cites Acebench: Who wins the match point in tool usage?arXiv preprint arXiv:2501.12851, 2025.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Acebench: Who wins the match point in tool usage?arXiv preprint arXiv:2501.12851, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.708055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.708055Z digest=sha256:b35bded1fb8ff8667a40295285af2a1f1a0dfb6cc3d1a8032fa212feebbf8124

Observation d2222a21-b736-4cfe-8275-ce8d485c1f75 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.712554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.712554Z digest=sha256:bb2423f774ce874050a37a1e25932310117c2a9fdcb4d9e486f3963ea6933aa1

Observation d9ac0475-4325-4b21-8e75-df4f7b354e95 · outbound

This paper cites Deepseek-v4: Towards highly efficient million-token context intelligence, 2026.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Deepseek-v4: Towards highly efficient million-token context intelligence, 2026

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.717827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.717827Z digest=sha256:93c404b361de135edfa2b09f678e28793847ac0452a1c71c973b2cae3cfd5291

Observation f80d6b94-f83a-4286-9116-df602d194120 · outbound

This paper cites Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.722166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.722166Z digest=sha256:35f3ca4a529b1ae1b3c626e478ce3ced6418f09bb10a2f4b1ad6ce13c670d519

Observation cfadd12e-87fa-42be-8641-d1fbb1574dde · outbound

This paper cites Gaia2: Benchmarking llm agents on dynamic and asynchronous environments.arXiv preprint arXiv:2602.11964, 2026.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Gaia2: Benchmarking llm agents on dynamic and asynchronous environments.arXiv preprint arXiv:2602.11964, 2026

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.726706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.726706Z digest=sha256:07ec8191b1da7e867d8d26966be5ce3301fadc99cc3db3bda54f0b7e0f6254b1

Observation c345dfa9-7688-4b25-b7ca-c55a6fe11a38 · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios GLM-5: from Vibe Coding to Agentic Engineering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.731165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.731165Z digest=sha256:740b414c1b43a9fbb465b953764886dca24f64eba67dcc02c165b63dc1f1dce7

Observation a1ad8539-97f9-47ff-874d-8edd8998011c · outbound

This paper cites Vitabench: Benchmarking llm agents with versatile interactive tasks in real-world applications.arXiv preprint arXiv:2509.26490, 2025.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Vitabench: Benchmarking llm agents with versatile interactive tasks in real-world applications.arXiv preprint arXiv:2509.26490, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.736318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.736318Z digest=sha256:e3cc57c8a090434beed90677a4734291f6307b87bafe4ad153d3332752b2e9b2

Observation d532e613-26c8-4670-99cf-ee37c8d0485f · outbound

This paper cites Swe-bench: Can language models resolve real-world github issues? In International Conference on Learning Representations, volume 2024, pages 54107–54157, 2024.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Swe-bench: Can language models resolve real-world github issues? In International Conference on Learning Representations, volume 2024, pages 54107–54157, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.740857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.740857Z digest=sha256:243f0738a5bf5f195ced374ca759d6447f26712d2348251b0fb0962ad49e1208

Observation de01f234-0ee2-40bb-a02f-ae99aa420a8e · outbound

This paper cites The tool decathlon: Bench- marking language agents for diverse, realistic, and long-horizon task execution.arXiv preprint arXiv:2510.25726, 2025.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios The tool decathlon: Bench- marking language agents for diverse, realistic, and long-horizon task execution.arXiv preprint arXiv:2510.25726, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.745842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.745842Z digest=sha256:aa96b4943d2d7a4613a8233b13c6f439b0e77bad7a0cb33e140f352af373bf50

Observation ae5b7737-562e-4b11-abc5-8bf9457f9a66 · outbound

This paper cites Dataflow: An llm-driven framework for unified data preparation and workflow automation in the era of data-centric ai, 2025.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Dataflow: An llm-driven framework for unified data preparation and workflow automation in the era of data-centric ai, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.750217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.750217Z digest=sha256:ced4e4c68d6274822a96bfa7aeaab7eb2afbc58bed6519dcd4332848ef69b43a

Observation dcec0a53-67d5-4c3d-bcb9-1ae873084715 · outbound

This paper cites Toolsandbox: A stateful, conversational, interactive evaluation benchmark for llm tool use capabilities.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Toolsandbox: A stateful, conversational, interactive evaluation benchmark for llm tool use capabilities

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.754402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.754402Z digest=sha256:b0f4d033b133410d9e7c4898a8009da91d3d63f538bfc242833343c99208d571

Observation 35d0ddb8-f4d3-4ac4-920f-1a718e7be0cb · outbound

This paper cites ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.758892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.758892Z digest=sha256:a2925a8fd1e5df6e4e0ea562d76267cad9d888a7671ca849bc7303fc1487ecd1

Observation 57e6884d-83e8-4a3e-aa7c-f3d8bac50212 · outbound

This paper cites Gorilla: Large language model connected with massive apis.Advances in Neural Information Processing Systems, 37:126544–126565, 2024.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Gorilla: Large language model connected with massive apis.Advances in Neural Information Processing Systems, 37:126544–126565, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.763968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.763968Z digest=sha256:e75a4e66bc7ef83e26943f27fa75901029d57c3446f6d4191c9083ee654953a3

Observation c13722e1-b478-40ba-a0d6-96a425633d9a · outbound

This paper cites The berkeley function calling leaderboard (bfcl): From tool use to agentic evaluation of large language models.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios The berkeley function calling leaderboard (bfcl): From tool use to agentic evaluation of large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.768484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.768484Z digest=sha256:50e2a95396dd597c65eb0d3da43cad30aed27ac122bc58ecb4311aebf5719153

Observation 146c6ecc-0f79-4997-83da-47bd6ce570af · outbound

This paper cites GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.772715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.772715Z digest=sha256:086498cdc1346e06a9b91140a1c37ff48e7617bb20b264608da88c199682cff3

Observation d90dc87d-f0b7-4e55-b017-cfc8530cc58c · outbound

This paper cites QwenClawBench: Real-user-distribution benchmark for openclaw agents, April.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios QwenClawBench: Real-user-distribution benchmark for openclaw agents, April

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.777037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.777037Z digest=sha256:9d97a40590f02fdc6510aa40c70f9c36c899a7d7b263faeef8d3fb497f60f689

Observation 8d291883-7561-47ed-b00e-4e53a70b11b8 · outbound

This paper cites One-eval: An agentic system for automated and traceable llm evaluation.arXiv preprint arXiv:2603.09821, 2026.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios One-eval: An agentic system for automated and traceable llm evaluation.arXiv preprint arXiv:2603.09821, 2026

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.785663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.785663Z digest=sha256:2b07945ed5cf07395439ed3c1ec70187938047ed021888aa6e734ae7ee51666b

Observation 8a43c760-0979-454d-9dcb-148f7ee0153a · outbound

This paper cites URLhttps://arxiv.org/abs/2603.04370.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios URLhttps://arxiv.org/abs/2603.04370

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.789506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.789506Z digest=sha256:bde12f7359a88c91bcb581fd68590905ca8b6d994f33b8fc17fda705575a9764

Observation e55cf75c-035a-4b25-a140-5a352ebe7b0d · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Kimi K2.5: Visual Agentic Intelligence

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.794297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.794297Z digest=sha256:0abf7a967874a76e94948904d3657fb544cc0b4715fe89764f2163ec6bbd276f

Observation bb12d6ba-b891-44d2-be04-6d568462a14a · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Systems, 37:95266–95290, 2024.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Systems, 37:95266–95290, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.799456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.799456Z digest=sha256:dd4852e7bf027eeca1da304242bbbd91e5ea0e3a002d2fc00488e0311bf47609

Observation af412506-04f8-471c-97d8-06db166c7e8e · outbound

This paper cites Swe-agent: Agent-computer interfaces enable automated software engineering.Advances in Neural Information Processing Systems, 37:50528–50652, 2024.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Swe-agent: Agent-computer interfaces enable automated software engineering.Advances in Neural Information Processing Systems, 37:50528–50652, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.804245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.804245Z digest=sha256:50c60165ac72933ab31350e392084db1bece6009bc1bfca78874c76169745013

Observation 380c0c35-124a-4edf-9d35-aae77bbfd9a5 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.809267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.809267Z digest=sha256:2147dc2fcb79687017c4044ce3dee4dd8d8214d565b52de6be6c3d3b337f2702

Observation f7303418-7fa0-46ef-9090-43c43d571a79 · outbound

This paper cites Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.814013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.814013Z digest=sha256:f9fa9f3fc8d6fc0de101543de1364ceb92da220aca12e04addc606b90d3c9a91

Observation 6fe38867-3a50-4928-8cc1-323641177e9f · outbound

This paper cites Deepplanning: Bench- marking long-horizon agentic planning with verifiable constraints.arXiv preprint arXiv:2601.18137, 2026.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Deepplanning: Bench- marking long-horizon agentic planning with verifiable constraints.arXiv preprint arXiv:2601.18137, 2026

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.818900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.818900Z digest=sha256:8080d88880daa726bef7ab606ed6a9d702d8e5128a79345fe3a4f4b39b417400

Observation 532400e7-adb4-410f-acfd-6b22622c9a7c · outbound

This paper cites an unresolved cited work.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.824559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.824559Z digest=sha256:8596b1dc57329b6a88fa4961fb8baa93feba38df000c869632500a3ae821b827

Observation 7099790e-86de-4bb8-89bc-b4636ce2f1e0 · outbound

This paper cites an unresolved cited work.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.828684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.828684Z digest=sha256:7239200657e3f8008e9d860fe62246ab9452b2a182bcf52767abba4f434e8b41

Observation 72246456-5f0d-4556-9518-d2adf7b427e4 · outbound

This paper cites an unresolved cited work.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.833186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.833186Z digest=sha256:83fc5acc786ee7bae975bfcac404575200adfa70aabf468b6936d72db676b03f

Observation 67998d9a-8708-4d4e-b59d-12469d9c2195 · outbound

This paper cites This protocol evaluates not only task execution, but also clarification, constraint tracking, adaptation to user feedback, and state maintenance across multiple turns.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios This protocol evaluates not only task execution, but also clarification, constraint tracking, adaptation to user feedback, and state maintenance across multiple turns

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.837265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.837265Z digest=sha256:5daa655f4726c27a0a2e9f558197f7ed79caf3350d16aa236aac8c707c440264

Observation 714e253d-0a79-4272-8875-05d696ed5080 · outbound

This paper cites an unresolved cited work.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.842384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.842384Z digest=sha256:16a71838107ed8c40a2746ca7754f4a372dbef156e131ba4c311300c8bc41ae1

Observation 77acbd66-6a2b-476f-bf2d-f0b21dcb3264 · outbound

This paper cites an unresolved cited work.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.846871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.846871Z digest=sha256:b09e2a18b4a615375e8dd77067b2324b8a9b33d3d752caabc95b8af8b6e01eae

Observation ef31ed00-4acd-458d-83e4-0b45f1df8868 · outbound

This paper cites 4.VerifyCodechecks the trajectory and final observation and returns a binary pass/fail result.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios 4.VerifyCodechecks the trajectory and final observation and returns a binary pass/fail result

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.852162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.852162Z digest=sha256:6815cda087bad6781b9ab48a54045a41ea5877fdb3305fbc5fb6fdf94e84af17

Observation e7af32ea-81ca-4c1b-bc37-92a81124aa3d · outbound

This paper cites [...additional description omitted ...].

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios [...additional description omitted ...]

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.856453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.856453Z digest=sha256:2c4f9f1635e40a77f4906b0173f034287bcf816b7c60c94b8226520841573fe3

Observation fcd424a9-9c6d-47a2-8ee4-9f4dde5f68d8 · outbound

This paper cites [...additional tools omitted ...].

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios [...additional tools omitted ...]

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.860557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.860557Z digest=sha256:706a5b4af598f1c317e282e75760402232172b92298983c22448b051b5ea413a

Observation 10c3a754-e51a-4626-9ef3-498b001c9a3b · outbound

This paper cites Your goal is to complete the user’s request in an interactive environment by gradually calling the available tools step by step, and to proactively communicate with the user.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Your goal is to complete the user’s request in an interactive environment by gradually calling the available tools step by step, and to proactively communicate with the user

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.865246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.865246Z digest=sha256:bdede0ebea2c57e6b9670cfbe77510229028f03e18f4804cf45d38ce108f25b8

Observation 85bd7733-9245-40e0-89a5-575d6b867983 · outbound

This paper cites an unresolved cited work.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.869397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.869397Z digest=sha256:f6a287f09a26dab7b11b8afc936a3dde92a9dad8f0c18cea839f91f7d9401793

Observation e457aacb-d23d-48fc-92db-2e1a7301f1e2 · outbound

This paper cites <Judge reason> The trajectory correctly resolvesLF-2024-PI-024, confirms Yunhe Foods / Lin Qiaoxue / He Shan, verifiesVER-004 as current, and flags the self-referentialLNK-003link.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios <Judge reason> The trajectory correctly resolvesLF-2024-PI-024, confirms Yunhe Foods / Lin Qiaoxue / He Shan, verifiesVER-004 as current, and flags the self-referentialLNK-003link

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.874308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.874308Z digest=sha256:4c33282242a8c55c6a5b09d4ad57f00f6f7876b5a712475e2ace9cc9a8c85e98

Observation 3b900804-95e7-4063-82de-3f3bce04cd52 · outbound

This paper cites store surveillance screenshots and incident timeline explanation.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios store surveillance screenshots and incident timeline explanation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.878606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.878606Z digest=sha256:9c516cd8b33476dd564763d58d7d8cd37b82957ced6732439ed91f4d1ebcdfd2

Observation 008c583d-9452-4fa5-936d-392a1db75110 · outbound

This paper cites an unresolved cited work.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.883109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.883109Z digest=sha256:53fa56162b11079ae30188238ab2f794cf130a9911bdc16d07858fdcd06ecf43

Observation 8db42b9d-0575-4d7e-9996-7269cca3ffab · outbound

This paper cites an unresolved cited work.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.887535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.887535Z digest=sha256:e4e5243e350ef729e81e44c317b80ad9f4964f7df4126e864e72572aad2cfe4a

Observation bd688285-357f-4f79-a60f-60a168ba8c39 · outbound

This paper cites an unresolved cited work.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.892243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.892243Z digest=sha256:8825fadb70509581d91027ffa8a4bbbf5fe0951908a2e6e21724ad70b9b1e722

Observation ecf144a3-b8d9-41c8-9c6e-b03b254a77ae · outbound

This paper cites an unresolved cited work.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.896432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.896432Z digest=sha256:4c9a99c33907a4f93679651b79d2e48ce6da383922a07dc3c3d378d564e6434b

Observation 2861883a-b8c8-47d8-b643-7ac88668d0c4 · outbound

This paper cites an unresolved cited work.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.781239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.781239Z digest=sha256:b4c232329c7c18a6934e74ecec739c411d6db60315300a99b7776464cbac5126

Pith citing papers

Observation 2bc7ea6a-b277-4acd-9cb3-21332b85d0d0 · inbound

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications cites this paper.

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T03:38:16.681296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:38:16.681296Z digest=sha256:d6e3fcffe1c76dd63f4dc75610aa320256b2aaddc9d76c6ffda8fee639b8b2a8