Pith. sign in

Paper Citation Record · LEDGER

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

As of 10 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 6 inbound Pith citation observations for arXiv:2604.23781.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.23781 v2

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T06:36:52.202847Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T00:36:15.758892Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T01:37:42.725645Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact12
  • verified fuzzy7
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c745d5cd-947d-45b1-b23e-3be81c733c16 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:11:10.841344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:b647bc8dfef5f5cf669f8ba39fba2d5ddff04a0185c313314d56bbbc16696502

Observation 43f1aa7d-cb1e-489a-9b73-05e2b84d18a4 · outbound

This paper cites WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:48:05.307585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:16e8872cf61a051a11f744fc7476086f8738112f35a7863e45997b93b10f7a68

Observation 5408acbf-6df5-4556-a7e8-0321276c6aab · outbound

This paper cites TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:39:31.128039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:4cd892c2cc2e09aff30249bcddfe4fc6efefb53bf33bef6006afd6a02d342cf9

Observation c3412cfb-7adf-412f-bf95-825aa7babade · outbound

This paper cites Visualwebarena: Evaluating multimodal agents on realistic visual web tasks.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Visualwebarena: Evaluating multimodal agents on realistic visual web tasks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T18:38:14.480542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:09c35abbaa4bbd9f1eb7c8da88af1ce66ecfe882854c847852380ebfb325f553

Observation 1396d387-47a3-4eb3-b6a0-a488517f3dad · outbound

This paper cites Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.Advances in Neural Information Processing Systems, 37:52040–52094.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.Advances in Neural Information Processing Systems, 37:52040–52094

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T18:38:14.472468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:efbee5f73d4d7bb1773cd90b5350fb9df6b62b578fc4763a9a9c7dfbd9929420

Observation e170de81-13c1-4c27-8264-99aa5a626e3e · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:11:10.918407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:877ddc9b38eea9716a67bb395fc77a16353ccf4ff26fb94421c191b216d1c4ed

Observation a4525fe0-769f-415d-9c41-d2dbc0339314 · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:11:10.854057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:33e48f7ba490642224cf31e8a10312a360bbeb4b19dd38958d4ec88735040473

Observation d9df3b27-bbb1-4b22-a4f1-1df1e37fed2a · outbound

This paper cites Mcpmark: A benchmark for stress-testing realistic and comprehensive mcp use.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Mcpmark: A benchmark for stress-testing realistic and comprehensive mcp use

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:11:10.898917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:7ca725840bf104b6554f55fd65bb2b7ee3d405b2a0970d5b8fdc1928c2ee31e5

Observation 889a55d8-3cc0-4ed2-a100-c62caa90a053 · outbound

This paper cites Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:11:10.795764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:354c5bae62f91bdacbdfb0da2b2d40816b4a2de2d4a34520609dd28af4acbc0b

Observation ec1f690c-0094-4a98-a97a-41eb93f7fb8a · outbound

This paper cites ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:11:10.876875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:5166a39465157d1118b9bad55cdfa9d29548ebd6cb75524e1e203c184f3c0821

Observation 77bd5cdf-05b7-48af-a50d-d0e1c6ece492 · outbound

This paper cites Mind2web: Towards a generalist agent for the web.Advances in Neural Information Processing Systems, 36:28091–28114.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Mind2web: Towards a generalist agent for the web.Advances in Neural Information Processing Systems, 36:28091–28114

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T18:38:14.484838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:124ae3d38ef5657b456d01653f540ccde6fa20aa009f999ad17f9e2e04137c1a

Observation df550c02-dd52-459a-aed6-c25951f65292 · outbound

This paper cites MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:11:10.779830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:4f922884de6480b3422eb6b3afaf6526367d79ce47808530ce0d23465dbeb835

Observation 2b1ea6ea-9559-4c50-b4e6-2ca1d7d62910 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:11:10.882978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:55988c2e60c3862b853495fb4fed05be413e40c96e059bf68bd7df3efce0957f

Observation 0ca1a61a-1aae-4786-ae12-899f465d57d1 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents AgentBench: Evaluating LLMs as Agents

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:11:10.925570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:4603f3f28166f365043bfd41d4507ab381e7e2b2bd86321bcd1e3f60629ae0e0

Observation 53ed90ca-aa7b-4bdc-a015-497c18def0b8 · outbound

This paper cites Gaia: a benchmark for general ai assistants.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Gaia: a benchmark for general ai assistants

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T18:38:14.476310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:9e1d14b3b909ddae0a5542560318ed1d03ec74821e26f97c1658572fd8c6c6b0

Observation b21eccad-d261-4c1d-aead-73eb8ef2d31e · outbound

This paper cites ClawArena: Benchmarking AI Agents in Evolving Information Environments.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents ClawArena: Benchmarking AI Agents in Evolving Information Environments

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:11:10.906594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:10d44031f3b3dbc9df1fc13736281ab3f36b05d8caa631e47469a245d5af1735

Observation 22252659-e486-4d29-934b-cc202fd46ae7 · outbound

This paper cites Swe-agent: Agent-computer interfaces enable automated software engineering.Advances in Neural Information Processing Systems, 37:50528–50652.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Swe-agent: Agent-computer interfaces enable automated software engineering.Advances in Neural Information Processing Systems, 37:50528–50652

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T18:38:14.488632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:c753743f53ce8008789999c1098d923940ab7ac095b4b70605acf6a61505188f

Observation 4aaa25e5-9a7c-4c13-a291-38dbfdc45c9f · outbound

This paper cites Autogen: Enabling next-gen llm applications via multi-agent conversations.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Autogen: Enabling next-gen llm applications via multi-agent conversations

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T18:38:14.465276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:98a04c754e0dea4b5a0b047d397aad2ba9922966699fd1a427150ee0e52c9a2d

Observation 79e1908f-7616-4582-af63-ef7707d9240f · outbound

This paper cites MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:11:10.806436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:6b7488a9867a571808a56b96eae8ad29e1e065cfbfc54721b29871857ce11be1

Observation 3f9702e2-a8e0-442e-94f1-e1f5dff331f7 · outbound

This paper cites Day 1 / Day 2 / Day 3.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Day 1 / Day 2 / Day 3

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T18:38:14.468512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:902656dd4739272e66fe6a903485f6a0e836e09b8801976def3d1214ef383546

Pith citing papers

Observation 8e670151-85ff-468d-a7c7-b5031919b87b · inbound

VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild cites this paper.

VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:23:28.378969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T13:14:07.487577Z digest=sha256:1fd6f70a19e61137eb2cad77b9e0e9ba91ebc91da18ba16876b272718fb4e0e1

Observation 903180f3-6562-4001-8539-c9c014c8b650 · inbound

Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks cites this paper.

Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-03T08:57:47.643579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T10:39:51.630229Z digest=sha256:55dad215119465a1cd3292ee80324517eb143ad040055bbfe6530534a77f6d3e

Observation 6b12aacc-012d-4920-bf42-26f09ecb7724 · inbound

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents cites this paper.

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-08T20:05:34.094932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-08T20:03:47.212663Z digest=sha256:e30cd2263ac2388dd806b8fdc4a18a4fc50566982137188cf36e683a3cb32741

Observation b3611f80-5451-40c1-96e3-10bec46c0c40 · inbound

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents cites this paper.

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:37:42.757410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T01:37:23.646537Z digest=sha256:b42dfe82ad17d0c72930465088672c5ab01d217ca9e7eb21b766e5b8d08456ff

Observation eade2889-5875-4bbc-859a-1e97167440ab · inbound

UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks cites this paper.

UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:46:41.270682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T01:39:57.741032Z digest=sha256:50bb077e44e5184fed61d7e97e086bcab1db3ebb2f7d93af57abf67207482d84

Observation 35d0ddb8-f4d3-4ac4-920f-1a718e7be0cb · inbound

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios cites this paper.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.758892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.758892Z digest=sha256:a2925a8fd1e5df6e4e0ea562d76267cad9d888a7671ca849bc7303fc1487ecd1