Pith. sign in

Paper Citation Record · LEDGER

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution

As of 10 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 0 inbound Pith citation observations for arXiv:2608.05573.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05573 v1

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T10:20:42.139206Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

87 of 87 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved70
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e1838b19-9ae6-4b53-9b74-849d764d42a1 · outbound

This paper cites 2023 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2023 , eprint=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.840481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.840481Z digest=sha256:04f63d8b47893482a8eb5d465f62443c33bf8cc7d35263f24fee3287092b152d

Observation d6746e21-fe57-4773-99be-8f88edde0e95 · outbound

This paper cites 2023 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2023 , eprint=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.844288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.844288Z digest=sha256:3daec4e432a8148038591155124e68259d9de4fb78d4c2b57552debc78094cbf

Observation 2dc189fb-9cbd-464f-827b-10e9c99b6d06 · outbound

This paper cites 2023 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2023 , eprint=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.847861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.847861Z digest=sha256:97fcf653fa5005da78ae4e9058539d3e0d916ae0d09a5509ffd4c880b63af467

Observation 115d3e34-87b1-4a40-8cff-2ca1193eacae · outbound

This paper cites 2023 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2023 , eprint=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.851210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.851210Z digest=sha256:5757e7bb2a59171e7d25c09a7de591830c9326c70746131224ef15f98a811533

Observation 36bc2a8c-a90a-44de-8ec9-87d43b5345fe · outbound

This paper cites 2023 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2023 , eprint=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.854672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.854672Z digest=sha256:640e43d989fc6a51de989e4cc95d933abc1a1541b995217c6511f2ae65c212c8

Observation 605723e7-9dd6-4b43-8317-2d4fd92ac338 · outbound

This paper cites 2023 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2023 , eprint=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.858096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.858096Z digest=sha256:5ffcf6b2ff15b3759ded66608beac625db6958ddb2544f63a62fd8f36e06f802

Observation 76a16a61-8e05-4305-80c2-a9854e4b31c8 · outbound

This paper cites 2024 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2024 , eprint=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.861913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.861913Z digest=sha256:3734d2ab2b661393da30604f9a498d46d65dde28e3d4092e70d618b13c9c5dfa

Observation 4124fb18-2c70-4e82-aaf4-b0c893cd52f8 · outbound

This paper cites 2024 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2024 , eprint=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.865188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.865188Z digest=sha256:d9b305c990bb384fabeb3c6d4085da65bd03712b76848745868e3ff3ea7b58d0

Observation c758cc47-ffcb-44dd-8a6a-0a3985549a56 · outbound

This paper cites 2025 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2025 , eprint=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.868568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.868568Z digest=sha256:ac8c5966f178c95c0357814337d8dcd0ddf136c8d654fb658d4c483c69014ca3

Observation 2f265538-767b-4323-b80f-6de9cbc71207 · outbound

This paper cites 2023 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2023 , eprint=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.871676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.871676Z digest=sha256:9f971fdd4160cf109b2fe41edd4de42653c9ad45efc00b633fde96be6f66532c

Observation d8f3d3ea-f70d-43b8-ad0d-60f8f493f376 · outbound

This paper cites 2023 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2023 , eprint=

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:43.886907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T10:20:41.874949Z digest=sha256:377757aa3a83b64ed4afa88af6cf98bf76b9594146868f00695e3b2b844dbf47

Observation fa0e07a9-5873-4a23-b1e9-4a0a38016745 · outbound

This paper cites 2024 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2024 , eprint=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.878151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.878151Z digest=sha256:bed1936202106baa71fa2f5f5392baf4e6a4bcd38a1a678fe19320f764b7f0ee

Observation c2d5d464-9c85-4661-b3a1-ce9ff103b120 · outbound

This paper cites 2024 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2024 , eprint=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.881504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.881504Z digest=sha256:e088b3d822f0388dad4b0e7a73601ade4b375bce99ed2f9f14221d54ad3c1d75

Observation 901690f8-2cfb-44f1-bfb8-63d6244c504a · outbound

This paper cites 2024 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2024 , eprint=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.885487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.885487Z digest=sha256:dea3a7249ae0548d7abfe0365fbc18398d0a8495bd7b6aa0cf6e7e08e43b78fc

Observation 71f7a4e3-e1fa-433a-8013-b1c38591f2d6 · outbound

This paper cites 2026 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.888670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.888670Z digest=sha256:dd9cd20af6a731987e1b4a50bc592675e37401c2f87a6f371969516bfe544693

Observation 1dd45944-d872-4523-a882-0b80e3f12839 · outbound

This paper cites American Journal of Physics , volume=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution American Journal of Physics , volume=

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:43.855529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T10:20:41.891952Z digest=sha256:8aad0a5fd9dfcf018da92c2e0a3ac527911c3a596db1ee9de714d99bd9666c0e

Observation 16f3e26a-adec-4f14-a051-53f9ee6a090e · outbound

This paper cites 2025 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2025 , eprint=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.895829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.895829Z digest=sha256:1250c76b4bc5740a394acbec5e21501379f0790e9949f7872f4cf7a73806e6f2

Observation f786dde9-5592-4bab-b3db-ae816f438117 · outbound

This paper cites 2026 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:43.840466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T10:20:41.899389Z digest=sha256:db94aebe258b22548d13d88741aa67c30ae8c970fc503187df4725041b3018dc

Observation 9a75888d-d688-494c-a431-133612daaca6 · outbound

This paper cites 2026 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.902619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.902619Z digest=sha256:1dec145172d007c2f6e50324692ddfc65486a267d159f317e054c574e45b7a8d

Observation 23fb6f24-093d-4c11-b0ae-cf3f32436750 · outbound

This paper cites 2025 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2025 , eprint=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.905914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.905914Z digest=sha256:420cc36519053dc266f9053d4b8daaf4d4f13bd3f80f190ad84d52a7361cb29b

Observation 20915c10-3a66-441e-a6ad-d8862cb4c515 · outbound

This paper cites 2026 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.909605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.909605Z digest=sha256:03cf0df4737f23c038b4d2470fad59b746726258b314cf88b22fd7af5b4c0338

Observation 73bba7e4-7e6a-4eac-a93b-536537643724 · outbound

This paper cites 2025 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2025 , eprint=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.912952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.912952Z digest=sha256:8861add133276e0c7f5512e8c9e5b175293ee720da4817f87906687bcebfc554

Observation 1814d7fd-ca8b-4c96-8dcd-0d9378370ca3 · outbound

This paper cites 2025 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2025 , eprint=

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:43.807975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T10:20:41.916102Z digest=sha256:d7dfc76ddec81797d58ef2cc5f35ffcfd4e08d2a6bcd8ab8e618d9c0d891552c

Observation aaff6c34-0d1b-4008-a4ef-01bca91a3c62 · outbound

This paper cites 2026 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:43.797956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T10:20:41.919673Z digest=sha256:baadb8b75f49d35d5b675602e114e5b1ba20d1a38f7b75a889eea878a9cdbdc2

Observation 4311a4f0-41f2-438b-928f-dc2c00302e75 · outbound

This paper cites 2025 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2025 , eprint=

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:43.787533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T10:20:41.923029Z digest=sha256:46104e65e94887da522d3f6574977a41bcdd8f05f5d94c080d50359bef673b3b

Observation 5d0f2b5c-1b21-472e-a54d-261d131047d2 · outbound

This paper cites 2026 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:43.777775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T10:20:41.926170Z digest=sha256:a6d8d855767ea5f4cfcd098a845d4a02bfdf48186269eaa740a48a027e5f16c5

Observation 46133c43-a6ae-4a56-a2c5-d36a7f450d01 · outbound

This paper cites 2026 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:43.767756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T10:20:41.929460Z digest=sha256:4d988fd52296c1aac39f7fe1b95855cb17edcba26d9f635dc9753d3e71d15565

Observation 46934b6c-ab7b-481b-9c2e-9d93de6f1191 · outbound

This paper cites 2026 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:43.757626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T10:20:41.932800Z digest=sha256:35b6917eedad1c7c8baac67cc704c97172a3a67b7b580dbb9c40e8ada0fad7ac

Observation 45cb4063-7c6f-4b5d-98c0-8e598c14ceed · outbound

This paper cites 2025 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2025 , eprint=

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:43.746771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T10:20:41.936365Z digest=sha256:937f2400318d4e4c200f6dcf5b10465d7a6272e2c6b198a41c78b5f286985e2f

Observation 2f72aca4-d14b-40ae-8a6d-d0c71cbc56bd · outbound

This paper cites 2024 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2024 , eprint=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.939386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.939386Z digest=sha256:be17d46f6ed4c53c2aa42fd6495e06326c08f7597a15795ecdc4b8639e1686c6

Observation 5f1d7a2f-0ba1-4b4b-b3ba-30cf16fedd09 · outbound

This paper cites 2024 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2024 , eprint=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.942919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.942919Z digest=sha256:e2025671907bea9bee9982d5bda590706c0d0417f6a8d952c614d3eb3654179e

Observation 3ccb06a9-53be-4771-9d5b-c493ad7efec8 · outbound

This paper cites an unresolved cited work.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-08T10:20:43.725545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T10:20:41.949776Z digest=sha256:978da757b00873a0ded0d6a8969a854ca25e986ffb116e7d54f10f510d436da2

Observation 0c2fca53-b17b-41f0-9807-9185cd688b15 · outbound

This paper cites 2026 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.952912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.952912Z digest=sha256:d063f1b0277ed657512fca314b0ece3abead518bd48b8db90d8b07e3f5a19691

Observation c7ce7d7d-5917-4404-827b-eef02ff86665 · outbound

This paper cites 2025 , month = dec, howpublished =.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2025 , month = dec, howpublished =

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.955776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.955776Z digest=sha256:eadc25f208e6cdb020b6ca77787f93bbee1095b2a1c8eb36a7ac59efc436edf4

Observation 3137ec12-ed27-466f-bda3-ec7535043a49 · outbound

This paper cites 2026 , month = feb, howpublished =.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , month = feb, howpublished =

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.958642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.958642Z digest=sha256:fc5c7a36251ea9b4a9c85c075f4c70bf47f600ddb103e1b746e14f6544e8b52a

Observation 9e023d98-eae1-4613-905d-3cd269a0cb94 · outbound

This paper cites an unresolved cited work.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.961697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.961697Z digest=sha256:8f47693ab55764491c148176b0edd552a273309b86f67ab44a244e89e7da7d20

Observation aded129b-ba23-485b-ab89-0f491432c2c7 · outbound

This paper cites 2026 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.964844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.964844Z digest=sha256:019380c6852170032dd1714abbd2ee51633d441a004090cd78ed3fcefa3be3a9

Observation ef923420-b94c-4741-9cb1-c5c198fa28f9 · outbound

This paper cites 2026 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:43.688386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T10:20:41.967762Z digest=sha256:7cc762ca4b264d738377566cfc2a2bb551f546ce4bd54aa0b1ba3a3b62ac093c

Observation 5c57f621-1b2f-4cd8-b98c-25ca48ea0b68 · outbound

This paper cites 2026 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:43.678640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T10:20:41.971179Z digest=sha256:475cf4b2ddbcd085b4b6655016287be342e2042bfe010065e287853d168a80f6

Observation 31d8c9d7-45c7-4681-9f31-d376959d28a3 · outbound

This paper cites 2026 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:43.668642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T10:20:41.974473Z digest=sha256:b1edf911477f9972e67754ef59825f23eb18dcd70208bef8e59adff678e2cbac

Observation 6c6fbaed-d074-42f9-b59a-6a94f937b1d1 · outbound

This paper cites 2026 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:43.659779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T10:20:41.977614Z digest=sha256:64adf173d2e588ca6407cab89c7f22faeedf799b6697df727f9b507acf27c211

Observation f1e1222e-e8f7-4d92-873e-4db3afc80033 · outbound

This paper cites 2024 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2024 , eprint=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.981047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.981047Z digest=sha256:ac97ac78dc8f028baf8534bf5e37e2c71ea1111cedc55959469e18282a88d61d

Observation a77bf0bc-47f4-4456-ab0a-70735985353c · outbound

This paper cites 2023 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2023 , eprint=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.984603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.984603Z digest=sha256:f8ac76b6bc1fe827c95ecf5bbe977a838813167a99486c8d5e9fb60d39ef95b6

Observation 60064129-dc20-4a03-ad04-bc0f3c5d9b3b · outbound

This paper cites 2025 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2025 , eprint=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:43.638229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T10:20:41.988036Z digest=sha256:377ede7540eac8cb7330d711e4f4ebaefd5f7930af42447760d7a43558301521

Observation 4467e8d7-58dd-454f-935c-f9a7b9504359 · outbound

This paper cites 2026 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:43.626085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T10:20:41.991087Z digest=sha256:478e84fe85e358bfae7bff7b5c44e5f5ae79427a9559d628232391bc21140f56

Observation 63e28695-0423-4b24-be9c-bb376043ed32 · outbound

This paper cites 2026 , eprint=.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.994447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.994447Z digest=sha256:15ea0a844822e2ac1972435603a80411f17b73900c0a02660437a7a71c54773d

Observation 00163140-ff00-42e1-993b-f26dc19b6bfd · outbound

This paper cites GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:41.997995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:41.997995Z digest=sha256:73de99c1d60edcf2fa62e266669cbd3822016c744f1ec0405bf26af033106aab

Observation 0db83985-a646-47f7-94c3-c790b4d5cb5c · outbound

This paper cites an unresolved cited work.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-08T10:20:43.607735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T10:20:42.001840Z digest=sha256:e15d552482bfc03a674cd0e6483e866dcae621a2dc7e5682955f3b32cfaefbce

Observation 03b43aa9-ea6a-472b-9b34-296cd4262595 · outbound

This paper cites an unresolved cited work.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.005062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.005062Z digest=sha256:7f34add0100cbfc4b083814514bb749861d1d0f6d60956cfadd06913fb99ff87

Observation a939bd00-543f-4d0a-9997-4734aa44cc00 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.008420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.008420Z digest=sha256:2e9fbe959ee9ce10486405326ef711e01b9c016c0ed10deda2b90f49199d64db

Observation 7e53c890-7176-4a94-be53-a6486cfcfd1d · outbound

This paper cites Skill-RM: Unifying Heterogeneous Evaluation Criteria via Agent Skill.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Skill-RM: Unifying Heterogeneous Evaluation Criteria via Agent Skill

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-08T10:20:43.390067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T10:20:42.012024Z digest=sha256:702b0bccafa021f3ebdfb8e397eaf37db2917c284303791f2f822bee43b74d96

Observation 99ce2704-f7f9-4a00-b166-14c7106acc03 · outbound

This paper cites SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.015720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.015720Z digest=sha256:983c8f3f72fce73d51a709c9288759e5389328da1b67ae38a94e570f714c0519

Observation a3e682bc-6a56-477b-8050-229d5f06caae · outbound

This paper cites TRAIL: Trace Reasoning and Agentic Issue Localization.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution TRAIL: Trace Reasoning and Agentic Issue Localization

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.019625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.019625Z digest=sha256:18e2c7191aab43ddc84df23e9e913743c099a4a4481fff1a37dcb43ee2aa0c84

Observation 77931dce-e424-4f4b-ba1b-cb2ba15ac6ab · outbound

This paper cites Agent Skill Evaluation and Evolution: Frameworks and Benchmarks.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Agent Skill Evaluation and Evolution: Frameworks and Benchmarks

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.023057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.023057Z digest=sha256:68c0043880d4e9e7ac616dbb80e6663572cb53e0c1edb2de4c3a8db20be8d86a

Observation f43a4304-974a-4340-a9b2-42e524b03d4b · outbound

This paper cites WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.027123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.027123Z digest=sha256:74560f53d7e77aecfb0672290b07258170ec6bf9ed3c8681843863bdeecc26a0

Observation 4027aed4-c4ec-422f-bf7b-b3e8db0287ca · outbound

This paper cites Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.030214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.030214Z digest=sha256:6759f8728c9f1c0c01f1314c8f3f42172376aa6b877b1e90bf6ed1d352f0419f

Observation 2b2b70a5-17a1-470a-b6ff-61047b3f4d51 · outbound

This paper cites an unresolved cited work.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.033779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.033779Z digest=sha256:48c47cac10366b9f00f323604f0f83a61abff636d7ee863637d05c8fd33f6508

Observation 509a09f3-8f4e-42c6-ba7c-3f437d975017 · outbound

This paper cites an unresolved cited work.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.036730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.036730Z digest=sha256:1085fc83aff4515cc59c4c42a92e9d27ee28c91cc163b6539423f056da303388

Observation 6be1e988-f0e5-4429-bf44-975420dc87b3 · outbound

This paper cites Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.039474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.039474Z digest=sha256:b2a010d60c262385370af5423017923ac8e554bbe51e10b3ef39e7a65b8c91c6

Observation a65af639-b90b-491a-b3b5-a2ed11e28685 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution RewardBench: Evaluating Reward Models for Language Modeling

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.042569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.042569Z digest=sha256:12cfdb3dd7894dd72cd3ee7687b2ffc56657159fb2f593452692a8ca9f15a5e5

Observation df2e98db-e610-41f0-8d71-f28ee9b8ca9f · outbound

This paper cites SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.046260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.046260Z digest=sha256:63492e2297730e4d11d96283959baa6c5dd141f887099f18a2acfede601e674d

Observation dc2babe8-a99b-403d-9c29-63d0d47ffee3 · outbound

This paper cites an unresolved cited work.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.049441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.049441Z digest=sha256:305e3173ff294949de2a9dd97b188ca5bf64e4cc60073847cb972030cff01840

Observation 57f586e4-c838-4d27-aa1c-84a528e5cf7c · outbound

This paper cites G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.052412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.052412Z digest=sha256:e8ec92c6e1bcf1d74e9a219b8909355d466a93b21e76cef815029fddef522f92

Observation 62a3a2b6-d3a9-4617-95f5-b5edce64e164 · outbound

This paper cites SkillRevise: Improving LLM-Authored Agent Skills via Trace-Conditioned Skill Revision.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution SkillRevise: Improving LLM-Authored Agent Skills via Trace-Conditioned Skill Revision

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.055705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.055705Z digest=sha256:d2b980c1b43e80c642293ac93ceb5ca3ef52c1895afc4f154658eeb32efe37da

Observation 4a4d8dca-9e04-46be-bc1c-95eef9990c49 · outbound

This paper cites Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.058799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.058799Z digest=sha256:09e431b315cca36cbdf6c98b41cb8dcfab54402389f362ce5f843d0c2521f142

Observation a7399c01-5fca-49af-a7de-a568a0a36584 · outbound

This paper cites H.; Kazemnejad, A.; Meade, N.; Patel, A.; Shin, D.; Zambrano, A.; Stańczak, K.; Shaw, P.; Pal, C.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution H.; Kazemnejad, A.; Meade, N.; Patel, A.; Shin, D.; Zambrano, A.; Stańczak, K.; Shaw, P.; Pal, C

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.061941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.061941Z digest=sha256:0a6f4f6318a19484a0de3c4844eadeceddd03b3983610b7baec2b4c8b6227474

Observation 76fe7547-9f91-4b51-be64-4965d5138f0a · outbound

This paper cites an unresolved cited work.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.064892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.064892Z digest=sha256:a3efaa7a8103e867e23298ee0e7aea1aa124939dbb93c4e9f3cde491ea382527

Observation d902ca15-77bb-4c25-ba82-86a530a72e0a · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.068370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.068370Z digest=sha256:07dfc133c6cb6a68ab65ab7b09c7ba7be2395a3ce3fd672ab9cfe65e4328b7d3

Observation a0be9b6c-63d0-4895-8b98-8901043a00ac · outbound

This paper cites an unresolved cited work.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-08T10:20:43.595893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T10:20:42.072021Z digest=sha256:6b8931c3aaccc8e5fc0564e46de6569e83fac74b75415e9a364864901750c374

Observation 414a0229-1d5c-4ef0-9ceb-e187ecd9b577 · outbound

This paper cites W.; Liu, J.; Chen, W.; Chen, Z.; and Lou, Y.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution W.; Liu, J.; Chen, W.; Chen, Z.; and Lou, Y

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.075461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.075461Z digest=sha256:91c914cbdaf22a7b4f57d85e894db39f58b6032b817145d549fb486b20e6bd74

Observation 75b3d18f-7b98-46be-b496-57586c529a14 · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.078991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.078991Z digest=sha256:3afa00645a2236409d247c09fb1b6881e48b99fd6eb8236af25ad938a91e80ed

Observation 1260d50e-e60a-4d04-aff8-03836dfcd4f3 · outbound

This paper cites Characterizing Faults in Agentic AI: A Taxonomy of Types, Symptoms, and Root Causes.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Characterizing Faults in Agentic AI: A Taxonomy of Types, Symptoms, and Root Causes

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.082598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.082598Z digest=sha256:0419a0a911ba01061a2f08750580aa8bc8387fffa2a5894fe00f7d67bd5667be

Observation 212e6fc4-a23a-4a68-83d9-b0dd0d1f55d8 · outbound

This paper cites AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.086849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.086849Z digest=sha256:a82b261b1bef2d159ddc0502708ac7e47285eb2f05c4722c7dba6b365bc1ea35

Observation 0e83ceb7-74c6-45c0-ab31-30937fb559cd · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.090825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.090825Z digest=sha256:af1211976f06e6324152879f9ba730ab4cc2805b9a1d6be3639d05a2a7736afa

Observation fe62e75e-1728-43a8-81e8-54901ceac076 · outbound

This paper cites Aegis: Taxonomy and Optimizations for Overcoming Agent-Environment Failures in LLM Agents.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Aegis: Taxonomy and Optimizations for Overcoming Agent-Environment Failures in LLM Agents

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.094789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.094789Z digest=sha256:a59714272b3d5ca993e3afae6a5746db0023dc5d7dc76a252de170d0878627ec

Observation 81177f5d-6c6b-4a08-bdec-73c47b27ea40 · outbound

This paper cites JudgeBench: A Benchmark for Evaluating LLM-based Judges.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.098785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.098785Z digest=sha256:ccfc8720212cac4cc864d506c6441da734a7bc46bfa770311cd4c57ebda18099

Observation b3b1ab29-11f5-49eb-89f4-f191c1ddb722 · outbound

This paper cites an unresolved cited work.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-08T10:20:43.584544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T10:20:42.102691Z digest=sha256:56b3f2815028f9fc57c177fcf428d645e8cb0d0fef3c4c897f40007ebce1b6ea

Observation 4998afc2-02ec-415c-8f4c-b9609f1d8b84 · outbound

This paper cites Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.106403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.106403Z digest=sha256:0448d8c11273d5093b7838c7917b3df05dc0db070afb7e7d84ee604766c70127

Observation b635b442-6671-4c56-87de-87b47a2168b1 · outbound

This paper cites Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.111062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.111062Z digest=sha256:a44a44b54f46bdcb3b5d2c062e2e420617021e543fcf662d175b56a51adce1ad

Observation 49cebb2b-7740-416a-a37f-6136e0684a27 · outbound

This paper cites Large Language Models as Optimizers.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Large Language Models as Optimizers

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.114847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.114847Z digest=sha256:a36a8dfbe68516372d0f8f82aaa93966c15460fef234286715fffe0e45d3d87f

Observation 0423ab39-9a53-47c7-8675-5b716629344d · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.118832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.118832Z digest=sha256:62654b0cbe27957225c004a2b1d4fab494b7d91548f2d1d162f2fc555209749f

Observation f645ac63-7a16-4883-bd6f-6f19ced31253 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution ReAct: Synergizing Reasoning and Acting in Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.122793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.122793Z digest=sha256:7443065fde850076a3b5b15bb565b7660fe1dd06f24ee3840f31139a91693f76

Observation 673bbefc-2f79-44a7-8952-0da07ea5c559 · outbound

This paper cites SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.126671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.126671Z digest=sha256:f9c91cbe3aa72e00b3ef3214ccba6342fab961d9eb0a8b4724c63f923aa7659e

Observation 3e914543-90f7-4f96-bbdc-5af9c64d2bb7 · outbound

This paper cites Scaling Test-time Compute for LLM Agents.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Scaling Test-time Compute for LLM Agents

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.130163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.130163Z digest=sha256:7c54d0e1ea62627ab83bf7f0981c15fccd1ba872ea7a7ce6a22c7d45ebfc6e01

Observation ba1d707d-f51f-4b5c-83a7-a5c9c0e1bbba · outbound

This paper cites an unresolved cited work.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Unresolved cited work

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.132968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.132968Z digest=sha256:b7c22ab06af0ea69df1898f203429a49690b90e3dd63216dc393cd3df383e942

Observation 84a52ccf-4e28-4d6c-8625-14447c1bc5e2 · outbound

This paper cites Where Agent Frameworks Fall Short: Examining Functional Challenges and Usability Concerns.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Where Agent Frameworks Fall Short: Examining Functional Challenges and Usability Concerns

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.136007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.136007Z digest=sha256:83e31ed0da2617364494421f6190d2c099afe4bee6c6fb8250331d4e71ec95b2

Observation 04bacc21-0956-4a45-b058-15d960ea72f2 · outbound

This paper cites Agent-as-a-Judge: Evaluate Agents with Agents.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Agent-as-a-Judge: Evaluate Agents with Agents

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.139206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.139206Z digest=sha256:1412bcdadff4d54c755f4732296e691ec5b65d40189db9a2b5c7864f3a4400b4

Pith citing papers

No inbound Pith citation observations are available.