Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T10:20:42.139206Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 0 inbound Pith citation observations for arXiv:2608.05573.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T10:20:42.139206Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
87 of 87 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e1838b19-9ae6-4b53-9b74-849d764d42a1 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2023 , eprint=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6746e21-fe57-4773-99be-8f88edde0e95 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2023 , eprint=
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dc189fb-9cbd-464f-827b-10e9c99b6d06 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2023 , eprint=
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 115d3e34-87b1-4a40-8cff-2ca1193eacae · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2023 , eprint=
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36bc2a8c-a90a-44de-8ec9-87d43b5345fe · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2023 , eprint=
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 605723e7-9dd6-4b43-8317-2d4fd92ac338 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2023 , eprint=
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76a16a61-8e05-4305-80c2-a9854e4b31c8 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2024 , eprint=
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4124fb18-2c70-4e82-aaf4-b0c893cd52f8 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2024 , eprint=
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c758cc47-ffcb-44dd-8a6a-0a3985549a56 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2025 , eprint=
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f265538-767b-4323-b80f-6de9cbc71207 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2023 , eprint=
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8f3d3ea-f70d-43b8-ad0d-60f8f493f376 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2023 , eprint=
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fa0e07a9-5873-4a23-b1e9-4a0a38016745 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2024 , eprint=
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2d5d464-9c85-4661-b3a1-ce9ff103b120 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2024 , eprint=
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 901690f8-2cfb-44f1-bfb8-63d6244c504a · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2024 , eprint=
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71f7a4e3-e1fa-433a-8013-b1c38591f2d6 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dd45944-d872-4523-a882-0b80e3f12839 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution American Journal of Physics , volume=
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 16f3e26a-adec-4f14-a051-53f9ee6a090e · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2025 , eprint=
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f786dde9-5592-4bab-b3db-ae816f438117 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9a75888d-d688-494c-a431-133612daaca6 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23fb6f24-093d-4c11-b0ae-cf3f32436750 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2025 , eprint=
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20915c10-3a66-441e-a6ad-d8862cb4c515 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73bba7e4-7e6a-4eac-a93b-536537643724 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2025 , eprint=
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1814d7fd-ca8b-4c96-8dcd-0d9378370ca3 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2025 , eprint=
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation aaff6c34-0d1b-4008-a4ef-01bca91a3c62 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4311a4f0-41f2-438b-928f-dc2c00302e75 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2025 , eprint=
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5d0f2b5c-1b21-472e-a54d-261d131047d2 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 46133c43-a6ae-4a56-a2c5-d36a7f450d01 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 46934b6c-ab7b-481b-9c2e-9d93de6f1191 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 45cb4063-7c6f-4b5d-98c0-8e598c14ceed · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2025 , eprint=
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2f72aca4-d14b-40ae-8a6d-d0c71cbc56bd · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2024 , eprint=
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f1d7a2f-0ba1-4b4b-b3ba-30cf16fedd09 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2024 , eprint=
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ccb06a9-53be-4771-9d5b-c493ad7efec8 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0c2fca53-b17b-41f0-9807-9185cd688b15 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7ce7d7d-5917-4404-827b-eef02ff86665 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2025 , month = dec, howpublished =
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3137ec12-ed27-466f-bda3-ec7535043a49 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , month = feb, howpublished =
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e023d98-eae1-4613-905d-3cd269a0cb94 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aded129b-ba23-485b-ab89-0f491432c2c7 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef923420-b94c-4741-9cb1-c5c198fa28f9 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5c57f621-1b2f-4cd8-b98c-25ca48ea0b68 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 31d8c9d7-45c7-4681-9f31-d376959d28a3 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6c6fbaed-d074-42f9-b59a-6a94f937b1d1 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f1e1222e-e8f7-4d92-873e-4db3afc80033 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2024 , eprint=
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a77bf0bc-47f4-4456-ab0a-70735985353c · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2023 , eprint=
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60064129-dc20-4a03-ad04-bc0f3c5d9b3b · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2025 , eprint=
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4467e8d7-58dd-454f-935c-f9a7b9504359 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 63e28695-0423-4b24-be9c-bb376043ed32 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution 2026 , eprint=
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00163140-ff00-42e1-993b-f26dc19b6bfd · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0db83985-a646-47f7-94c3-c790b4d5cb5c · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 03b43aa9-ea6a-472b-9b34-296cd4262595 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Unresolved cited work
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a939bd00-543f-4d0a-9997-4734aa44cc00 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e53c890-7176-4a94-be53-a6486cfcfd1d · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Skill-RM: Unifying Heterogeneous Evaluation Criteria via Agent Skill
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 99ce2704-f7f9-4a00-b166-14c7106acc03 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3e682bc-6a56-477b-8050-229d5f06caae · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution TRAIL: Trace Reasoning and Agentic Issue Localization
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77931dce-e424-4f4b-ba1b-cb2ba15ac6ab · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Agent Skill Evaluation and Evolution: Frameworks and Benchmarks
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f43a4304-974a-4340-a9b2-42e524b03d4b · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4027aed4-c4ec-422f-bf7b-b3e8db0287ca · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b2b70a5-17a1-470a-b6ff-61047b3f4d51 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Unresolved cited work
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 509a09f3-8f4e-42c6-ba7c-3f437d975017 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Unresolved cited work
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6be1e988-f0e5-4429-bf44-975420dc87b3 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a65af639-b90b-491a-b3b5-a2ed11e28685 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution RewardBench: Evaluating Reward Models for Language Modeling
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df2e98db-e610-41f0-8d71-f28ee9b8ca9f · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc2babe8-a99b-403d-9c29-63d0d47ffee3 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Unresolved cited work
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57f586e4-c838-4d27-aa1c-84a528e5cf7c · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62a3a2b6-d3a9-4617-95f5-b5edce64e164 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution SkillRevise: Improving LLM-Authored Agent Skills via Trace-Conditioned Skill Revision
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a4d8dca-9e04-46be-bc1c-95eef9990c49 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7399c01-5fca-49af-a7de-a568a0a36584 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution H.; Kazemnejad, A.; Meade, N.; Patel, A.; Shin, D.; Zambrano, A.; Stańczak, K.; Shaw, P.; Pal, C
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76fe7547-9f91-4b51-be64-4965d5138f0a · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Unresolved cited work
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d902ca15-77bb-4c25-ba82-86a530a72e0a · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0be9b6c-63d0-4895-8b98-8901043a00ac · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Unresolved cited work
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 414a0229-1d5c-4ef0-9ceb-e187ecd9b577 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution W.; Liu, J.; Chen, W.; Chen, Z.; and Lou, Y
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75b3d18f-7b98-46be-b496-57586c529a14 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Toolformer: Language Models Can Teach Themselves to Use Tools
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1260d50e-e60a-4d04-aff8-03836dfcd4f3 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Characterizing Faults in Agentic AI: A Taxonomy of Types, Symptoms, and Root Causes
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 212e6fc4-a23a-4a68-83d9-b0dd0d1f55d8 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e83ceb7-74c6-45c0-ab31-30937fb559cd · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe62e75e-1728-43a8-81e8-54901ceac076 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Aegis: Taxonomy and Optimizations for Overcoming Agent-Environment Failures in LLM Agents
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81177f5d-6c6b-4a08-bdec-73c47b27ea40 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3b1ab29-11f5-49eb-89f4-f191c1ddb722 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Unresolved cited work
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4998afc2-02ec-415c-8f4c-b9609f1d8b84 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b635b442-6671-4c56-87de-87b47a2168b1 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49cebb2b-7740-416a-a37f-6136e0684a27 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Large Language Models as Optimizers
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0423ab39-9a53-47c7-8675-5b716629344d · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f645ac63-7a16-4883-bd6f-6f19ced31253 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution ReAct: Synergizing Reasoning and Acting in Language Models
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 673bbefc-2f79-44a7-8952-0da07ea5c559 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e914543-90f7-4f96-bbdc-5af9c64d2bb7 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Scaling Test-time Compute for LLM Agents
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba1d707d-f51f-4b5c-83a7-a5c9c0e1bbba · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Unresolved cited work
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84a52ccf-4e28-4d6c-8625-14447c1bc5e2 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Where Agent Frameworks Fall Short: Examining Functional Challenges and Usability Concerns
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04bacc21-0956-4a45-b058-15d960ea72f2 · outbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Agent-as-a-Judge: Evaluate Agents with Agents
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.