Pith. sign in

Paper Citation Record · LEDGER

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise

As of 9 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2608.02372.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02372 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T08:32:49.575837Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved48
  • parse uncertain0
  • malformed identifier4
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 504de7a3-c67a-47a9-8bcc-9fc7f327b4a1 · outbound

This paper cites Claude Haiku 4.5 System Card.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Claude Haiku 4.5 System Card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:43.743272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:43.743272Z digest=sha256:806fb38fbaf7abca5d5840c83eb6ca4d31f14781ae412926d8fe01e0c0fa5d6c

Observation d6651767-5d6b-4e87-9e57-7be3754334cf · outbound

This paper cites Claude Opus 4.7 System Card.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Claude Opus 4.7 System Card

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:43.873998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:43.873998Z digest=sha256:c59d721bf4b5f72fa0450c37f85d0d80a4ec5877505277cf9cb147f74c97ceb7

Observation cb68e588-dc12-4889-a02c-ae1a043ca196 · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:43.971429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:43.971429Z digest=sha256:a2330f0f0f90234001c887b5d1f093720f9ade3f75e1b25dff0b87188df05ef9

Observation c61c0bf5-ac1a-4ccb-8fc9-43d31a4b96a6 · outbound

This paper cites MultiWOZ - a large-scale multi-domain Wizard-of-Oz dataset for task-oriented dialogue modelling.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise MultiWOZ - a large-scale multi-domain Wizard-of-Oz dataset for task-oriented dialogue modelling

Reference 4

Resolution
malformed identifier
no resolver link, observed 2026-08-04T08:32:44.072039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:44.072039Z digest=sha256:82b6b440bf88c2937cc54614793253f78e1f56fbbcd588a03547b728f4804e7c

Observation e8cba22b-efea-4878-94e8-fe83925c57f7 · outbound

This paper cites Timebench: A comprehensive evaluation of temporal reasoning abilities in large language models.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Timebench: A comprehensive evaluation of temporal reasoning abilities in large language models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:44.169178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:44.169178Z digest=sha256:5fb0a80606fb1633ebc3cbba51cdb1d7c30a087223cca2775ed847ca7c2a8fe5

Observation 8fef0fec-1a1b-4150-a7fc-761f7621088e · outbound

This paper cites Timer: Temporal instruction modeling and evaluation for longitudinal clinical records.npj Digital Medicine, 8(1):577, 2025.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Timer: Temporal instruction modeling and evaluation for longitudinal clinical records.npj Digital Medicine, 8(1):577, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:44.284896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:44.284896Z digest=sha256:33bdfaf107e11dbb3f27b3842cf88e36ed3a1e774bf5261975b0b1fc58f01cde

Observation b4ce5431-deca-4e9f-a547-6dd8fb684c4c · outbound

This paper cites Deepseek-v4: Towards highly efficient million-token context intelligence, 2026.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Deepseek-v4: Towards highly efficient million-token context intelligence, 2026

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:44.458523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:44.458523Z digest=sha256:671e8694f0563873284d52308a5a4548a3ad4bf9142bf3a7cb06cd5d28be43e1

Observation 3cca429d-0fc0-4c08-a55c-1c4a31f629a2 · outbound

This paper cites Mind2web: Towards a generalist agent for the web.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Mind2web: Towards a generalist agent for the web

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:44.629181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:44.629181Z digest=sha256:b04aacb53d32b63174baa06b884d19cb35e7ddefeffa8fcfcbe9ad176d923ef0

Observation 8ce20688-edbe-45a7-9bb0-5f444207b151 · outbound

This paper cites MultiWOZ 2.1: A consol- idated multi-domain dialogue dataset with state corrections and state tracking baselines.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise MultiWOZ 2.1: A consol- idated multi-domain dialogue dataset with state corrections and state tracking baselines

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:44.770225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:44.770225Z digest=sha256:dbf938d3b4ba9c7f6c2a68fc1b90a326c4ae9e339b7e097a29baa18ea2d179c4

Observation 72ffb4de-1d67-4d89-8d48-91273adfc39b · outbound

This paper cites Verification of forecasts expressed in terms of probability.Monthly weather review, 78(1):1–3, 1950.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Verification of forecasts expressed in terms of probability.Monthly weather review, 78(1):1–3, 1950

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:44.861482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:44.861482Z digest=sha256:42944c4f1a182cb3cbfd678f21b7d211b4e5a5409c56641b1ebd73b069134b54

Observation 00a40c3e-78a1-42fe-ac92-2c25ac880ac6 · outbound

This paper cites Gemini 3 Flash Model Card.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Gemini 3 Flash Model Card

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:45.014107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:45.014107Z digest=sha256:e3305d200521a48efc25f7f77095c7a3574ba9008bbb14936702c7b85491a8f6

Observation a3bf996f-f749-4ae0-859c-fc4493a44e5b · outbound

This paper cites Gemini 3.1 Pro Model Card.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Gemini 3.1 Pro Model Card

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:45.136049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:45.136049Z digest=sha256:0ef7729cc89fdc592194f2e9f2d2a50a30753c1099722ecd5a4f322f69813b05

Observation b067a85c-c5c5-46a1-8cb9-b206d13a74bb · outbound

This paper cites Weinberger.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Weinberger

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:45.233219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:45.233219Z digest=sha256:b5e894703c852cef95cd02bbe056e474812834b610d4714d188a316336adefb6

Observation 07f4bd87-66ad-4a6f-a044-5613a91275f4 · outbound

This paper cites Multiwoz 2.3: A multi-domain task-oriented dialogue dataset enhanced with annotation corrections and co-reference annotation.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Multiwoz 2.3: A multi-domain task-oriented dialogue dataset enhanced with annotation corrections and co-reference annotation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:45.318824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:45.318824Z digest=sha256:6c54ae427ee7caff963ec15104a723e13d721486d6583aed8417c0ba1dabcdf1

Observation 830417f9-cd75-43b6-afd5-a1ca3fe3b38c · outbound

This paper cites Towards explainable temporal reasoning in large language models: A structure-aware generative framework.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Towards explainable temporal reasoning in large language models: A structure-aware generative framework

Reference 15

Resolution
verified exact
doi, observed 2026-08-04T08:33:23.343722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T08:32:45.449837Z digest=sha256:65dd4b9ede54309d943d0d308c9e20e0b0831873511a469c12fd4ea5c1c7efcd

Observation 08f3d5c5-c7f0-42cc-bb66-d06d2a73ae06 · outbound

This paper cites Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:45.556145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:45.556145Z digest=sha256:f94996d69b8ac871d522b57065c1446ae9e17538f073e43963671ab1e0a218ec

Observation 284f6816-524c-4673-b9d1-b91056d127c5 · outbound

This paper cites Beyond perfect apis: A comprehensive evaluation of llm agents under real-world api complexity, 2026.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Beyond perfect apis: A comprehensive evaluation of llm agents under real-world api complexity, 2026

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:45.656984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:45.656984Z digest=sha256:d5d6c07d81c7bbaff9a8ef6ec0d71a9aa08deb2dd9b7a2f075b1feb3b37e0a00

Observation 2e060976-f4d5-4e0a-952c-fb79fe3dd40c · outbound

This paper cites Counterfactual-consistency prompting for relative tem- poral understanding in large language models.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Counterfactual-consistency prompting for relative tem- poral understanding in large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:45.761300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:45.761300Z digest=sha256:ffecdd0651a9494250a2bc04a7603545569a53a9ccb55d3ba2e4bffd9bd89539

Observation 84088aa6-4cc8-4cc2-ac77-616a30bf0104 · outbound

This paper cites Open university learning analytics dataset.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Open university learning analytics dataset

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:45.875768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:45.875768Z digest=sha256:aa7c9015c5961c908ca9cc245318b8755b3ef8c9b09b4f9d8d80bee292336f60

Observation 26a899b7-ad39-479d-b78e-e5d8038d5f50 · outbound

This paper cites Prefix: Understand and adapt to user preference in human-agent interaction, 2026.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Prefix: Understand and adapt to user preference in human-agent interaction, 2026

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:45.975785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:45.975785Z digest=sha256:e3eddcb29c8dc5196b9e7cdfc67813e945ddb4c3442e1a4bdf9119ba7be9cf0e

Observation 66332d06-a20d-41c4-afb5-9cb5b9bed7b4 · outbound

This paper cites Ministral 3.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Ministral 3

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:46.082440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:46.082440Z digest=sha256:af3810af52f82a783d5b9696dc5863ea5bd8604ff8d4ea0dad2e4b45642f1bd4

Observation 255100c5-8742-4475-b396-e1400a260400 · outbound

This paper cites ER-Reason: A Benchmark Dataset for LLM Clinical Reasoning in the Emergency Room.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise ER-Reason: A Benchmark Dataset for LLM Clinical Reasoning in the Emergency Room

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:46.218986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:46.218986Z digest=sha256:b74ab20e02f8702f2247a3d4dad8b7f4c100ef37679a3e408488dfdf43ac7025

Observation b1e86b4c-fcf0-4065-85c7-3749efa01303 · outbound

This paper cites Mistral-Small-24B-Instruct-2501.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Mistral-Small-24B-Instruct-2501

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:46.348786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:46.348786Z digest=sha256:211777c3ddcd608921f5d6b5984b5585a916aca884d995c712aab931f8126ec8

Observation 71686b62-e7c6-459a-bd70-7a6636cb7f99 · outbound

This paper cites Time is encoded in the weights of finetuned language models.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Time is encoded in the weights of finetuned language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:46.464972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:46.464972Z digest=sha256:851d2617c24751a8e66bf5beda86a811d16e53fa32a2ae3a418177a8f627d177

Observation a706882c-f725-4e17-bd48-abea13598413 · outbound

This paper cites GPT-4o mini: Advancing cost-efficient intelligence.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise GPT-4o mini: Advancing cost-efficient intelligence

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:46.582534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:46.582534Z digest=sha256:3029b0b556b38c6d093b88131dcaaec8a8b6970c3939275458423cb07d03f637

Observation 72914927-0152-422b-bfff-28cb4effaeb2 · outbound

This paper cites Introducing GPT-5.4 mini and nano.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Introducing GPT-5.4 mini and nano

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:46.659446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:46.659446Z digest=sha256:7226f6d486aefdc80bf2fd3326165714d00b3fc1c7c3e3e19cbd006b4f91f675

Observation a40f4512-4bf9-4399-97a5-9c030786c193 · outbound

This paper cites GPT-5.5 System Card.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise GPT-5.5 System Card

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:46.815954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:46.815954Z digest=sha256:e2a8884ce4e6e554ae1d2cfc44d399ac14cff34b7535ecc35172d7202753377f

Observation d3f37d3c-eed8-47bf-8043-126e87bb8ea7 · outbound

This paper cites Patil, Tianjun Zhang, Xin Wang, and Joseph E.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Patil, Tianjun Zhang, Xin Wang, and Joseph E

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:46.878354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:46.878354Z digest=sha256:b3618b29a4d4991007030d8fb664cb172c0d477cf6c223cf10e4a6a57d78dcca

Observation e01096d4-b7c7-4943-bd83-118cb1a920f0 · outbound

This paper cites Gonzalez.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Gonzalez

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:46.926362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:46.926362Z digest=sha256:0c1a6aef94780ad02841a33b41294fb80cbe56fccd742f1985f11029518f42a8

Observation f5c0ab94-5970-4b58-ad47-ffba373e0c5e · outbound

This paper cites LLMD: A Large Language Model for Interpreting Longitudinal Medical Records.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise LLMD: A Large Language Model for Interpreting Longitudinal Medical Records

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:46.996990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:46.996990Z digest=sha256:0ccc9b59b21736c4fc497642d55c01ed62790d25163d90187bbf5dae6a0511a8

Observation 32d9be73-7da3-4619-8016-bb71a2b0e0db · outbound

This paper cites ToolLLM: Facilitating large language models to master 16000+ real-world APIs.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise ToolLLM: Facilitating large language models to master 16000+ real-world APIs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:47.081849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:47.081849Z digest=sha256:ee4f871e065bc68d5e8557a31988c0692a092d9ee786f4bc17737a96a0add528

Observation c10d19e3-2d72-4e75-9188-f1ad4d526830 · outbound

This paper cites Qwen3.5: Towards native multimodal agents, February 2026.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Qwen3.5: Towards native multimodal agents, February 2026

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:47.213183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:47.213183Z digest=sha256:5b99e115df71d7ef0e107bc67656f7e40d5d7c89613059e71d5c840461f4957e

Observation 84b448c8-e92b-47fb-af8c-57308fb56607 · outbound

This paper cites Towards scalable multi-domain conversational agents: The schema-guided dialogue dataset.Proceedings of the AAAI Conference on Artificial Intelligence, 34(05):8689–8696, Apr.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Towards scalable multi-domain conversational agents: The schema-guided dialogue dataset.Proceedings of the AAAI Conference on Artificial Intelligence, 34(05):8689–8696, Apr

Reference 33

Resolution
malformed identifier
no resolver link, observed 2026-08-04T08:32:47.333976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:47.333976Z digest=sha256:990704bf0bae933a13b2ace0443f617bbf029c199ca2a202315686e0b0cd81ba

Observation e40a3819-c6ba-4e12-a8cd-b8fdf4d9aa35 · outbound

This paper cites Appropriate reliance on ai advice: Conceptualization and the effect of explanations.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Appropriate reliance on ai advice: Conceptualization and the effect of explanations

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:47.394889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:47.394889Z digest=sha256:d1a07888ce0399a3e2f2f05f6822d880b13abab8323c0a48dbc67902650ff909

Observation 47de76a1-0fa2-4b7c-942f-79974f0ab4a2 · outbound

This paper cites Timo: Towards better temporal reasoning for language models.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Timo: Towards better temporal reasoning for language models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:47.507523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:47.507523Z digest=sha256:ba8c99854eef848a4714d2b128fad95e2abd9fd2865ee04a97b24cfe33a6b8c0

Observation 89d6230b-03b7-4c1a-80fd-b2298631812a · outbound

This paper cites Paladin: Self-correcting language model agents to cure tool-failure cases, 2025.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Paladin: Self-correcting language model agents to cure tool-failure cases, 2025

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:47.600137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:47.600137Z digest=sha256:8d7d36f7eae904e6375f5d73bbde560a04f84497da06924523f8ec8adf21faff

Observation 47e98ff6-7aba-4a92-b24f-eb9de741ed8f · outbound

This paper cites Agentnoisebench: Benchmarking robustness of tool-using llm agents under noisy condition, 2026.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Agentnoisebench: Benchmarking robustness of tool-using llm agents under noisy condition, 2026

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:47.762337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:47.762337Z digest=sha256:c2c56111595cb417279334f2ffabfeec89c6bfdffd274dce929027b1acf14afd

Observation b08e8d73-0612-4b52-9e9b-f23719076039 · outbound

This paper cites Butterfly effects in toolchains: A comprehensive analysis of failed parameter filling in LLM tool-agent systems.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Butterfly effects in toolchains: A comprehensive analysis of failed parameter filling in LLM tool-agent systems

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:47.874010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:47.874010Z digest=sha256:7c56d4047588f278dd457a4a193c653f7f7670985eb006931cd87b79f3c20157

Observation a1028e53-6f80-4a14-af3b-5d0589a0e7ac · outbound

This paper cites Reducing Tool Hallucination via Reliability Alignment.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Reducing Tool Hallucination via Reliability Alignment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:48.011938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:48.011938Z digest=sha256:32ea15024882e4969f88c71c2c2b0c59e0eba5cda64766c9c434d140f22563b2

Observation a016a723-7571-439b-88df-58fd08e72c43 · outbound

This paper cites Can Tool-augmented Large Language Models be Aware of Incomplete Conditions?.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Can Tool-augmented Large Language Models be Aware of Incomplete Conditions?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:48.154098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:48.154098Z digest=sha256:8c2319fc38c0c4fe72db7e607112d068395551311a02516641fecc623dd529ce

Observation 44df9c9e-d438-4cc8-9a43-8a7c522550a6 · outbound

This paper cites τ-bench: A benchmark for Tool-Agent-User interaction in real-world domains.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise τ-bench: A benchmark for Tool-Agent-User interaction in real-world domains

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:48.264525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:48.264525Z digest=sha256:8a8d0f2bf58a0276ccbbd45ef92c31e5cd855b8290a6d68623928e8b399678ff

Observation d1d09bbd-ee5a-40d9-ade5-c71ec30779f5 · outbound

This paper cites MultiWOZ 2.4: A multi-domain task-oriented dialogue dataset with essential annotation corrections to improve state track- ing evaluation.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise MultiWOZ 2.4: A multi-domain task-oriented dialogue dataset with essential annotation corrections to improve state track- ing evaluation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:48.388211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:48.388211Z digest=sha256:c956cea31252c6398691cf95a7d0bb5f62d81a1d5fe7721c3c54a375785fdd9e

Observation 8ff6ca94-4e10-49a4-95c4-83d46d8e6326 · outbound

This paper cites MultiWOZ 2.2 : A dialogue dataset with additional annotation corrections and state tracking baselines.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise MultiWOZ 2.2 : A dialogue dataset with additional annotation corrections and state tracking baselines

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:48.511445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:48.511445Z digest=sha256:be9a2ca10076aebcd107372ee3b9a61f37dc4d3bcdcac304881b1bb29398272a

Observation 0ee22073-1f93-49c2-bb72-19f0b9bcd373 · outbound

This paper cites From allies to adversaries: Manipulating LLM tool-calling through adversarial injection.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise From allies to adversaries: Manipulating LLM tool-calling through adversarial injection

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:48.612316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:48.612316Z digest=sha256:2c67210360b053630db9a5091b9ec8ba1aab2850b44695de715e73610bbe27f9

Observation 26473219-2caa-4716-8338-478878987eef · outbound

This paper cites CLAMBER: A benchmark of identifying and clarifying ambiguous information needs in large language models.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise CLAMBER: A benchmark of identifying and clarifying ambiguous information needs in large language models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:48.760073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:48.760073Z digest=sha256:50eea57a3570fbc835400d62c9a08eec7618e94f862f8abebb1cb52599bad206

Observation b893b0a0-8e23-4c5e-b412-4f229cb1024f · outbound

This paper cites ToolBeHonest: A multi-level hallucination diagnostic benchmark for tool-augmented large language models.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise ToolBeHonest: A multi-level hallucination diagnostic benchmark for tool-augmented large language models

Reference 46

Resolution
malformed identifier
no resolver link, observed 2026-08-04T08:32:48.911714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:48.911714Z digest=sha256:31bb197860fd2b4971047e949bcc6ffae8041064a07dabeaa9b8faa29f50390a

Observation b7205e0c-2a3a-410a-b1b7-6f0cb03fe58f · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:49.005953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:49.005953Z digest=sha256:539e47ac8403e2d3e3d7055fec6037b3e1699835e9556c4cc96e88aa02c19e25

Observation ce43522e-37e3-4a92-be93-0966cec3ac4a · outbound

This paper cites an unresolved cited work.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:49.105357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:49.105357Z digest=sha256:bc25cb860779236c672c0404b5273d4b20d45d20fc630731e32a16c6c86ac006

Observation 101634f9-4680-4160-80c2-014b0440f502 · outbound

This paper cites an unresolved cited work.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:49.193660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:49.193660Z digest=sha256:d7652bd46291bc8d65a03cb0f758d0c163e34837eea7af58f8feccc730709f72

Observation 998d2d4c-86fd-4029-90a9-bc50bc628c88 · outbound

This paper cites an unresolved cited work.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:49.300953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:49.300953Z digest=sha256:4343b595f6c5126320c4bbb555a0c3f5257636a244964fd3a3955635433708be

Observation 3363725e-05f2-415f-aa54-fe35f671204a · outbound

This paper cites an unresolved cited work.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:49.405786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:49.405786Z digest=sha256:a28bbc43c79c0cdd2b5c7ffcae602792f63fda2fa2e61bd92fd4884222d04d24

Observation 69870187-8700-4e82-b1ee-fbb3c57269df · outbound

This paper cites an unresolved cited work.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:49.476122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:49.476122Z digest=sha256:e2f15b1663920253d4e766d95b9dd7dc33bfcfbb50291f6e4f40bc7f8184944f

Observation e5e5948d-c24b-4d20-a769-f1996a9946d4 · outbound

This paper cites Failure Risk: Critical.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Failure Risk: Critical

Reference 53

Resolution
malformed identifier
no resolver link, observed 2026-08-04T08:32:49.575837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:49.575837Z digest=sha256:8e5526c6b6417ec012d97d280a95c42d979b0ecfdc060715effba833be71be86

Pith citing papers

No inbound Pith citation observations are available.