Pith. sign in

Paper Citation Record · LEDGER

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues

As of 8 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2506.22853.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22853 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:02:41.100115Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact5
  • verified fuzzy3
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 67af86c5-852a-4eab-9f36-fc04dcc5593f · outbound

This paper cites Granite-Function Calling Model: Introducing Function Calling Abilities via Multi-task Learning of Granular Tasks.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Granite-Function Calling Model: Introducing Function Calling Abilities via Multi-task Learning of Granular Tasks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:37.650566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:37.650566Z digest=sha256:f3f52a684263567b72b16178bcba3a3f794497b61e7121f1e7acde51fd185f71

Observation fb98c640-b9ef-4d10-b882-d412b61048b5 · outbound

This paper cites Phi-4 Technical Report.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Phi-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:37.731201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:37.731201Z digest=sha256:7386565871f89f6426db19f706ec6ca8f80ceff0c49f12875fd2490359132f7c

Observation 1263e47b-b410-4cc9-862a-b4c46cd4ebdd · outbound

This paper cites Can a Single Model Master Both Multi-turn Conversations and Tool Use? CoALM: A Unified Conversational Agentic Language Model.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Can a Single Model Master Both Multi-turn Conversations and Tool Use? CoALM: A Unified Conversational Agentic Language Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:37.849489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:37.849489Z digest=sha256:13f9740379aea537b83530d5c8efd1c49043d4c28ff4f3210901b3c4336eb3de

Observation ace2d624-af41-4043-99fc-7f8c256c5689 · outbound

This paper cites API-BLEND: A Comprehensive Corpora for Training and Benchmarking API LLMs.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues API-BLEND: A Comprehensive Corpora for Training and Benchmarking API LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:37.949602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:37.949602Z digest=sha256:da84581cdc0c6dbbf89f1388b87c9e87125af333cf3a3c8d962a6240beded33b

Observation 07f59721-6407-409e-955d-616765d4661a · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:02:45.244990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:02:38.066276Z digest=sha256:4e7b16fc7ac055c64e36f878fcb5a965f0408db4f27109916cf17c9603ce66b5

Observation ee157638-d7a2-4999-a40a-7d55726c6ca0 · outbound

This paper cites Genie: A Generator of Natural Language Semantic Parsers for Virtual Assistant Commands.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Genie: A Generator of Natural Language Semantic Parsers for Virtual Assistant Commands

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:02:43.088698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:02:38.215072Z digest=sha256:628fafbfbd50f94d6ef888832ae5f0302fbd4462e6073d87b4e189dd03a88086

Observation 6512ffbd-b7dd-481e-88e9-49d88c10eb49 · outbound

This paper cites T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.371492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.371492Z digest=sha256:911096c0017607a2cb4e3e38c25c6b5935371b30ceb5efbd28a3be6859bfb142

Observation eaa1891e-46d6-412f-9683-9fdaeac463ce · outbound

This paper cites TinyAgent: Function Calling at the Edge.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues TinyAgent: Function Calling at the Edge

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.481560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.481560Z digest=sha256:f4ee4012363b080496e3b67ffa29be27ecf2bd9346caad1ebc927999c51c919a

Observation b7820be4-1746-493c-bf47-d3385e69a83e · outbound

This paper cites ToolTalk: Evaluating Tool-Usage in a Conversational Setting.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ToolTalk: Evaluating Tool-Usage in a Conversational Setting

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.534202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.534202Z digest=sha256:b880024b340f3135a83ed969a52d3c712fa1a1079b0cbff89bb4e1cecc27d06e

Observation f8a22051-a2c3-4fc7-9b9b-e797ebfe2b16 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.574014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.574014Z digest=sha256:b342c1e1ee2eeb89a34c40f24686ddd7d80d88c342c6686264b9c265f53d079a

Observation f7cbef17-dec1-4b55-b70c-7db2756b1bc5 · outbound

This paper cites Is It Really Long Context if All You Need Is Retrieval? Towards Genuinely Difficult Long Context NLP.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Is It Really Long Context if All You Need Is Retrieval? Towards Genuinely Difficult Long Context NLP

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.627939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.627939Z digest=sha256:5d2dd854c06d2875815b5ab6011752c49844eb789f5de7ea67ceffd60a200092

Observation 140e6b49-c3e1-4f90-99d9-53875517b35e · outbound

This paper cites CoSearchAgent: A Lightweight Collaborative Search Agent with Large Language Models.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues CoSearchAgent: A Lightweight Collaborative Search Agent with Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.725718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.725718Z digest=sha256:cb4c790732b49e17bfd8b7bebefd52aad76161a269bc1edd4204e54c7e3049f0

Observation 64c732c0-1389-4afa-8e90-b22d0b70dc3b · outbound

This paper cites Intelligent Virtual Assistants with LLM-based Process Automation.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Intelligent Virtual Assistants with LLM-based Process Automation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.797381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.797381Z digest=sha256:12100ccbe62ace3e15819c5b3d3c45aadc1a41ab55f3fb937e1ecd9e2d32182f

Observation 20756517-6596-4bec-9837-d139ff5bdd0b · outbound

This paper cites MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.861022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.861022Z digest=sha256:2770018ac3f02c8bedbb39c7cbed7d55d346ac9b0810675d3e988c80e036fadc

Observation 17b10d36-0521-4cdb-9bc0-99c1e690b5a6 · outbound

This paper cites An LLM Benchmark for Addressee Recognition in Multi-modal Multi-party Dialogue.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues An LLM Benchmark for Addressee Recognition in Multi-modal Multi-party Dialogue

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:02:42.730281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:02:38.931012Z digest=sha256:a5a4de7cd3ff6c282079ea8a380557cfba58e06a836e8f316e4fe0a9ce01a4a3

Observation 1890a7c7-d408-4896-9b3f-b6f43833978d · outbound

This paper cites Mistral 7B.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Mistral 7B

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.975216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.975216Z digest=sha256:81ff67838878823176977012fa8cb6a80aea18a91d79fbd46158cc3ad9b0d064

Observation 9920ca20-b5c2-400f-a140-a0c3904cd7f0 · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:02:45.073905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:02:39.071046Z digest=sha256:49385ec9764e4f6787068b6f99557327332e9babd1bb71baef5a2b4f263e77b8

Observation abb7d8da-005f-4b43-89b6-c9ed2ea968b6 · outbound

This paper cites Why and When LLM-Based Assistants Can Go Wrong: Investigating the Effectiveness of Prompt-Based Interactions for Software Help-Seeking.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Why and When LLM-Based Assistants Can Go Wrong: Investigating the Effectiveness of Prompt-Based Interactions for Software Help-Seeking

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.143917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.143917Z digest=sha256:5814cb14fdca4d33b89ef2945babf399ed20ce89940a2262542a396098beb612

Observation cce720a0-115d-4a8e-88d7-e34a0c75fab4 · outbound

This paper cites SEAL: Suite for Evaluating API-use of LLMs.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues SEAL: Suite for Evaluating API-use of LLMs

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:02:42.398827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:02:39.212803Z digest=sha256:2b339b7601a0be4b62f8da5eeeade325e32b1e7bca151905391b25b251e5574b

Observation 6ac36575-d17c-4731-8683-e28c29e9355e · outbound

This paper cites ETHIC: Evaluating Large Language Models on Long-Context Tasks with High Information Coverage.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ETHIC: Evaluating Large Language Models on Long-Context Tasks with High Information Coverage

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:02:42.211387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:02:39.273106Z digest=sha256:b4c4659f5dbe2548fdfe12967acd45e4c1834a69dbaf4742c23b0f25b33c6db3

Observation a4d59512-941d-4138-8afd-e64b0a39d3d4 · outbound

This paper cites API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.336321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.336321Z digest=sha256:7ee49fbbbe739a636145082d46d2b2f951dee408be5bce428a83b0803c4a75d4

Observation 2f8e3684-0fda-440d-884d-98e4ea610dc6 · outbound

This paper cites ToolACE: Winning the Points of LLM Function Calling.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ToolACE: Winning the Points of LLM Function Calling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.409173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.409173Z digest=sha256:bfac53f0c5ebd476d94526119200f4e2257c89190c3fccaceca8822a1ebc1a35

Observation 657ffe74-1b1a-4030-b055-d8fcbe58a610 · outbound

This paper cites G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.463112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.463112Z digest=sha256:63b0654ad9f0896ac6dca19a1143584957c70074ee1c05f15d610e138d7eb86b

Observation 5a800961-81c6-4fc3-9866-d9d3e3ff30c3 · outbound

This paper cites Manning and Hinrich Schütze.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Manning and Hinrich Schütze

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:02:44.883790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:02:39.552815Z digest=sha256:b182894c7bdcda94ee60ce31cfeb0e6f7394f987c000c8a49436a031741859b7

Observation f0c6393b-284a-4b21-9170-17b734c0e26e · outbound

This paper cites GPT-4 Technical Report.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues GPT-4 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.596859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.596859Z digest=sha256:7a299836405dca3d39d5da1a2b9d1b98c0f6d6de7f49c3c27a823e6a3a0d6052

Observation 2eaa2706-e189-4aef-9f14-8d5dd678067d · outbound

This paper cites Bernstein.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Bernstein

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:02:44.719098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:02:39.688152Z digest=sha256:024f806c9f8c9561788b7c3e13f3a6dd0c085adffec4cc2f68d940fd27693ae9

Observation ecc55d3b-2e3a-4235-9a5c-e8ae82c90c84 · outbound

This paper cites Gorilla: Large Language Model Connected with Massive APIs.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Gorilla: Large Language Model Connected with Massive APIs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.734802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.734802Z digest=sha256:4f554ddcbed5c104aecaafc2949113919e2f3e09c9c7d0accbc6746a41638273

Observation 966f52ad-265e-4f2f-9823-e73d8ad708d8 · outbound

This paper cites Tool Learning with Foundation Models.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Tool Learning with Foundation Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.781162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.781162Z digest=sha256:a034ef140d7b154b9fd3d8b94ee909bd53d5c3c7b974812f3487a3f75043dbe2

Observation d8eefb7e-6129-43c4-9966-2846c830af1d · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.832816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.832816Z digest=sha256:dbabaaeaef37142542745cbe9ede27f32b5cc97e4786932ef2b8956731d4d6b9

Observation ebd8a6f6-34b6-4430-8e74-939421a62030 · outbound

This paper cites Tool Learning with Large Language Models: A Survey.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Tool Learning with Large Language Models: A Survey

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.878785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.878785Z digest=sha256:6f0b4588db47224ebed5d5a2347a9ed06ad7af0345c16a290c17d6e152e22ef6

Observation 32c161d1-5d57-4b99-b867-84df2740b0c3 · outbound

This paper cites Qwen2.5 Technical Report.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Qwen2.5 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.913760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.913760Z digest=sha256:dcc75d563161ffab3258b0943113e987a2b5847826142bccf0b79aadede3406d

Observation 194f2059-bec4-4236-aa48-eaba2aac480e · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:02:44.547233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:02:39.971592Z digest=sha256:a145a657e54c5bb488d7226b3b89905cd71c029963879df7a2fd01d1fbaa2771

Observation 5ad492a3-1457-44b2-90e9-cdee7a3a7618 · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.010427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.010427Z digest=sha256:e8d91f8e39c646b2a0f5a9a6e89338aa43bf73efb9cdc41ba2bd589d865f0a02

Observation 2fa50049-4b1e-41f4-801b-d5d2a67165f4 · outbound

This paper cites Bridging HCI and AI Research for the Evaluation of Conversational SE Assistants.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Bridging HCI and AI Research for the Evaluation of Conversational SE Assistants

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.051068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.051068Z digest=sha256:f510cc42b1a5c885c0bf49df09ab557a5022f520384190ea824e3689af93767a

Observation 7584eeff-89c0-452b-a71c-7ac4129b577f · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:02:44.405014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:02:40.104042Z digest=sha256:8ec61e85ca612d3a326a3f494080c4140ceeffa1eb169d587d51b39c40b14e29

Observation ddb17d27-535d-41e0-9462-2883dd5cd10f · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.144671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.144671Z digest=sha256:0160a2233382ee64debc8b6cd2e08b50dd18f71db2351b1926db44b67cb44715

Observation 598b52fb-ac15-4bd5-a9b2-6d0320e92621 · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:02:44.220824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:02:40.206383Z digest=sha256:5c0e7d2fdb94086be3b2033442a56ae9ba64026580cda3a94194b987d46d6015

Observation ab839f1b-689b-47f7-a359-3fff8177c7b7 · outbound

This paper cites TaskBench: Benchmarking Large Language Models for Task Automation.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues TaskBench: Benchmarking Large Language Models for Task Automation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.257193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.257193Z digest=sha256:f2cd62bdaeb4333417a24b98a36f22ad7fcd0a82f8838bb7cc116bdcea181e07

Observation da63ca7c-fa85-4fa8-903e-c1c44a098e9e · outbound

This paper cites ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.291583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.291583Z digest=sha256:588909837b049bfa9978e9cae5fa27b9d406168874702e0608f9a013f84bfc86

Observation 4c1a3027-16e2-4f09-8e63-345342b65479 · outbound

This paper cites Language Models are Few-Shot Learners.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Language Models are Few-Shot Learners

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.342829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.342829Z digest=sha256:7e98265c30f214c7943be80eafe7dabb11deeabb7d4ea5f1eabc55360b3ddfe4

Observation bfc23798-d9ac-4d89-8eed-a0dc20e9b9e5 · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:02:44.037471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:02:40.389822Z digest=sha256:aeb5756e88f1132e7dec1024aab2b9dc3f3881f25fabbcfe45d66333622189ad

Observation 0c97b5fb-b823-42e4-a498-5d6d718bb907 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues LLaMA: Open and Efficient Foundation Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.421187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.421187Z digest=sha256:633b45e858d6fa0fc031974790cf6be025fc7d0fa14e9bc3fd9f59070dda4a2b

Observation 56f80d99-8a0d-437b-bc30-9babfb099cdb · outbound

This paper cites GPTVoiceTasker: Advancing Multi-step Mobile Task Efficiency Through Dynamic Interface Exploration and Learning.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues GPTVoiceTasker: Advancing Multi-step Mobile Task Efficiency Through Dynamic Interface Exploration and Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.459154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.459154Z digest=sha256:67c39a22b456bfae914e0c20845c28e80cbf296cb984deeea29284bd41cc887b

Observation 65dc56bb-2c0b-442d-b3db-66e711f76d3d · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:02:43.889198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:02:40.521310Z digest=sha256:43b04e5ba12df7b50f36a8b42d6732525305d2557ec2e42c449a19f8da5b2aca

Observation f8e6265f-9550-45e8-9f25-2c6b399afcd5 · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:02:43.705152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:02:40.565257Z digest=sha256:db5fa4dcd97c9099abbd31e22e755fce368a08bc5d539bfb84e3d93793695df6

Observation 5909157e-e86c-44ea-9604-ca43eb156c05 · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.606739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.606739Z digest=sha256:d22d786d714ef42241498ef1d429c51a850302278146c9e26abc7ed2c0328ca7

Observation 223e56c2-0253-48b2-8fd1-41e74a31a625 · outbound

This paper cites MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.651835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.651835Z digest=sha256:af3d3a059720dbf25730a766fba87c1d4181fcd88106b0bbb30f25fd4fa9df14

Observation 97428b19-a228-4f2d-af5d-9c1cf2c353fc · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:02:43.580860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:02:40.699899Z digest=sha256:f28a91935e0c62a40232c354a983353bf79989b488a87b9132346f791d5496f4

Observation 5321c14a-8d40-4db5-a641-3352980a350c · outbound

This paper cites On the Tool Manipulation Capability of Open-source Large Language Models.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues On the Tool Manipulation Capability of Open-source Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.749933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.749933Z digest=sha256:5fa1d9aaa953e85a23be169017ad8cee2ec0c86f81333e25f2de0a79f300e844

Observation aa42965a-feaf-4205-9eb2-6acc7788c794 · outbound

This paper cites ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.814433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.814433Z digest=sha256:4bcef452ab2cd9a5b748c3806933a2962ad644436fd72eb4f31b18f6b71150f2

Observation 14228d02-f6a2-4b96-bd90-60acba2b90c2 · outbound

This paper cites RoTBench: A Multi-Level Benchmark for Evaluating the Robustness of Large Language Models in Tool Learning.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues RoTBench: A Multi-Level Benchmark for Evaluating the Robustness of Large Language Models in Tool Learning

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:02:41.368776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:02:40.848146Z digest=sha256:78893ea06aaa63a96639df727ee5102498c1f0c2c314e51687f3aa97115a11e7

Observation 23c8b628-30fe-47e5-a5ad-8bcf7a1084c4 · outbound

This paper cites Schweitzer, and Alison Wood Brooks.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Schweitzer, and Alison Wood Brooks

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:02:43.458492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:02:40.893204Z digest=sha256:85df115ffbf12c2e88ce9a1b04871f941d70564408760c90205f8834090720df

Observation 381dc53e-59e5-4fd9-b493-50ebe7b64b8d · outbound

This paper cites A Survey on Multi-Turn Interaction Capabilities of Large Language Models.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues A Survey on Multi-Turn Interaction Capabilities of Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.944596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.944596Z digest=sha256:70b81910858eef864d3138651f224f85c96f88812a3371f9a792773c4618eb61

Observation 3e5d6cd4-2439-4fcc-9d6c-a6c20b2f61bf · outbound

This paper cites ToolQA: A Dataset for LLM Question Answering with External Tools.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ToolQA: A Dataset for LLM Question Answering with External Tools

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.986835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.986835Z digest=sha256:dea8629801f3c7de4884a42e48f78f2e9b883256a9383c9b36c4092d914888b6

Observation e7d363ed-31e3-499e-9178-35297422add4 · outbound

This paper cites online" 'onlinestring :=.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues online" 'onlinestring :=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:41.035264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:41.035264Z digest=sha256:3a88695ce5c2c5f6530b21cd852ba5221e1b0c1857f99af21589e347d2873ad8

Observation 6c7cba1e-c6ed-4806-8dfb-91e3dfc4747b · outbound

This paper cites write newline.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues write newline

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:41.100115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:41.100115Z digest=sha256:a25a0709d095ac1ddc90f85620b1bf5462d0f0fefb1ef6ad267fb9f98ab8507c

Pith citing papers

No inbound Pith citation observations are available.