Pith. sign in

Paper Citation Record · LEDGER

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues

As of 18 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2506.22853.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22853 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:02:41.100115Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact5
  • verified fuzzy3
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 67af86c5-852a-4eab-9f36-fc04dcc5593f · outbound

This paper cites Granite-Function Calling Model: Introducing Function Calling Abilities via Multi-task Learning of Granular Tasks.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Granite-Function Calling Model: Introducing Function Calling Abilities via Multi-task Learning of Granular Tasks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:37.650566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:37.650566Z digest=sha256:288fa5bbd3e3645454194fa2cd905e60b155a8d0e9b621414ecf94a3dfc99f2e

Observation fb98c640-b9ef-4d10-b882-d412b61048b5 · outbound

This paper cites Phi-4 Technical Report.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Phi-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:37.731201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:37.731201Z digest=sha256:753c825e0c01af82b1c33368457c36e6990ae73c326892f356e74e5b3f044083

Observation 1263e47b-b410-4cc9-862a-b4c46cd4ebdd · outbound

This paper cites Can a Single Model Master Both Multi-turn Conversations and Tool Use? CoALM: A Unified Conversational Agentic Language Model.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Can a Single Model Master Both Multi-turn Conversations and Tool Use? CoALM: A Unified Conversational Agentic Language Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:37.849489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:37.849489Z digest=sha256:3d23e135900c47c682073172446e8d95190c88ff459f45fcd76839cbea498d6a

Observation ace2d624-af41-4043-99fc-7f8c256c5689 · outbound

This paper cites API-BLEND: A Comprehensive Corpora for Training and Benchmarking API LLMs.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues API-BLEND: A Comprehensive Corpora for Training and Benchmarking API LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:37.949602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:37.949602Z digest=sha256:6bf56bc197ffbb81cc535404509d91439a2606b8a0d5a45bfd68c6b0a39383b8

Observation 07f59721-6407-409e-955d-616765d4661a · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:02:45.244990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:02:38.066276Z digest=sha256:59707299ccf5b11edac57b7077ace9908ffbb65e665f580846c9e94d29c3a735

Observation ee157638-d7a2-4999-a40a-7d55726c6ca0 · outbound

This paper cites Genie: A Generator of Natural Language Semantic Parsers for Virtual Assistant Commands.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Genie: A Generator of Natural Language Semantic Parsers for Virtual Assistant Commands

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:02:43.088698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:02:38.215072Z digest=sha256:ab2e82534d7985f15d740374ca900fb2598ca7dbc3df9166ee9b35a795c59885

Observation 6512ffbd-b7dd-481e-88e9-49d88c10eb49 · outbound

This paper cites T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.371492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.371492Z digest=sha256:abc00187fa6c8cf86906ee2683918c8e5a50cc88e7a38af270f9579c112b7579

Observation eaa1891e-46d6-412f-9683-9fdaeac463ce · outbound

This paper cites TinyAgent: Function Calling at the Edge.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues TinyAgent: Function Calling at the Edge

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.481560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.481560Z digest=sha256:21567def7d0435ccb2f63af9a10fa752c29d9500797542a4b6317c2ff9593025

Observation b7820be4-1746-493c-bf47-d3385e69a83e · outbound

This paper cites ToolTalk: Evaluating Tool-Usage in a Conversational Setting.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ToolTalk: Evaluating Tool-Usage in a Conversational Setting

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.534202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.534202Z digest=sha256:61468982c239c4d4d9f6a0bd8d46d64c84218b895f6c8b8b9a6a518f33dddf94

Observation f8a22051-a2c3-4fc7-9b9b-e797ebfe2b16 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.574014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.574014Z digest=sha256:716031a4e791ecb3bcdd2e1f102afb397106b5490adafd123937e76de7dcca7c

Observation f7cbef17-dec1-4b55-b70c-7db2756b1bc5 · outbound

This paper cites Is It Really Long Context if All You Need Is Retrieval? Towards Genuinely Difficult Long Context NLP.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Is It Really Long Context if All You Need Is Retrieval? Towards Genuinely Difficult Long Context NLP

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.627939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.627939Z digest=sha256:5a36f9dbffb5e40cf60d76e3513010757e4ffacd6eddd88a3ef2b2bb36cbf667

Observation 140e6b49-c3e1-4f90-99d9-53875517b35e · outbound

This paper cites CoSearchAgent: A Lightweight Collaborative Search Agent with Large Language Models.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues CoSearchAgent: A Lightweight Collaborative Search Agent with Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.725718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.725718Z digest=sha256:0f2fe140d9527d30f9bec37abff01e01f7f2896792beac0b758d079e576093f9

Observation 64c732c0-1389-4afa-8e90-b22d0b70dc3b · outbound

This paper cites Intelligent Virtual Assistants with LLM-based Process Automation.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Intelligent Virtual Assistants with LLM-based Process Automation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.797381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.797381Z digest=sha256:0e72c62701822a541fb44097f858e8e0eca9038d0419f214f8c062b4f22da635

Observation 20756517-6596-4bec-9837-d139ff5bdd0b · outbound

This paper cites MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.861022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.861022Z digest=sha256:6b12eeade2d2de0447b025aef59d52b6f97db923b315c5f549c7387447594384

Observation 17b10d36-0521-4cdb-9bc0-99c1e690b5a6 · outbound

This paper cites An LLM Benchmark for Addressee Recognition in Multi-modal Multi-party Dialogue.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues An LLM Benchmark for Addressee Recognition in Multi-modal Multi-party Dialogue

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:02:42.730281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:02:38.931012Z digest=sha256:83b1be314043cb9b5ea5e09eea240a6362862b053a61b951ec4e6dacb2cf28ca

Observation 1890a7c7-d408-4896-9b3f-b6f43833978d · outbound

This paper cites Mistral 7B.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Mistral 7B

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.975216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.975216Z digest=sha256:1d6f073a7c561fd0966ec208ac4402bf876ef3c875b105ec3508a5809576e82c

Observation 9920ca20-b5c2-400f-a140-a0c3904cd7f0 · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:02:45.073905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:02:39.071046Z digest=sha256:85c8125aec27d026a19c7222b4c844e9c5c06c84711c07fa5b8a6f42bddbc191

Observation abb7d8da-005f-4b43-89b6-c9ed2ea968b6 · outbound

This paper cites Why and When LLM-Based Assistants Can Go Wrong: Investigating the Effectiveness of Prompt-Based Interactions for Software Help-Seeking.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Why and When LLM-Based Assistants Can Go Wrong: Investigating the Effectiveness of Prompt-Based Interactions for Software Help-Seeking

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.143917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.143917Z digest=sha256:18be152489731c9f7b0c75eb305a6d87f467bacab3603bcfb0eb161b4839c5e5

Observation cce720a0-115d-4a8e-88d7-e34a0c75fab4 · outbound

This paper cites SEAL: Suite for Evaluating API-use of LLMs.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues SEAL: Suite for Evaluating API-use of LLMs

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:02:42.398827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:02:39.212803Z digest=sha256:f91e534f4f820e5de7d94fec8ab1a936180362bf867b1cc8c668c8bb11e8ed40

Observation 6ac36575-d17c-4731-8683-e28c29e9355e · outbound

This paper cites ETHIC: Evaluating Large Language Models on Long-Context Tasks with High Information Coverage.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ETHIC: Evaluating Large Language Models on Long-Context Tasks with High Information Coverage

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:02:42.211387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:02:39.273106Z digest=sha256:3e413a42334e28d7c30997407a83b8f81f1055fb4f5ac5c68257db0ff253e76a

Observation a4d59512-941d-4138-8afd-e64b0a39d3d4 · outbound

This paper cites API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.336321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.336321Z digest=sha256:120b0f529f0a66fdc4076e3515a8fcee5f9437a8673713f1590999876847d14d

Observation 2f8e3684-0fda-440d-884d-98e4ea610dc6 · outbound

This paper cites ToolACE: Winning the Points of LLM Function Calling.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ToolACE: Winning the Points of LLM Function Calling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.409173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.409173Z digest=sha256:2050e3e1c022dbd397caec55291703179d798942d10c2991a9634fc9f3be83b5

Observation 657ffe74-1b1a-4030-b055-d8fcbe58a610 · outbound

This paper cites G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.463112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.463112Z digest=sha256:d6dfe0f311e619514edfaad2b8a01d99535424fd38beb0106be1be97134454ab

Observation 5a800961-81c6-4fc3-9866-d9d3e3ff30c3 · outbound

This paper cites Manning and Hinrich Schütze.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Manning and Hinrich Schütze

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:02:44.883790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:02:39.552815Z digest=sha256:91c27ad985c1cacad8cbf291f161305cf84d92ab0bf4ef40368df2c62b5a4490

Observation f0c6393b-284a-4b21-9170-17b734c0e26e · outbound

This paper cites GPT-4 Technical Report.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues GPT-4 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.596859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.596859Z digest=sha256:a58903028f6e32bb45004b9c7b2c6c7c06e957faf45d78265d8d5a553ee84807

Observation 2eaa2706-e189-4aef-9f14-8d5dd678067d · outbound

This paper cites Bernstein.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Bernstein

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:02:44.719098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:02:39.688152Z digest=sha256:41729f328d303e017fd9f63c1b36df8a73c0c927d871bb3b2fd66c72beb9ab25

Observation ecc55d3b-2e3a-4235-9a5c-e8ae82c90c84 · outbound

This paper cites Gorilla: Large Language Model Connected with Massive APIs.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Gorilla: Large Language Model Connected with Massive APIs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.734802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.734802Z digest=sha256:35cbb02533bdd2fd469103c2477e28ecfc0bd2128603dc83b4c6c7e56137a94e

Observation 966f52ad-265e-4f2f-9823-e73d8ad708d8 · outbound

This paper cites Tool Learning with Foundation Models.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Tool Learning with Foundation Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.781162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.781162Z digest=sha256:9fedd9f5b7b4ca436ea9ee5f0cf00243bc705d65aaca3ee5336b46310ca4cd3f

Observation d8eefb7e-6129-43c4-9966-2846c830af1d · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.832816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.832816Z digest=sha256:999ff9aa77482ea3e4371ffc2f72589cb88b9bab958c0a40470ec4515af1abf0

Observation ebd8a6f6-34b6-4430-8e74-939421a62030 · outbound

This paper cites Tool Learning with Large Language Models: A Survey.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Tool Learning with Large Language Models: A Survey

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.878785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.878785Z digest=sha256:1ed7beae38f6bbf18d2971b11bdbc53e4b67c49c9a64b3e914ccdca4ed606c7a

Observation 32c161d1-5d57-4b99-b867-84df2740b0c3 · outbound

This paper cites Qwen2.5 Technical Report.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Qwen2.5 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.913760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.913760Z digest=sha256:2b34aafdd196501371acf123359215997a4841a819aee3596c353b1276603196

Observation 194f2059-bec4-4236-aa48-eaba2aac480e · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:02:44.547233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:02:39.971592Z digest=sha256:a886151151d95d56d8289de0a5c79f887ccacc460be480cb3ad20590a3e38437

Observation 5ad492a3-1457-44b2-90e9-cdee7a3a7618 · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.010427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.010427Z digest=sha256:c521dfa961eac03800b5be2faee687d058309a0d431c71684fca00587869a108

Observation 2fa50049-4b1e-41f4-801b-d5d2a67165f4 · outbound

This paper cites Bridging HCI and AI Research for the Evaluation of Conversational SE Assistants.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Bridging HCI and AI Research for the Evaluation of Conversational SE Assistants

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.051068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.051068Z digest=sha256:ceded515bfbe48a1c9b359eb3404482247f50de80064a9d5bee02dd41fe91431

Observation 7584eeff-89c0-452b-a71c-7ac4129b577f · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:02:44.405014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:02:40.104042Z digest=sha256:35222aad0ef26d94918d18764746b09b9a07e1f4e46550ba6108e831efa40e0c

Observation ddb17d27-535d-41e0-9462-2883dd5cd10f · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.144671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.144671Z digest=sha256:d135266b24c10fc005aa45edc96888e5824ef5579b0da0b1655942b61952b285

Observation 598b52fb-ac15-4bd5-a9b2-6d0320e92621 · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:02:44.220824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:02:40.206383Z digest=sha256:049d7c2db75b5e97af83cdec88965d6341503c25776f0723d600f80482d588d7

Observation ab839f1b-689b-47f7-a359-3fff8177c7b7 · outbound

This paper cites TaskBench: Benchmarking Large Language Models for Task Automation.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues TaskBench: Benchmarking Large Language Models for Task Automation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.257193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.257193Z digest=sha256:8cd5358b0b183b7c5bc2c47525deb0ca61a2027dd7b74cde05a79bb50e99bb92

Observation da63ca7c-fa85-4fa8-903e-c1c44a098e9e · outbound

This paper cites ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.291583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.291583Z digest=sha256:4dcd6c3d3d876c8562f63652b058955403bcee2888bde9c57beab5c9d5bac8a3

Observation 4c1a3027-16e2-4f09-8e63-345342b65479 · outbound

This paper cites Language Models are Few-Shot Learners.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Language Models are Few-Shot Learners

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.342829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.342829Z digest=sha256:6edaa9782c7a6be6263690808f34b2a4f958327ed44c29ef128620832b6c4913

Observation bfc23798-d9ac-4d89-8eed-a0dc20e9b9e5 · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:02:44.037471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:02:40.389822Z digest=sha256:4b8255ab71d29008a3243f75ffd1f06f4c8b18d0d2edd18bce8419b271cee795

Observation 0c97b5fb-b823-42e4-a498-5d6d718bb907 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues LLaMA: Open and Efficient Foundation Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.421187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.421187Z digest=sha256:2fdbffff63f1bfd33b59f564edb585f6c901919d3ad0d577c674e7ea147fcb7e

Observation 56f80d99-8a0d-437b-bc30-9babfb099cdb · outbound

This paper cites GPTVoiceTasker: Advancing Multi-step Mobile Task Efficiency Through Dynamic Interface Exploration and Learning.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues GPTVoiceTasker: Advancing Multi-step Mobile Task Efficiency Through Dynamic Interface Exploration and Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.459154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.459154Z digest=sha256:181596b49bf61c40b8d34dc2bc69595b0e42afa64b7f38dbf9d6cedcef23d196

Observation 65dc56bb-2c0b-442d-b3db-66e711f76d3d · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:02:43.889198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:02:40.521310Z digest=sha256:64e16e79a62c0636a4903558073eab0053db6359278bf80218c1ce09301cd018

Observation f8e6265f-9550-45e8-9f25-2c6b399afcd5 · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:02:43.705152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:02:40.565257Z digest=sha256:b52d5bfb2e573ad6263a8e6cf15cdaf7b9f0a35a634e75401c151eb1682929b9

Observation 5909157e-e86c-44ea-9604-ca43eb156c05 · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.606739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.606739Z digest=sha256:dc7a5274c32d6140c616e3e53a17084a6b56fc1f85fc5c8289395939a0218ab1

Observation 223e56c2-0253-48b2-8fd1-41e74a31a625 · outbound

This paper cites MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.651835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.651835Z digest=sha256:3e94d5e19d88c1cf0ca2044ed7c1e97e68943e007367a2132b1e768e2af848ec

Observation 97428b19-a228-4f2d-af5d-9c1cf2c353fc · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:02:43.580860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:02:40.699899Z digest=sha256:b74b4b62fb92e031332a8c6b8a32dcd26d3f2b6f169e9f6687e766c662bb95a7

Observation 5321c14a-8d40-4db5-a641-3352980a350c · outbound

This paper cites On the Tool Manipulation Capability of Open-source Large Language Models.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues On the Tool Manipulation Capability of Open-source Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.749933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.749933Z digest=sha256:a341b531cc4fde37bd83cbfcc43fe5ec1000884fa410e23d4791a4236016d29f

Observation aa42965a-feaf-4205-9eb2-6acc7788c794 · outbound

This paper cites ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.814433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.814433Z digest=sha256:29481ff014882c94317976f38e501a7373ef5dd634cb63f999049ae713b98340

Observation 14228d02-f6a2-4b96-bd90-60acba2b90c2 · outbound

This paper cites RoTBench: A Multi-Level Benchmark for Evaluating the Robustness of Large Language Models in Tool Learning.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues RoTBench: A Multi-Level Benchmark for Evaluating the Robustness of Large Language Models in Tool Learning

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:02:41.368776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:02:40.848146Z digest=sha256:23686ef5d593e566e251466291f1bbcc7516001bf6710d9b47a88a22026770c1

Observation 23c8b628-30fe-47e5-a5ad-8bcf7a1084c4 · outbound

This paper cites Schweitzer, and Alison Wood Brooks.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Schweitzer, and Alison Wood Brooks

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:02:43.458492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:02:40.893204Z digest=sha256:ad44be19f99278ab1dc20f70a3e719c26cbabf076a9319d2eacc3f4255d7c3bb

Observation 381dc53e-59e5-4fd9-b493-50ebe7b64b8d · outbound

This paper cites A Survey on Multi-Turn Interaction Capabilities of Large Language Models.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues A Survey on Multi-Turn Interaction Capabilities of Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.944596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.944596Z digest=sha256:fad5c16922982bcee10cfabd42c2f7bc59892420f3af2c92277d46442738a4ba

Observation 3e5d6cd4-2439-4fcc-9d6c-a6c20b2f61bf · outbound

This paper cites ToolQA: A Dataset for LLM Question Answering with External Tools.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ToolQA: A Dataset for LLM Question Answering with External Tools

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.986835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.986835Z digest=sha256:f3bacac770ec01d0ff4bcd4f45890280552c32f26a229ff86bcf32502a7db422

Observation e7d363ed-31e3-499e-9178-35297422add4 · outbound

This paper cites online" 'onlinestring :=.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues online" 'onlinestring :=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:41.035264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:41.035264Z digest=sha256:144b85ade0b619ca97f084517965a3adcfd65c4639d699927d2d220d85453936

Observation 6c7cba1e-c6ed-4806-8dfb-91e3dfc4747b · outbound

This paper cites write newline.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues write newline

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:41.100115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:41.100115Z digest=sha256:81e9c048a40783a0150e814c667f2a7e98286413fb9f0cb7f9e8acde2be11047

Pith citing papers

No inbound Pith citation observations are available.