Pith. sign in

Paper Citation Record · LEDGER

MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2309.10691.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.10691 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:12:01.390439Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

16
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cc51e12d-9b0c-4636-b45d-2e8181925a91 · inbound

MORTAR: Multi-turn Metamorphic Testing for LLM-based Dialogue Systems cites this paper.

MORTAR: Multi-turn Metamorphic Testing for LLM-based Dialogue Systems MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T11:22:56.721158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:22:56.721158Z digest=sha256:09d0a4da5b80471110bc6f2fc325da968828592a575130a87a4379ee7a54d597

Observation fefc07c8-f462-46ea-bdda-247b2719279a · inbound

Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks cites this paper.

Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 276

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:01.390439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:12:01.390439Z digest=sha256:72a6d43a065c21a5f29620c8fae1e46e57912ba68679487ef5f44793d882a3e8

Observation dd0d9f5f-27ce-4fbb-81b7-6f27b5f70e55 · inbound

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production cites this paper.

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-22T15:44:58.119606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T15:42:05.266854Z digest=sha256:7e0fad233d8c3079377354c11254fd5423ee367edb13fb7fbf93f676cd260001

Observation 2099b230-e2c8-494b-8a8f-8fc1a85b6b63 · inbound

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges cites this paper.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.362158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.362158Z digest=sha256:a897e63f3377b599c4c89209cc0789e3c85e4007b699b98e01b04cddf9e7a036

Observation dbc2158c-70d8-4b05-a080-05b7abb71823 · inbound

ChemGraph: An Agentic Framework for Computational Chemistry Workflows cites this paper.

ChemGraph: An Agentic Framework for Computational Chemistry Workflows MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:08.786395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:08.786395Z digest=sha256:10d480e7d03234ff5c68f3e9ce6e9d308afde9a2dac2b0ebf70e46f86ec25169

Observation fe819c0f-4e05-4e74-b792-821941c5a143 · inbound

A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy cites this paper.

A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T04:52:00.986893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:52:00.986893Z digest=sha256:a48da68ac76143d231b3f57b20e61567076f871cd478081f4ceb7cab510c798d

Observation be19d469-391a-47f3-8e42-b4a525f8d81e · inbound

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey cites this paper.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 190

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.896513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.896513Z digest=sha256:c5cfbf6996f0ac5efb5970b30b99e1baceaf9118faf7fb4953e52fe3eb0351d2

Observation 27374b34-1684-49b9-9c3e-101ca3db1365 · inbound

Automating Financial Statement Audits with Large Language Models cites this paper.

Automating Financial Statement Audits with Large Language Models MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:15.549331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:53:15.549331Z digest=sha256:43ea55384c566f7328ca9937efe33c6e5d8244c074fe909ad05b8abe665082ac

Observation 1b717073-e197-45fc-ad3f-e3a5e158c86c · inbound

DrafterBench: Benchmarking Large Language Models for Tasks Automation in Civil Engineering cites this paper.

DrafterBench: Benchmarking Large Language Models for Tasks Automation in Civil Engineering MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T17:11:33.977624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:11:33.977624Z digest=sha256:968a9d103729da7ca4b847077dc6461a5df60730102bb9c70de910125b07cce1

Observation 7e0cb956-51fa-47cc-be48-1410d9573aba · inbound

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning cites this paper.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:04.918820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:04.918820Z digest=sha256:cd60912028cd1077f99aa321f55cf482cce09e5d8fb256c12bc9f5e275c016b4

Observation 18e33ee5-6707-45b0-b00b-a5fa0cd1ceac · inbound

GEM-Bench: A Benchmark for Ad-Injected Response Generation within Generative Engine Marketing cites this paper.

GEM-Bench: A Benchmark for Ad-Injected Response Generation within Generative Engine Marketing MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T15:55:33.882467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:55:33.882467Z digest=sha256:9b4968dafb170541912f07dada51e9f6042ca55b2add8ea1d100a85228295268

Observation 2dce9817-e5db-4adb-a203-2e56ebe70506 · inbound

An Empirical Study of Testing Practices in Open Source AI Agent Frameworks and Agentic Applications cites this paper.

An Empirical Study of Testing Practices in Open Source AI Agent Frameworks and Agentic Applications MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:12:39.816366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T14:12:08.776876Z digest=sha256:947d0d9f7d510bf1c3553ac37c2addd5abea41ae4a0d3e933dd42db6f44226df

Observation 54ca4c07-5857-4c29-adaa-0985a18cbca1 · inbound

Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live cites this paper.

Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:00:39.797394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T01:58:23.234348Z digest=sha256:4dbc68663e428ae8eed0040092dccaff503ef1d930593382538738a9334629ba

Observation f4751383-d045-463c-be3b-321ec00feebb · inbound

Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live cites this paper.

Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T00:19:21.129190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:19:21.129190Z digest=sha256:84d04d7f14e1d18322980a851332ebb2d9c9fbe2d6a503cf71cbf810c104f432

Observation 0164b675-c62d-424f-afc9-63de5f2c3494 · inbound

Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models cites this paper.

Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T11:44:07.135108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:44:07.135108Z digest=sha256:480d8a67688d0f74941661e358466ac76ec3604aeded939f77afab4dd2590899

Observation a29eb7da-c3fa-4131-9675-0118d26be51c · inbound

LLM-Based Agentic Systems for Software Engineering: Challenges and Opportunities cites this paper.

LLM-Based Agentic Systems for Software Engineering: Challenges and Opportunities MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T10:30:41.526303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:30:41.526303Z digest=sha256:cd6e46c6061e2e8da15b1a086cd0f67d8c0e32492957e6b9cc0ba9e737398c99

Observation d32cf886-8fce-4c05-b174-d866c7b8c931 · inbound

Qualixar OS: A Universal Operating System for AI Agent Orchestration cites this paper.

Qualixar OS: A Universal Operating System for AI Agent Orchestration MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:00:55.622769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T18:44:25.389675Z digest=sha256:6ac25d26cec67f0b33e6d451f7f97330214934244598aedc5c276a0a3bfe466b

Observation f05026b1-8666-411f-8889-4c4da5e839d3 · inbound

Evaluating Temporal Consistency in Multi-Turn Language Models cites this paper.

Evaluating Temporal Consistency in Multi-Turn Language Models MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:36:12.942416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T11:37:10.330674Z digest=sha256:2915ff29765aae9dac0e135428d5274f9a4324fd775c2b03cf5cd342a6c34008

Observation fc8facb9-335a-4d59-a160-5a2a4d7691a5 · inbound

AgentFloor: How Far Up the tool use Ladder Can Small Open-Weight Models Go? cites this paper.

AgentFloor: How Far Up the tool use Ladder Can Small Open-Weight Models Go? MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:21:10.587173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-09T20:09:11.566825Z digest=sha256:63732a51717356ed1ca9a504e867ea861eaafd8a709b3d76c24f4f7d7eed3010

Observation 29dad227-b50f-4304-88af-6fe780563325 · inbound

To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling cites this paper.

To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:36:10.316387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-09T19:32:57.054584Z digest=sha256:d668b57892c4cd4902cf22ecb81519c752a630500df993fd94a0ac3f01571f2a

Observation 3ac429c7-71f4-4082-8f33-8e3c45480718 · inbound

A Language for Describing Agentic LLM Contexts cites this paper.

A Language for Describing Agentic LLM Contexts MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:21:09.182239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-09T17:18:16.256307Z digest=sha256:cbeb8985b6560876a47c95847ea73e5ee640b3731085e24dcbfa4b1377e3f592

Observation 077c04f4-2078-4930-8337-d5cb9dfc6c8d · inbound

Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability cites this paper.

Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:41:21.833327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T04:41:15.286881Z digest=sha256:bdf1f5f4f6cbe210e242c295f58e21058f00838680b3d492ca297577977ce60c

Observation 310ebe67-b272-46ba-a4de-a4da00752f82 · inbound

The Scaling Laws of Skills in LLM Agent Systems cites this paper.

The Scaling Laws of Skills in LLM Agent Systems MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:13:37.568810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T18:10:08.737710Z digest=sha256:3fef9ee9bf038a52abb5e01b19b08ae2686da22188bffc4936ff854c0c28b749

Observation 23514e88-d818-4ff8-81da-582867fa3c7a · inbound

Interactive Evaluation Requires a Design Science cites this paper.

Interactive Evaluation Requires a Design Science MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:58:14.000188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T10:55:08.135630Z digest=sha256:00e4e84c9babe6945267af137308f5344bf994bc0ea274cd6a9ce46e0b37d09d

Observation 90999811-e646-45b4-a6bc-ed75fca53284 · inbound

SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science cites this paper.

SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:38:12.096119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T10:36:09.724234Z digest=sha256:9e5534d3bd8e708d00499ff30454eec2519be807ff8a83884d02b2adafb2138f

Observation 63860cd3-a857-4659-a905-39e5599c0da3 · inbound

SeDT: Sentence-Transformer Decision-Transformer Conditioning for Multi-Turn Conversation Reliability cites this paper.

SeDT: Sentence-Transformer Decision-Transformer Conditioning for Multi-Turn Conversation Reliability MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:53:51.659118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T18:45:40.697579Z digest=sha256:4ba4b6b0bf5efca1343493897cb1f2c5b21d17f54c40778a5b2875e40d03d2c3

Observation 6cb1c69f-d869-4e53-bca4-624b15c3ff28 · inbound

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns cites this paper.

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:19:23.870272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T19:47:59.090249Z digest=sha256:d003b593a3265cb4b6176cb936b33bc180190a8a19812dc2cf25e6237346ad36

Observation 3fa3b650-be81-43d9-b8f2-967f5167901d · inbound

What Drives Interactive Improvement from Feedback? cites this paper.

What Drives Interactive Improvement from Feedback? MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T12:05:43.465485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-01T02:36:20.649687Z digest=sha256:340197811990e3c04439994b0a3c4524dee65587a228e87fce762c34ada25460

Observation 7761b0a8-c03b-43d6-9ee7-f924f4de7fff · inbound

DataClawEval: A Benchmark for Data Engineering Agents in Real Industrial Harness cites this paper.

DataClawEval: A Benchmark for Data Engineering Agents in Real Industrial Harness MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-31T19:54:08.546745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T19:54:08.546745Z digest=sha256:7a0d5eb8dfe1e71b92f74937f3c6e068ff9d8991138115e24f35af52c44987eb