Pith. sign in

Paper Citation Record · LEDGER

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems

As of 10 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 6 inbound Pith citation observations for arXiv:2501.11067.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.11067 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T18:43:53.758609Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:44:21.636050Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T10:58:14.019840Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2ba4f3c6-1653-4bd7-afc4-d98a3676eadb · outbound

This paper cites Abbasian, I.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Abbasian, I

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:43:54.519211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:43:53.585622Z digest=sha256:2c8b8e71f738fcc7296dc6e55c5689e5f70f5c21a1d490ac87d390822d3e61bf

Observation 8e19b320-5490-44c1-8ef8-1c6a7d42025a · outbound

This paper cites Alberts, B.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Alberts, B

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:43:54.503628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:43:53.591274Z digest=sha256:f21e07f3be1213739bb401afd09f12bd04add72f02b572ff7899abb16afc7dd1

Observation 120dbcbf-2850-4000-904c-80afd79c3454 · outbound

This paper cites Automated test generation to evaluate tool-augmented LLMs as conversational AI agents.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Automated test generation to evaluate tool-augmented LLMs as conversational AI agents

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T18:43:53.601467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:43:53.601467Z digest=sha256:401f5e46a3b393717cbe61b7c1165b582a6c25e77a68cb70fe5a54ca2b6067b9

Observation ff441c3c-d61e-4cf7-ad9f-4c9654b71a64 · outbound

This paper cites Banerjee, P.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Banerjee, P

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:43:54.484345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:43:53.606452Z digest=sha256:5db120dffe87770e6f9fdc5af07fe8c79808396b762715cc67391d6344c344e8

Observation 864af797-095f-4193-b175-c7128ecd25c5 · outbound

This paper cites Betker, G.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Betker, G

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:43:54.463365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:43:53.610857Z digest=sha256:2a8fb510720513dba4b42b2331dd0a767729a11e2a8fd838fe60b41b55544bb0

Observation d74ba6b3-ff60-4065-8a6e-9d7813c0634e · outbound

This paper cites Beyond Prompts: Dynamic Conversational Benchmarking of Large Language Models.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Beyond Prompts: Dynamic Conversational Benchmarking of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T18:43:53.615356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:43:53.615356Z digest=sha256:11dc687750de51e1f3d81f2ecbf724078269d5920b42c65f3bb3162c5999a938

Observation 61e06fcb-7f7a-4d5d-9385-cd696583e928 · outbound

This paper cites an unresolved cited work.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:43:54.443124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:43:53.620346Z digest=sha256:0d41876908c0e26bab0e08d1908f13970b823962db783da6c4876f99104dc212

Observation b049339d-d601-4158-ba9d-c8d33a785c1a · outbound

This paper cites an unresolved cited work.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:43:54.424920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:43:53.625397Z digest=sha256:36b772b9a3aaeb23252f675a2a3c40d478d483e55bea3d02759a97f92a73cabc

Observation fa874ec5-b020-4832-bf09-0fe198584fcb · outbound

This paper cites an unresolved cited work.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:43:54.405510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:43:53.629554Z digest=sha256:2a266516abf9406209218abf6eec843d1482737e220a1697be039ec59a072727

Observation 9b3e003c-402f-4947-8683-5c397b4e3539 · outbound

This paper cites From Local to Global: A Graph RAG Approach to Query-Focused Summarization.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems From Local to Global: A Graph RAG Approach to Query-Focused Summarization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T18:43:53.633747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:43:53.633747Z digest=sha256:57098e2585e0d2a81df8253b986dd75c0521da508e932c9537670be836a6100a

Observation 01d5335a-76e4-4e3c-86dc-a0e35381fdef · outbound

This paper cites Ragas: Automated Evaluation of Retrieval Augmented Generation.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Ragas: Automated Evaluation of Retrieval Augmented Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T18:43:53.638756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:43:53.638756Z digest=sha256:a5ef1a5643245dd3f5c9fe6ee1c1b109b087c2a2321125a01173e0d798c92bff

Observation 36365a3f-426d-433c-a472-f8802ac8287a · outbound

This paper cites Textbooks Are All You Need.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Textbooks Are All You Need

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T18:43:53.643993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:43:53.643993Z digest=sha256:83dc820355434886bda8379269d8f708b04cdb3d620be13aaa2028c4a5c83ce0

Observation 64c6518c-741b-4a71-b6d3-7fa10a004813 · outbound

This paper cites Honovich, T.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Honovich, T

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:43:54.372346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:43:53.648623Z digest=sha256:8d342ba81c932f364058aafed2245da7af8d629b6a6be037a31a31a911606616

Observation 94f1e61d-d4f0-4048-af9a-6dde17e90960 · outbound

This paper cites Huang, A.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Huang, A

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:43:54.350687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:43:53.652863Z digest=sha256:c590ce5185f6e0db65cbb8fc31b8bd9184ac33bfa72ac3ca0c217c6e6bf36e9e

Observation 1265742e-3035-4112-a623-6c0a5fbcc07d · outbound

This paper cites an unresolved cited work.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:43:54.332770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:43:53.657300Z digest=sha256:bc94e636a006d0d9035ffbe484009e22339ff1403a9a47baf052b2edfd6fd46b

Observation e3ecfb34-d7d3-4499-958f-bde26121fc88 · outbound

This paper cites an unresolved cited work.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:43:54.306040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:43:53.666256Z digest=sha256:127251f085b65e9c44d3df104ea2e164d80f4ea17d5ba5e0ce073a7c2c675555

Observation 488ce502-6b20-4306-870b-e4c6c61fdb25 · outbound

This paper cites an unresolved cited work.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:43:54.287378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:43:53.670747Z digest=sha256:0845269cd3f9996da78636d95b93819d97f40317c1c66d109cbebce61a57179c

Observation 60f0898c-7fb5-4aa1-afed-a6897b43d296 · outbound

This paper cites WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T18:43:53.675025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:43:53.675025Z digest=sha256:2b679fa8b331aff266393da1ac800c3030dffdcafe25ed9edf69a2ef2e36e75b

Observation 49aa8bb6-05a1-417e-a40a-f159e73c3385 · outbound

This paper cites an unresolved cited work.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:43:54.269950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:43:53.680513Z digest=sha256:43031f0ab6713d4685433ce527eb0cc057c20bef4458fcfaa341e8dcc8faf42f

Observation 3d32d8a9-f40d-489c-bbc1-35ab0242d00c · outbound

This paper cites Mehandru, B.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Mehandru, B

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:43:54.243261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:43:53.685350Z digest=sha256:ce8c8236c25743ff297d22f144d608967f759e544ec3e7141efb7f1ae387aa3a

Observation d04fcc82-1da6-4221-98a9-ce0c1b7707df · outbound

This paper cites ask me anything.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems ask me anything

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:43:54.226019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:43:53.690380Z digest=sha256:b3fda42609bde2aad9463145227c642c8fd2c329e66e1b1ebd5ce86c26bd290d

Observation 1f7f464f-9d8a-4775-93b9-3c8f27541608 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Code Llama: Open Foundation Models for Code

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T18:43:53.695662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:43:53.695662Z digest=sha256:2fe6242d466c5f3e4e0684d2ffba6a9389f712dab43ce5b875e363d5dd22845a

Observation daa89ba4-b4a8-4baa-9627-b0be7f0553fa · outbound

This paper cites an unresolved cited work.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:43:54.209531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:43:53.701862Z digest=sha256:a6a114eb81cca69d8c85ecbf2311ab51a48444642569923c09c52ebae2138ec3

Observation f9da4632-408c-4d34-a357-239a4c69b480 · outbound

This paper cites Improving Text Embeddings with Large Language Models.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Improving Text Embeddings with Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T18:43:53.707147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:43:53.707147Z digest=sha256:4cc27b14687a0d099806ab1b9ff1f46e50ecebf3c7027cff7e2f04a107356045

Observation 0b9b9dba-1484-4a94-bfb2-aba6f1354e9e · outbound

This paper cites an unresolved cited work.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:43:54.186295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:43:53.712319Z digest=sha256:bd8d411dae0cca4ddd92c9c74e568b93de9e45d176d6477997976f7bbcd74dfb

Observation 3da4173e-3865-4609-b6e3-fb97e98a4151 · outbound

This paper cites an unresolved cited work.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:43:54.162347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:43:53.717373Z digest=sha256:a9463b48094661015654de892fa2b22d7ea1f5c8e5931825f16bc41977574344

Observation 5d967670-e111-45fe-b511-3f2f4a436c41 · outbound

This paper cites an unresolved cited work.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:43:54.138827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:43:53.722808Z digest=sha256:c02bb8c7081eff3c66473276b602f6631121c9a102d3d654f228022b51ef3492

Observation 40742cb0-949e-4874-af72-80a351e560f8 · outbound

This paper cites Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language Models.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T18:43:53.727559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:43:53.727559Z digest=sha256:7cece1a57e4db379855b9ff24b6db08bcbb0796db19ab3915923f4f48aa5b4e4

Observation 14f01197-372c-4fd4-8b67-f22acea2ab26 · outbound

This paper cites an unresolved cited work.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:43:54.117460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:43:53.732284Z digest=sha256:ff855231ce05138a1221b5c043b159bf0e0fa26f5418abc4e0eb454d3f8b619d

Observation 9eb16bd0-ee2d-45ca-bfb8-b2a633bf0965 · outbound

This paper cites an unresolved cited work.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:43:54.100916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:43:53.736454Z digest=sha256:1b0a1306cf761407856542f2bdd70cf03cd8cb572dbfa8e3cd19ce2b6f663c96

Observation 6a5161df-fd22-4128-9b46-865d102df32d · outbound

This paper cites an unresolved cited work.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:43:54.084321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:43:53.740603Z digest=sha256:01dae5fd0e88b1547f499dad3ae1e5e4ec743956723ac8b9bd2a58a1956deb36

Observation 0de91527-b830-48fc-8c69-50cfb4d73481 · outbound

This paper cites Decoding Data Quality via Synthetic Corruptions: Embedding-guided Pruning of Code Data.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Decoding Data Quality via Synthetic Corruptions: Embedding-guided Pruning of Code Data

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T18:43:53.744625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:43:53.744625Z digest=sha256:2b410a0109cc3111a46a329a5815e6f59a5778ce734e589408984122bc6177d9

Observation 9a10d60b-cd69-4c72-92e0-2e03143d1744 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T18:43:53.749066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:43:53.749066Z digest=sha256:7d6a68f61e8e86b708d7ff72efaa8b8a104327c51f3b4318ca0be7ffeab9dec8

Observation 5000dc8e-e6bf-4ded-beee-a4497e09c9c3 · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T18:43:53.753467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:43:53.753467Z digest=sha256:f374e56101c50d9d39fc7a7d681e5dbd72ebefa0fd62ee3748717320c994a070

Observation 3fec2bc5-ec90-40d2-bac9-07c4bc745788 · outbound

This paper cites emma_jones_4829.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems emma_jones_4829

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:43:54.065905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:43:53.758609Z digest=sha256:2b6371457bcc2f7bb7c93c7b93aa9b246db89327bb0fd0f5cead45f75d72b076

Observation d89cc873-7c03-45fc-90e6-c5305e40d624 · outbound

This paper cites an unresolved cited work.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T18:43:53.661698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:43:53.661698Z digest=sha256:39c2dc00fe01facddcd72ec4ab41e63c8cd26333422331f864a3dd8d2c304a71

Observation f286962b-0d7c-4bec-9432-9990d2373b5b · outbound

This paper cites CURATe: Benchmarking Personalised Alignment of Conversational AI Assistants.

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems CURATe: Benchmarking Personalised Alignment of Conversational AI Assistants

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T18:43:53.596538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:43:53.596538Z digest=sha256:fe0df5ad2117bd35929c78be4c27b56abb52f8c2bc1a3ce1b5691b7621b5da55

Pith citing papers

Observation 86576c80-50e8-48f5-b106-d7eb39242f96 · inbound

$\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment cites this paper.

$\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:52:17.378438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:52:17.174347Z digest=sha256:4845e7cd3ff7d2f04c029b0c346826d16d956e96cf3dd749bf12f728f7929adb

Observation a7a1f55d-5608-463f-ba76-40eaeb9259b8 · inbound

Evaluation and Benchmarking of LLM Agents: A Survey cites this paper.

Evaluation and Benchmarking of LLM Agents: A Survey IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.636050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.636050Z digest=sha256:edca6be36cb803b0c8239ff350008ea2e2c3ca24d6a1168c75cfc113bc5e3fb4

Observation 988949c2-74aa-4222-9d9f-bdbecac438b3 · inbound

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory cites this paper.

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:13:15.534657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-14T23:13:15.016486Z digest=sha256:22f47a476c8b5edff6e925a754a29dcc3c3fee41e34a3b5e2914b5de9d06b17b

Observation 2c5c510e-7a6e-4cbd-98c4-77944c47aee0 · inbound

AgentCE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments cites this paper.

AgentCE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:30:49.458088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T19:07:46.077831Z digest=sha256:8d9986205e7beb65ae6783abdaf67f8edff702b15e90856e6c1e028674f62cb5

Observation 83caedcb-f771-4e56-98fc-588ed22b31f6 · inbound

MANTRA: Synthesizing SMT-Validated Compliance Benchmarks for Tool-Using LLM Agents cites this paper.

MANTRA: Synthesizing SMT-Validated Compliance Benchmarks for Tool-Using LLM Agents IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:06:12.706281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T10:18:08.444296Z digest=sha256:2a70a26a00e8ee531caaac2f49578c75a16185345a4f86c19911f275bd519016

Observation 09cbdbfa-b8e1-48e5-93ca-fd187250d11e · inbound

Interactive Evaluation Requires a Design Science cites this paper.

Interactive Evaluation Requires a Design Science IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:58:14.021687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T10:55:08.135630Z digest=sha256:a74f1e1a676b97c8ccd156fdc8a6abc646f775f51decc75bb7375bf144674c7b