Pith. sign in

Paper Citation Record · LEDGER

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents

As of 20 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2607.06873.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06873 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T00:02:58.383840Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact9
  • verified fuzzy50
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8c9d1e7c-c348-422f-946b-62e9568cbc46 · outbound

This paper cites ReAct: Synergizing reasoning and acting in language models.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents ReAct: Synergizing reasoning and acting in language models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.388221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:ebe16b281725f942faa2e525107f1e668b90e53b5d02aa41d79b10cec6fc8ada

Observation 72f10df0-056f-4523-8e0e-6aca88842e4b · outbound

This paper cites Toolformer: Language models can teach themselves to use tools,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Toolformer: Language models can teach themselves to use tools,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.389906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:b970055bf418b998d4df0fb989d964f411975759488c42dd8f0f9b14d174c621

Observation 4ea94b3d-f9b9-4888-9852-dd8a529e6290 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:06:37.954013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:18c60aa9ee501667715f592525fb80462e1a378183c5b04a7458e10164494b04

Observation 4a4dfd04-19c2-4c43-a2fe-eafb3ce7c78b · outbound

This paper cites Preventing repeated real world AI failures by cataloging incidents: The AI incident database,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Preventing repeated real world AI failures by cataloging incidents: The AI incident database,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.393408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:c701e7745abe45fc4e6a209fe7450f3cc0638eb827a075854cb8127918242cce

Observation 4ffa99de-f735-4887-b6f9-e719cb61aae6 · outbound

This paper cites RealHarm: A collection of real-world language model application failures,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents RealHarm: A collection of real-world language model application failures,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.353168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:671c43a2fab203b4f82e8ba68dbafc59a7da0a71654acd45c36e251e2bfdbef2

Observation f27b15c7-3df5-4d10-bf94-f07e62d6107d · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:06:37.972449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:186e415337ca417d4deb8884627e68e4614ee3cb7fc4117bd77fba437f9ac5eb

Observation e592fa41-3f52-480f-a16a-9c01898fae04 · outbound

This paper cites WebArena: A realistic web environment for building autonomous agents,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents WebArena: A realistic web environment for building autonomous agents,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.384707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:3a13ed49a2f0b12a69e9784ac4c9899ca842ff02d24ca9248f531e662ca66390

Observation 0f3b2126-34aa-4a65-826f-ade350a7d323 · outbound

This paper cites GAIA: A benchmark for general AI assistants,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents GAIA: A benchmark for general AI assistants,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.381201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:2a353638cc5c857c0b0ab565719eafbe7ae90f9b314f358b23974220a6048755

Observation 99a03855-90bb-4410-81dd-38744782f243 · outbound

This paper cites AppWorld: A controllable world of apps and people for benchmarking interactive coding agents,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents AppWorld: A controllable world of apps and people for benchmarking interactive coding agents,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.377808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:3218fc6cf530bb9a14b132a0c515dcf9b6a0c25b4dd6cb6a8ac6d6a300582920

Observation c9221d46-79f0-4923-8502-f38c9021ec18 · outbound

This paper cites SWE-bench: Can language models resolve real-world GitHub issues?.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents SWE-bench: Can language models resolve real-world GitHub issues?

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.401263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:3af4626867dd0b4d373e96a4e956fe6de6f3d5148c0c9c849051e099444785a5

Observation 1f0725ff-4b9e-447f-b1c5-f27a6c1c8728 · outbound

This paper cites an unresolved cited work.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-07-10T00:06:38.398826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:c827db396d442b2d27ef00a532efea9abc8aae06e83b527e9d7a3937189b81ce

Observation 63c1db74-4e01-44c4-8c8c-8f522b8203c8 · outbound

This paper cites SpecOps: A fully automated AI agent testing framework in real-world GUI environments,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents SpecOps: A fully automated AI agent testing framework in real-world GUI environments,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.347539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:b2f223d2ef03825eb593d03a24167e5dafa69e36beba372d6d6660412ecda8c5

Observation daea19d1-e896-4fbd-9314-21ce8a47e832 · outbound

This paper cites STELLAR: A search- based testing framework for large language model applications.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents STELLAR: A search- based testing framework for large language model applications

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-10T00:06:37.967164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:3536cc569e962bd8fc041bc9fe62c069a4b7b348e38138c3011e497b36e60724

Observation 6143456a-537e-4313-8166-7a7f273e6ff7 · outbound

This paper cites an unresolved cited work.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-07-10T00:06:38.420670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:c57fd225586490595450ce336f74017e4b9d71388668e7de8d80a1bb5e75b17c

Observation aec1170e-470f-4be9-8296-7486691ec429 · outbound

This paper cites A practitioner’s guide to process mining: Limitations of the directly-follows graph,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents A practitioner’s guide to process mining: Limitations of the directly-follows graph,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.422481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:427ac49d098334816f4ba684d6d356157ea71b477d403acffe08b012740776c4

Observation 57d39246-2796-4bff-b036-68890eadf922 · outbound

This paper cites τ 3-bench: From text-only to multimodal, knowledge- aware agent evaluation,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents τ 3-bench: From text-only to multimodal, knowledge- aware agent evaluation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.429430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:0f567d5ccf0a16428bcfde42ff1378314fd5ad465b25e3b950dedaa788930dbe

Observation e5e93c59-a7ef-4463-9592-d07d6a3f5697 · outbound

This paper cites Event abstraction for process mining using supervised learning techniques,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Event abstraction for process mining using supervised learning techniques,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.424048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:4d930c0160e39a55880a56d90486d0c97e653c82f1525de95fc6ec3f34a383de

Observation 94bedfb8-7f23-49ae-8e03-092a82ad01ac · outbound

This paper cites Event abstraction in process mining: Literature review and taxonomy,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Event abstraction in process mining: Literature review and taxonomy,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.408812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:11ff2b86029f329a84e6d3f2959fcc11829ea2d7631a05d415212a0e992eac85

Observation a3e665c4-f53a-48ac-830b-15ce92ec3fb1 · outbound

This paper cites Boundary value exploration for software analysis,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Boundary value exploration for software analysis,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.413741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:72c0d794807d24728d36b8447b41cdf8120e7e5fd102b81403f693e9313bb04e

Observation 94441f16-9357-421f-bc16-0c2ed032bfb8 · outbound

This paper cites Automated robustness testing of off-the-shelf software components,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Automated robustness testing of off-the-shelf software components,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.427615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:ea9f3e6563ec7f700ed23ed11a45bef48bd9b0fe9863b2c43397c3c716b4ebd8

Observation 954f9e5c-be91-4ff4-8818-3fd3c6f4b45d · outbound

This paper cites τ- knowledge: Evaluating conversational agents over unstructured knowl- edge,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents τ- knowledge: Evaluating conversational agents over unstructured knowl- edge,

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-10T00:06:37.976247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:be7ce8030ccd157134ee04c6285ad991923f544f086d49c1e4ef6844c794fc29

Observation 528726a0-4bbc-43f0-82da-02e1076c96fc · outbound

This paper cites ToolLLM: Facilitating large language models to master 16000+ real-world APIs,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents ToolLLM: Facilitating large language models to master 16000+ real-world APIs,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.391348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:4d1818213ba155a555437f7f0178e94698dd3b92e89ca5e569f2e41d67f2911a

Observation 313d6d59-d163-401c-9181-18e859c0f448 · outbound

This paper cites Towards self-evolving benchmarks: Synthesizing agent trajectories via test-time exploration under validate-by-reproduce paradigm.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Towards self-evolving benchmarks: Synthesizing agent trajectories via test-time exploration under validate-by-reproduce paradigm

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-10T00:06:37.963929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:7477c69c653b4b4ec8c7474fb410bfaabf496051595b110dc89b41b57af01eca

Observation 97062318-6ecb-476d-8c13-da94c9bc37b9 · outbound

This paper cites Revisiting benchmark and assessment: An agent-based exploratory dynamic evaluation framework for llms.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Revisiting benchmark and assessment: An agent-based exploratory dynamic evaluation framework for llms

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-10T00:06:37.975111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:a1879dbed5a56ca452f01d1a663e626a66ef57190bef1e13a8dc816ada7fd327

Observation 86574250-54b4-41bf-a838-365082b4f56f · outbound

This paper cites Graph2Eval: Automatic multimodal task generation for agents via knowledge graphs,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Graph2Eval: Automatic multimodal task generation for agents via knowledge graphs,

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-10T00:06:37.973251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:446b1cca3b2638e07fdb5d2d7950441983c8886a5de44bd30b35813c201eec6a

Observation 5bddc95b-cf75-4d8b-bd2c-3374077b922a · outbound

This paper cites Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:06:37.956842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:bdee6481ec2b5a599ff63ee40bd707cb7a136bb863c68d4ab74e62c4561dd27e

Observation bdd38177-ce3c-4b83-9b07-36928981b4da · outbound

This paper cites Mining specifications,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Mining specifications,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.372426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:fcb183db491d447561cb412a3792411618a032cc12b5d98cfeac9ce32ba41b1d

Observation 84bb6543-3b44-4012-a793-c6147a4fb1f0 · outbound

This paper cites Discovering models of software processes from event-based data,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Discovering models of software processes from event-based data,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.395137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:8c57244706c404e77d5ef7ba0195bfbfbae2be34463e8b1032094448cdba478c

Observation f89dc7ec-340b-4e20-b62f-18d00ca50c54 · outbound

This paper cites Automatic generation of software behavioral models,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Automatic generation of software behavioral models,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.403140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:13c3911840361e18f08bc7a86f0f1fd27b6a4196295a06837b8cccd216fb198b

Observation 2dc6c576-bf93-4be9-bcc0-4dc77de44aad · outbound

This paper cites Inferring models of concurrent systems from logs of their behavior with CSight,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Inferring models of concurrent systems from logs of their behavior with CSight,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.410430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:32608b7718e958a46ef95651e62680641336f25ac692b3ff42015d1a2ad11dd3

Observation a0c731b5-5311-457b-9da7-7e14474b33d1 · outbound

This paper cites Workflow mining: Discovering process models from event logs,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Workflow mining: Discovering process models from event logs,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.412159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:925967bdcd307fbb559afc30990d3907497c1de5a8caeb54ebcc75b55146c41e

Observation 18aa5651-0695-48e9-822c-646501335467 · outbound

This paper cites Discov- ering block-structured process models from event logs—a constructive approach,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Discov- ering block-structured process models from event logs—a constructive approach,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.425884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:5f3469acc951eb6980f234ca913b6546e34b6cc147e741c1072e8a7aec5a4ca7

Observation 1f2e5127-a857-461b-abc0-1219399aa20c · outbound

This paper cites Applying graph reduction techniques for identifying structural conflicts in process models,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Applying graph reduction techniques for identifying structural conflicts in process models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.407049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:e71b31215a894c0f2ec338bdc98ece2175735b854bd0ebaf4ada01b2f8673909

Observation d81ef0ce-87de-4eb1-a2a2-1871b5a90b68 · outbound

This paper cites Learning regular sets from queries and counterexamples,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Learning regular sets from queries and counterexamples,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.431693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:a4b9d761245362c4afbec9d931a47a62479b7bb20dcef939b9f79e942b34c702

Observation 9e496994-d97f-432f-b6da-538d382c9579 · outbound

This paper cites Unsupervised dialog structure learning,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Unsupervised dialog structure learning,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.415418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:a3857404fb73ddda9b80bcea7893047e3eac23c20ba827f97fa2f1fb0759d158

Observation 8b21782c-e4f6-4cf1-8c76-f9e32d4497df · outbound

This paper cites Dialog2Flow: Pre-training soft-contrastive action-driven sentence embeddings for automatic dia- log flow extraction,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Dialog2Flow: Pre-training soft-contrastive action-driven sentence embeddings for automatic dia- log flow extraction,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.393236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:b56c89c3f9f1315bca22047e05759e85bed2370be5097f108b7036339c2450bc

Observation 47d02b58-9eb9-4c9d-8561-ca85fee5e5e5 · outbound

This paper cites Agenda-based user simulation for bootstrapping a POMDP dialogue system,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Agenda-based user simulation for bootstrapping a POMDP dialogue system,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.404999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:48b2a471afa672ed225ed64d148abd669985b26281e856d68e6f17aae6973a83

Observation 8b43514d-7a55-415b-94e7-19673430fb4d · outbound

This paper cites ConvLab-2: An open-source toolkit for building, evaluating, and diagnosing dialogue systems,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents ConvLab-2: An open-source toolkit for building, evaluating, and diagnosing dialogue systems,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.407241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:e03d4e4d7c70f83805663f4eaba66197b6a529e7fe1c5d25bbe757176d63fd64

Observation 2fc55fc5-d849-4b04-aa23-93c0f960bcdb · outbound

This paper cites A taxonomy of model- based testing approaches,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents A taxonomy of model- based testing approaches,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.385993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:921289927cb38a8ad5f3a1d86451f89cd4c7492c2661209b030f4d0bb0c49696

Observation 75ce28d8-4ae0-40ca-ba6d-2e5f3c4d9ca9 · outbound

This paper cites Principles and methods of testing finite state machines—a survey,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Principles and methods of testing finite state machines—a survey,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.389665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:be19e47efd98c498ad2ba5842d23bf0240d813fdbcbf793875c2fe3e7e268062

Observation f73ca176-6c76-43e6-9d34-280c4bf33ab5 · outbound

This paper cites Testing software design modeled by finite-state machines,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Testing software design modeled by finite-state machines,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.396696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:115190deba8c5331e336ca1d6d13b701248d77177ab18deeb7b49e118da52482

Observation 25e0e7b8-c0d7-4dfd-bdc5-0355e81b625d · outbound

This paper cites RESTler: Stateful REST API fuzzing,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents RESTler: Stateful REST API fuzzing,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.384367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:3bbac655fca691aa9fd2dfa86af72c663508f6de95aa37796128d944ccb00051

Observation 413c8bf1-30d6-41dd-9201-7823baa165db · outbound

This paper cites RESTful API automated test case generation with Evo- Master,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents RESTful API automated test case generation with Evo- Master,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.373884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:02f0a6fdabd9eab81abe861b2e8bfb4d9d4f42cd7f816fe3fbcdd57d80b945fc

Observation 9c71df55-6f0d-40e5-86cb-c486edd2c362 · outbound

This paper cites Morest: Model-based RESTful API testing with execution feedback,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Morest: Model-based RESTful API testing with execution feedback,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.351133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:658fd90f5b7eeb3d1dd3ae960b80293c975ce853745ccfc73e061742bbefc612

Observation fdfc078e-4aac-477b-a422-f5bbcb06dd86 · outbound

This paper cites KAT: Dependency-aware automated API testing with large language models,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents KAT: Dependency-aware automated API testing with large language models,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.375692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:fff8df10f63849b4addb6bfbe9c0fa710c6de431b3ac8f1dd61096805c2721c5

Observation 5ccd9764-a37b-474b-b12d-7afdcb0079dd · outbound

This paper cites Testing RESTful APIs: A survey,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Testing RESTful APIs: A survey,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.408976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:9947e693d27991c7c7e9c99e0ba9d8651d4e7c58eb79264321b00523c4c4fc60

Observation 2b3b289f-f2e8-4e78-98d5-f07c0cd671c2 · outbound

This paper cites An empirical evaluation of using large language models for automated unit test generation,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents An empirical evaluation of using large language models for automated unit test generation,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.412345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:3efb32067ef1a4d78c19920464501cf8ed8555d7884f9909eb167d72942e8046

Observation 5352081e-274f-4b47-bb01-3889218a77e5 · outbound

This paper cites CoverUp: Effective high coverage test generation for Python,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents CoverUp: Effective high coverage test generation for Python,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.417481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:b96ea19d065703140c6c48f5bafa36a818a9dd372fd59102103cf83fa1f22fb6

Observation a19faacc-39ec-42e5-9c38-01a65f71e8cd · outbound

This paper cites Evaluating and improving ChatGPT for unit test generation,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Evaluating and improving ChatGPT for unit test generation,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.417180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:554ec2fe9bf18bf1b4e5326cdb661bce26559859681f819694e3f658f8281880

Observation c6508104-bd2e-4132-b3b5-e0026d11b7f7 · outbound

This paper cites CodaMosa: Escaping coverage plateaus in test generation with pre-trained large language models,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents CodaMosa: Escaping coverage plateaus in test generation with pre-trained large language models,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.410662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:2b3d4c17c36086941d6d414f0b55ea3675c2dadd9395052b60fdab5880d7dc83

Observation 74a8218e-fa88-4365-9149-1b2c1b93a272 · outbound

This paper cites TestART: Improving LLM-based Unit Testing via Co-evolution of Automated Generation and Repair Iteration.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents TestART: Improving LLM-based Unit Testing via Co-evolution of Automated Generation and Repair Iteration

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:06:37.970137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:be7e95319edcb0712f4d0b2d1f455af962e833d69a10f5d3e03bbd68d6ea21ce

Observation d2ee0e6d-23e2-41d1-865d-8836ec6aaeef · outbound

This paper cites The oracle problem in software testing: A survey.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents The oracle problem in software testing: A survey

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.362723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:0112df785698c602ad577480ac7afd6c1b245cdc8c112cec90e6e4cd215527dc

Observation 91b24357-a3aa-40bb-a95e-781f191ad93b · outbound

This paper cites Pseudo-oracles for non-testable programs,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Pseudo-oracles for non-testable programs,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.368635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:55ad1ad67c579e7bc645c980622c6990f22d740c3a771f9d69df0690e1c5c787

Observation e755b481-bc90-4bad-9642-d31ac5d8b987 · outbound

This paper cites On testing non-testable programs,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents On testing non-testable programs,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.356276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:fe82ed785aa4719139510812d5c76e078588be055f0cfe6c9c4402d6af4387ec

Observation 5280139e-b263-4a46-9432-063d6ccc5e10 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.379148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:1dcc97bb9481c5ca49caa1f4a5054dd7fee9a85e103f863ad773925f0ce95939

Observation 8919763f-e676-4082-b6e1-3691f34f621c · outbound

This paper cites Large language models cannot self-correct reasoning yet.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Large language models cannot self-correct reasoning yet

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.415797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:6bf1642649c5dfedd0fac6a52ed4d90c12afe15effc1227f33a050fdaf9e65c5

Observation 56f5958e-caa8-4320-9086-ead614e1fd1c · outbound

This paper cites On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:06:37.970034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:c6d17684e291e0123063dcf7b94b3c5d524a887772007e05c6e9aa6c5d4d02ba

Observation fc6e1223-e3df-467e-90f8-89dcda18765c · outbound

This paper cites Self-Refine: Iterative refinement with self-feedback,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Self-Refine: Iterative refinement with self-feedback,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.372182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:de1fbb95562561de83a4c86f30c0400d001ec3171f4da33f4f9033284a52a41e

Observation ffd8ae57-7bb3-41f1-a318-d4876be7b2df · outbound

This paper cites CRITIC: Large language models can self-correct with tool-interactive critiquing,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents CRITIC: Large language models can self-correct with tool-interactive critiquing,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.377391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:4a186032c9b0769859e768bf43b9df84a827b9728b44f527431c7048ea023d88

Observation 4b153557-3da1-4ffe-b1b5-bd2cff717884 · outbound

This paper cites Metamorphic testing: A review of challenges and opportunities,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Metamorphic testing: A review of challenges and opportunities,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.398672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:e8f0ccaed5e69ef192dd55374b511dc91096b13acdce427458a1efdbe5c922c3

Observation 45b02c00-7cdd-4acc-b72b-106b5381a916 · outbound

This paper cites Not what you’ve signed up for: Compromising real-world LLM- integrated applications with indirect prompt injection,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Not what you’ve signed up for: Compromising real-world LLM- integrated applications with indirect prompt injection,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.403085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:c438c86ac5c19a6100fe02f62b993d52fb890b4daa150f658d8d4b85b25f5434

Observation 587a72f6-11a9-45de-b57e-fa05e453fc52 · outbound

This paper cites Ignore previous prompt: Attack techniques for language models.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Ignore previous prompt: Attack techniques for language models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.419028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:fa3f252e2a5ad715b32038f3a0e079c562ac3100548e74206a4c74a4c5884de0

Pith citing papers

No inbound Pith citation observations are available.