Pith. sign in

Paper Citation Record · LEDGER

Configurable multi-agent framework for scalable and realistic testing of llm-based agents

As of 9 August 2026, this Paper Citation Record lists 4 of 4 outbound references and 2 inbound Pith citation observations for arXiv:2507.14705.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.14705 v1

Coverage vector

measured 4 of 4 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:53:37.742865Z

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T15:14:55.551856Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T15:15:16.673584Z

Reference resolution

4 of 4 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bd4ecba1-9834-461e-b2b9-05f032675470 · outbound

This paper cites Generative agents: Interactive simulacra of human behavior.

Configurable multi-agent framework for scalable and realistic testing of llm-based agents Generative agents: Interactive simulacra of human behavior

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:53:38.681889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:53:37.248828Z digest=sha256:f4ed4e456ebe0f2e97151695d44985f51eca7a27af442fcc3f22302150426711

Observation d5c14cd0-84ff-461f-9acb-51485b8cafd0 · outbound

This paper cites A survey of statistical user simulation techniques for reinforcement-learning of dialogue management.

Configurable multi-agent framework for scalable and realistic testing of llm-based agents A survey of statistical user simulation techniques for reinforcement-learning of dialogue management

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:53:38.312972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:53:37.393521Z digest=sha256:b040cc13bfe0ffbb2bc0ec59c50fe7182ca7130132f0265df1482924a80bf003

Observation 5dc81a3b-8e1f-4c2c-9d5a-816b4206a179 · outbound

This paper cites User simulation for reinforcement learning of dialogue management policies: Initial progress report.

Configurable multi-agent framework for scalable and realistic testing of llm-based agents User simulation for reinforcement learning of dialogue management policies: Initial progress report

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:53:38.024252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:53:37.534694Z digest=sha256:16cb35caf2e381eb619e54311481a52728e9923701a5889e36d70d52772fe9d2

Observation 66e57e5f-cdb9-4967-b15c-286dece8a6b0 · outbound

This paper cites Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation.

Configurable multi-agent framework for scalable and realistic testing of llm-based agents Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:53:37.742865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:53:37.742865Z digest=sha256:967161334ccf3003f25238c5be31f82bf16ffe9a1d9085658f7550884b67802f

Pith citing papers

Observation 5f770e9d-d13e-4274-81dd-94b175fe2c88 · inbound

MirrorBench: A Benchmark to Evaluate Conversational User-Proxy Agents for Human-Likeness cites this paper.

MirrorBench: A Benchmark to Evaluate Conversational User-Proxy Agents for Human-Likeness Configurable multi-agent framework for scalable and realistic testing of llm-based agents

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:15:16.675455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T15:14:55.551856Z digest=sha256:48854a69a4102c399dcf45a3bf8bdb439f702a7d663080738284d7fe92540195

Observation 001b6ff6-591a-409e-8fcb-15c30b43ed12 · inbound

SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems cites this paper.

SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems Configurable multi-agent framework for scalable and realistic testing of llm-based agents

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:48:01.660508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T08:45:54.303143Z digest=sha256:10d0a3f3b71d926d799ba573cc473bd22aee789f3db020096cc6612310a42c1a