Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:35:09.898417Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 5 inbound Pith citation observations for arXiv:2505.13546.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:35:09.898417Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:26:27.695830Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-22T00:40:51.097386Z
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c697badb-72c8-47df-a2cc-4420684a0b8e · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems LLM-AutoDiff: Auto-Differentiate Any LLM Workflow
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccc6489c-72a6-4789-806a-3377495c30af · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems Efficient multi-prompt evaluation of LLMs
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1ed9ba4-16bb-4231-b6f8-117de28f23ba · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems Concentration Inequalities: A Nonasymptotic Theory of Independence
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 20b32132-251e-481d-83f8-b0a8db913791 · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in HuggingFace
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e91246e0-2bb3-46dd-879c-c8d65d7c9fce · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb9dd25a-877e-42ed-b689-abfbb54239e4 · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems Generative Agents: Interactive Simulacra of Human Behavior
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f15d743f-8d84-4215-845e-c75b19d56866 · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversa- tions
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 76022750-cdac-4aca-a471-b031e8010743 · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems Language Models are Few-Shot Learners
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a12b8a66-eae0-4a93-a281-764ce15165d1 · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems Exploiting Cloze Questions for Few Shot Text Classification and Natural Language Inference
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc637efd-1499-44a6-be07-8b1219fbc4b0 · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems Cross-Task Generalization via Natural Language Crowdsourcing Instructions
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2e8aae8-5035-479a-9873-90518bac8511 · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e43520a-9e14-40f0-b0f0-9367b6ddb3f7 · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems The Power of Scale for Parameter-Efficient Prompt Tuning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3214cee0-406f-4d75-a880-32d2a7a46216 · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems Prefix-Tuning: Optimizing Continuous Prompts for Generation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56ce45e0-a765-44b4-bf90-8e2368fca8c8 · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems RLPrompt: Optimizing Discrete Text Prompts with Reinforcement Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19b486d3-9fde-4d80-8c8a-af868126f49a · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems Large Language Models Are Human-Level Prompt Engineers
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bca9ce95-38b6-40b6-8229-c76562f9b215 · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 533d78ec-c8a0-4e07-b2dd-5ef81a2b117e · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems Kullback and R
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 95731e7f-666a-4f52-bf4b-7b74ae1aa365 · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems BERTScore: Evaluating Text Generation with BERT
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0365e017-ddbf-446f-b3b4-0710dd3e3afe · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems Sentence-BERT: Sentence Embeddings using Siamese BERT- Networks
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation eadf666d-721f-4c3f-96e8-54c37315dbeb · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems Universal Sentence Encoder for English
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fb94a454-ee5a-4405-b260-2a913a344ca2 · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems Non-Determinism of "Deterministic" LLM Settings
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 561a50ee-e1fc-4ab2-ab3f-3f3af366bbda · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 18609012-e8fb-4cd0-8dd2-25bb696bdd85 · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems Are Large Language Models Consistent over Value-laden Questions?
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff87c08f-2265-441b-a084-6faa22a7c303 · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems Beyond Accuracy: Behavioral Testing of NLP Models with CheckList
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c9602a3a-3515-46ef-9855-5baab69a7af7 · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06305638-1c26-4d0b-9f9d-439bcd62a37f · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems Data Interpreter: An LLM Agent For Data Science
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47239252-dd1a-4cad-92d8-9872b0894fcf · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22b345db-2f41-4855-a51a-ec92ef626598 · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems Self-Evolving Multi-Agent Collaboration Networks for Software Development
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa2b6d2b-96e6-40ee-994d-989ba15e8a2c · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems Evaluating Large Language Models Trained on Code
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b30973ca-36d1-4fee-8ab2-5a959560b561 · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems DA-Code: Agent Data Science Code Generation Benchmark for Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 671fd378-cb34-4757-8c2c-4c0b6310715d · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems Measuring Mathematical Problem Solving With the MATH Dataset
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0804f88e-e849-480f-8acd-9ef3f452df19 · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb89272f-b359-44ef-9c7b-8ba0aa52b632 · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems CellAgent: An LLM-driven Multi-Agent Framework for Automated Single-cell Data Analysis
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c311e2b2-7b34-47bd-b4be-832b6515aa67 · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems A Multimodal Foundation Agent for Financial Trading: Tool-Augmented, Diversified, and Generalist
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a078fce0-5191-4245-ab95-3ffa43c90b28 · outbound
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems ChemLLM: A Chemical Large Language Model
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4f2e7ee-463c-497a-b3a0-43119574eb64 · inbound
PREMISE: Scalable and Strategic Prompt Optimization for Efficient Mathematical Reasoning in Large Models Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8541d78-6217-4e5f-a137-fa8d260877a7 · inbound
Audit, Alignment, and Optimization of LM-Powered Subroutines with Application to Public Comment Processing Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72c8c0cc-ca66-467e-97f5-ca387eab47cd · inbound
GenoMAS: A Multi-Agent Framework for Scientific Discovery via Code-Driven Gene Expression Analysis Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d92a3d6c-5eba-4411-b26d-8f6d47172579 · inbound
Knowing How to Edit: Reliable Evaluation Signals for Diagnosing and Optimizing Prompts at Query Level Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 226d97fc-a631-467a-a820-6e0c50f09089 · inbound
PEEM: Prompt Engineering Evaluation Metrics for Interpretable Joint Evaluation of Prompts and Responses Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.