Pith. sign in

Paper Citation Record · LEDGER

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation

As of 18 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2607.14256.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.14256 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T02:44:17.432243Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 52876c55-4ea9-4615-bc02-c38753137410 · outbound

This paper cites Security in LLM-as-a-Judge: A Comprehensive SoK.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Security in LLM-as-a-Judge: A Comprehensive SoK

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:13.210092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:13.210092Z digest=sha256:577050cff098350ccfab4ba7c773a215c6e213a893f9025989e597ab7a57989c

Observation 58511f56-6fb5-48e6-81c5-87a22d890e21 · outbound

This paper cites Reference-guided verdict: Llms-as-judges in automatic evaluation of free-form qa.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Reference-guided verdict: Llms-as-judges in automatic evaluation of free-form qa

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:13.302235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:13.302235Z digest=sha256:55e77e28ddc7532c162d4520ea89748d68d5ee733dbb23e679a5043420e64f7a

Observation 25ec6ad5-7450-430d-8bc1-36f5a9a20c14 · outbound

This paper cites Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:13.401198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:13.401198Z digest=sha256:6400f445d2ee8efe6f168a65714a8727d0625839c5fbb3a24523414c0ebdb489

Observation 7c0bc200-52d7-411b-b170-f737ba754229 · outbound

This paper cites Rethinking fine- tuning when scaling test-time compute: Limiting confidence improves mathematical reasoning.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Rethinking fine- tuning when scaling test-time compute: Limiting confidence improves mathematical reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:13.549266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:13.549266Z digest=sha256:7bc38e3f19de499634d365e0a4b83c06b24d7f49af0b106963bfa4d92db780b3

Observation d430814e-b2b2-4e82-9c3d-bf34a6ec2e5f · outbound

This paper cites Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:13.644542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:13.644542Z digest=sha256:26bd3401637773ab1105813cd284d588e5e65e5155211c9f52bcc8730da08ad1

Observation 4697ea5a-7ce8-48f8-a8cb-ac68ac45c49e · outbound

This paper cites The side effects of being smart: Safety risks in mllms’ multi-image reasoning.arXiv preprint arXiv:2601.14127, 2026.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation The side effects of being smart: Safety risks in mllms’ multi-image reasoning.arXiv preprint arXiv:2601.14127, 2026

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:13.742784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:13.742784Z digest=sha256:1aa439a4fbbf5f263248fc3f6cb015ed94243e8a6aa8ba8f36fd41fc17567c58

Observation 9550f4db-fcb5-44ca-aa70-6971f60b464a · outbound

This paper cites Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:13.880726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:13.880726Z digest=sha256:aeb6cf4ec0594a08c2fc2924b0fb4f2a8ae1564804bea2d39276cf4fae6ae658

Observation 0e4b58c6-1b82-4790-b85a-10630cd29029 · outbound

This paper cites Jailbreaking llms & vlms: Mechanisms, evaluation, and unified defense.arXiv preprint arXiv:2601.03594, 2026.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Jailbreaking llms & vlms: Mechanisms, evaluation, and unified defense.arXiv preprint arXiv:2601.03594, 2026

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:13.955654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:13.955654Z digest=sha256:19bcbeafbdbf887418b47aaae97fbdb6df17fb54b845720253f861be2eeec219

Observation 756d73ff-2ec1-4636-81fd-ecb371ac8b24 · outbound

This paper cites Courtroom-Style Multi-Agent Debate with Progressive RAG and Role-Switching for Controversial Claim Verification.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Courtroom-Style Multi-Agent Debate with Progressive RAG and Role-Switching for Controversial Claim Verification

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:14.030820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:14.030820Z digest=sha256:bd1057b6b2a3c2afa922f2bc19db29e40a2b7b60880c98a2dafb0eee0c78a865

Observation 482ba05c-36e7-4bbd-96b1-5fffbe61dc7b · outbound

This paper cites Improv- ing factuality and reasoning in language models through multiagent debate.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Improv- ing factuality and reasoning in language models through multiagent debate

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:14.131820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:14.131820Z digest=sha256:08dbaea8a09e58427c31a39d3ead0b46fbc3689a79714c8055793a2f52f54994

Observation 2d181dc7-b008-4baf-9ebb-65978dae824b · outbound

This paper cites Bad students make great teachers: Active learning accelerates large-scale visual understanding.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Bad students make great teachers: Active learning accelerates large-scale visual understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:14.201724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:14.201724Z digest=sha256:cdb90a578b0508cac2c4ebbe0e6261812f3d4e36db33ea5b873542bdede4d81a

Observation 15ff5940-23fc-4a59-830b-8f83a5e745a9 · outbound

This paper cites Contextnav: Towards agentic multimodal in-context learning.arXiv preprint arXiv:2510.04560, 2025.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Contextnav: Towards agentic multimodal in-context learning.arXiv preprint arXiv:2510.04560, 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:14.292307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:14.292307Z digest=sha256:6dbb0b760bb0125e44b472ef39eb3f051690ec72e662ddb492f3df8490b2a1ef

Observation 41636004-8125-44d0-958a-94928a86f81f · outbound

This paper cites Adversarial defense in vision-language models: An overview.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Adversarial defense in vision-language models: An overview

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:14.406045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:14.406045Z digest=sha256:58d96478acd9b174759d1fe0357a72491f0ccbb9df070c634a9dd102ec3f13dd

Observation 64785b11-9b35-4244-b252-0f932846d848 · outbound

This paper cites DetPO: In-Context Learning with Multi-Modal LLMs for Few-Shot Object Detection.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation DetPO: In-Context Learning with Multi-Modal LLMs for Few-Shot Object Detection

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:14.501395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:14.501395Z digest=sha256:23daabf120455477fb553965b8974fb0cdd715fda22a2bf0a0dc52c57d2a1cfc

Observation dd6e3a12-bcfc-4e2a-bc28-fdd9f16717ec · outbound

This paper cites Debate, deliberate, decide (d3): A cost-aware adversarial framework for reliable and interpretable llm evaluation.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Debate, deliberate, decide (d3): A cost-aware adversarial framework for reliable and interpretable llm evaluation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:14.602511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:14.602511Z digest=sha256:3c7abdb4fa1b75c4f7c81458449ab2747e871c5192839a0726a8a90236ff6879

Observation ca22c62f-9c53-4427-ae90-2a3130017d0f · outbound

This paper cites Aetheria: A multimodal interpretable content safety framework based on multi-agent debate and collaboration.arXiv preprint arXiv:2512.02530, 2025.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Aetheria: A multimodal interpretable content safety framework based on multi-agent debate and collaboration.arXiv preprint arXiv:2512.02530, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:14.693049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:14.693049Z digest=sha256:568d4c01f7529181fdef61f7148e8829e7c99440010fc86bcac0bca7b3a02c15

Observation 13e9c64d-93e5-40cb-8bae-88ef7353e82a · outbound

This paper cites LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:14.769734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:14.769734Z digest=sha256:a7a78a1d901a52247b8d6b696d0b19149cd8e4132e70e57fbdc82b7d63887edc

Observation d069afdb-8d62-4b5b-bf68-da6366714a31 · outbound

This paper cites A Judge-free LLM Open-ended Generation Benchmark Based on the Distributional Hypothesis.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation A Judge-free LLM Open-ended Generation Benchmark Based on the Distributional Hypothesis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:14.843081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:14.843081Z digest=sha256:58ce529356bda7c30da6b4b440480697759cf7ff1c417be15db65c14a4271091

Observation a35ee8e3-d5dd-4d57-99ce-0c7df1dfd0e4 · outbound

This paper cites AI safety via debate.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation AI safety via debate

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:14.904515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:14.904515Z digest=sha256:8f71dd490d9e9ce4e0f1636a92b45e06706d9d55fa7f036bd07e3a6f61d41d04

Observation 91538c90-cdc5-4d1d-be9d-95236c6607c7 · outbound

This paper cites Adversarial attacks on multimodal large language models: A comprehensive survey.arXiv preprint arXiv:2603.27918, 2026.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Adversarial attacks on multimodal large language models: A comprehensive survey.arXiv preprint arXiv:2603.27918, 2026

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:14.990323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:14.990323Z digest=sha256:464f74da0d643e5bab0b8362b9df898167ce2eb7e669fb96a37b8bf7b3de9a81

Observation 83e048fd-179d-4526-9624-52b5e4a6387b · outbound

This paper cites Curriculum guided massive multi agent system solving for robust long horizon tasks.arXiv preprint arXiv:2512.08545, 2025.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Curriculum guided massive multi agent system solving for robust long horizon tasks.arXiv preprint arXiv:2512.08545, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:15.074317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:15.074317Z digest=sha256:0e171f4f33020172c9b2a11f9bb79e6e8107719e74d77db48dcb08c3c755c2a4

Observation ddf61226-84f2-4a8b-98c8-c4d3febcff2b · outbound

This paper cites Evaluating nova 2.0 lite model under amazon’s frontier model safety framework.arXiv preprint arXiv:2601.19134, 2026.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Evaluating nova 2.0 lite model under amazon’s frontier model safety framework.arXiv preprint arXiv:2601.19134, 2026

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:15.141754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:15.141754Z digest=sha256:ffe1554b6b24ce0ebbad42efa648e221be8d96a8a17ef70bea7b185eb09d70ad

Observation a5eae34c-4ea6-4f95-971b-48f20d8970f9 · outbound

This paper cites Biasscope: Towards automated detection of bias in llm-as-a-judge evaluation.arXiv preprint arXiv:2602.09383, 2026.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Biasscope: Towards automated detection of bias in llm-as-a-judge evaluation.arXiv preprint arXiv:2602.09383, 2026

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:15.193060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:15.193060Z digest=sha256:2286ae2858dd61a2e0854af16dfc47e97026d6ffe3ba67e51332a52a5463b839

Observation 30ae7cc6-cf76-4203-828a-0d13922aa6b6 · outbound

This paper cites T-map: Red-teaming llm agents with trajectory-aware evolutionary search.arXiv preprint arXiv:2603.22341, 2026.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation T-map: Red-teaming llm agents with trajectory-aware evolutionary search.arXiv preprint arXiv:2603.22341, 2026

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:15.275900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:15.275900Z digest=sha256:b15a53306d265fb0b8648414495e6742dbe8ad9857d43241299e936ba755d43c

Observation 9137ca88-9e26-4c79-86cf-143d26060fda · outbound

This paper cites THINKSAFE: Self-Generated Safety Alignment for Reasoning Models.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation THINKSAFE: Self-Generated Safety Alignment for Reasoning Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:15.360226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:15.360226Z digest=sha256:8e770b48fc98ac15743d8608c0ac0cfcfeb9fe617275f5f1bc56462cddb64d28

Observation 923205e8-ed21-4f3d-8ddd-1697e24ec7a1 · outbound

This paper cites Holisafe: Holistic safety benchmarking and modeling for vision- language model.arXiv preprint arXiv:2506.04704, 2025.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Holisafe: Holistic safety benchmarking and modeling for vision- language model.arXiv preprint arXiv:2506.04704, 2025

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:15.415061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:15.415061Z digest=sha256:9dac768915497467b7698348a0781b4ae335ac070849ea519a8399081306b3ed

Observation 07572d23-b8af-4145-a2de-79170a537b03 · outbound

This paper cites an unresolved cited work.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:15.502229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:15.502229Z digest=sha256:9e0f91bc4da3a97986852b54b31d2f62b333defd11806d14a51d04cd9eeaf2b9

Observation 9594631e-637b-42b6-9e7e-8911513f7531 · outbound

This paper cites From generation to judg- ment: Opportunities and challenges of llm-as-a-judge.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation From generation to judg- ment: Opportunities and challenges of llm-as-a-judge

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:15.559113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:15.559113Z digest=sha256:31c23f48266471558a1361d2e5feacf67f58a91925ba43c7479deeb132c05281

Observation f808fac0-18f9-4a75-be1a-ba4060ee7ead · outbound

This paper cites Evaluating Scoring Bias in LLM-as-a-Judge.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Evaluating Scoring Bias in LLM-as-a-Judge

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:15.637387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:15.637387Z digest=sha256:24de1d3e63937f440e4f3947540b1da926ae4e603471c78a5513b0c43677e22c

Observation 9d216852-3fe2-4f35-8201-37ac60540127 · outbound

This paper cites Who judges the judge? llm jury-on-demand: Building trustworthy llm evaluation systems.arXiv preprint arXiv:2512.01786, 2025.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Who judges the judge? llm jury-on-demand: Building trustworthy llm evaluation systems.arXiv preprint arXiv:2512.01786, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:15.711786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:15.711786Z digest=sha256:e4d9219f9ac30df552a0c90ee60e02746d41d5aa280dad41bcec0208f8f33bde

Observation 67fad8d2-9f74-4a5c-82a7-14a406e4c347 · outbound

This paper cites Benchmark test-time scaling of general llm agents.arXiv preprint arXiv:2602.18998, 2026.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Benchmark test-time scaling of general llm agents.arXiv preprint arXiv:2602.18998, 2026

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:15.789634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:15.789634Z digest=sha256:fc35674d62464d6fff67257c26b98885ea95d2c33145950f15f332c319f9f924

Observation 1df74600-1e78-46b9-8622-439d85ef6787 · outbound

This paper cites Elhplan: Efficient long-horizon task planning for multi-agent collaboration.arXiv preprint arXiv:2509.24230, 2025.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Elhplan: Efficient long-horizon task planning for multi-agent collaboration.arXiv preprint arXiv:2509.24230, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:15.871821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:15.871821Z digest=sha256:3e16097e978e91d84d2a6b26671db08147490a3d1b49e90558560280c9328907

Observation 8ab350e0-6b97-41c7-ab24-217c230c7269 · outbound

This paper cites Examining LLMs' Uncertainty Expression Towards Questions Outside Parametric Knowledge.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Examining LLMs' Uncertainty Expression Towards Questions Outside Parametric Knowledge

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:15.960327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:15.960327Z digest=sha256:2ef7c7a1a533f40fa9245ac97c0ab7fe614d0ab369cdabc7a2dc3fcbbc0c8235

Observation dd68d04f-a297-440d-af26-260c3a2cfa9d · outbound

This paper cites WebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation WebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.036683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.036683Z digest=sha256:c61f2b9679bd7103e2dfe3a4a871d97f4e7d8bb6eb0388b4e709554fe3823782

Observation 7ce210d1-3829-4b1c-812a-83df63e4f122 · outbound

This paper cites Mosaic: Modeling social ai for content dissemination and regulation in multi-agent simulations.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Mosaic: Modeling social ai for content dissemination and regulation in multi-agent simulations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.085510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.085510Z digest=sha256:ee515e3c912217be11b77dab47d6be6a74f7399cb5967b0774ce70563b8d7f52

Observation 5ab79f7c-a3bc-497f-8223-45b369e101bb · outbound

This paper cites Mtmcs-bench: Evaluating contextual safety of multimodal large language models in multi-turn dialogues.arXiv preprint arXiv:2601.06757, 2026.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Mtmcs-bench: Evaluating contextual safety of multimodal large language models in multi-turn dialogues.arXiv preprint arXiv:2601.06757, 2026

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.135485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.135485Z digest=sha256:814f37e9c7a14e617942b6e64e90f791aaf66d90857b69c69024210e1c6d1d6f

Observation 531fa067-92e2-407c-a07f-12b8772125dc · outbound

This paper cites Ai debate aids assessment of controversial claims.arXiv preprint arXiv:2506.02175, 2025.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Ai debate aids assessment of controversial claims.arXiv preprint arXiv:2506.02175, 2025

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.185073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.185073Z digest=sha256:c5bb949e1d4a514930161ad35b1f03dcad005da05c0ceed6773a543a6c7aacd3

Observation 9236d518-02f5-4c78-86c7-a3e914e97b47 · outbound

This paper cites X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.255412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.255412Z digest=sha256:c74322a1b1689a545bad101a26b2206febd11dbcc0a1dd1eadf99b7dfc783846

Observation ffb7c359-7219-48e5-bcf7-2e1325c7a3cd · outbound

This paper cites Disc-amc: Token-and parameter-efficient discretized statistics in-context automatic modulation classification.arXiv preprint arXiv:2510.00316, 2025.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Disc-amc: Token-and parameter-efficient discretized statistics in-context automatic modulation classification.arXiv preprint arXiv:2510.00316, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.307323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.307323Z digest=sha256:28237b99131500c1566f438b5c77ff5375290c36f1224d760fad7c05c9be74e4

Observation 7058828b-0ee1-4b46-a2c9-16c53fc97cb7 · outbound

This paper cites Interaction Protocol Shapes Moral Judgment in Multi-Agent Debate.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Interaction Protocol Shapes Moral Judgment in Multi-Agent Debate

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.371761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.371761Z digest=sha256:cf1a93536ab63de6eb8477cd4d129441eaedb02e3ba12e54e44b890401def9a6

Observation 95cefedf-3041-4473-b764-6c41c76f0066 · outbound

This paper cites Assessment of Multimodal Large Language Models in Alignment with Human Values.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Assessment of Multimodal Large Language Models in Alignment with Human Values

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.420847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.420847Z digest=sha256:5baa646c8d14a659c9d5412105d492c3ffeadca94d8ebdd99900a19a5d632780

Observation b3ab628d-ad3d-427c-a6a4-17e7ae4e68e0 · outbound

This paper cites Llm-as-a-judge for time series explanations.arXiv preprint arXiv:2604.02118, 2026.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Llm-as-a-judge for time series explanations.arXiv preprint arXiv:2604.02118, 2026

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.446615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.446615Z digest=sha256:46c58e97ac96ff0a92e3c718bcaf745ea5a0305df6eb1170f5aa1adeef42fa21

Observation 120bdebe-f285-4bb1-bb70-5a21c75d58a0 · outbound

This paper cites Unigame: Turning a unified multimodal model into its own adversary.arXiv preprint arXiv:2511.19413, 2025.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Unigame: Turning a unified multimodal model into its own adversary.arXiv preprint arXiv:2511.19413, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.513676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.513676Z digest=sha256:4288f27cfee4188caa7ea66762ae702c13f96150d412038d0574ec1ab9a81c5e

Observation b5435b84-4cfe-4289-a100-9e21c2b2aba3 · outbound

This paper cites Supporting human raters with the detection of harmful content using large language models.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Supporting human raters with the detection of harmful content using large language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.574250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.574250Z digest=sha256:1a51fcc513a1b3ebb7fc874e09552d18b70c71f5d3b2e00d5e7d5c720a6c38ca

Observation 26823bb7-34c3-4c03-b006-82a95b9ab134 · outbound

This paper cites Automated concept discovery for llm-as-a-judge preference analysis.arXiv preprint arXiv:2603.03319, 2026.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Automated concept discovery for llm-as-a-judge preference analysis.arXiv preprint arXiv:2603.03319, 2026

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.631763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.631763Z digest=sha256:eba1554a37c07fed26382601d8ad0cfa26b3b7fc5127a06fedf9e30446181cee

Observation ae09ccf1-033f-4661-b6d9-4cdeb3e18042 · outbound

This paper cites Can llm agents really debate? a controlled study of multi-agent debate in logical reasoning.arXiv preprint arXiv:2511.07784, 2025.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Can llm agents really debate? a controlled study of multi-agent debate in logical reasoning.arXiv preprint arXiv:2511.07784, 2025

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.691039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.691039Z digest=sha256:dabc60d81c501554189234b27206f77cd9d8a80717af0586260accfa7fa787f6

Observation 9921a8aa-53d0-4ed6-8c6d-f66a804b1077 · outbound

This paper cites OutSafe-Bench: A Benchmark for Multimodal Offensive Content Detection in Large Language Models.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation OutSafe-Bench: A Benchmark for Multimodal Offensive Content Detection in Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.755987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.755987Z digest=sha256:a1e3cfcee4cfc95627065daa03f5611b262bb7cbd7634ec794e8b3205159e544

Observation 808993bb-8fe5-42ae-b424-758395746a7d · outbound

This paper cites Llama-3.1-foundationai- securityllm-reasoning-8b technical report.arXiv preprint arXiv:2601.21051, 2026.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Llama-3.1-foundationai- securityllm-reasoning-8b technical report.arXiv preprint arXiv:2601.21051, 2026

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.818828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.818828Z digest=sha256:7222021c27166f5d69b9fe611c1901edf5ebd378db459dfab5447e11542fbed7

Observation 243d2632-a2f4-454d-a540-90894987e1cb · outbound

This paper cites When AIs Judge AIs: The Rise of Agent-as-a-Judge Evaluation for LLMs.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation When AIs Judge AIs: The Rise of Agent-as-a-Judge Evaluation for LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.888129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.888129Z digest=sha256:c75e2397099cc653180a5d36f9e0abbc68af977c387241491007672f6b837695

Observation d7e7d1ef-c352-43da-8b8c-b317f6c69775 · outbound

This paper cites Evolving contextual safety in multi-modal large language models via inference-time self-reflective memory.arXiv preprint arXiv:2603.15800, 2026.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Evolving contextual safety in multi-modal large language models via inference-time self-reflective memory.arXiv preprint arXiv:2603.15800, 2026

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.956146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.956146Z digest=sha256:af9b35f42f0a063eef87f9cb98022dc380ba47e0d5b22f46c55d2d9210475480

Observation 91a72551-1176-4d1c-9b5d-d29132eb03b7 · outbound

This paper cites Visual exclusivity attacks: Automatic multimodal red teaming via agentic planning.arXiv preprint arXiv:2603.20198, 2026.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Visual exclusivity attacks: Automatic multimodal red teaming via agentic planning.arXiv preprint arXiv:2603.20198, 2026

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:17.016819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:17.016819Z digest=sha256:b75957f6c2253e2c5eebd5e853425e9c55567bf8048462739163305be7da20db

Observation b2f95ed4-09d2-4d0e-a9f8-f48b33a53814 · outbound

This paper cites Amulet: ReAlignment During Test Time for Personalized Preference Adaptation of LLMs.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Amulet: ReAlignment During Test Time for Personalized Preference Adaptation of LLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:17.054063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:17.054063Z digest=sha256:c9f1d456ff39b51d8d8e83aafa282a96d708a7a6cb633f554a8b2bcf1f0679c7

Observation bc460a0f-f52e-4af6-8b41-e9d10122f13f · outbound

This paper cites inhumane conditions.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation inhumane conditions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:17.106150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:17.106150Z digest=sha256:8e59584766f88fb3dd845b0c171d1a0397473e9d1e608d15064b69fa86eb92dc

Observation 24c00459-75db-4a1e-8d9b-515db75a8ca2 · outbound

This paper cites {policy_text}.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation {policy_text}

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:17.183685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:17.183685Z digest=sha256:587fc90bcba1f535bab65815ffca16ada4866dbd4d96e1aa680810d093d58cd2

Observation 4aaef1b5-e16b-45b9-b010-037a7d8986f6 · outbound

This paper cites This synthesizes a complete context block explicitly documenting the distinct arguments for and against each candidate classification label.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation This synthesizes a complete context block explicitly documenting the distinct arguments for and against each candidate classification label

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:17.271516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:17.271516Z digest=sha256:a29659b6530a5478ef158524f06fc11ef07cebe3c02a1a22ca9ea828ba60aed2

Observation 0af11b24-c32d-4db8-bc87-8fd0451384ff · outbound

This paper cites {policy_text}.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation {policy_text}

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:17.347594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:17.347594Z digest=sha256:3851a3fd0d7fadacd8393d6411e0d46cff8c3eaf5eb3fa2df356825a3090f8fa

Observation a4c2016a-a3d2-46bc-b8c6-7d70ce6315a9 · outbound

This paper cites If any dissent continues even after the debate round, the system automatically categorizes the instance as a deadlock and escalates it to the Level II Jury Committee.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation If any dissent continues even after the debate round, the system automatically categorizes the instance as a deadlock and escalates it to the Level II Jury Committee

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:17.432243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:17.432243Z digest=sha256:54c71d61e93ab53e9a8e0739b22b5d58672217f6c5c5adb599466cdc415ea18c

Pith citing papers

No inbound Pith citation observations are available.