Pith. sign in

Paper Citation Record · LEDGER

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows

As of 9 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 2 inbound Pith citation observations for arXiv:2506.03332.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03332 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:10:45.767347Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-22T03:53:03.807297Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T03:54:34.003611Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved47
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8b8395ea-c8c5-4af7-a3da-a1ff96abb61e · outbound

This paper cites ReConcile: Round-table conference improves reasoning via consensus among diverse LLMs.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows ReConcile: Round-table conference improves reasoning via consensus among diverse LLMs

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:48.042471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.517449Z digest=sha256:6175e1f983f07f25ad2c7830fe537bb25e9ff5c96c5e78d3fb6223a4935ed56f

Observation 1361cfdc-fabe-4716-b7f6-8fefba645e89 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.521483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.521483Z digest=sha256:db3cbc5f6eb2402c9b2e8b0a5f70221276612eb0affa0e772b5f754b5f6288d4

Observation 1676adec-811d-4096-b08f-16bcf22e89b1 · outbound

This paper cites Synthetic disinformation attacks on automated fact verification systems.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Synthetic disinformation attacks on automated fact verification systems

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:47.859384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.524917Z digest=sha256:8a6d9537372380175685b8b2be9a590239d4fa7fa48ac61e70940c380b1a5d11

Observation b8f5c337-7361-4ea6-bcfd-9825eb220d40 · outbound

This paper cites Improving Factuality and Reasoning in Language Models through Multiagent Debate.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Improving Factuality and Reasoning in Language Models through Multiagent Debate

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.528343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.528343Z digest=sha256:c80ed4061a9d95656b2eb3de736d7faeae4071f1979d2a8ab45e66fd2f096e53

Observation 3aba0d3b-023e-48d5-8fb5-9a0f63d0116c · outbound

This paper cites DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.531959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.531959Z digest=sha256:3350d13b8b35ffa0a0cb39ad1d3400ecade2512be3ad5f5e2bc130fec4d970c1

Observation f9c8f781-43ed-4575-98e9-6d9d5f450882 · outbound

This paper cites SearchQA: A New Q&A Dataset Augmented with Context from a Search Engine.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows SearchQA: A New Q&A Dataset Augmented with Context from a Search Engine

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.536054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.536054Z digest=sha256:a8240bbd3ff79863e8fb399a58bb35b61965f70732352539a052d44483caa160

Observation c46706d3-b975-4a89-9062-bb9564179cb3 · outbound

This paper cites Baldur: Whole-Proof Generation and Repair with Large Language Models.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Baldur: Whole-Proof Generation and Repair with Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.540047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.540047Z digest=sha256:c6b916e64ade8a9b8bb1b78b44492a3bc9430360295e7f4ed3f4637416d8fd01

Observation a53f184f-325a-4c2d-87ab-0386c753d6f8 · outbound

This paper cites CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.544006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.544006Z digest=sha256:6c4439e53a63a34f3382389aa5c0b32debfbcb9edaff531d104c01c2766cc1dd

Observation 014c86a3-04a1-4df9-b5b7-e1cb9428fb57 · outbound

This paper cites The larger the better? improved LLM code-generation via budget reallocation.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows The larger the better? improved LLM code-generation via budget reallocation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:47.713786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.547462Z digest=sha256:a75864d8454eb0c6eb28e6fc34c99f16f316eb9d7a36d1e149d13010e150915a

Observation 29a76a2f-4925-44d7-8582-e6a5c9d92acb · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Measuring Massive Multitask Language Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.550880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.550880Z digest=sha256:718076f33e33bf79a3db85459239595bfbfe7cef9c52188842b6127ec575a794

Observation 424f8645-1142-4952-b467-92cd4ba4e026 · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Large Language Models Cannot Self-Correct Reasoning Yet

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.554308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.554308Z digest=sha256:82b6e61082644899ea79133375d82b9e3342706f4052eec11264010a98d36157

Observation 9e438a02-8bcd-42db-ad95-d3711708e956 · outbound

This paper cites GPT-4o System Card.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows GPT-4o System Card

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.557635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.557635Z digest=sha256:475f1727b22bceb8295abd80445888007bc15573b36ed5f109308e57590aa544

Observation 369e756e-bf62-4a64-828a-cc97e895e1f9 · outbound

This paper cites OpenAI o1 System Card.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows OpenAI o1 System Card

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.561129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.561129Z digest=sha256:2f889b68e49f963491297842fd9e60652ac2a26190efe20c06b8e576d5024a46

Observation 0d8abfcb-76f5-464c-ac2c-af43ba4c354c · outbound

This paper cites TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.564319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.564319Z digest=sha256:40f0d7bc7d58e44168193a8f5833a614f8582312e77d4daa404c307cfc360445

Observation 65cd85af-c690-480d-b6df-a537f449e9fd · outbound

This paper cites When can LLMs actually correct their own mistakes? a critical survey of self-correction of LLMs.Transactions of the Association for Computational Linguistics, 12:1417–1440, 2024.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows When can LLMs actually correct their own mistakes? a critical survey of self-correction of LLMs.Transactions of the Association for Computational Linguistics, 12:1417–1440, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:47.579967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.567712Z digest=sha256:4faeb5907a96ee1dae9b7b8d5e0161cd7e17c81d7536e31e79eb5183fda4ca5d

Observation 81f7c4be-d452-422d-9d8f-d5bd2c12d249 · outbound

This paper cites AI Agents That Matter.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows AI Agents That Matter

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.571170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.571170Z digest=sha256:c5e9868eb82879c2cd0a6260ba842253c8791dd68261c728e0289f090c53c201

Observation ec74cd6d-d6ba-485c-be06-83020276e60d · outbound

This paper cites Bowman, Tim Rocktäschel, and Ethan Perez.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Bowman, Tim Rocktäschel, and Ethan Perez

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:47.375433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.574364Z digest=sha256:f93d343c0c183f977619a59e2ca5dc8fbcaecde8d2e9f0434fc26a1926e9e659

Observation 52cdb31c-8fde-43c1-9c84-baa153c6daca · outbound

This paper cites Natural questions: a benchmark for question answering research.Transactions of the Association for Computational Linguistics, 7:453–466, 2019.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Natural questions: a benchmark for question answering research.Transactions of the Association for Computational Linguistics, 7:453–466, 2019

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.577818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.577818Z digest=sha256:149a3c7aea8dfa1c690fb254aa6d39bd3f1dbf1b00d85499666efddc99d700e4

Observation 70b948cd-baec-4553-8abd-ef315209b3b4 · outbound

This paper cites Making language models better reasoners with step-aware verifier.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Making language models better reasoners with step-aware verifier

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.581968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.581968Z digest=sha256:7f5605e12277068dfa9ce91f877f6e1436dd55d2807cb4480f453d97be56df01

Observation 3e3d13e9-9249-4e03-8909-d85da7e37c06 · outbound

This paper cites Encouraging divergent thinking in large language models through multi- agent debate.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Encouraging divergent thinking in large language models through multi- agent debate

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:47.161684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.585154Z digest=sha256:ba1f286eeefb9da8086cbb3a8eb60c41de17360c2254b5b58e5f8b13426890d0

Observation b58eb157-5c5c-40ba-b168-9470b7f9814c · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Self-Refine: Iterative Refinement with Self-Feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.589027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.589027Z digest=sha256:1e84f220427255dcbf74a4c7e1bf4bd3e610ca94c36b7521721ed0d1f83211cb

Observation f6843d3e-326c-48dd-9049-be2657dc59ad · outbound

This paper cites Debate Helps Supervise Unreliable Experts.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Debate Helps Supervise Unreliable Experts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.592227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.592227Z digest=sha256:2f9bcc6b6d3ceaa627003269c5c68ded22c6442a19a57dd96a1d3b0eabbb84d7

Observation b140cf55-d037-4299-a99d-98c319d50c95 · outbound

This paper cites Faitheval: Can your language model stay faithful to context, even if ”the moon is made of marshmallows”.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Faitheval: Can your language model stay faithful to context, even if ”the moon is made of marshmallows”

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.993915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.595523Z digest=sha256:63829c46e0a7a49a3caac2efa721f1cd7d8dffb2d8cd3b4d73de6c8e95aababc

Observation 0e540f63-cb26-4645-bd04-120ec0f23f93 · outbound

This paper cites Lever: Learning to verify language-to-code generation with execution.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Lever: Learning to verify language-to-code generation with execution

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.847888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.598687Z digest=sha256:20c4062011012240e6cabca28ed0a54b338ad4b7ad12d6995cd893af4b801902

Observation a94a1d61-8e0d-470a-8a80-68e877edd8e6 · outbound

This paper cites On the Risk of Misinformation Pollution with Large Language Models.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows On the Risk of Misinformation Pollution with Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.601797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.601797Z digest=sha256:3d4ec4229f7fdaac6f001b37e7227b62d036784b4d535212d797b3332a4c2761

Observation 56114032-d189-4cae-8ef5-06c0cdd39060 · outbound

This paper cites OffsetBias: Leveraging Debiased Data for Tuning Evaluators.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.605360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.605360Z digest=sha256:7b89beff9162a469441aa83f03c6f0398c63cfa3d1f3ece2dace5d9b26552ba1

Observation 8c3f49e0-dbb3-4f3a-828b-3ad7cfbcd84d · outbound

This paper cites Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.671505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.608726Z digest=sha256:4ae69e814f7f61f97f2a318532e594c0f3c51809363e37ddffddb2389586a22f

Observation a25700a9-0615-4de9-a465-29cc43230d9a · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Gpqa: A graduate-level google-proof q&a benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.611812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.611812Z digest=sha256:c42c6d7a074010a323223ae875ddd8fa717649f80e118454e407e2c52198301d

Observation f78f1aca-f91e-42a5-b10a-e3c0dfaf3c88 · outbound

This paper cites Archon: An Architecture Search Framework for Inference-Time Techniques.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Archon: An Architecture Search Framework for Inference-Time Techniques

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.615625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.615625Z digest=sha256:839d6bfaf8a796220c2c7098f9c5ec574aaa7f6f6ff13b6014d128a83ecf12e6

Observation 78054db9-5947-4319-9573-31d6d4e47596 · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106, 2021.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Winogrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106, 2021

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.618993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.618993Z digest=sha256:38357fceb083a4df122046a0abd4b729afd7f8f1c16ca13c5a2fa8c913ab14be

Observation 41c39a13-2c28-46cd-89d4-ccd9cc75da88 · outbound

This paper cites Battling Misinformation: An Empirical Study on Adversarial Factuality in Open-Source Large Language Models.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Battling Misinformation: An Empirical Study on Adversarial Factuality in Open-Source Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.622334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.622334Z digest=sha256:dff2851b0085c244e0e341da3645ef82defde5e22d9043bc8d7131af35a51576

Observation 31e67fda-94ee-45ba-bbfe-e651cdfd53fb · outbound

This paper cites an unresolved cited work.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:10:46.553778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.625638Z digest=sha256:47c9d6f4036e9065b0b9a0db20944c1ae5aaa11f3cbb4f55af141b7e48c65fdb

Observation 842dfd1d-5e0f-46ac-b200-c9efd47354d8 · outbound

This paper cites Practices for governing agentic ai systems.Research Paper, OpenAI, December, 2023.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Practices for governing agentic ai systems.Research Paper, OpenAI, December, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.535097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.628595Z digest=sha256:5f72c0cb25b5109a7b01ce6dea4cd9d1baf316350cfadc460522a02920a1c530

Observation 5e5cffbc-a1a3-4d63-a9b6-61944c43f564 · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.632368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.632368Z digest=sha256:bc61eb4821e2b95dfb6743d3a280e3ccaf6bd2b34a3cf0ee178da7d1cfbb8271

Observation e7aa50e3-d305-4872-9050-3b414728ed19 · outbound

This paper cites Inference scaling flaws: The limits of llm resampling with imperfect verifiers.arXiv preprint arXiv:2411.17501, 2024.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Inference scaling flaws: The limits of llm resampling with imperfect verifiers.arXiv preprint arXiv:2411.17501, 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.636111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.636111Z digest=sha256:d748235916766868cde455e0cf420f3c9efe5e19b77b939d7d7a02f24f762772

Observation 659746c9-1a21-40c3-9810-e03b8f945461 · outbound

This paper cites Gemma 3 Technical Report.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Gemma 3 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.639268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.639268Z digest=sha256:b630999ea20d15bc436473a5dc8bd00851feb321b7c2243c5b11d3f602d9f1e5

Observation 9b1b878d-ba31-4c7f-bde1-3703747c52e4 · outbound

This paper cites An in- context learning agent for formal theorem-proving.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows An in- context learning agent for formal theorem-proving

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.513148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.643196Z digest=sha256:449447e4c549742cee1aa3a7993b03c76c9e4469aed7a6696c20e64e317e3254

Observation 2c8095f3-cbcc-4b22-9706-b030bf3a72e6 · outbound

This paper cites Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.650713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.650713Z digest=sha256:a880e97662dae0815a79a69464a99a7ef94d9c32f81f6eaad16839ec6cf4ca37

Observation df1619c1-a459-40f1-b5a0-6510583c4d5c · outbound

This paper cites Toward self-improvement of llms via imagination, searching, and criticizing.Advances in Neural Information Processing Systems, 37:52723–52748, 2024.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Toward self-improvement of llms via imagination, searching, and criticizing.Advances in Neural Information Processing Systems, 37:52723–52748, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.654664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.654664Z digest=sha256:cd43a5a9d7f2df8831afa7d0f8a1f8dcfb6042da2541ec0a7dd7150338f2921e

Observation ca280cb7-08fe-4757-b276-fb36c5ca5541 · outbound

This paper cites NewsQA: A Machine Comprehension Dataset.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows NewsQA: A Machine Comprehension Dataset

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.657768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.657768Z digest=sha256:e5ac2f6985bf3f49c349d955580e59b77719e399e19d52c328763535122b2ccd

Observation 02705d8d-9e81-4e3e-a8e6-e51996d1acca · outbound

This paper cites LEGO-prover: Neural theorem proving with growing libraries.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows LEGO-prover: Neural theorem proving with growing libraries

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.462307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.661889Z digest=sha256:201695cfa4918e5605bf9bb750462b6e320bcb9d2c55846c461d2c5fee51182a

Observation 304c27ca-984c-4766-aaad-e5d1b58d3881 · outbound

This paper cites Resolving Knowledge Conflicts in Large Language Models.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Resolving Knowledge Conflicts in Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.665136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.665136Z digest=sha256:133dab4d4537685307d2d0ce34b832533da73923a41d094ff2bec37c9984d0da

Observation ce2080a2-9371-4e44-8d81-267562a7ef10 · outbound

This paper cites Measuring short-form factuality in large language models.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Measuring short-form factuality in large language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.669134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.669134Z digest=sha256:669e286234e2690331acdf13961941cdb31fba8c1a050dad3d9c8175d17d4604

Observation 8160064c-d682-4a19-81db-c13f2bdcfb9e · outbound

This paper cites Simple synthetic data reduces sycophancy in large language models.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Simple synthetic data reduces sycophancy in large language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.672950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.672950Z digest=sha256:8bac6fb6b6ee9acc4642b5206f661d832720d9ba498b7cec69e0429537b6c417

Observation 257dc43e-dd17-4bd5-a5ea-9a73874ed609 · outbound

This paper cites Examining inter-consistency of large language models collaboration: An in-depth analysis via debate.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Examining inter-consistency of large language models collaboration: An in-depth analysis via debate

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.422874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.676470Z digest=sha256:4db91f6d25255041396f17665950e83903e7dd7514d35cad973446ff02a9df39

Observation 3503e66b-4cd2-435a-bd88-7619c505b8ca · outbound

This paper cites Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.679960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.679960Z digest=sha256:58afdf39d2d046fad7abdcedd4a1ef95d6c280941b428b8093e0c682cd18561b

Observation 58d5fa96-f9fb-404b-884f-7bb3d5894e24 · outbound

This paper cites The Earth is Flat because...: Investigating LLMs' Belief towards Misinformation via Persuasive Conversation.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows The Earth is Flat because...: Investigating LLMs' Belief towards Misinformation via Persuasive Conversation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.684166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.684166Z digest=sha256:5e707fc9af407f96886b67c7df20ec315e30ef0f09dc429aea5c26cf2d9e8bf1

Observation 2452f366-8686-4d5c-acce-2afacc8d01fa · outbound

This paper cites Knowledge Conflicts for LLMs: A Survey.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Knowledge Conflicts for LLMs: A Survey

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.688349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.688349Z digest=sha256:39ba80c1fdd44b333c43b2f2c046406f7cac9539b92b3832c46d93f999a86f57

Observation 03d51244-6295-4ae4-a1b1-348b5a65cd5e · outbound

This paper cites Qwen2.5 Technical Report.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Qwen2.5 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.691942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.691942Z digest=sha256:8f1d8678d32d0369341a4305ce7e4da353770d4da07c27f60727b55b3206ea58

Observation 2c4865fa-1973-4da3-a5e3-4519ccad2170 · outbound

This paper cites Generating Natural Language Proofs with Verifier-Guided Search.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Generating Natural Language Proofs with Verifier-Guided Search

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.695856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.695856Z digest=sha256:6dbe91c131b1a3cdc400de3851f3d34635c8cd576cb89c42e205928a45611bc5

Observation 4df5328b-2433-4d1d-b33b-c9d76a162aa4 · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.699426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.699426Z digest=sha256:b799ada1b4d0fbcb81865d19e322c2aaec3e752c948eb14121f37bf0b3086145

Observation ee76ef9b-3d3e-4419-8144-19406373b589 · outbound

This paper cites LLMCrit: Teaching large language models to use criteria.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows LLMCrit: Teaching large language models to use criteria

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.383357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.702959Z digest=sha256:638b6103b7db9e1eaab6cc46383b6f41991f6f14baa81d194803e3d1c9892581

Observation 85143046-42f5-45bf-a161-4eab3dea9688 · outbound

This paper cites AFlow: Automating agentic workflow generation.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows AFlow: Automating agentic workflow generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.706946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.706946Z digest=sha256:7f4adcf2b1d1aee8ef674fea4a8c0ef03f6ed951fd7fb88dec9890858b6f3377

Observation 83e44c2b-a4f8-483e-92bb-2e3918ca5c43 · outbound

This paper cites SituatedQA: Incorporating extra-linguistic contexts into QA.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows SituatedQA: Incorporating extra-linguistic contexts into QA

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.710982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.710982Z digest=sha256:abd68ebae336dae5c880fd651f02fc28ad5cb97c960b5ee2a2c65498ad1e3654

Observation 7036bc86-9c64-44ad-9f76-3a3870d4ac62 · outbound

This paper cites Position- aware attention and supervised data improve slot filling.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Position- aware attention and supervised data improve slot filling

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.364770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.714283Z digest=sha256:88fd0b71da2c8a17432f6a538f7bca0a259334bce3837cecce7f6a3674705932

Observation 9dc70918-abe1-4e64-99f7-dbe0d6ade062 · outbound

This paper cites Merging Generated and Retrieved Knowledge for Open-Domain QA.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Merging Generated and Retrieved Knowledge for Open-Domain QA

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T11:10:45.811721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.717582Z digest=sha256:b10d964edb940bbb7fc38def91c214d37983b1c28fd27f5e414311baa42cd382

Observation fb463a37-ff55-4bdb-8faa-15fda3878e57 · outbound

This paper cites an unresolved cited work.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:10:46.355427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.722403Z digest=sha256:c2091c10ba7510dee45d46b130c8f13f25698ace85ed73a38076ef12d0ae3d76

Observation ee76e201-3308-4478-ad17-b3b5c3a75b4e · outbound

This paper cites an unresolved cited work.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:10:46.346548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.725664Z digest=sha256:57485e8be179a8e90a027e28ce441a6c0bc6325d627d028e6cd10346fc688e04

Observation 572d83ec-8ba1-4651-90bf-354c801e2a6b · outbound

This paper cites studies” or “statistics.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows studies” or “statistics

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.336953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.729976Z digest=sha256:b99a46250f239f74806e87cb146419f3145ed2e586ae1938b819ecefbec2d016

Observation 035fbc01-6134-4025-922f-234761dc69a8 · outbound

This paper cites an unresolved cited work.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:10:46.326832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.734594Z digest=sha256:778b7ec217a6d5a3eaf4893c0f9f060511c9d416ea376dc06abf6fe58903a086

Observation 4d7d511f-fe48-42b9-b46d-cb6b044086aa · outbound

This paper cites expert opinions.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows expert opinions

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.316400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.738482Z digest=sha256:fe751a794c6d601a162c7b13ab56f7886980ee1c4a75517cda14fec02d67398a

Observation 10b453a9-9b0f-49ce-820a-680c620ab5b3 · outbound

This paper cites an unresolved cited work.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:10:46.307058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.741757Z digest=sha256:ffed5274b7f947b4cc6b88ace683e135652980d99cc6ad06428ebf3e907f3bc0

Observation 876b9632-643b-4d15-9a95-236dea98d87f · outbound

This paper cites an unresolved cited work.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:10:46.297727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.745532Z digest=sha256:ad879366734e0a760eb423d14594ef28f5e0ea033ca7f9180b852c73d95f6b10

Observation abeea410-1d77-4113-921e-e055787fb5c6 · outbound

This paper cites an unresolved cited work.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:10:46.287737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.749904Z digest=sha256:ed526da1b996f5000bc7149f4282076a54b4afb7a44006f22fa70c56eacf4b6e

Observation 9e052475-cf78-4112-8276-42ded6e7cda6 · outbound

This paper cites Be creative and ruthless in your criticism.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Be creative and ruthless in your criticism

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.277779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.753651Z digest=sha256:c14ea3f1cd314e9ab05ec149c40b10e341d188eb2ed677b519155582e26728d8

Observation a6288341-3502-4487-97de-6f57e5dd6779 · outbound

This paper cites Are you sure about this? I don’t think this answer is correct because.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Are you sure about this? I don’t think this answer is correct because

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.266954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.756990Z digest=sha256:3067458206173f7847c2c70af9cb615e98621d404f8f263076c3f57a76f45613

Observation a240727e-2006-44d5-9ecd-1242c28fe83e · outbound

This paper cites This conclusion seems hasty. What if.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows This conclusion seems hasty. What if

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.256950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.760176Z digest=sha256:6016f6c76cbd08fe3047eef8b542f379658480b1b918b0d1229e107c2db854c8

Observation bffffb8b-2a6b-4d8d-a43b-f864042d6cca · outbound

This paper cites I don’t think this follows logically because.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows I don’t think this follows logically because

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.245063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.763239Z digest=sha256:cce442e0c6e85f7986e5da6f33afb0e9643d92fec6776bd1f2ffa822ecdb4001

Observation 5eb77fa0-ba39-4b9c-a33d-d088f8b89736 · outbound

This paper cites an unresolved cited work.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Unresolved cited work

Reference 70

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:10:46.235130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.767347Z digest=sha256:57fb50521eaa0ab3470f94df943d5d70c7f8b2a91d959590492a660687ad1edf

Observation 8e63d4c5-bcba-478c-b046-9e21234a13b8 · outbound

This paper cites an unresolved cited work.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:10:46.492842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:10:45.647597Z digest=sha256:aca9bd453f332966fba289a092c9aa25e86acc59e519e5306241b5358ec5d8df

Pith citing papers

Observation 7789b71c-de9e-4cfe-a038-da6c13ecd995 · inbound

The Authorization-Execution Gap Is a Major Safety and Security Problem in Open-World Agents cites this paper.

The Authorization-Execution Gap Is a Major Safety and Security Problem in Open-World Agents Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:22:01.851015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T01:19:48.189694Z digest=sha256:72365d19ee080c9f24b50e109a705082befed8e5097a78a8a551c261b09e5dee

Observation 984043c9-0be9-42eb-b35b-bf409fdce334 · inbound

Why Are Agentic Pull Requests Merged or Rejected? An Empirical Study cites this paper.

Why Are Agentic Pull Requests Merged or Rejected? An Empirical Study Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T03:54:34.006576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T03:53:03.807297Z digest=sha256:139705bc6f172ec6610e01f2b41965d98a0a8c63c943597c6079c2a30de80389