Pith. sign in

Paper Citation Record · LEDGER

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows

As of 7 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 2 inbound Pith citation observations for arXiv:2506.03332.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03332 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:10:45.767347Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-22T03:53:03.807297Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T03:54:34.003611Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved47
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8b8395ea-c8c5-4af7-a3da-a1ff96abb61e · outbound

This paper cites ReConcile: Round-table conference improves reasoning via consensus among diverse LLMs.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows ReConcile: Round-table conference improves reasoning via consensus among diverse LLMs

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:48.042471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.517449Z digest=sha256:54e692d206cd5a65eed1e859e55e241f90f207c0120c8772dfb150b1418215bd

Observation 1361cfdc-fabe-4716-b7f6-8fefba645e89 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.521483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.521483Z digest=sha256:db3cbc5f6eb2402c9b2e8b0a5f70221276612eb0affa0e772b5f754b5f6288d4

Observation 1676adec-811d-4096-b08f-16bcf22e89b1 · outbound

This paper cites Synthetic disinformation attacks on automated fact verification systems.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Synthetic disinformation attacks on automated fact verification systems

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:47.859384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.524917Z digest=sha256:f5f297b2ecbca43ad18ba2767ca85400a49aa9b43d6b263fb531c1a807d7bbd8

Observation b8f5c337-7361-4ea6-bcfd-9825eb220d40 · outbound

This paper cites Improving Factuality and Reasoning in Language Models through Multiagent Debate.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Improving Factuality and Reasoning in Language Models through Multiagent Debate

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.528343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.528343Z digest=sha256:c80ed4061a9d95656b2eb3de736d7faeae4071f1979d2a8ab45e66fd2f096e53

Observation 3aba0d3b-023e-48d5-8fb5-9a0f63d0116c · outbound

This paper cites DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.531959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.531959Z digest=sha256:3350d13b8b35ffa0a0cb39ad1d3400ecade2512be3ad5f5e2bc130fec4d970c1

Observation f9c8f781-43ed-4575-98e9-6d9d5f450882 · outbound

This paper cites SearchQA: A New Q&A Dataset Augmented with Context from a Search Engine.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows SearchQA: A New Q&A Dataset Augmented with Context from a Search Engine

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.536054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.536054Z digest=sha256:6fe8d73c57431dcece6eb7017f015c4b36abdfa0f03fb5deadcef4e5f2f31b05

Observation c46706d3-b975-4a89-9062-bb9564179cb3 · outbound

This paper cites Baldur: Whole-Proof Generation and Repair with Large Language Models.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Baldur: Whole-Proof Generation and Repair with Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.540047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.540047Z digest=sha256:c6b916e64ade8a9b8bb1b78b44492a3bc9430360295e7f4ed3f4637416d8fd01

Observation a53f184f-325a-4c2d-87ab-0386c753d6f8 · outbound

This paper cites CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.544006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.544006Z digest=sha256:6c4439e53a63a34f3382389aa5c0b32debfbcb9edaff531d104c01c2766cc1dd

Observation 014c86a3-04a1-4df9-b5b7-e1cb9428fb57 · outbound

This paper cites The larger the better? improved LLM code-generation via budget reallocation.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows The larger the better? improved LLM code-generation via budget reallocation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:47.713786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.547462Z digest=sha256:747787dfe5521b682be6afee93cfb17d4f638fc60767ceba16d55c2628607f07

Observation 29a76a2f-4925-44d7-8582-e6a5c9d92acb · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Measuring Massive Multitask Language Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.550880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.550880Z digest=sha256:1b3b2819ae2e3a59cd4c33c8d2e4dba8dd623631cacad8b7d11165c350a951c4

Observation 424f8645-1142-4952-b467-92cd4ba4e026 · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Large Language Models Cannot Self-Correct Reasoning Yet

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.554308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.554308Z digest=sha256:82b6e61082644899ea79133375d82b9e3342706f4052eec11264010a98d36157

Observation 9e438a02-8bcd-42db-ad95-d3711708e956 · outbound

This paper cites GPT-4o System Card.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows GPT-4o System Card

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.557635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.557635Z digest=sha256:475f1727b22bceb8295abd80445888007bc15573b36ed5f109308e57590aa544

Observation 369e756e-bf62-4a64-828a-cc97e895e1f9 · outbound

This paper cites OpenAI o1 System Card.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows OpenAI o1 System Card

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.561129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.561129Z digest=sha256:2f889b68e49f963491297842fd9e60652ac2a26190efe20c06b8e576d5024a46

Observation 0d8abfcb-76f5-464c-ac2c-af43ba4c354c · outbound

This paper cites TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.564319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.564319Z digest=sha256:40f0d7bc7d58e44168193a8f5833a614f8582312e77d4daa404c307cfc360445

Observation 65cd85af-c690-480d-b6df-a537f449e9fd · outbound

This paper cites When can LLMs actually correct their own mistakes? a critical survey of self-correction of LLMs.Transactions of the Association for Computational Linguistics, 12:1417–1440, 2024.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows When can LLMs actually correct their own mistakes? a critical survey of self-correction of LLMs.Transactions of the Association for Computational Linguistics, 12:1417–1440, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:47.579967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.567712Z digest=sha256:aa8645790714cdcee607e1a446cf27203913b24fa5a1bb68f4facfb30951efb5

Observation 81f7c4be-d452-422d-9d8f-d5bd2c12d249 · outbound

This paper cites AI Agents That Matter.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows AI Agents That Matter

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.571170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.571170Z digest=sha256:c5e9868eb82879c2cd0a6260ba842253c8791dd68261c728e0289f090c53c201

Observation ec74cd6d-d6ba-485c-be06-83020276e60d · outbound

This paper cites Bowman, Tim Rocktäschel, and Ethan Perez.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Bowman, Tim Rocktäschel, and Ethan Perez

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:47.375433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.574364Z digest=sha256:1b5bc8111663c19c95f1e3cf861f11ba860f1861b25b6ec2b45dfd67d71672f3

Observation 52cdb31c-8fde-43c1-9c84-baa153c6daca · outbound

This paper cites Natural questions: a benchmark for question answering research.Transactions of the Association for Computational Linguistics, 7:453–466, 2019.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Natural questions: a benchmark for question answering research.Transactions of the Association for Computational Linguistics, 7:453–466, 2019

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.577818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.577818Z digest=sha256:149a3c7aea8dfa1c690fb254aa6d39bd3f1dbf1b00d85499666efddc99d700e4

Observation 70b948cd-baec-4553-8abd-ef315209b3b4 · outbound

This paper cites Making language models better reasoners with step-aware verifier.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Making language models better reasoners with step-aware verifier

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.581968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.581968Z digest=sha256:7f5605e12277068dfa9ce91f877f6e1436dd55d2807cb4480f453d97be56df01

Observation 3e3d13e9-9249-4e03-8909-d85da7e37c06 · outbound

This paper cites Encouraging divergent thinking in large language models through multi- agent debate.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Encouraging divergent thinking in large language models through multi- agent debate

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:47.161684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.585154Z digest=sha256:7e83d629c84100611551fdbea3d6bf5297ad54f25d9725633096807467cf9594

Observation b58eb157-5c5c-40ba-b168-9470b7f9814c · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Self-Refine: Iterative Refinement with Self-Feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.589027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.589027Z digest=sha256:1e84f220427255dcbf74a4c7e1bf4bd3e610ca94c36b7521721ed0d1f83211cb

Observation f6843d3e-326c-48dd-9049-be2657dc59ad · outbound

This paper cites Debate Helps Supervise Unreliable Experts.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Debate Helps Supervise Unreliable Experts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.592227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.592227Z digest=sha256:2f9bcc6b6d3ceaa627003269c5c68ded22c6442a19a57dd96a1d3b0eabbb84d7

Observation b140cf55-d037-4299-a99d-98c319d50c95 · outbound

This paper cites Faitheval: Can your language model stay faithful to context, even if ”the moon is made of marshmallows”.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Faitheval: Can your language model stay faithful to context, even if ”the moon is made of marshmallows”

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.993915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.595523Z digest=sha256:ece53542cc47304781f2198f31102e167a1415a379d33a95e4631f671abc2e9b

Observation 0e540f63-cb26-4645-bd04-120ec0f23f93 · outbound

This paper cites Lever: Learning to verify language-to-code generation with execution.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Lever: Learning to verify language-to-code generation with execution

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.847888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.598687Z digest=sha256:7d736572a3dbcbbd157dca0dd6710d996769e79f7878b2bd3dfde39b80ae34de

Observation a94a1d61-8e0d-470a-8a80-68e877edd8e6 · outbound

This paper cites On the Risk of Misinformation Pollution with Large Language Models.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows On the Risk of Misinformation Pollution with Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.601797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.601797Z digest=sha256:3d4ec4229f7fdaac6f001b37e7227b62d036784b4d535212d797b3332a4c2761

Observation 56114032-d189-4cae-8ef5-06c0cdd39060 · outbound

This paper cites OffsetBias: Leveraging Debiased Data for Tuning Evaluators.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.605360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.605360Z digest=sha256:7b89beff9162a469441aa83f03c6f0398c63cfa3d1f3ece2dace5d9b26552ba1

Observation 8c3f49e0-dbb3-4f3a-828b-3ad7cfbcd84d · outbound

This paper cites Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.671505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.608726Z digest=sha256:0cd44781485a495c9ff1f55699ab141e4957da7bd64c328e3efa5e4ebad4e07f

Observation a25700a9-0615-4de9-a465-29cc43230d9a · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Gpqa: A graduate-level google-proof q&a benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.611812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.611812Z digest=sha256:c42c6d7a074010a323223ae875ddd8fa717649f80e118454e407e2c52198301d

Observation f78f1aca-f91e-42a5-b10a-e3c0dfaf3c88 · outbound

This paper cites Archon: An Architecture Search Framework for Inference-Time Techniques.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Archon: An Architecture Search Framework for Inference-Time Techniques

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.615625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.615625Z digest=sha256:839d6bfaf8a796220c2c7098f9c5ec574aaa7f6f6ff13b6014d128a83ecf12e6

Observation 78054db9-5947-4319-9573-31d6d4e47596 · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106, 2021.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Winogrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106, 2021

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.618993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.618993Z digest=sha256:38357fceb083a4df122046a0abd4b729afd7f8f1c16ca13c5a2fa8c913ab14be

Observation 41c39a13-2c28-46cd-89d4-ccd9cc75da88 · outbound

This paper cites Battling Misinformation: An Empirical Study on Adversarial Factuality in Open-Source Large Language Models.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Battling Misinformation: An Empirical Study on Adversarial Factuality in Open-Source Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.622334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.622334Z digest=sha256:dff2851b0085c244e0e341da3645ef82defde5e22d9043bc8d7131af35a51576

Observation 31e67fda-94ee-45ba-bbfe-e651cdfd53fb · outbound

This paper cites an unresolved cited work.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:10:46.553778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.625638Z digest=sha256:d1cd9775be1327672f48a3bea04ab6b1a3c6e51eeb4a1661982ffadcc67941c7

Observation 842dfd1d-5e0f-46ac-b200-c9efd47354d8 · outbound

This paper cites Practices for governing agentic ai systems.Research Paper, OpenAI, December, 2023.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Practices for governing agentic ai systems.Research Paper, OpenAI, December, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.535097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.628595Z digest=sha256:cf669d1718199a506f8897d5603846dbcb9b7abdbc5ce17273339c7fb25c404a

Observation 5e5cffbc-a1a3-4d63-a9b6-61944c43f564 · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.632368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.632368Z digest=sha256:bc61eb4821e2b95dfb6743d3a280e3ccaf6bd2b34a3cf0ee178da7d1cfbb8271

Observation e7aa50e3-d305-4872-9050-3b414728ed19 · outbound

This paper cites Inference scaling flaws: The limits of llm resampling with imperfect verifiers.arXiv preprint arXiv:2411.17501, 2024.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Inference scaling flaws: The limits of llm resampling with imperfect verifiers.arXiv preprint arXiv:2411.17501, 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.636111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.636111Z digest=sha256:d748235916766868cde455e0cf420f3c9efe5e19b77b939d7d7a02f24f762772

Observation 659746c9-1a21-40c3-9810-e03b8f945461 · outbound

This paper cites Gemma 3 Technical Report.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Gemma 3 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.639268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.639268Z digest=sha256:b630999ea20d15bc436473a5dc8bd00851feb321b7c2243c5b11d3f602d9f1e5

Observation 9b1b878d-ba31-4c7f-bde1-3703747c52e4 · outbound

This paper cites An in- context learning agent for formal theorem-proving.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows An in- context learning agent for formal theorem-proving

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.513148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.643196Z digest=sha256:4c4f3e7e6d4f59da41a37e182e45be267b270fcb4c82d05b254ed63bacf0126f

Observation 2c8095f3-cbcc-4b22-9706-b030bf3a72e6 · outbound

This paper cites Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.650713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.650713Z digest=sha256:a880e97662dae0815a79a69464a99a7ef94d9c32f81f6eaad16839ec6cf4ca37

Observation df1619c1-a459-40f1-b5a0-6510583c4d5c · outbound

This paper cites Toward self-improvement of llms via imagination, searching, and criticizing.Advances in Neural Information Processing Systems, 37:52723–52748, 2024.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Toward self-improvement of llms via imagination, searching, and criticizing.Advances in Neural Information Processing Systems, 37:52723–52748, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.654664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.654664Z digest=sha256:cd43a5a9d7f2df8831afa7d0f8a1f8dcfb6042da2541ec0a7dd7150338f2921e

Observation ca280cb7-08fe-4757-b276-fb36c5ca5541 · outbound

This paper cites NewsQA: A Machine Comprehension Dataset.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows NewsQA: A Machine Comprehension Dataset

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.657768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.657768Z digest=sha256:e5ac2f6985bf3f49c349d955580e59b77719e399e19d52c328763535122b2ccd

Observation 02705d8d-9e81-4e3e-a8e6-e51996d1acca · outbound

This paper cites LEGO-prover: Neural theorem proving with growing libraries.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows LEGO-prover: Neural theorem proving with growing libraries

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.462307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.661889Z digest=sha256:9746cda57b4f439260dc5391cc54d35a59adf9160a6c8ec1d45430234bb5e728

Observation 304c27ca-984c-4766-aaad-e5d1b58d3881 · outbound

This paper cites Resolving Knowledge Conflicts in Large Language Models.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Resolving Knowledge Conflicts in Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.665136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.665136Z digest=sha256:133dab4d4537685307d2d0ce34b832533da73923a41d094ff2bec37c9984d0da

Observation ce2080a2-9371-4e44-8d81-267562a7ef10 · outbound

This paper cites Measuring short-form factuality in large language models.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Measuring short-form factuality in large language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.669134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.669134Z digest=sha256:669e286234e2690331acdf13961941cdb31fba8c1a050dad3d9c8175d17d4604

Observation 8160064c-d682-4a19-81db-c13f2bdcfb9e · outbound

This paper cites Simple synthetic data reduces sycophancy in large language models.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Simple synthetic data reduces sycophancy in large language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.672950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.672950Z digest=sha256:8bac6fb6b6ee9acc4642b5206f661d832720d9ba498b7cec69e0429537b6c417

Observation 257dc43e-dd17-4bd5-a5ea-9a73874ed609 · outbound

This paper cites Examining inter-consistency of large language models collaboration: An in-depth analysis via debate.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Examining inter-consistency of large language models collaboration: An in-depth analysis via debate

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.422874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.676470Z digest=sha256:d31f17a1c0b3fd8db8b2f4f76516752aeed3d8c72cdffcc497ec2779b8256bb7

Observation 3503e66b-4cd2-435a-bd88-7619c505b8ca · outbound

This paper cites Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.679960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.679960Z digest=sha256:58afdf39d2d046fad7abdcedd4a1ef95d6c280941b428b8093e0c682cd18561b

Observation 58d5fa96-f9fb-404b-884f-7bb3d5894e24 · outbound

This paper cites The Earth is Flat because...: Investigating LLMs' Belief towards Misinformation via Persuasive Conversation.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows The Earth is Flat because...: Investigating LLMs' Belief towards Misinformation via Persuasive Conversation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.684166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.684166Z digest=sha256:5e707fc9af407f96886b67c7df20ec315e30ef0f09dc429aea5c26cf2d9e8bf1

Observation 2452f366-8686-4d5c-acce-2afacc8d01fa · outbound

This paper cites Knowledge Conflicts for LLMs: A Survey.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Knowledge Conflicts for LLMs: A Survey

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.688349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.688349Z digest=sha256:39ba80c1fdd44b333c43b2f2c046406f7cac9539b92b3832c46d93f999a86f57

Observation 03d51244-6295-4ae4-a1b1-348b5a65cd5e · outbound

This paper cites Qwen2.5 Technical Report.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Qwen2.5 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.691942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.691942Z digest=sha256:8f1d8678d32d0369341a4305ce7e4da353770d4da07c27f60727b55b3206ea58

Observation 2c4865fa-1973-4da3-a5e3-4519ccad2170 · outbound

This paper cites Generating Natural Language Proofs with Verifier-Guided Search.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Generating Natural Language Proofs with Verifier-Guided Search

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.695856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.695856Z digest=sha256:6dbe91c131b1a3cdc400de3851f3d34635c8cd576cb89c42e205928a45611bc5

Observation 4df5328b-2433-4d1d-b33b-c9d76a162aa4 · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.699426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.699426Z digest=sha256:a6a059309c730230bac116d678110abc364a8e81df4f1825a005b73d4b47cc62

Observation ee76ef9b-3d3e-4419-8144-19406373b589 · outbound

This paper cites LLMCrit: Teaching large language models to use criteria.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows LLMCrit: Teaching large language models to use criteria

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.383357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.702959Z digest=sha256:b61410caab7f23fd9d7100095f21a9337b5a8040247e5a339f6a448e2c1591ed

Observation 85143046-42f5-45bf-a161-4eab3dea9688 · outbound

This paper cites AFlow: Automating agentic workflow generation.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows AFlow: Automating agentic workflow generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.706946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.706946Z digest=sha256:7f4adcf2b1d1aee8ef674fea4a8c0ef03f6ed951fd7fb88dec9890858b6f3377

Observation 83e44c2b-a4f8-483e-92bb-2e3918ca5c43 · outbound

This paper cites SituatedQA: Incorporating extra-linguistic contexts into QA.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows SituatedQA: Incorporating extra-linguistic contexts into QA

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:45.710982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:10:45.710982Z digest=sha256:abd68ebae336dae5c880fd651f02fc28ad5cb97c960b5ee2a2c65498ad1e3654

Observation 7036bc86-9c64-44ad-9f76-3a3870d4ac62 · outbound

This paper cites Position- aware attention and supervised data improve slot filling.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Position- aware attention and supervised data improve slot filling

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.364770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.714283Z digest=sha256:cc68c3e1bce6435a21057d8bbc33ddb56fbac286472ba90d83a1bdb052f4764d

Observation 9dc70918-abe1-4e64-99f7-dbe0d6ade062 · outbound

This paper cites Merging Generated and Retrieved Knowledge for Open-Domain QA.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Merging Generated and Retrieved Knowledge for Open-Domain QA

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T11:10:45.811721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.717582Z digest=sha256:397e2c0f37c2afacae99957b7b646db56fec04cf78499dc85da3dbebcdc1bae4

Observation fb463a37-ff55-4bdb-8faa-15fda3878e57 · outbound

This paper cites an unresolved cited work.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:10:46.355427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.722403Z digest=sha256:193f4344e67ccca60f523b380a84804cdb09dffa60ad0e83174b94d7db2edd47

Observation ee76e201-3308-4478-ad17-b3b5c3a75b4e · outbound

This paper cites an unresolved cited work.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:10:46.346548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.725664Z digest=sha256:fd37af466920c939f79e9d12b7f944ce4e675af559859b0769fdd60461684a76

Observation 572d83ec-8ba1-4651-90bf-354c801e2a6b · outbound

This paper cites studies” or “statistics.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows studies” or “statistics

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.336953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.729976Z digest=sha256:023654293c595d548ee95cc26588e3eddedc3c8d18cad66efd45e6996c849770

Observation 035fbc01-6134-4025-922f-234761dc69a8 · outbound

This paper cites an unresolved cited work.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:10:46.326832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.734594Z digest=sha256:5bddecf3bfcbb1733cc48710b91690626f037dd02cabb41ad210b1f5fcafd04b

Observation 4d7d511f-fe48-42b9-b46d-cb6b044086aa · outbound

This paper cites expert opinions.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows expert opinions

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.316400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.738482Z digest=sha256:7b74e07a4ca6ecf8f8dc241129b2b39e27eb150ce3c35e44805c98ac3c9ea7e1

Observation 10b453a9-9b0f-49ce-820a-680c620ab5b3 · outbound

This paper cites an unresolved cited work.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:10:46.307058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.741757Z digest=sha256:7759966a3c0e753c25724d6b3893c1fe128d2b8a5a2364fd94415042c64e8c17

Observation 876b9632-643b-4d15-9a95-236dea98d87f · outbound

This paper cites an unresolved cited work.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:10:46.297727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.745532Z digest=sha256:817d8f29b68d673a9726539c2a2dfb02f104f0fd8ac53cf2685aa8c1f37570fa

Observation abeea410-1d77-4113-921e-e055787fb5c6 · outbound

This paper cites an unresolved cited work.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:10:46.287737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.749904Z digest=sha256:b41d556d267058c758decbe2f4a30a9fc29de2f99983d9932a3ffd644014d9b8

Observation 9e052475-cf78-4112-8276-42ded6e7cda6 · outbound

This paper cites Be creative and ruthless in your criticism.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Be creative and ruthless in your criticism

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.277779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.753651Z digest=sha256:466518b549f6f2aab86eb15a65ca466aefd4d00fb7b59bcfcce070a6b0352e3c

Observation a6288341-3502-4487-97de-6f57e5dd6779 · outbound

This paper cites Are you sure about this? I don’t think this answer is correct because.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Are you sure about this? I don’t think this answer is correct because

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.266954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.756990Z digest=sha256:6f8648cbb8d03b531379412f3f0e978adf1b7a7d85333b225aaac0f1791430ea

Observation a240727e-2006-44d5-9ecd-1242c28fe83e · outbound

This paper cites This conclusion seems hasty. What if.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows This conclusion seems hasty. What if

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.256950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.760176Z digest=sha256:4a275ccbd95462860650a8c00fa13508b68e02f920042bdb79ceb7c72e5ddc5f

Observation bffffb8b-2a6b-4d8d-a43b-f864042d6cca · outbound

This paper cites I don’t think this follows logically because.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows I don’t think this follows logically because

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:10:46.245063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.763239Z digest=sha256:17345deaf53d04746bc0804e6d363d36a5dcb46665168bfe3525ce89fb02c902

Observation 5eb77fa0-ba39-4b9c-a33d-d088f8b89736 · outbound

This paper cites an unresolved cited work.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Unresolved cited work

Reference 70

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:10:46.235130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.767347Z digest=sha256:a7f29298e25d2036525d178b006ed291a07ac2f60bb59970cfac3c3ebefdaced

Observation 8e63d4c5-bcba-478c-b046-9e21234a13b8 · outbound

This paper cites an unresolved cited work.

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:10:46.492842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:10:45.647597Z digest=sha256:132c118de556a78a3214f0c9bba6147dd86661134f6276aeb73a2537f88ec63b

Pith citing papers

Observation 7789b71c-de9e-4cfe-a038-da6c13ecd995 · inbound

The Authorization-Execution Gap Is a Major Safety and Security Problem in Open-World Agents cites this paper.

The Authorization-Execution Gap Is a Major Safety and Security Problem in Open-World Agents Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:22:01.851015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:48.189694Z digest=sha256:99acdde5cd5eb69bfe09340ad29df087313051056edd7e495b35f115dd14a368

Observation 984043c9-0be9-42eb-b35b-bf409fdce334 · inbound

Why Are Agentic Pull Requests Merged or Rejected? An Empirical Study cites this paper.

Why Are Agentic Pull Requests Merged or Rejected? An Empirical Study Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T03:54:34.006576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T03:53:03.807297Z digest=sha256:4b1560bbbc9f4b0c915f0b029907dce1a8ad5e29c60c74d932c70da430a4dbe3