Pith. sign in

Paper Citation Record · LEDGER

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs

As of 5 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 1 inbound Pith citation observation for arXiv:2604.22937.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.22937 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T11:33:21.391661Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-10T19:01:21.650333Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T19:07:35.241005Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact20
  • verified fuzzy3
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c46bb44e-3c1a-4045-b196-572805592936 · outbound

This paper cites online" 'onlinestring :=.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs online" 'onlinestring :=

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T14:07:51.807847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:82762ed4be6ae82c2d9292e4ee6947c1218ceeded724effb0f6e7061b0be7fb2

Observation 36aafd43-bd3d-4c22-830f-d6076901e1b1 · outbound

This paper cites write newline.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs write newline

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T14:07:51.829409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:4a531c5592f456fdc0465e6fdc1f7ea0486a5130124bcaa14a705a9bc7922db1

Observation 44cdc814-62bf-4033-9e5f-a12af7dc4de3 · outbound

This paper cites MathArena: Evaluating LLMs on Uncontaminated Math Competitions.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs MathArena: Evaluating LLMs on Uncontaminated Math Competitions

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:10:14.938194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:82d1b0d5c83dbcc21a9945b7774922b8fa40c8843300217f38ee22583234902e

Observation 7407d28d-99e4-4b14-b14a-af2bcd5a30c8 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:36:13.750346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:e169c066dc7074694d627641dad99fd694b05e1dff0a46cce29f0cac4c40ac36

Observation 1ebfef7f-f252-4e21-aae1-29f8a5981718 · outbound

This paper cites Scoring Verifiers: Evaluating Synthetic Verification for Code and Reasoning.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Scoring Verifiers: Evaluating Synthetic Verification for Code and Reasoning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:36:13.787691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:2aeb7076ec8e5baac79695aed54037af40deb7eeb32f444fa986ba79ac9738b4

Observation d8ea1117-5b30-456d-8dac-147a32ff1d04 · outbound

This paper cites an unresolved cited work.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-26T14:07:51.832178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:42d8cc68a34922061d545e94dca141f62f7f362153a3d6d0552124c403b7e369

Observation 7ecbb288-1d78-4a7c-89a9-b47f84a32448 · outbound

This paper cites an unresolved cited work.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-26T14:07:51.820152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:f066428c7f9a06e1b8f9286d2609e21877bd0b612bc2d3a1ec5cffa9acc4ac71

Observation 21cfe2ab-2b31-4a46-97d2-828433491170 · outbound

This paper cites Beyond oracle: Verifier-supervision for instruction hierarchy in reasoning and instruction-tuned llms.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Beyond oracle: Verifier-supervision for instruction hierarchy in reasoning and instruction-tuned llms

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T14:07:51.813014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:3ab769eb41db6b6b94911e73a0aa8553de46567a2e751cc193e51f2076e6ba14

Observation 48029d7e-f467-4632-9309-499a21e736a7 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:36:13.731528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:de7d77f5241710da73bffad9ef4a7812101a218a22259f78e4ff4328bb41d9bd

Observation 2ac75d33-6d5b-4bf5-b07b-3fe8b882cb8c · outbound

This paper cites an unresolved cited work.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-26T14:07:51.835109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:ea259f23efc79a4e796d2a97325674e9383d89ca7bf2e01b74c4eb57cab1be59

Observation 45661ab3-f3cb-4ea9-a9c9-efe335a0f535 · outbound

This paper cites Process reward models that think.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Process reward models that think

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:36:13.762894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:b776e4a9627d73db1f170d0bb121b9074546c78873640da48f91750b3f216054

Observation f128a48d-731f-46f1-8d57-728379350875 · outbound

This paper cites How to Correctly Report LLM-as-a-Judge Evaluations.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs How to Correctly Report LLM-as-a-Judge Evaluations

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-02T03:04:02.584259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:b7f3b9e344d0139518144a22c34a5c304d97a833df346af91cbfe9f9169778cc

Observation e763ad97-8017-4944-a0c3-90190430f9c1 · outbound

This paper cites an unresolved cited work.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-26T14:07:51.815193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:bd06d042107daff59aacc0a236ae778150afe98bc4931bf022cc8418900d2abd

Observation ec19ef6e-2772-44da-aae4-b49188f278c2 · outbound

This paper cites LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:36:13.714408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:7ba32a73168c7d6917d7da62b3e52f1d4016e6a2d33adbd195a295f0dda3064e

Observation 8ad72153-e9aa-4dd1-a290-1e3c637781a2 · outbound

This paper cites an unresolved cited work.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-26T14:07:51.805284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:aced7559971512becc6df3e30697847d90a1fb87a03cb11f5c84548fc891bf97

Observation 768b5491-8865-4ee8-963c-b45f37ef4ee0 · outbound

This paper cites Autoharness: improving llm agents by automatically synthesizing a code harness.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Autoharness: improving llm agents by automatically synthesizing a code harness

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:36:13.775751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:9f9bad63f33f7c7a4ff2a41ec6141e41094147c13cbeb554b157bd035fe43a0d

Observation a75a7856-c5f8-409c-b872-eaec0a77c1f8 · outbound

This paper cites an unresolved cited work.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-26T14:07:51.799298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:4fe96dc2e4cac2246fcd7fc28ee88e423b1c3d19d523b89d67d3400ee272face

Observation 3a597ab0-7276-4732-b31c-d285b5cfc3a8 · outbound

This paper cites an unresolved cited work.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-26T14:07:51.810638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:321879a9db5590028bd64a7953728851c2938641789981e9415004abc338eed6

Observation 91b8ae1e-4a0e-402a-a76c-98d6d5cb03de · outbound

This paper cites Natural-Language Agent Harnesses.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Natural-Language Agent Harnesses

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:04:30.082400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:3c1942404e58d5acd12a2aad3bfdaf27d1bb71f1edc6b7930668b9eb10610210

Observation 89ed779a-d700-45f5-963a-a3e25281d343 · outbound

This paper cites an unresolved cited work.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-26T14:07:51.822974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:03424d4b2126c8687804cf656cbe3a9b516ad788911d41d2f5c79c9026100e97

Observation 67002705-f036-44ce-95a4-3a7e5675ba47 · outbound

This paper cites Beyond outcome verification: Verifiable process reward models for structured reasoning.arXiv preprint arXiv:2601.17223.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Beyond outcome verification: Verifiable process reward models for structured reasoning.arXiv preprint arXiv:2601.17223

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:36:13.725542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:3deba40c38645c0bd910044c707be5451c1b0216dfa5ebac5f6e0e7482dd831d

Observation bc17e442-b7b8-4bf7-a2e3-e26fa3ebcc59 · outbound

This paper cites Generalizing Verifiable Instruction Following.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Generalizing Verifiable Instruction Following

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:36:13.699446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:9dde7c2106e047fbdfc794370a8b09190efcbcfadc6750220e074df53aa49dc8

Observation 85b937d4-e414-4c45-83ff-e793b734385e · outbound

This paper cites OpenAI GPT-5 System Card.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs OpenAI GPT-5 System Card

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:36:13.679536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:72eb0d3b0cba755bc3d11c13d0a52555a8abc8983dff81cac041af1cbd85cf7f

Observation 119e44ff-95f2-4253-a83a-9b506b9a545c · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:36:13.660771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:c9253f0eaf7c45e6065cc0e14afcd0b0d7416b74235b744b095bca3b9e331043

Observation 145e1f44-a3ec-43c5-a78b-146fd6307c3c · outbound

This paper cites Trust- judge: Inconsistencies of LLM-as-a-judge and how to alleviate them.arXiv preprint arXiv:2509.21117,.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Trust- judge: Inconsistencies of LLM-as-a-judge and how to alleviate them.arXiv preprint arXiv:2509.21117,

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:36:13.649003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:02f1342bf7d04524f7ef694a5fe2e04ca4f9aee3309dcd1f996de83852a405b4

Observation 51f8e7e6-b35a-42e7-a489-d3107768fd20 · outbound

This paper cites an unresolved cited work.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-05-26T14:07:51.826464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:7a70d78c456acd690f88266bbaf21f2e19ed1bcfd5bf55bfceff05df6bf1f233

Observation 1017f8af-d607-4e1b-b968-50192fc8f6da · outbound

This paper cites StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:36:13.704509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:b67b86fa6bee0c933bf2f7fc15b03cb6a79aecc6577efb073fa55be5a7e36001

Observation 7875a516-346b-4f3a-8e20-5c64e9321aca · outbound

This paper cites an unresolved cited work.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-05-26T14:07:51.802944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:7a66201fecf341e85b37b131a526b8560883d242cdecc5c2c30d1eeabd97d774

Observation 099524b1-63dd-425c-b722-87a3f0b40fdd · outbound

This paper cites MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:10:22.017517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:b747e6aba4128055fb0488000b28caa6e9b96491309e1ea427dd4c073899d6f6

Observation 92666ef6-494a-4a0a-be5b-9eb7e62098bd · outbound

This paper cites AgentV-RL: Scaling Reward Modeling with Agentic Verifier.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs AgentV-RL: Scaling Reward Modeling with Agentic Verifier

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:36:13.694509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:fa31808505142c2d412374c7f82b889a61c1d055b0db4b75ec356bbff27da573

Observation fd5e532e-81b4-4da3-aca3-7a62b4d7a490 · outbound

This paper cites Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T16:44:08.223014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:89be985f1db59811ed7ca987aa6148cdea615d7d0e70fb9eb2822a434ecc323e

Observation 00d1a3ea-39d5-4299-b112-6b2d5ca445a7 · outbound

This paper cites an unresolved cited work.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-05-26T14:07:51.817477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:138e51e94f6b71cb8acb53027d954abdf19354da336a60436752bf28d0d35141

Observation f3f8c785-c41e-426c-b7d3-cbd8aa881af9 · outbound

This paper cites Available: https://arxiv.org/abs/2603.11445.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs Available: https://arxiv.org/abs/2603.11445

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:36:13.668075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:81985793cf371ec53b3eb5fab54332d356163777eeb21aee12111c72364e88c9

Observation c90841e7-ff2f-4029-b6d1-7a368a59c826 · outbound

This paper cites ComplexFuncBench: Exploring Multi-Step and Constrained Function Calling under Long-Context Scenario.

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs ComplexFuncBench: Exploring Multi-Step and Constrained Function Calling under Long-Context Scenario

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:36:13.672811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:33:21.391661Z digest=sha256:4ea084df862be914132a2cf6845b4b31214ea8b2214ff9145425755788121633

Pith citing papers

Observation 939fb746-b159-48e3-a0cc-4d79b29405b3 · inbound

The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents cites this paper.

The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-10T19:07:35.244603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T19:01:21.650333Z digest=sha256:69452ecc2e7e7f96e8e0ca7c635be7de0efa3b3e7772ea3ea2467b3caf4254e2