Pith. sign in

Paper Citation Record · LEDGER

Training Agents with Weakly Supervised Feedback from Large Language Models

As of 18 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2411.19547.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.19547 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:08:13.270416Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b0164537-341b-4af9-9018-a9ad87aa6898 · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

Training Agents with Weakly Supervised Feedback from Large Language Models Dota 2 with Large Scale Deep Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.103779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.103779Z digest=sha256:96571c53f44a0f0bcea0b7531a83757754f3fe461d9567da5470d51cde046617

Observation 3911caca-315b-48dc-ab0f-151a45476f58 · outbound

This paper cites FireAct: Toward Language Agent Fine-tuning.

Training Agents with Weakly Supervised Feedback from Large Language Models FireAct: Toward Language Agent Fine-tuning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.110427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.110427Z digest=sha256:f1419d65f71365c855742cca90b2a76cf8b414b9b4350e1d61b7b3822e54d1ea

Observation ef04eb93-7dd7-4bca-bd77-6e1efd7c23fd · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

Training Agents with Weakly Supervised Feedback from Large Language Models Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.115910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.115910Z digest=sha256:c340673205eb2d7fa95ff53c9d625b36eaa3c347a96e0d1a651c0db42842371d

Observation 6f793bbd-7632-4028-b1b5-75ba93f4de6d · outbound

This paper cites Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models.

Training Agents with Weakly Supervised Feedback from Large Language Models Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.122617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.122617Z digest=sha256:97359bd2ae7a69bd6ec224703373a3b2ed97da74aa2c05cec4513dfcab8959c3

Observation 28587567-42b0-4924-987c-a60391b68760 · outbound

This paper cites an unresolved cited work.

Training Agents with Weakly Supervised Feedback from Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:08:13.766325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-12T10:08:13.128743Z digest=sha256:ed391ad7288450e53e0f64e5b67c3763b25cabbb300e462b58db2d2218eb49ec

Observation 718e2d56-b39d-4f36-be58-647422554ef5 · outbound

This paper cites A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis.

Training Agents with Weakly Supervised Feedback from Large Language Models A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.133623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.133623Z digest=sha256:b0e14c540f7601608743b39c00d1df91f1ba859997367286ca52f5263b06406e

Observation db37de30-47ce-433c-a044-113fc4eb932f · outbound

This paper cites Large Language Models Can Self-Improve.

Training Agents with Weakly Supervised Feedback from Large Language Models Large Language Models Can Self-Improve

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.139576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.139576Z digest=sha256:af1dcb048bc5b2ae21253bfa03542ebba0022236cc6b50b8dec99af55e83d3e3

Observation 4d8aa783-376a-4892-b90d-fe4e86b880c8 · outbound

This paper cites SelfEvolve: A Code Evolution Framework via Large Language Models.

Training Agents with Weakly Supervised Feedback from Large Language Models SelfEvolve: A Code Evolution Framework via Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.144370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.144370Z digest=sha256:cdec37a1512c4bd5c35059ee745835ed787f20f491ed5a2e9a45ee9ad3eee8f3

Observation 9fd38620-f1e9-4e65-b8e7-b22ecbfcf6a4 · outbound

This paper cites Self-training language models in arithmetic reasoning.

Training Agents with Weakly Supervised Feedback from Large Language Models Self-training language models in arithmetic reasoning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:08:13.751304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-12T10:08:13.149088Z digest=sha256:04237129745f20ddd3bc225b507130ce4c9e91250da9366ca509b91aeec88882

Observation 8b844dc2-c8b8-4dc0-9edb-81bc36c5c626 · outbound

This paper cites an unresolved cited work.

Training Agents with Weakly Supervised Feedback from Large Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:08:13.734739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-12T10:08:13.154116Z digest=sha256:109981bfeb8e3136f66f5c1a595eb3a917098b3c295f1ac7919a29d8b60445f7

Observation 28d9802b-9719-4342-a050-5e0f7222eac1 · outbound

This paper cites MARIO: MAth Reasoning with code Interpreter Output -- A Reproducible Pipeline.

Training Agents with Weakly Supervised Feedback from Large Language Models MARIO: MAth Reasoning with code Interpreter Output -- A Reproducible Pipeline

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.158944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.158944Z digest=sha256:390dbf1aea238f329acb026050ed973ebcf7018861b626c4579fd960916b7a1c

Observation 7d6a8eb0-7b47-468f-a64e-8bce11b4c6ed · outbound

This paper cites an unresolved cited work.

Training Agents with Weakly Supervised Feedback from Large Language Models Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.164434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.164434Z digest=sha256:b0004510ba7c2dd3e4b0d08f6ec5c5d6cf3385b0c4e98973df541f9cdabd0e82

Observation b8711026-5546-4613-a603-0378ef86e67a · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Training Agents with Weakly Supervised Feedback from Large Language Models WebGPT: Browser-assisted question-answering with human feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.169686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.169686Z digest=sha256:a666d6dcf530eac7d59a4c562ad053f3002d6b1c29af83773ae4445e6dbfbc7b

Observation 22770d27-8490-46bf-b7e4-ddb966e6a351 · outbound

This paper cites an unresolved cited work.

Training Agents with Weakly Supervised Feedback from Large Language Models Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.175010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.175010Z digest=sha256:f095d86c42c9167042e3fa4679c30dcda9a4c040ccaf984ac90743ae5bc55fb4

Observation fc839b6f-707d-4c4d-b329-ff278a5b0229 · outbound

This paper cites Gorilla: Large Language Model Connected with Massive APIs.

Training Agents with Weakly Supervised Feedback from Large Language Models Gorilla: Large Language Model Connected with Massive APIs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.179868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.179868Z digest=sha256:af3f39358cdb28aa2d0ab07018e9525a11875c47971ba69560b0e7fc59d1438d

Observation 8342b52c-e514-4a8a-9361-2ac05e32dcba · outbound

This paper cites an unresolved cited work.

Training Agents with Weakly Supervised Feedback from Large Language Models Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.184620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.184620Z digest=sha256:7ab739e9391a786190dc513eedd6ee3c40dc0b750aee1621ffec69672ed6890d

Observation 464f3854-a11b-4c8a-9b4f-92286fd17e73 · outbound

This paper cites an unresolved cited work.

Training Agents with Weakly Supervised Feedback from Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.189463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.189463Z digest=sha256:fab6d98b40426692c2b90e3cc864eb53e7b376cc57b7ebda422e0986d4c0775a

Observation 3e5d3a35-5ba9-467f-afd8-0d04ad5bd38a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Training Agents with Weakly Supervised Feedback from Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.194394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.194394Z digest=sha256:1689625a52671a964c76e8569d973c3309204e91f8ff8b0978061f0c4617ca52

Observation 93d9b338-0bdf-4cd7-ae47-22bb95f431e7 · outbound

This paper cites an unresolved cited work.

Training Agents with Weakly Supervised Feedback from Large Language Models Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.199373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.199373Z digest=sha256:c02c78b6eee5fbf42975a89b945f806d787783090314f60dc0ce339b98208288

Observation ca4bf90c-b752-47e8-ab71-18f14afc5c18 · outbound

This paper cites an unresolved cited work.

Training Agents with Weakly Supervised Feedback from Large Language Models Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.204838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.204838Z digest=sha256:02fda7724deb36b97edfff585b66e1619329a8805ebd30f74a29f462998d21f4

Observation 30f9efe6-4235-41ab-80fe-fae95410b525 · outbound

This paper cites ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases.

Training Agents with Weakly Supervised Feedback from Large Language Models ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.209522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.209522Z digest=sha256:c48fcf91e132ed9b79d6d3ea623eff83a7d3850d9109811b73172009de5ea5ce

Observation 961843a0-f1d2-4a48-a538-4833fad57d09 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Training Agents with Weakly Supervised Feedback from Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.214607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.214607Z digest=sha256:7ace466313e01a52be08bfb039040499bbb2757825750bf0a765719e7e654f5e

Observation 10a9cc14-2510-4922-9bd9-38ad76b28956 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Training Agents with Weakly Supervised Feedback from Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.219817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.219817Z digest=sha256:1d96b569229aa9e8d5bafb03485bf05329d539e93cd20a679d30e7afa32bd942

Observation d4bca0c1-da40-4e88-a251-c4d224b5cad8 · outbound

This paper cites an unresolved cited work.

Training Agents with Weakly Supervised Feedback from Large Language Models Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:08:13.662093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-12T10:08:13.225241Z digest=sha256:e4bdf01750bae985913877847beedea8a243d2ad1a98efc99209a43f8cb758ec

Observation 9ee9c83f-1709-4c35-82f4-122a2c3981aa · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Training Agents with Weakly Supervised Feedback from Large Language Models Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.230991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.230991Z digest=sha256:a62503f76e89b8fbc79b49a2b7af4a2a7bfd0cc8fdfc90e8bcf429278b58fca4

Observation 66349724-564e-4ced-804f-1e789f21684b · outbound

This paper cites The Rise and Potential of Large Language Model Based Agents: A Survey.

Training Agents with Weakly Supervised Feedback from Large Language Models The Rise and Potential of Large Language Model Based Agents: A Survey

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.235856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.235856Z digest=sha256:8ed24fb47429a48eb0be7ebf22a659a530a033ea52999d71a58bac1d848e9eda

Observation c40f4bc0-521e-4a6c-aaa8-3033dc969e04 · outbound

This paper cites Lemur: Harmonizing Natural Language and Code for Language Agents.

Training Agents with Weakly Supervised Feedback from Large Language Models Lemur: Harmonizing Natural Language and Code for Language Agents

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:08:13.380613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-12T10:08:13.240119Z digest=sha256:25773b5559ea7a2c7eb6570193b4ebb2e3dc892459c36f18e10d9a7140489fd9

Observation 56c74e4a-5805-45d7-af8c-af24eb062a9e · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Training Agents with Weakly Supervised Feedback from Large Language Models ReAct: Synergizing Reasoning and Acting in Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.244892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.244892Z digest=sha256:92b736478eb5e7cf4deacc9ed0884ea8b82eebe51e6820d03e922a284e74a293

Observation 2894738e-c67c-4bff-8d49-7d942bb8df60 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Training Agents with Weakly Supervised Feedback from Large Language Models Yi: Open Foundation Models by 01.AI

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.249386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.249386Z digest=sha256:1c46b3bcc7ee52cc672258974fbdfd9d320378ae14b2be9937fce7fb4c8873f9

Observation 8030f6ea-7901-4745-9615-cce2475dd0d9 · outbound

This paper cites AgentTuning: Enabling Generalized Agent Abilities for LLMs.

Training Agents with Weakly Supervised Feedback from Large Language Models AgentTuning: Enabling Generalized Agent Abilities for LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.254686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.254686Z digest=sha256:799fc7bcbe6aa23fff564da9a0964257dd46ff01d9c3655d515a6287692aa689

Observation af16f172-6a01-4b14-89ea-2613ee381586 · outbound

This paper cites Agent-Pro: Learning to Evolve via Policy-Level Reflection and Optimization.

Training Agents with Weakly Supervised Feedback from Large Language Models Agent-Pro: Learning to Evolve via Policy-Level Reflection and Optimization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.259699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.259699Z digest=sha256:48803a5588ed1ed9fed52e40d8e4ad0d9d9b3d5d5ae70109f633b158e0895c01

Observation 2395cfca-8a4e-477b-8f87-0948d7c215ed · outbound

This paper cites online" 'onlinestring :=.

Training Agents with Weakly Supervised Feedback from Large Language Models online" 'onlinestring :=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.264817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.264817Z digest=sha256:5457d973683d3f36b10fafae3d80f4773914d18ee0457950c6c110fef4a93a31

Observation 159d06ae-b93f-4975-ad0e-a008cba38128 · outbound

This paper cites write newline.

Training Agents with Weakly Supervised Feedback from Large Language Models write newline

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:13.270416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:08:13.270416Z digest=sha256:93ac69b06745436e13f71bb4f5c4637a3e551071a6a14a345fc5ab6744cacb5c

Pith citing papers

No inbound Pith citation observations are available.