Pith. sign in

Paper Citation Record · LEDGER

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization

As of 10 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2602.11351.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.11351 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:13:30.397515Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4cf34756-4688-4a94-a3f3-eb8df2a296f7 · outbound

This paper cites Consistently simulating human personas with multi-turn reinforcement learning.arXiv preprint arXiv:2511.00222,.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Consistently simulating human personas with multi-turn reinforcement learning.arXiv preprint arXiv:2511.00222,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:24.996133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:24.996133Z digest=sha256:bc9ec28205a785e92e02a0423c943347ad456d6b88ae9d33f1f105b4766146e6

Observation 925a5ce1-d2d7-4178-bfe5-2fb732e95e7c · outbound

This paper cites Reinforcement Learning for Long-Horizon Interactive LLM Agents.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:25.232880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:25.232880Z digest=sha256:8dee5329d3ead37541c2d3becafc0a412e8ca75759faf02e7d8ac02e03309f59

Observation 5d47239c-cc22-4a5b-98c7-0fd887ce84b7 · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:25.386643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:25.386643Z digest=sha256:a823d3a0e13acfcb4094cd109f4b5af9d1342a333b0ce485ff483dcf83bf5da9

Observation 2f6680bd-5aac-44f1-a237-74399e3fa718 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:25.542401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:25.542401Z digest=sha256:861aa429b4f301db5361e7ecc49cbaef00b878cb368a54bafc6209ec6dea62aa

Observation de66711d-3d1d-4196-a7cb-7a50aa09c23d · outbound

This paper cites Contextual Markov Decision Processes.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Contextual Markov Decision Processes

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:25.745024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:25.745024Z digest=sha256:76896b96506087272cd0a16efe49c4fa5a89ce371efe8a283fb1f023fef87b6f

Observation d3cd077b-4b50-4363-b7c5-5ba09127493d · outbound

This paper cites R., He, J., Yu, H., et al.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization R., He, J., Yu, H., et al

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:26.242813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:26.242813Z digest=sha256:8f33c10c51a7a7e2e7ba242504f6e082e819e1eb323cb6220ade1460651044c3

Observation ed6abef3-9eaa-4136-af20-fc317702da7d · outbound

This paper cites Quagmires in sft-rl post- training: When high sft scores mislead and what to use instead.arXiv preprint arXiv:2510.01624,.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Quagmires in sft-rl post- training: When high sft scores mislead and what to use instead.arXiv preprint arXiv:2510.01624,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:26.329894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:26.329894Z digest=sha256:6efe3a892bd42ec75cbd73e56609c81bf35fb6a9a3768414ac59259694564603

Observation fdafd0cb-8359-4589-80fa-a54d0b1c4e9f · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:26.456107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:26.456107Z digest=sha256:0f60a0443a15adb8ff8d0c8c9da698c3dcc9ebdf597f5b5038a7d05bcadd0745

Observation 1dbddb93-2803-4ca7-88e5-63fe4e100703 · outbound

This paper cites NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:26.575467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:26.575467Z digest=sha256:a95f7494e52c1f2c952c3f69bebb838b6dfffe99dc9a5cdcb25e1727b8e70b29

Observation bd3124c4-1797-4d42-9b63-7bc0125cb811 · outbound

This paper cites A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:26.701496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:26.701496Z digest=sha256:674da4b11d9314d08f4f59fa6f7910dc4da21c4c962db71adb170f8903739b9a

Observation f75ca2ed-f3e2-4825-8c23-754567efa4a3 · outbound

This paper cites A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:26.854490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:26.854490Z digest=sha256:ef294dc110b1c1c19400e8147fc4ea7260099e1b3255d449f619f079287b5f82

Observation 21967e0c-e5f6-42a9-a2a2-d93f2d683900 · outbound

This paper cites UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:27.015676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:27.015676Z digest=sha256:075f219df1380776d7672092429cc7f55ac8135e46d616eafe69d5826adb8f21

Observation c097a627-c7e7-408b-87bf-6b323c113e63 · outbound

This paper cites GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:27.189438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:27.189438Z digest=sha256:2c46789929fed44da9abf95dfd226c9e08a1cb2785331a8e1b51f3c4e0b376aa

Observation ab94e218-f79c-409f-8062-ecaef6c54f67 · outbound

This paper cites HierTOD: A Task-Oriented Dialogue System Driven by Hierarchical Goals.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization HierTOD: A Task-Oriented Dialogue System Driven by Hierarchical Goals

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:27.347739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:27.347739Z digest=sha256:a8ee460bb12cad4e0d1fefc09b388f9f83ae53870abb9f7de48da185961216f3

Observation a8a8e0be-0c63-4936-a4fb-ca7f1444b804 · outbound

This paper cites Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:27.574670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:27.574670Z digest=sha256:c49b3e020c18312838afb5a5f4fb65ab165b2dd542a922cf274afae29e18ff29

Observation 97c2fc39-6466-4cf3-a527-32042fbe3344 · outbound

This paper cites BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:27.793264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:27.793264Z digest=sha256:d2978b13bce95205e56d08782810d6c3d0e221b640e822de79d0f58f546c296f

Observation 670737dc-3c3a-42ab-a041-0bf224cc8170 · outbound

This paper cites APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:27.863872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:27.863872Z digest=sha256:f8350f489cd4ba9fdaebcb22eed357b20e75bcfaba29aa5c918becb0f829f2ea

Observation 1f86f967-51b6-442c-a056-e36475a4f012 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization ToolRL: Reward is All Tool Learning Needs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:27.985409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:27.985409Z digest=sha256:21e63423c3c03f8ea199d3679bbd44778d63f0f406e6cba3ea4b0cec86cad651

Observation df72b7ae-2db8-4bda-985b-47de9f1a2166 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:28.112926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:28.112926Z digest=sha256:50d04266f9391a889b8e9a3158a04e82cfcfeea0c6bf25737a8323f1da24ee10

Observation 2465e9b8-0ca7-4897-9838-026eacf3cd3d · outbound

This paper cites Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:28.211988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:28.211988Z digest=sha256:16e9d1f7147a57850d84c0aa22ee98bce470572e6512719d16cdc094a053f58e

Observation fbaf1417-b5da-400e-9a9f-40f7dd5c0742 · outbound

This paper cites Language Model Personalization via Reward Factorization.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Language Model Personalization via Reward Factorization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:28.352379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:28.352379Z digest=sha256:7ff8608caa773d3e0d5aba50ceb2dae6813d29593bc64933786c6257d9758dbf

Observation fc728a6f-6756-4464-bffe-ece9910977f7 · outbound

This paper cites Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:28.480694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:28.480694Z digest=sha256:e015d0882a63ffa320c700da83d2f7a41e7f96bc2c67d8a16016e500925c24a3

Observation d746fc4d-08f7-4f89-804f-0866a12d889f · outbound

This paper cites Training proactive and per- sonalized llm agents.arXiv preprint arXiv:2511.02208,.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Training proactive and per- sonalized llm agents.arXiv preprint arXiv:2511.02208,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:28.610154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:28.610154Z digest=sha256:b410d52b58b7935bcd25702d0c342eb61e60231dcf3d791583605c633dc32f6f

Observation 0b738638-e145-4a6d-a552-07f26960a010 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Gemini: A Family of Highly Capable Multimodal Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:28.684575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:28.684575Z digest=sha256:5bcf66195021503c858963ca3e68ab9c74050b1b95feca316b150ae29caa5189

Observation 10028b63-61cd-4504-8a77-e75ff9eb10a6 · outbound

This paper cites En- hancing personalized multi-turn dialogue with curiosity reward.arXiv preprint arXiv:2504.03206,.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization En- hancing personalized multi-turn dialogue with curiosity reward.arXiv preprint arXiv:2504.03206,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:28.984404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:28.984404Z digest=sha256:7331840c2febbce4b463c864edee27fd82a899992d5ffb0628b97947f6b76ea2

Observation 9391ca7b-429b-4994-a220-a8f2c6f0f8b9 · outbound

This paper cites OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:29.080919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:29.080919Z digest=sha256:a829e51879fe6415a3616f3c651876a82e6b2fa80c8b50d9678e3b2efaad4fcf

Observation b279e7df-c54c-4fe5-a995-6298419c2e5b · outbound

This paper cites Boad: Discovering hierarchi- cal software engineering agents via bandit optimization.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Boad: Discovering hierarchi- cal software engineering agents via bandit optimization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:29.386885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:29.386885Z digest=sha256:b9f4ed361e2aa21b421d75066e3711ad1331c1e35f274b947bd954fdb3313791

Observation 43f0cd01-802e-484c-b23a-15c30fe2faac · outbound

This paper cites Qwen3 Technical Report.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Qwen3 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:29.423336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:29.423336Z digest=sha256:aa13bda4847c53765952ace7a9755f70cb533b12207ce476a427f69d6d83e1aa

Observation 3bd0f250-4516-44ba-9c1a-1f3650c12246 · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:29.528540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:29.528540Z digest=sha256:c1ef8dffb9a79c7a4c34f5be0f8af70b084cbc27f177ca87edb891d2bf3bc551

Observation 23557dd4-11bf-41ca-905a-c1c1bfb8f343 · outbound

This paper cites Demysti- fying reinforcement learning in agentic reasoning.arXiv preprint arXiv:2510.11701,.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Demysti- fying reinforcement learning in agentic reasoning.arXiv preprint arXiv:2510.11701,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:29.699949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:29.699949Z digest=sha256:3ce9902e4ea5d99a92a4caaa7dee30111c045a47c96376d01ba53af59251e046

Observation 91f78ddb-4cbf-4659-9b2a-0d815ea1ca68 · outbound

This paper cites Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:29.838143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:29.838143Z digest=sha256:f4c499c17bff188ff8e8000e7367a7240bcdddc4c14a49e556e16295cca24e34

Observation 10def8fe-8f92-4e75-af87-e71342188605 · outbound

This paper cites Teaching language models to evolve with users: Dynamic profile modeling for personalized alignment.arXiv preprint arXiv:2505.15456,.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Teaching language models to evolve with users: Dynamic profile modeling for personalized alignment.arXiv preprint arXiv:2505.15456,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:30.021924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:30.021924Z digest=sha256:0368d6131009ae1047872cfb604d1c5d22a796c79d029a049fa2d255d9bb3bb1

Observation e3738c32-c27f-4543-aa82-ea31e4711187 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:30.206953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:30.206953Z digest=sha256:833df93c99c5de81a0a8fc37f103218f8533bd92cf69c5e78d0228a0e5c4b965

Observation 9a4cb23e-93d5-4e7a-b7a3-f78585c61635 · outbound

This paper cites E., and Zhou, W.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization E., and Zhou, W

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:30.321333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:30.321333Z digest=sha256:a39676aeaa7d0411969319e3d0aee8fd867a9962c8d45c9e5397b77a7a5dda59

Observation 0f994dda-d6a8-42fd-964a-a8595da18906 · outbound

This paper cites Yes”, “No.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Yes”, “No

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:30.397515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:30.397515Z digest=sha256:e818b7931593f490b2178daec5701976c21d0d04e017b22d61270c65a4f85287

Observation 5a0a8997-cc09-4a0b-a013-6621dda5a016 · outbound

This paper cites GPT-4o System Card.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization GPT-4o System Card

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:25.873768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:25.873768Z digest=sha256:a691611fd3fd7cb218b1696e10cda8f2c423a20639528cd80e0e11f1342069d8

Observation 11359bac-b2ad-4c72-8e1a-64e34cd6f14f · outbound

This paper cites ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:27.661295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:27.661295Z digest=sha256:e7882af40e73a1ca225997fdef9a6441f8dbd9035cb4aa1cf1923bb6d33d8738

Observation 8ce43951-02aa-4c32-9da1-fd317ea5175d · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:28.882418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:28.882418Z digest=sha256:de2fa7816c6d07cca219e5dd59c40e7fcc6bf841bccf04d43819f1a7e06f31c0

Observation d3b8c592-c102-4c97-a2f1-41c7335db57c · outbound

This paper cites OpenAI o1 System Card.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization OpenAI o1 System Card

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:26.072593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:26.072593Z digest=sha256:d53cb851217455e353b07c2f547897979267723303f3b3cd4f7cd450a3c07e20

Observation f09582dc-a9eb-4fe9-9b35-23ec975f43e3 · outbound

This paper cites Behavior injection: Preparing language models for reinforcement learning.arXiv preprint arXiv:2505.18917,.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Behavior injection: Preparing language models for reinforcement learning.arXiv preprint arXiv:2505.18917,

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:25.140784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:25.140784Z digest=sha256:ebb040a90b98a8120c64bb32642f82137a20c3bdc1121f65c2a283484352e3d7

Observation a279625d-8e07-44d8-a331-6aec4580110f · outbound

This paper cites SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:29.208730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:29.208730Z digest=sha256:ac87c245590b5be7fd88388f6afa468ad67a4b135986484a469a40e3ed6fb526

Pith citing papers

No inbound Pith citation observations are available.