Pith. sign in

Paper Citation Record · LEDGER

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

As of 20 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 33 inbound Pith citation observations for arXiv:2505.00024.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00024 v2

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:31:35.336899Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 33 of 33 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:52:52.497186Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 49749bbb-d435-4f3b-af0a-4434aec63f2e · outbound

This paper cites Chemcrow: Augmenting large-language models with chemistry tools.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Chemcrow: Augmenting large-language models with chemistry tools

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:31:36.245643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:31:35.142072Z digest=sha256:3943974f7a1652bc81f5b938a0806b66e8bff53fd2a12df98dbd6564ec91e878

Observation e0c1508c-1359-4b61-9251-4d11471cc246 · outbound

This paper cites Acebench: Who wins the match point in tool learning?arXiv preprint arXiv:2501.12851, 2025.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Acebench: Who wins the match point in tool learning?arXiv preprint arXiv:2501.12851, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.146688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.146688Z digest=sha256:f2b69d0dee5c938f2af140e0557d8b8d4b95b90dfda02927f87e3c3196bcb605

Observation e3a7c3c3-8352-4c95-aae5-36d9bd1e102d · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.150743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.150743Z digest=sha256:4510b82ca087fced5c009803e47d115976024af532652c3bc51d5ff6f016e8d6

Observation 0ea275ae-b33d-49c0-a79b-45caae5b7a50 · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.155015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.155015Z digest=sha256:9386b741850e91c2ebb2c3cf225be16fa2a19d32e959bb5c080bc61acb141cc5

Observation 233de737-a27b-4396-ad4d-a4ccd6a9d90c · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.158993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.158993Z digest=sha256:772ccd4c06797b731cc754c2f2319b3276a9d1b2352ad57ec3eba11de18e4370

Observation d2778d79-002b-4933-a14e-a60da02af70b · outbound

This paper cites On Designing Effective RL Reward at Training Time for LLM Reasoning.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning On Designing Effective RL Reward at Training Time for LLM Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.163114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.163114Z digest=sha256:2c29de1072d5dd9e5cf4cdfbd0fe2505aa323fe72941eccf68ad53583f1e42d7

Observation 535983c5-7f87-4d25-b01f-24368c4205ed · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.167122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.167122Z digest=sha256:a51a0ba18da1c60648513fae1e856dd28d9d2af9fbc61b82e09d02ab9cc5bff9

Observation c07a2ff4-1df8-422a-84ae-0d7bdc74ad41 · outbound

This paper cites Visual sketchpad: Sketching as a visual chain of thought for multimodal language models.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Visual sketchpad: Sketching as a visual chain of thought for multimodal language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:31:36.233585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:31:35.171643Z digest=sha256:357ea50a0e11d5526f980cdbf085a9ef004902c2afcb0d4bb1310c8f2e8fd599

Observation 7bf87a75-a25f-4b96-a628-2d61f2f66f0b · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.175548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.175548Z digest=sha256:43125388319579731709e5b559bf329e81f24ca190fc3c256d62232b3b1ecc55

Observation 101f98bd-0676-4c5b-9f69-8b4381246a80 · outbound

This paper cites Language models can solve computer tasks.Advances in Neural Information Processing Systems, pages 39648–39677, 2023.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Language models can solve computer tasks.Advances in Neural Information Processing Systems, pages 39648–39677, 2023

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:31:36.220802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:31:35.180136Z digest=sha256:db03de9d741de682ef93139cb8e2fda6d80cff25566bb67f6c42868ff7369d41

Observation 7514f049-cbad-4e6d-ae99-1dec9d45f589 · outbound

This paper cites Internet-Augmented Dialogue Generation.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Internet-Augmented Dialogue Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.183794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.183794Z digest=sha256:9ee091026b5b96a39881cacb1d70bd5ddc0bd24c69f72fc351599ec926338b38

Observation 7763e7e6-1f61-44b9-8beb-4508c1dc541d · outbound

This paper cites Internet-augmented language models through few-shot prompting for open-domain question answering.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Internet-augmented language models through few-shot prompting for open-domain question answering

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.187574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.187574Z digest=sha256:a7254b352573040d76f2577c159c72b730356d4ab89500c9e96cdac335221987

Observation 9fd8d46e-a389-4b22-8027-c2ee02c3e820 · outbound

This paper cites API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.191591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.191591Z digest=sha256:8931c8b6c4e72579f8263a2ee47f7e8544ab268a2a9fb26201b394316f9a7179

Observation 9041ee50-4f15-4d4f-84a2-55162dffceb1 · outbound

This paper cites Process Reward Model with Q-Value Rankings.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Process Reward Model with Q-Value Rankings

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.195492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.195492Z digest=sha256:54606b644b5b708ac886ea8c8b86935f812df0822be5cc8a2667808e19d49fc4

Observation 272b3cae-8532-4a53-bd03-4fe608b7a982 · outbound

This paper cites Hammer: Robust function-calling for on-device language models via function masking.International Conference on Learning Representations, 2024.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Hammer: Robust function-calling for on-device language models via function masking.International Conference on Learning Representations, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:31:36.197189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:31:35.203176Z digest=sha256:631f96f1a0415ce56226c980322ed743b06b497f2410cfb85e8181cb029b8d7e

Observation 1944ec4a-afe2-4d42-9903-f0914591b028 · outbound

This paper cites Code-r1: Reproducing r1 for code with reliable rewards.arXiv preprint arXiv:2503.18470, 2025.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Code-r1: Reproducing r1 for code with reliable rewards.arXiv preprint arXiv:2503.18470, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.206844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.206844Z digest=sha256:f1608ff1f7f885cb3e78b6008b07641e82cca36216b41d5254b6997fd2271c58

Observation 13ad4f4c-1f8e-4401-9356-bf7838496cf8 · outbound

This paper cites Toolace: Winning the points of llm function calling.International Conference on Learning Representations, 2024.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Toolace: Winning the points of llm function calling.International Conference on Learning Representations, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:31:36.186319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:31:35.210586Z digest=sha256:9d1cf67f42479335d05876194950fdd5ec7e7d763d82ded25ff0161db5eb6853

Observation 60be5adb-7108-48dc-b2f9-01f1d30f23d7 · outbound

This paper cites Apigen: Automated pipeline for generating verifiable and diverse function-calling datasets.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Apigen: Automated pipeline for generating verifiable and diverse function-calling datasets

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:31:36.175189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:31:35.213951Z digest=sha256:c081b5c3d5c5e233fa752ce50375309cda3413fe55af69bf433bfe59e184ae3d

Observation 1da50eb4-e3ec-48ad-ad90-dae696ae5724 · outbound

This paper cites UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.217426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.217426Z digest=sha256:1df008e9b1066bf68c223e7b162ea199882150c7936ed41ba8cb527e9101e8d6

Observation 77fe5306-4e2c-41da-b142-3e1d5292b172 · outbound

This paper cites Sql-r1: Training natural language to sql reasoning model by reinforcement learning.arXiv preprint arXiv:2504.08600, 2025.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Sql-r1: Training natural language to sql reasoning model by reinforcement learning.arXiv preprint arXiv:2504.08600, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.221075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.221075Z digest=sha256:63522a98ede07cf1fe27df7b465ceba1756b35b038585757e3099be25b9b3e3f

Observation 905e18bf-2716-4731-a106-7f46488d7e63 · outbound

This paper cites m & m’s: A benchmark to evaluate tool-use for m ulti-step m ulti-modal tasks.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning m & m’s: A benchmark to evaluate tool-use for m ulti-step m ulti-modal tasks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:31:36.164101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:31:35.224520Z digest=sha256:1bbd9997d098fabc9d3354799e69e78778419cd9b0afaa75b0868fa9f29d82fd

Observation b219db24-bc17-4e84-82ce-f8be6fd346fe · outbound

This paper cites Taco: Learning multi-modal action models with synthetic chains-of-thought-and-action.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Taco: Learning multi-modal action models with synthetic chains-of-thought-and-action

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.227973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.227973Z digest=sha256:be2c386305ce87162fe1b825aa8c40acc93a5adb96c4f6c57646a3a70c77657b

Observation d10fe966-0042-4d22-bfb4-d7275ffeb6bf · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.231595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.231595Z digest=sha256:4f254d6799efd733c66d9580899a933da80e7b9e765ca43a48e04563514c9fde

Observation df8f7afb-1578-4c15-80c7-6888bd3599b1 · outbound

This paper cites s1: Simple test-time scaling.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning s1: Simple test-time scaling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.235425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.235425Z digest=sha256:a5a39225bdf7fcf086d738b90dca86c95fc1ecab28d95b147b7a06af67912bf6

Observation b69517a6-469d-4b12-8ad0-33e86ad191f9 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning WebGPT: Browser-assisted question-answering with human feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.239129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.239129Z digest=sha256:30339e50f2a9529c8596643ab4f7555e9d54e7714ad0e2d37cc399cd7e2c4ee8

Observation 151820d6-d8fa-4189-afef-79e1bae26c1a · outbound

This paper cites Feedback loops with language models drive in-context reward hacking.International Conference on Machine Learning, 2024.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Feedback loops with language models drive in-context reward hacking.International Conference on Machine Learning, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:31:36.152794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:31:35.242745Z digest=sha256:ec0c1dd6f328795343e937f6e3c7b317177e3a5b8e6c45e168050bb70cdfd4d9

Observation 16ae09ac-c8c9-48a9-93d3-6c205c64ec3a · outbound

This paper cites ART: Automatic multi-step reasoning and tool-use for large language models.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning ART: Automatic multi-step reasoning and tool-use for large language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.246109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.246109Z digest=sha256:ef8f41753a531f8e2b7cb2a5676667476646e7d48fb31a64ff8059b3ef7a6e2a

Observation 5aa30e77-1943-4b17-8665-b839f1c73df8 · outbound

This paper cites APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.249740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.249740Z digest=sha256:e694b8e8ec2feeadb0e9cda19c7be8613dfa5c96b0abcf56bb8bb5d44a44d8fe

Observation 633fddce-3d8c-4473-8097-49657e00cc80 · outbound

This paper cites Toolllm: Facilitating large language models to master 16000+ real-world apis, 2023.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Toolllm: Facilitating large language models to master 16000+ real-world apis, 2023

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:31:36.141124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:31:35.253495Z digest=sha256:cf44b68d30b42b2ce028e245b04be3e2780dc78313f775e733a4f318be3da2cb

Observation 75ca8db0-98b9-4e1b-9543-1eb05663aa5b · outbound

This paper cites Tool learning with large language models: A survey.Frontiers of Computer Science, 2025.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Tool learning with large language models: A survey.Frontiers of Computer Science, 2025

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:31:36.129251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:31:35.256742Z digest=sha256:179a5a704fa273c25dae510856c40d5291d3fe42c6f61c8125b66311bd65a661

Observation fbdde0be-f158-490f-9cee-af2009785ad6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.260272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.260272Z digest=sha256:4dadb53fe6cc36c5db8b0fbadbce9033f309b1aa9f288b663501b96b127bbdea

Observation 7d33f48c-7904-4f69-b690-7013b99c95b4 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.264007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.264007Z digest=sha256:ccf06ccb15eb51b03234fade37a37ed6bdd375c4898749985e8fc95c1226801a

Observation a3ee43dd-0922-4b87-9fa1-0b67ad618721 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.267902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.267902Z digest=sha256:671a7f00f84f6aa0e1daf003560be23f02b7567a54d7d28587ec6e6c3d8658c1

Observation 1b06be77-10f9-4c4e-9d48-f770fc1a75a5 · outbound

This paper cites BlenderBot 3: a deployed conversational agent that continually learns to responsibly engage.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning BlenderBot 3: a deployed conversational agent that continually learns to responsibly engage

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.271641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.271641Z digest=sha256:51f2f303461a872182939b87ffe9d9d093c4d254682104148dd4379e7b521df3

Observation 4b016cdb-4648-4cab-9972-1a7e7b2cd76b · outbound

This paper cites Adaptive In-conversation Team Building for Language Model Agents.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Adaptive In-conversation Team Building for Language Model Agents

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.275414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.275414Z digest=sha256:84cd336b336781d537c90409482a423dd4b0d263ea75edd5774ebd870422914f

Observation 04bd80a6-9b36-45d6-9e89-f09f70f91a4e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning LLaMA: Open and Efficient Foundation Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.279709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.279709Z digest=sha256:06bc766d2a9fa45e1baeada2ba8aa046537d2233798673c19c0b1e968e59d0fe

Observation 814848fb-9b3f-45c1-ab47-0bc0be7e120a · outbound

This paper cites Executable code actions elicit better llm agents.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Executable code actions elicit better llm agents

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:31:36.117489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:31:35.283091Z digest=sha256:7af26edee91dea638051caee8de5b7079ef8ff0fb1597244a4442dcdfc3206b6

Observation 70f6df23-82ac-48fd-8278-005d01441b6c · outbound

This paper cites What Are Tools Anyway? A Survey from the Language Model Perspective.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning What Are Tools Anyway? A Survey from the Language Model Perspective

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.286404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.286404Z digest=sha256:811105a537db3727ccd8f390fa411869eb0ac6fb7b9b61d370012d54098dc18f

Observation 2d97cc87-ebe6-48ca-bd82-4a63ab76ea4e · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 2022.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 2022

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:31:36.106426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:31:35.289994Z digest=sha256:5a773302b06ce00fa751e07f7b0de98848e5d96488c1e1df4d15b3b80c249fad

Observation b7e559d7-b39d-4c84-a680-acb075286350 · outbound

This paper cites MathChat: Converse to Tackle Challenging Math Problems with LLM Agents.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning MathChat: Converse to Tackle Challenging Math Problems with LLM Agents

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.293284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.293284Z digest=sha256:0ab90c5310ff4d020b11fa2e351970c9a511700a804ca80292df10981383fade

Observation 4793b878-d277-4c8f-962d-6ce7ce5bbbba · outbound

This paper cites Patil, Ion Stoica, and Joseph E.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Patil, Ion Stoica, and Joseph E

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:31:36.095636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:31:35.297032Z digest=sha256:9ee6f2df09a920f9102593985d43af77560d2ceeaca1b643e6e8debd316d0c73

Observation 3fb9dbfc-9aa7-4579-9c72-e4671ebebffc · outbound

This paper cites Qwen2.5 Technical Report.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Qwen2.5 Technical Report

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.300396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.300396Z digest=sha256:a272fe32a237fcd722fe1fc639d7e034f9f9dfb9e8921058dd8a60fc77d943f8

Observation bf48dc31-ad7a-420b-860e-11655884a4e7 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.Advances in neural information processing systems, 2023.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Tree of thoughts: Deliberate problem solving with large language models.Advances in neural information processing systems, 2023

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:31:36.083647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:31:35.304027Z digest=sha256:a53b2823d42ff2c623d95d2de6331026c1514320ff0d6bd01af80c1f5a2819ab

Observation 271d900e-186c-4826-9ed1-cc464b547e53 · outbound

This paper cites Re- act: Synergizing reasoning and acting in language models.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Re- act: Synergizing reasoning and acting in language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:31:36.071753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:31:35.307454Z digest=sha256:93417f4615bf0d7684d925d429fb17f738748b0f614b3dd28aac71f254189024

Observation cd4235b7-8e28-4ca2-af53-50a5108d722f · outbound

This paper cites Magnet: Multi-turn Tool-use Data Synthesis and Distillation via Graph Translation.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Magnet: Multi-turn Tool-use Data Synthesis and Distillation via Graph Translation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.310996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.310996Z digest=sha256:1f037f8374a60a04525f9e9ca2436078413da045357a2aa1d1104438eac748aa

Observation 7c451a0a-5028-41b7-878f-c237507484d4 · outbound

This paper cites StepTool: Enhancing Multi-Step Tool Usage in LLMs via Step-Grained Reinforcement Learning.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning StepTool: Enhancing Multi-Step Tool Usage in LLMs via Step-Grained Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.315086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.315086Z digest=sha256:920c9c751dc56bbbb2a58046d4590a9157fdf8c1bff93fde99efcc870c1dfcde

Observation 32818700-c092-4cb1-86dc-6b1ddbc3c202 · outbound

This paper cites Boosting tool use of large language models via iterative reinforced fine-tuning.arXiv preprint arXiv:2501.09766, 2025.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Boosting tool use of large language models via iterative reinforced fine-tuning.arXiv preprint arXiv:2501.09766, 2025

Reference 47

Resolution
verified exact
raw_fallback, observed 2026-08-16T10:31:35.449180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:31:35.318603Z digest=sha256:edab881ff61cde3ec09a9fd16b4ebdce31249a8c57473b04f89a5f89d97791ea

Observation df9dcf89-bf8d-460b-9b34-6bbe979f00ff · outbound

This paper cites Data-centric artificial intelligence: A survey.ACM Computing Surveys, 2025.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Data-centric artificial intelligence: A survey.ACM Computing Surveys, 2025

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:31:36.058847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:31:35.322487Z digest=sha256:fd5f42f01fc42f76de77c6271f35542541ffa0b580e7c8783afdccea2f5ea6a7

Observation 8960663d-00cd-47ba-9e1a-6ebaf8e7c4ef · outbound

This paper cites xlam: A family of large action models to empower ai agent systems.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning xlam: A family of large action models to empower ai agent systems

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:31:36.047237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:31:35.326325Z digest=sha256:ac675e0a9acdf5d4a456d11709e538add9096adf40e7d236d1934fbc917b1114

Observation e03dbd94-d63e-4575-8306-aff5f9d98a15 · outbound

This paper cites EcoAct: Economic Agent Determines When to Register What Action.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning EcoAct: Economic Agent Determines When to Register What Action

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.329671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.329671Z digest=sha256:5245478875edb34648e3c49599a15385944cebe4c4160a314a0049f8e3718a1a

Observation 979a8588-a0a2-438f-9d2f-02c0b15156ae · outbound

This paper cites Training languagemodelagentswithoutmodifyinglanguagemodels.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Training languagemodelagentswithoutmodifyinglanguagemodels

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:31:36.035983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:31:35.333279Z digest=sha256:4817a1c8bd1cdcf52744adc5f12ef03d0711c51c617a2820063fcd3aeb10d5b7

Observation 45d4a086-6334-4ca7-bd61-bffce55f1fb8 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.336899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.336899Z digest=sha256:cc6d0bc7646fa6f7b9ab5abcb787ab4ce801fbecf9800d2f964a5ebdfa1ecdf4

Observation 14039814-2878-4ee8-8886-41acc98ad8cb · outbound

This paper cites an unresolved cited work.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:31:36.209155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:31:35.199339Z digest=sha256:02c6827803aebdf0ce067c1926e5d3d50e89755b672f5247b1a67517c6648b41

Pith citing papers

Observation 55fba8a4-b624-42d2-84ac-ee7f9f39423a · inbound

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems cites this paper.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:52.497186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:52:52.497186Z digest=sha256:15059bd7efa982a4cdbe9b5e4ace4b9c48456f738ec00c7ed5b4bfccd64c1ee6

Observation f32edaa0-bd51-4cb7-81a7-533956c446c1 · inbound

The Hallucination Tax of Reinforcement Finetuning cites this paper.

The Hallucination Tax of Reinforcement Finetuning Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:36.354899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:36.354899Z digest=sha256:192f30218e079cfff195c195db380139c80f8fe538f15692b02ab129521fe52a

Observation f62d10de-d64c-42de-b2d7-77681a765934 · inbound

Visual Agentic Reinforcement Fine-Tuning cites this paper.

Visual Agentic Reinforcement Fine-Tuning Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:32.736707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:32.736707Z digest=sha256:806f0a491915061affa925583e124c69c3cb4c010e656af3bd28a8a04c9f257b

Observation 6fc74151-5eff-4081-a180-57d2aad2b823 · inbound

WebDancer: Towards Autonomous Information Seeking Agency cites this paper.

WebDancer: Towards Autonomous Information Seeking Agency Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:02.846595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:02.846595Z digest=sha256:d1cbb9472993530e89eb8d7431fd2351adfd9d73fad8ca9557a1b5c432a3d93c

Observation 7932a883-2bfe-4888-803c-20a7cea739a8 · inbound

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents cites this paper.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:26.754466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:26.754466Z digest=sha256:cbdb7d92e8f96215a85096ba7186f506222aded262cecd0bf52370d827708d7f

Observation 1d3b4cfe-bc69-4358-96c9-0d690ce2e958 · inbound

StepFun-Prover Preview: Let's Think and Verify Step by Step cites this paper.

StepFun-Prover Preview: Let's Think and Verify Step by Step Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T13:47:38.076693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:47:38.076693Z digest=sha256:c8c4fb70f920a2473036790dd6adc479b16213a98eb659ac02aa1de6c233e4e0

Observation 159b4990-e5c7-4400-b315-9e5dbff73feb · inbound

MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use cites this paper.

MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T16:22:51.631063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:22:51.631063Z digest=sha256:c2565681b9945faafae67e5498cc71164375fb8cad298c1e618ab03c7d3ca4ed

Observation d2a6fe6a-5e9d-4f0b-9f1d-f754a154fbe8 · inbound

How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on $\tau$-bench cites this paper.

How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on $\tau$-bench Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T14:49:00.746594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:49:00.746594Z digest=sha256:140470b5136a3a72fb643154e89505f597eda7865ac37cd153d70662ef59e6c3

Observation 8d24ae44-7827-42b1-bf90-4ea7efc56487 · inbound

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model cites this paper.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.953253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.953253Z digest=sha256:82f9c13ad9ca53a70d8d468a2f3b69a9e67c4d12a00e2d6d569511579908a400

Observation d928ccf2-7034-4c4f-a90f-1b64febeffe8 · inbound

Reinforced Visual Perception with Tools cites this paper.

Reinforced Visual Perception with Tools Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:05.083989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:27:05.083989Z digest=sha256:b9b755a7a42318e4e86ac43e984eb03efe659fe84584c1b8030df42850a2871b

Observation 6a2fa7a9-9b77-4d60-865b-50ae03e35d12 · inbound

Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels cites this paper.

Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:41:08.319271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T08:39:54.747656Z digest=sha256:3a308dc6ed17d57d226968b6cef2e32f3734bf79a91b2a8ff0ad852491ef33a9

Observation bb5db8fd-4415-4d2c-bb58-9cc652e33de5 · inbound

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation cites this paper.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:15:26.752887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:2826464e8d8a2a8435b1a83c18c9e167e1b3dbb102f4a85bdda00432f04e80f7

Observation 9a615701-ee8e-4f43-88ba-0606415905e3 · inbound

Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models cites this paper.

Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T05:28:16.047540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:28:16.047540Z digest=sha256:2ce92fd6b77743ef61f1220c6a00e4758e7516405247685747e51d338c2b388e

Observation e642e121-706e-41e9-a0b6-4f77905bd761 · inbound

LAST: Leveraging Tools as Hints to Enhance Spatial Reasoning for Multimodal Large Language Models cites this paper.

LAST: Leveraging Tools as Hints to Enhance Spatial Reasoning for Multimodal Large Language Models Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:00:47.686992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T19:26:23.660206Z digest=sha256:36d6540d22f8c853b4070ae249ed52e7abe1f745cf675044914f4f6266d5099c

Observation f7167414-ff93-4c04-9fb4-de6f59864fb9 · inbound

Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning cites this paper.

Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:15:58.363701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T17:42:57.596073Z digest=sha256:61ac698c5d7bf2ec9c5d69a54f6bd5d6f0f395fd5b205f680f5c438773460ce6

Observation 142182ba-3d01-4ab8-8105-7d86ff9c2db3 · inbound

Democratizing Tool Learning with Environments Fully Simulated by a Free 8B Language Model cites this paper.

Democratizing Tool Learning with Environments Fully Simulated by a Free 8B Language Model Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:25:20.381487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T04:52:37.197183Z digest=sha256:bd583c5cb3d9a007718019b75ba514fd8561873534e3c80cdf03d0e6dc0cc410

Observation 97593b0d-1f44-478e-8c0e-abbda9ea3e8c · inbound

R2IF: Aligning Reasoning with Decisions via Composite Rewards for Interpretable LLM Function Calling cites this paper.

R2IF: Aligning Reasoning with Decisions via Composite Rewards for Interpretable LLM Function Calling Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:05.598895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T00:31:31.327074Z digest=sha256:d9fa78202a843f2f1b3d5640e903cc5dd576fa70cd1b646b892f94495f4fd189

Observation 5287c186-a7de-47a5-a21d-5b2c0e598478 · inbound

R2IF: Aligning Reasoning with Decisions via Composite Rewards for Interpretable LLM Function Calling cites this paper.

R2IF: Aligning Reasoning with Decisions via Composite Rewards for Interpretable LLM Function Calling Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-05T04:40:40.695576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-05T04:32:42.150335Z digest=sha256:ec34fbc55da1638fb97af54ef81343f1a26e8e1d1f175c010a6df3c8d90ea68b

Observation 50d2d743-3da8-4f94-95d3-0edc9010f587 · inbound

CuraView: A Multi-Agent Framework for Medical Hallucination Detection with GraphRAG-Enhanced Knowledge Verification cites this paper.

CuraView: A Multi-Agent Framework for Medical Hallucination Detection with GraphRAG-Enhanced Knowledge Verification Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:56:30.586469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-07T04:07:33.610288Z digest=sha256:225fc3c438e772008355d03f86773e58bf8d4017e948ebc2d13b9d9eb23e4238

Observation 60d09fb1-7a28-4c2c-9a49-5cb1d35291c4 · inbound

RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement cites this paper.

RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:26.747358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-12T03:37:09.000442Z digest=sha256:fe12c6728d64f35feeaac32aeb5136f533db17a9a99d878ce4d626dadb058192

Observation 1691925d-e988-412a-b22d-2a65c095f354 · inbound

RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement cites this paper.

RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:29:48.025544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-15T05:25:43.890367Z digest=sha256:c6fae235e1bd8155fb104d1c1eece311c09cab276dcfafa972713803da1148ac

Observation e24e6cda-8f49-44bc-8708-0ee6cce84a0f · inbound

RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement cites this paper.

RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:13:47.042443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-20T22:12:23.680155Z digest=sha256:d9f9a801748edfaa914a231406598ef223c4207214ae5071a4fecea3e392b932

Observation b610b465-0527-41fb-bae3-78b36d5cee19 · inbound

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control cites this paper.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:32:29.756995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-13T07:30:41.399083Z digest=sha256:507aa83cf36d1f9e34bea57e223308617c220c632d4dabc8243393f74d347e06

Observation 43e4fe24-08dc-4499-af74-b9b8f36dbb38 · inbound

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control cites this paper.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:05:06.427969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:44369fb8a463fb38091d19af8f0ef17a499b33df6390c43f673fd8a0919e0834

Observation e15b3f73-7bf1-4501-ab0a-b930c5a74237 · inbound

Reinforcement Learning for Tool-Calling Agents in Fast Healthcare Interoperability Resources (FHIR) cites this paper.

Reinforcement Learning for Tool-Calling Agents in Fast Healthcare Interoperability Resources (FHIR) Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:59:44.534655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-15T04:56:21.439473Z digest=sha256:26651dbbdccbdff23073c9540ff74348fb28682ff8ccd3df83d24bb5f81e6df6

Observation 131cda4a-c012-41cc-bbb6-824cbd1c1fa7 · inbound

Entropy-KL Divergence-based Token Masking: A Novel Approach for Selective Fine-tuning of Large Language Models cites this paper.

Entropy-KL Divergence-based Token Masking: A Novel Approach for Selective Fine-tuning of Large Language Models Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:13.511091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-29T08:01:39.412431Z digest=sha256:0c4affd759eeaba0d3220a7ed54efcfb0155b8f2cd793a0331391f5f6d31a5e8

Observation 623b7556-ad61-4315-ad77-6ff632f35973 · inbound

On Effectiveness and Efficiency of Agentic Tool-calling and RL Training cites this paper.

On Effectiveness and Efficiency of Agentic Tool-calling and RL Training Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:43:15.374233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-29T08:38:52.411671Z digest=sha256:e3883a758fe4590da1b04b63946eace1d241a6e439e1afcb0c6a14c897fecd41

Observation fa3b5909-2615-42e9-8579-37aa52a39b5e · inbound

Synthesize and Reward -- Reinforcement Learning for Multi-Step Tool Use in Live Environments cites this paper.

Synthesize and Reward -- Reinforcement Learning for Multi-Step Tool Use in Live Environments Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:36:29.699536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T09:54:00.111238Z digest=sha256:54f30a7716cf303b6e42440795c95857ca9de3875b4711b96eeab3acbd87d9c1

Observation 53b32c7b-1cbe-468d-aa76-2e2eed45add2 · inbound

Capability-Aligned Hierarchical Learning for Tool-Augmented LLMs cites this paper.

Capability-Aligned Hierarchical Learning for Tool-Augmented LLMs Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:17:30.701824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T16:43:00.259139Z digest=sha256:ddc8c571a365a0db07ddc86bbdc48b1acca4574393687771288d94a1402f43e6

Observation 96f730d3-8d75-45fa-b278-cbec30e2f639 · inbound

Pushing the Limits of LLM Tool Calling via Experiential Knowledge Integration and Activation cites this paper.

Pushing the Limits of LLM Tool Calling via Experiential Knowledge Integration and Activation Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 4

Resolution
malformed identifier
arxiv_id, observed 2026-07-03T05:07:39.282163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T13:23:42.745788Z digest=sha256:d58b5168de7391c2d9a917d0f8f0f48670a6cba0d97e993607d31ef9e5e69b31

Observation 8ded1493-716c-4061-a89a-18136fec6c84 · inbound

TCPO: Turn-Level Credit Policy Optimization cites this paper.

TCPO: Turn-Level Credit Policy Optimization Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.686054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.686054Z digest=sha256:af637354c5a710a300ff88b005e432e8c5e1058f9edd27b21765d14a4a2c673b

Observation a641ada2-d208-4a43-987a-a284cdd29e67 · inbound

TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning cites this paper.

TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T04:18:15.310447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:18:15.310447Z digest=sha256:e4199ee06064ea1e523b1cc84407ea30d28627d4e019a6215c5fd86551b9cbd6

Observation 7ff4c1bc-fd62-4dd1-a8fc-b299bc0ae9d3 · inbound

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents cites this paper.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.161918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.161918Z digest=sha256:9256a07c2d2e5b0a32b40be44c190adf3ac167fcee091d99f883ccc68164fdd5