Pith. sign in

Paper Citation Record · LEDGER

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning

As of 16 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2607.08647.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.08647 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T04:00:47.056185Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact6
  • verified fuzzy1
  • unresolved1
  • parse uncertain0
  • malformed identifier4
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 208e9110-1738-43dc-8668-0ea9739d0691 · outbound

This paper cites URLhttps: //doi.org/10.1007/978-3-642-00982-2_1.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning URLhttps: //doi.org/10.1007/978-3-642-00982-2_1

Reference 1

Resolution
verified exact
doi, observed 2026-07-10T04:06:44.494310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:8571bf57106aa1bf58ce4401ac64ac9f9ae4ed7deee250ea0d60171f753365e7

Observation 724bce8f-8f9f-4345-bdfe-4a3af69c8bed · outbound

This paper cites Understanding the Power and Limitations of Teaching with Imperfect Knowledge.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Understanding the Power and Limitations of Teaching with Imperfect Knowledge

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-10T04:06:44.795817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:a06b370b8e9cd0089a3bc240dd4106ca07b0b2da51b02626e34502e9ae5479a1

Observation 89ae9dcb-9b29-4a60-bfdc-084100c778b2 · outbound

This paper cites Learning Robust Rewards with Adversarial Inverse Reinforcement Learning.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Learning Robust Rewards with Adversarial Inverse Reinforcement Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-10T04:06:44.799707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:f542752389b1e96a4b5a0031f5131c69f7d70bda5a119cd48e04953a9cdfeb09

Observation 0cfe838c-0d66-413a-bbbf-cf1ecf6df5d8 · outbound

This paper cites The effect of modeling human rationality level on learning rewards from multiple feedback types.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning The effect of modeling human rationality level on learning rewards from multiple feedback types

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T04:06:45.255728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:d0bac1b31dbeacd0b1664e95b65f99ad1a8432509a53511e7e07bdfe09b661c9

Observation 4bbd38d5-4249-47c0-bb47-2893f642f869 · outbound

This paper cites Assisted Robust Reward Design.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Assisted Robust Reward Design

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-10T04:06:44.786673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:dcfe5b718310a3705209c82301e849b3e2fac15aa0963ae16d4e22dc44bcc692

Observation eab54314-22fd-4d18-8bc3-a2d7e765b742 · outbound

This paper cites Interactive Teaching Algorithms for Inverse Reinforcement Learning.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Interactive Teaching Algorithms for Inverse Reinforcement Learning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-10T04:06:44.792911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:8834f4e070247dcbbc8a1edf8467813cd7c391293bdb8b57378fc0121193e399

Observation 165abf34-3469-4981-a613-e4638b6b1fc0 · outbound

This paper cites Mehta and Dylan P.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Mehta and Dylan P

Reference 7

Resolution
metadata mismatch
doi, observed 2026-07-10T04:06:44.496762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:f60764d0996c96b2d2219b22d9e03716960ff7f4a7b0828ce315998a6e366dab

Observation f31d1d4c-ea9e-4b6a-856b-c57a00822c25 · outbound

This paper cites Effects of Robot Competency and Motion Legibility on Human Correction Feedback.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Effects of Robot Competency and Motion Legibility on Human Correction Feedback

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-10T04:06:44.789465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:443f96bb307f7a43e5d2188e45d255730c2a2c4b1f5a24735681194d14c3d5eb

Observation 05e0804c-f802-43d4-803d-94241f7b81be · outbound

This paper cites An Overview of Machine Teaching.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning An Overview of Machine Teaching

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T04:06:44.803111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:4a70b32a7dd8d836dc50c06da1879c1613e1e71e0735aa5ba1acd899ba524506

Observation 395203d5-155b-44ed-8db4-dd4d596c1fd9 · outbound

This paper cites an unresolved cited work.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Unresolved cited work

Reference 10

Resolution
malformed identifier
raw_fallback, observed 2026-07-10T04:06:45.253692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:09e026383a3b2b66f054193f78cefd4260bcfb8ef1dc7360699d8e8cae94b795

Observation db90686e-f1e6-4046-9935-59cf7cb9866c · outbound

This paper cites an unresolved cited work.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-07-10T04:06:45.250108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:3c7603aa9793bfab4a617131d5eb7d865bc9a5362200b4ff8e417d8214e80e22

Observation f4b88810-7a0b-417d-ab89-bd2594532659 · outbound

This paper cites We use2×3gridworlds with two cell features (drawn gray and white) and a randomly placed terminal cellTthat may occupy either feature.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning We use2×3gridworlds with two cell features (drawn gray and white) and a randomly placed terminal cellTthat may occupy either feature

Reference 12

Resolution
malformed identifier
raw_fallback, observed 2026-07-10T04:06:45.246599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:4e9d510c1adf00257e77375dcf5b077fbf221715bd1dbe571138f6472a8af6f9

Observation 8a7cab3b-aa36-43cc-a292-133c09631019 · outbound

This paper cites an unresolved cited work.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Unresolved cited work

Reference 13

Resolution
malformed identifier
raw_fallback, observed 2026-07-10T04:06:45.248495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:cc2fcd795285d4e7444ee1ede7ddf89cc2be5b15531db7eb86303fb400786ef1

Observation 5a762f50-e6a9-4847-9796-477fb4079fb4 · outbound

This paper cites S1) 2:Restrict candidate atoms to those in environmentsK 3:D←Greedy Atom Selection(K,U)(Alg.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning S1) 2:Restrict candidate atoms to those in environmentsK 3:D←Greedy Atom Selection(K,U)(Alg

Reference 14

Resolution
malformed identifier
raw_fallback, observed 2026-07-10T04:06:45.252086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:90526efd4e1933dc30760b366788e9f84fb08002e359177a40e7e4a727475a1d

Pith citing papers

No inbound Pith citation observations are available.