Pith. sign in

Paper Citation Record · LEDGER

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning

As of 16 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2607.08647.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.08647 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T04:00:47.056185Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact6
  • verified fuzzy1
  • unresolved1
  • parse uncertain0
  • malformed identifier4
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 208e9110-1738-43dc-8668-0ea9739d0691 · outbound

This paper cites URLhttps: //doi.org/10.1007/978-3-642-00982-2_1.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning URLhttps: //doi.org/10.1007/978-3-642-00982-2_1

Reference 1

Resolution
verified exact
doi, observed 2026-07-10T04:06:44.494310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:4635d50138c0f7914cdb2f71fc12e3bee9fe3a1a7fdfa760e5c1af5513fab03b

Observation 724bce8f-8f9f-4345-bdfe-4a3af69c8bed · outbound

This paper cites Understanding the Power and Limitations of Teaching with Imperfect Knowledge.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Understanding the Power and Limitations of Teaching with Imperfect Knowledge

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-10T04:06:44.795817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:0ca8a89d6dd68c842cedf1da4fbaa46117f765cc40f161f0b982b37d0644cf36

Observation 89ae9dcb-9b29-4a60-bfdc-084100c778b2 · outbound

This paper cites Learning Robust Rewards with Adversarial Inverse Reinforcement Learning.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Learning Robust Rewards with Adversarial Inverse Reinforcement Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-10T04:06:44.799707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:02903cf652a7523b9a84fe0a75c9375645b33d9b47148a1ea8fa3823466d0aad

Observation 0cfe838c-0d66-413a-bbbf-cf1ecf6df5d8 · outbound

This paper cites The effect of modeling human rationality level on learning rewards from multiple feedback types.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning The effect of modeling human rationality level on learning rewards from multiple feedback types

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T04:06:45.255728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:5d583ec206402346967ac5fabbf202f247e8a4b0c81255c4d05168d3e8acf59d

Observation 4bbd38d5-4249-47c0-bb47-2893f642f869 · outbound

This paper cites Assisted Robust Reward Design.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Assisted Robust Reward Design

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-10T04:06:44.786673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:507c9cfe117d0db987e4dd5994b6fc02d91dba08064f45b05ae84ca2c312b4be

Observation eab54314-22fd-4d18-8bc3-a2d7e765b742 · outbound

This paper cites Interactive Teaching Algorithms for Inverse Reinforcement Learning.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Interactive Teaching Algorithms for Inverse Reinforcement Learning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-10T04:06:44.792911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:4e189edb532a27b0f9344205578aa4449a7d91ba566343346506a0a94c4b5c72

Observation 165abf34-3469-4981-a613-e4638b6b1fc0 · outbound

This paper cites Mehta and Dylan P.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Mehta and Dylan P

Reference 7

Resolution
metadata mismatch
doi, observed 2026-07-10T04:06:44.496762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:a10bef3fcf50fed89c5436818862320330f580054b315da9b5d6e8d5bfa7fa87

Observation f31d1d4c-ea9e-4b6a-856b-c57a00822c25 · outbound

This paper cites Effects of Robot Competency and Motion Legibility on Human Correction Feedback.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Effects of Robot Competency and Motion Legibility on Human Correction Feedback

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-10T04:06:44.789465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:5bcaae8e729b6d1d99c637b2831e2605152bceb6951a1c93b9fad4cda5037c58

Observation 05e0804c-f802-43d4-803d-94241f7b81be · outbound

This paper cites An Overview of Machine Teaching.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning An Overview of Machine Teaching

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T04:06:44.803111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:63e7c794fa020ad54e7526849d6961dadf419e98b665d32ab2ad18fbf56ea0a3

Observation 395203d5-155b-44ed-8db4-dd4d596c1fd9 · outbound

This paper cites an unresolved cited work.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Unresolved cited work

Reference 10

Resolution
malformed identifier
raw_fallback, observed 2026-07-10T04:06:45.253692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:866eb656dc613246bdbff299383e798b95f3dc30544ecb669e9746de7ea34e9e

Observation db90686e-f1e6-4046-9935-59cf7cb9866c · outbound

This paper cites an unresolved cited work.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-07-10T04:06:45.250108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:0b1a41902ed39990adfec00971e7ed1784e99ac2da04f1042df27eab08afe190

Observation f4b88810-7a0b-417d-ab89-bd2594532659 · outbound

This paper cites We use2×3gridworlds with two cell features (drawn gray and white) and a randomly placed terminal cellTthat may occupy either feature.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning We use2×3gridworlds with two cell features (drawn gray and white) and a randomly placed terminal cellTthat may occupy either feature

Reference 12

Resolution
malformed identifier
raw_fallback, observed 2026-07-10T04:06:45.246599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:f9281fec3bfab67ead08029fd9e48159472404aecc3377ad065da73bdaef42f3

Observation 8a7cab3b-aa36-43cc-a292-133c09631019 · outbound

This paper cites an unresolved cited work.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Unresolved cited work

Reference 13

Resolution
malformed identifier
raw_fallback, observed 2026-07-10T04:06:45.248495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:ee753ac98012c669554d388914e20aa833690f6b0ea96756309778848ada012d

Observation 5a762f50-e6a9-4847-9796-477fb4079fb4 · outbound

This paper cites S1) 2:Restrict candidate atoms to those in environmentsK 3:D←Greedy Atom Selection(K,U)(Alg.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning S1) 2:Restrict candidate atoms to those in environmentsK 3:D←Greedy Atom Selection(K,U)(Alg

Reference 14

Resolution
malformed identifier
raw_fallback, observed 2026-07-10T04:06:45.252086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:fde24690fa188e60d44a2d5bb5711dd70bc0d8204bb8f9d641cda904a5f89b0b

Pith citing papers

No inbound Pith citation observations are available.