Pith. sign in

Paper Citation Record · LEDGER

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents

As of 12 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2412.10410.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10410 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:41:45.249886Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f50f39f2-e052-4655-b5da-a369fe0d3d19 · outbound

This paper cites write newline.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:44.943729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:44.943729Z digest=sha256:c3cffe0a8c7a83cbaaa70a113a2aad8b28cd36765b6ed17359462c5638e73d0c

Observation 28369869-51de-4e3f-b658-9b9c64591727 · outbound

This paper cites Multi-objective latent space optimization of generative molecular design models.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Multi-objective latent space optimization of generative molecular design models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:46.001240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:44.949822Z digest=sha256:9899d4d863ed7e970d6ce70d41754d2019c03d964005366df8cd348e2fd18008

Observation 14e1e029-8980-4bd3-a986-8f2bf97615fc · outbound

This paper cites Imitating Interactive Intelligence.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Imitating Interactive Intelligence

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:44.960832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:44.960832Z digest=sha256:0a3b32ac2be816f907d3d090da2086e06d923b25ff5ee88fe79dd3b113413425

Observation 218ff899-930d-4e8f-a422-d73ca4dcb064 · outbound

This paper cites An optimistic perspective on offline reinforcement learning.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents An optimistic perspective on offline reinforcement learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:44.965648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:44.965648Z digest=sha256:09b70f313c7e87f60c3c8da9de3dd70d8425e6cd5680bfe75a9ed43fa580a260

Observation e7a7730c-45b8-4919-b357-713eb95f3b6b · outbound

This paper cites OPAL: Offline Primitive Discovery for Accelerating Offline Reinforcement Learning.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents OPAL: Offline Primitive Discovery for Accelerating Offline Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:44.970264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:44.970264Z digest=sha256:d11741d1b44c87c8b374f0cb49d92095bd01062814085263ff7382ce6ae27c6c

Observation 5edea55d-0b14-4aca-a234-78064cce8fe4 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Flamingo: a Visual Language Model for Few-Shot Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:44.976022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:44.976022Z digest=sha256:c8c609fb90532a9fef4305d8b3d2b4b7fc9f8d08a7cbc209ef894bf2a8fbf8a0

Observation 8744f31c-5e5a-4fa7-9fa7-7aaeed675001 · outbound

This paper cites Fixing a broken elbo.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Fixing a broken elbo

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:44.981536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:44.981536Z digest=sha256:0459afd06b3e54805fdf809178251d2fdd13ad2cac3e384480d68dac6e02acb8

Observation 69ca765b-5454-4b5d-9ab9-0dba13e472d5 · outbound

This paper cites Hindsight Experience Replay.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Hindsight Experience Replay

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:44.985917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:44.985917Z digest=sha256:ef875c5caea7187ba76c94030fc3dfdcf9c09f25620eb587d44c738f7b085285

Observation 7c984fe0-9b1a-41e1-9940-fc603d54ba1f · outbound

This paper cites Agent57: Outperforming the atari human benchmark.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Agent57: Outperforming the atari human benchmark

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.972738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:44.990370Z digest=sha256:f6452ebf6d62d72f95984fad1e9c9cd4c6df18882206caa3b13417c6400ddb84

Observation 06676397-6534-4398-8f3b-c47ad9508a58 · outbound

This paper cites Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:44.993876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:44.993876Z digest=sha256:a93562b7871709274577e4e5db21b84e6a38690355ba791e16952031f3cb9780

Observation 500ce3f9-1759-4886-9837-0aaeff9386bb · outbound

This paper cites The arcade learning environment: An evaluation platform for general agents.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents The arcade learning environment: An evaluation platform for general agents

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:44.999032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:44.999032Z digest=sha256:d01ca828724993eb7c34c9ae11ecbb8a119f4ad98b391ff1d2d1c637f3d08770

Observation 2cf34361-3637-4dd8-8133-e34fa7774007 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents RT-1: Robotics Transformer for Real-World Control at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.003192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.003192Z digest=sha256:a628134b86f64113c2088496093ef81881950f0146efe3b4819f17b33637b474

Observation 1732c799-21f6-4b38-b370-1f9ef5864090 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.013831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.013831Z digest=sha256:72d27391163676ecee6c3a14151597ab1995fb0a154f7e68cfb3e3604ad8107f

Observation 8da9e833-c68d-4395-b509-f4bd10d0700b · outbound

This paper cites Language Models are Few-Shot Learners.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Language Models are Few-Shot Learners

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.021602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.021602Z digest=sha256:04829883b5509a1f2cc020ffdd132013cedc84ad5c05a931b20e3cc5ada99bb4

Observation 70444ac8-ee54-4e16-946a-323d725855cd · outbound

This paper cites Open-world multi-task control through goal-aware representation learning and adaptive horizon prediction.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Open-world multi-task control through goal-aware representation learning and adaptive horizon prediction

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.950291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.026106Z digest=sha256:b0b99384387047c32e4b66054076903f964045bb199e562f651f76e21e2c68b9

Observation f24ea4a2-6aa5-4c12-90fb-e4beb55d40b1 · outbound

This paper cites Groot: Learning to follow instructions by watching gameplay videos.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Groot: Learning to follow instructions by watching gameplay videos

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.937746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.030662Z digest=sha256:33753457a51a64b52dbd16b6053f45a33f273790152f0b78c9e0a2573fe30890

Observation 83b80df0-fc92-467f-b030-94ed50058ad0 · outbound

This paper cites Abbeel, A.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Abbeel, A

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.923865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.034180Z digest=sha256:7d7bb2644825c92e6dd01c39691009519a90c100cbe0a42407a5ae3ce837a32d

Observation 18ad4ab7-8f55-4320-bc0c-aa46bd7932f0 · outbound

This paper cites Transformer-xl: Attentive language models beyond a fixed-length context.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Transformer-xl: Attentive language models beyond a fixed-length context

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.042834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.042834Z digest=sha256:5c39e7ef420d94b683263ae6a21e1c0cf30ca9ff6a1454134d9b369351451db8

Observation 4f7cf927-0b2e-4ad3-abc6-6678d4951c82 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.046948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.046948Z digest=sha256:1f6682c715a1bb22e36e4cbbfb8ac623f6903c8d513738ab7e9eadef8dd1ed39

Observation bc4d0dcc-4b45-404c-bb45-3bccd86d4211 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.049894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.049894Z digest=sha256:1499f163ebd855a7e73542fcf02c7671138b7c443f5cbb051d1965ee9f567730

Observation 2c6542bb-2217-4c5d-8b02-0beca396f1c0 · outbound

This paper cites One-shot imitation learning.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents One-shot imitation learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.911107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.053932Z digest=sha256:2f073bd3546072d84a81903dd93c96d29f2bfffc866187677a87c756cf95ed21

Observation d26f0261-4177-45a8-8f12-8f61498869d1 · outbound

This paper cites Implicit Deep Latent Variable Models for Text Generation.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Implicit Deep Latent Variable Models for Text Generation

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:41:45.530666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.057282Z digest=sha256:c6bb23759471d45479160b8ba36ad008fb8fe144fe6d8b9f62a4d65d5c0b4816

Observation 293ec0b5-6d0c-43aa-8f62-93c0d783b7d5 · outbound

This paper cites Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.061059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.061059Z digest=sha256:b57662fb6cef7531a0eca37a665dad4e084a29b8e23a7d7afe81f6b72f5a2373

Observation 13f6abe7-be3c-4463-8429-2a2b75e76ab5 · outbound

This paper cites Guss, Brandon Houghton, Nicholay Topin, Phillip Wang, Cayden Codel, Manuela M.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Guss, Brandon Houghton, Nicholay Topin, Phillip Wang, Cayden Codel, Manuela M

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.064991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.064991Z digest=sha256:7f718da84cd2109d0ae86da638146e19054cef843020e999dae456b98028e0df

Observation 02419236-e53b-41a7-8067-4f6d8b769539 · outbound

This paper cites Vln-bert: A recurrent vision-and-language bert for navigation.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Vln-bert: A recurrent vision-and-language bert for navigation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.891844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.068745Z digest=sha256:9303926ba5ee3ca5c6e0fe2428769aac7e6059bc78725a8ef4e364aab17b06c1

Observation 14dc7441-235e-4a76-a095-56bdd444ded0 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents An Embodied Generalist Agent in 3D World

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.074200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.074200Z digest=sha256:5e343348c7fba468358401405e94b132ad58a2802344ac71df1eeaebac8eae4a

Observation 5fafca65-a57c-47b7-ab54-a0b2091c12a5 · outbound

This paper cites Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.078763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.078763Z digest=sha256:27d84eb6dea743da16c67643c7ddd943838be55acefc8c51918bc804d7029950

Observation 61b1740d-0afc-4bfd-9177-f8d85f0c2d28 · outbound

This paper cites BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.083397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.083397Z digest=sha256:97d362c01112ac3730aa18a7816c540f64428ba9fd5f209dd71807a447f189a4

Observation 0352c213-bfdc-4930-bc33-46e87ae24b0f · outbound

This paper cites The malmo platform for artificial intelligence experimentation.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents The malmo platform for artificial intelligence experimentation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.878744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.088113Z digest=sha256:1d995e90f0c8c669d98b609ce349c5658435c631096e5368456a2777e4d4a1a3

Observation 1b0a3218-9245-42d3-960d-a904105392fc · outbound

This paper cites Auto-Encoding Variational Bayes.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Auto-Encoding Variational Bayes

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.092438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.092438Z digest=sha256:5e0179dba6408451d2daf3dcdacd9b2daffd3ffe84badfc90a085090d76ca7b6

Observation 3f602080-664e-463e-a314-8bd7086f9af2 · outbound

This paper cites Multi-game decision transformers.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Multi-game decision transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.096021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.096021Z digest=sha256:46e5a464171ccd7f48c967118541a451383794a7c76ccaaa91aa6ea0fa85407f

Observation 3daa7d74-c326-499c-82d9-52b066f61ea0 · outbound

This paper cites Evaluating Real-World Robot Manipulation Policies in Simulation.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Evaluating Real-World Robot Manipulation Policies in Simulation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.100715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.100715Z digest=sha256:d8f10d279d155983797d60af83b4ba284f6905009bef6adac4c727678e78292d

Observation 158adb09-896f-45fb-9640-3dd60c27624c · outbound

This paper cites STEVE-1: A Generative Model for Text-to-Behavior in Minecraft.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents STEVE-1: A Generative Model for Text-to-Behavior in Minecraft

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.105155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.105155Z digest=sha256:9565ef62dd3fc41b6dac0e90bc5ceb60420904630c48536286a087375b080302

Observation 4bb96c3a-44ef-4fd2-8a33-8135c6256bc9 · outbound

This paper cites Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.116529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.116529Z digest=sha256:01050ba303e289047a6dc66c2c5e2531db3354142635447046390e983871ea42

Observation 151077ff-6bb6-4603-adb4-be779ff5523e · outbound

This paper cites Language Conditioned Imitation Learning over Unstructured Data.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Language Conditioned Imitation Learning over Unstructured Data

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.120926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.120926Z digest=sha256:bdf5ca3eb5d892397558fa9ab20bb972bb2aa2729afc3a560dfed325c997764b

Observation 8b7ab2c4-e7d8-40eb-acd9-3de1f078aaf5 · outbound

This paper cites Learning latent plans from play.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Learning latent plans from play

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.854733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.124914Z digest=sha256:26217affc0acd433ff00d741a0d3a54e3cf9b20b9f4a85e16d6d5333ac2d2e73

Observation 8f0b4f59-1616-491c-9527-c1e1dc268d6d · outbound

This paper cites Learning latent plans from play.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Learning latent plans from play

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.841767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.128533Z digest=sha256:41dc22f4fa04807d6c02c1e1edc59d2d41a8d037fcdb9da765828127f346b0dd

Observation 85dae5c4-4642-4d51-9bf8-e5e2390ec52a · outbound

This paper cites Learning latent plans from play.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Learning latent plans from play

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.827698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.132613Z digest=sha256:5166ea9e6e7b1a839db1e03dd74315dfbb59c517a0a79bd6a3150df424fa4508

Observation 53c39c27-ffb2-4a5c-a38b-200f01ec7b37 · outbound

This paper cites Interactive language: Talking to robots in real time.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Interactive language: Talking to robots in real time

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.136889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.136889Z digest=sha256:59958e75030b051e2fbcd16a3ba529c6716b1954c2b7226dc7b319f0e2df1ce6

Observation 1c85f61f-d6cc-400c-9029-c29b0d88694e · outbound

This paper cites ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal Embeddings.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal Embeddings

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.140755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.140755Z digest=sha256:13e2e57e887f562d394afaaa55ad380a1cb69bc909c1a452d31a66baef51b690

Observation ee0692f7-bc58-4de7-9fd4-0870847b5b14 · outbound

This paper cites What matters in language conditioned robotic imitation learning over unstructured data.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents What matters in language conditioned robotic imitation learning over unstructured data

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.807648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.145567Z digest=sha256:2b71efe7ab38575e8ec506857801eaa0bb6da2b9e29d83faa36015d6da038481

Observation f5b071b7-f4f6-432b-978a-2fe318553e90 · outbound

This paper cites Rusu, Joel Veness, Marc G.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Rusu, Joel Veness, Marc G

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.149036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.149036Z digest=sha256:b063f8f5e18a5371a4a9eb911cba850f55edb211c79accbaa51b964a31ed8a70

Observation 58d8e0f2-56fc-4a86-b104-0c53e31d9b77 · outbound

This paper cites Goal representations for instruction following: A semi-supervised language interface to control.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Goal representations for instruction following: A semi-supervised language interface to control

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.787144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.152603Z digest=sha256:dc4c08226faad1508f161014c17e9277cfa13bbe43111c12e18f27a312f82f0b

Observation f89879f9-987a-4a0c-99f1-cc009537a33d · outbound

This paper cites Octo: An open-source generalist robot policy.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Octo: An open-source generalist robot policy

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.156190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.156190Z digest=sha256:2d13b177547b3182c671e1849c8ab5cdf8a0c4f6e0b22f9088e57f1b4dd5beaf

Observation 152c4e05-a8ed-4815-983d-9f22f37cce05 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.159279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.159279Z digest=sha256:f002848778f1cc2ec12d45012075381235d7feabc26372c36392ac7a8e29d28f

Observation 0bb8bfe7-c9e8-4958-aef8-a31f933bc552 · outbound

This paper cites Conditional Variational Autoencoder for Neural Machine Translation.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Conditional Variational Autoencoder for Neural Machine Translation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.162510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.162510Z digest=sha256:2524956716ba498751c2d5f25a54f0f9ecc8a8245e42074aad07f410ad294462

Observation 84e935d3-5426-47a3-b218-c758916988de · outbound

This paper cites Episodic transformer for vision-and-language navigation.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Episodic transformer for vision-and-language navigation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.766547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.166517Z digest=sha256:3dd1696487353bd888f453fe449e37d891af26b162cf72cf3e5a2ba832e309f5

Observation 87dd3bdf-20cf-41ee-badf-f7e0df4ffe25 · outbound

This paper cites Accelerating reinforcement learning with learned skill priors.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Accelerating reinforcement learning with learned skill priors

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.749383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.169963Z digest=sha256:db19c785353a93452a6ee7479280b36038f850ede6077e12e235fe2baf32bbae

Observation b0412b2c-a674-4c5b-812e-05dd181154cb · outbound

This paper cites Scaling Instructable Agents Across Many Simulated Worlds.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Scaling Instructable Agents Across Many Simulated Worlds

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.173544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.173544Z digest=sha256:54713195851c16c766fd35927672bc6353b3453aa9f84b6fe88a895a5ebb3720

Observation ea1ff3cd-f5d9-45a3-b93d-d96a13eb9d6c · outbound

This paper cites Improving language understanding by generative pre-training.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Improving language understanding by generative pre-training

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.177377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.177377Z digest=sha256:8075b5933fe3c47d0855279e05e6086afeaeb150cb4a3005c59ea3eacff7555e

Observation 35cf3de0-21f3-4e86-848e-2cd6bc0f9f24 · outbound

This paper cites Language models are unsupervised multitask learners.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Language models are unsupervised multitask learners

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.181411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.181411Z digest=sha256:77b0ce6ae5c10329add5f8d48a821d3a767dbf9ef287c5cbceacbf3ce566ecce

Observation 4c021881-6925-4123-a7b0-f3c5c08f3fee · outbound

This paper cites Learning transferable visual models from natural language supervision.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Learning transferable visual models from natural language supervision

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.186191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.186191Z digest=sha256:cdfbe9c211ddcd58b3b6583293150bfcd7d17c52f9ffacf25a96fa9c781570e2

Observation 407263d1-4278-4042-b2b2-63265791bec9 · outbound

This paper cites A Generalist Agent.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents A Generalist Agent

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.191304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.191304Z digest=sha256:de90454cb9ffe65910a3e5368f48c4a40d19444cd91d07b78311f354b63d0830

Observation e417dc5c-f1c3-4089-a3b8-ec46ce4657ab · outbound

This paper cites Habitat: A platform for embodied ai research.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Habitat: A platform for embodied ai research

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.195401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.195401Z digest=sha256:713ca3d1722103d112ccd95463caeef67af6667b8fdfc37740cea40a58d8d529

Observation c31d09c1-597f-45b4-a07a-4debc3ede1aa · outbound

This paper cites Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.198736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.198736Z digest=sha256:75e42498359e4b84f6950c12ae671410853cac576e9c92789fdbb0577394538f

Observation f5f0f9ab-0452-4a05-859e-96f9a2094da0 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Gemini: A Family of Highly Capable Multimodal Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.203599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.203599Z digest=sha256:06b0706976dcb73778979ea29bc20256e9d9973b1c681ee3fb63075716a0f1e8

Observation 3b518a75-3d8d-4552-bceb-8f71998c5f2b · outbound

This paper cites Describe, explain, plan and select: interactive planning with llms enables open-world multi-task agents.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Describe, explain, plan and select: interactive planning with llms enables open-world multi-task agents

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.706557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.213308Z digest=sha256:f47612804315505fbf3eb1c77350a84a22b5edd464e02aaebc5ab864dd637eb3

Observation e80437eb-59ba-462c-83a0-4a437aeed3ba · outbound

This paper cites JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.217706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.217706Z digest=sha256:ca823ff70766ab3415d6a25957117de5355d8d58dfb786aafc1215b34708fd3f

Observation 3b12a3f0-374d-4c8a-91ef-fbd89547d5b3 · outbound

This paper cites Xskill: Cross embodiment skill discovery.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Xskill: Cross embodiment skill discovery

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.695643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.223484Z digest=sha256:fe0f5f66945cdda0818ded25db5e9640c7a39dc4ed217a7f70b050d0bfcbc02f

Observation cdb9dadc-fcbe-496d-b6d0-89fe6d0c414c · outbound

This paper cites Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.228028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.228028Z digest=sha256:943440e572454c87a4c194d9de85e3c29588db2489808b15cd1708f403f316b0

Observation a5e56cbd-a5b0-421b-bd47-db18e53ca0e6 · outbound

This paper cites Learning to drive by watching youtube videos: Action-conditioned contrastive policy pretraining.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Learning to drive by watching youtube videos: Action-conditioned contrastive policy pretraining

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:45.685235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:41:45.232668Z digest=sha256:59ddc109a93181f7190aa6160a3afab2cb88ebea28e7004a55999d11acf1277a

Observation 6a3966b8-a3b1-4ae6-b34e-daebfa857382 · outbound

This paper cites Hierarchical Task Learning from Language Instructions with Unified Transformers and Self-Monitoring.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Hierarchical Task Learning from Language Instructions with Unified Transformers and Self-Monitoring

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.236842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.236842Z digest=sha256:4d9c1cf297eff4196c4f7f3b9ebc199d8ece972135a05fee5c9e30eaeefc1fa2

Observation 09bb8d68-c8de-434a-bcf7-6a0cf4209639 · outbound

This paper cites @esa (Ref.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents @esa (Ref

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.240583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.240583Z digest=sha256:0fa84dce362a9bd46a8c68c08d82ef7341a0fbf71ffc714f29a424dd3d3dfd6a

Observation 83a22687-78ff-4f75-b957-fa99edf90c46 · outbound

This paper cites an unresolved cited work.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.245534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.245534Z digest=sha256:0c920064eea7c610ae50c5156e322a582f8a47555e56838058c3a4149f39ec9c

Observation 6b90dc01-faa4-4423-835c-a6af1e088068 · outbound

This paper cites MCU: An Evaluation Framework for Open-Ended Game Agents.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents MCU: An Evaluation Framework for Open-Ended Game Agents

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.249886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.249886Z digest=sha256:e42cafebcbc4cdde97700deeb065a0f193f5448bd7d699d1ed1290170fc9a741

Pith citing papers

No inbound Pith citation observations are available.