Pith. sign in

Paper Citation Record · LEDGER

ViSTa Dataset: Do vision-language models understand sequential tasks?

As of 16 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2411.13211.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13211 v2

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:45:47.406893Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fcc889a0-0f25-4e85-9f8f-3a7b480f2842 · outbound

This paper cites Mas- tering the game of go with deep neural networks and tree search.

ViSTa Dataset: Do vision-language models understand sequential tasks? Mas- tering the game of go with deep neural networks and tree search

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.161355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.161355Z digest=sha256:c76bf5475ec973f02a112f67266e8e518cb0de6f940c657e67d5399c77955e1b

Observation 534fdd08-0d02-4d4d-bc0a-d5ea9d2d8de7 · outbound

This paper cites Deep Reinforcement Learning for Robotics: A Survey of Real-World Successes.

ViSTa Dataset: Do vision-language models understand sequential tasks? Deep Reinforcement Learning for Robotics: A Survey of Real-World Successes

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.168907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.168907Z digest=sha256:f20ff1ad994afabc2677e3cd5a5e4f90522064e96379a6ed8fdef3872db91f26

Observation e7042915-8ccb-4735-8104-11b6b7592550 · outbound

This paper cites Specification gaming: the flip side of ai ingenuity.

ViSTa Dataset: Do vision-language models understand sequential tasks? Specification gaming: the flip side of ai ingenuity

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:48.155246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T16:45:47.176949Z digest=sha256:f4c77ac6d68d9035efc3e2e6bc0623751af3712fbd23538fe4688d8bfdd3558b

Observation e93a6f76-1f30-4f21-868b-f11ca058f76c · outbound

This paper cites Defining and characterizing reward gaming.

ViSTa Dataset: Do vision-language models understand sequential tasks? Defining and characterizing reward gaming

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:48.138168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T16:45:47.184902Z digest=sha256:fe5feb207b3541c4bcf2f0d4146e5162096ea3a249629a18f6cb764c0df5fee4

Observation ae624301-8431-4ccc-9a5a-72ae93f62973 · outbound

This paper cites On the importance of hyperparameter optimization for model-based reinforcement learning.

ViSTa Dataset: Do vision-language models understand sequential tasks? On the importance of hyperparameter optimization for model-based reinforcement learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:48.120939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T16:45:47.191282Z digest=sha256:aff517f6700759951497645cfa20f7e66f8bd7704f681cc45b5e40bfc431da5c

Observation ee53c557-2a5a-4e63-a00a-7aada7a3959c · outbound

This paper cites Variational inverse control with events: A general framework for data-driven reward definition.

ViSTa Dataset: Do vision-language models understand sequential tasks? Variational inverse control with events: A general framework for data-driven reward definition

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:48.101686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T16:45:47.198072Z digest=sha256:b549218df084d9e75f7e18e87bb216be0268654e5d573f253e0d870f2aae08cd

Observation b13ec86d-3df7-422f-bb34-42b565521e38 · outbound

This paper cites Deep reinforcement learning from human preferences.

ViSTa Dataset: Do vision-language models understand sequential tasks? Deep reinforcement learning from human preferences

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.204429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.204429Z digest=sha256:5ac0564667a7d8b867eab9496858975e420ab55ec828ea84a28680d5093e7ede

Observation 1c616136-d84a-46b6-ad8c-e02bdf1af33c · outbound

This paper cites Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning.

ViSTa Dataset: Do vision-language models understand sequential tasks? Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.211087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.211087Z digest=sha256:cc790459dd5bd7a6502651c27638f0e432275f8e3d29ca3cc64ddbd516fe93e2

Observation ae1fa5bf-97f7-4b05-8bdb-1c3086adaa46 · outbound

This paper cites RoboCLIP: One Demonstration is Enough to Learn Robot Policies.

ViSTa Dataset: Do vision-language models understand sequential tasks? RoboCLIP: One Demonstration is Enough to Learn Robot Policies

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.217346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.217346Z digest=sha256:0bbfbdf484bfc98dd8a8df21bb16d23aa71c9fbc121817057e05b3e4fb43de8b

Observation aa9aa0c6-3229-493d-a173-15357e8c13cb · outbound

This paper cites Task Success is not Enough: Investigating the Use of Video-Language Models as Behavior Critics for Catching Undesirable Agent Behaviors.

ViSTa Dataset: Do vision-language models understand sequential tasks? Task Success is not Enough: Investigating the Use of Video-Language Models as Behavior Critics for Catching Undesirable Agent Behaviors

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.222909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.222909Z digest=sha256:a8105a98610de7c34d779863f6addeb665cc6c7340e027e6b6e97845e4f96183

Observation a0169577-c9f9-4b4f-aace-dcf08d21b2bf · outbound

This paper cites Learning Generalizable Robotic Reward Functions from "In-The-Wild" Human Videos.

ViSTa Dataset: Do vision-language models understand sequential tasks? Learning Generalizable Robotic Reward Functions from "In-The-Wild" Human Videos

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.228810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.228810Z digest=sha256:029d34a57b519fa8de56b47e4e8f7ab4033b829d2f510cbd54bae20d3fda8e4c

Observation 6718bc60-6efa-4b08-99ee-8934da57ee5c · outbound

This paper cites The unsur- prising effectiveness of pre-trained vision models for control.

ViSTa Dataset: Do vision-language models understand sequential tasks? The unsur- prising effectiveness of pre-trained vision models for control

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:48.070961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T16:45:47.233963Z digest=sha256:2c254ca455ef481f9134dfa3ddc61a7e74ce9c6d3bc6ad5281b40df926f3ec1f

Observation e8819983-5ecb-496f-bbce-4df5d85468f5 · outbound

This paper cites Reward Design with Language Models.

ViSTa Dataset: Do vision-language models understand sequential tasks? Reward Design with Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.239421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.239421Z digest=sha256:38f6bb15af59fc940f0fa5234616a70fe0ecf4d20bafe33e7b831a1feccb135b

Observation ed1eaabf-97a8-4d83-ba2c-94c173cec3a2 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

ViSTa Dataset: Do vision-language models understand sequential tasks? RewardBench: Evaluating Reward Models for Language Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.244847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.244847Z digest=sha256:434b337caa564e30034c8386a537f3639142da91087ca311174695ee896431f3

Observation 2d89699d-a082-48fc-9303-2e05cfd4779b · outbound

This paper cites Reward learning from narrated demonstrations.

ViSTa Dataset: Do vision-language models understand sequential tasks? Reward learning from narrated demonstrations

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:48.053066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T16:45:47.251372Z digest=sha256:4f6ffec7c5f822a671ffe5fd9404673fc301158ce87e53dcf185949b6f7fd4cd

Observation 710e80fd-f955-423f-8ebe-7018aaa971bc · outbound

This paper cites Zero-shot reward specification via grounded natural language.

ViSTa Dataset: Do vision-language models understand sequential tasks? Zero-shot reward specification via grounded natural language

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:48.035370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T16:45:47.256190Z digest=sha256:c072bc189af7e418536bb1c168debacd40a21197daf04c6a4fbe5c53831bb7dd

Observation cc57d9c5-0e98-4c8f-95e1-cf4e4564c3f0 · outbound

This paper cites RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback.

ViSTa Dataset: Do vision-language models understand sequential tasks? RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.260990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.260990Z digest=sha256:38423fe242e6ee3692d3e4bd1943e18b8c5a607e4bf1c6702c88ac0892618705

Observation e93e662b-5681-4cfe-b64a-ed702c29e75b · outbound

This paper cites PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain.

ViSTa Dataset: Do vision-language models understand sequential tasks? PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.265925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.265925Z digest=sha256:6c39cf8c6bc97d9fa2673d63d1da1ae2d8f3f7067e4b698f88b47277016ef3de

Observation bd236a1c-d166-4044-b0fc-81fa4e9e125e · outbound

This paper cites Paxion: Patching Action Knowledge in Video-Language Foundation Models,.

ViSTa Dataset: Do vision-language models understand sequential tasks? Paxion: Patching Action Knowledge in Video-Language Foundation Models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:48.018311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T16:45:47.271210Z digest=sha256:dd5bf157e33544201017996005fc9421bd9417cca9aa00dba0044b2bf71b325d

Observation 12d3fae1-9b50-49d4-b082-a0c025bed755 · outbound

This paper cites BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation.

ViSTa Dataset: Do vision-language models understand sequential tasks? BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.281865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.281865Z digest=sha256:42feea9ba841fffd6f9e041a9ced1fc50d4088f4c7a4fc5352be7d71515411c8

Observation 5874afd9-ede9-4663-bd7a-ac3b4119de37 · outbound

This paper cites VirtualHome: Simulating Household Activities via Programs.

ViSTa Dataset: Do vision-language models understand sequential tasks? VirtualHome: Simulating Household Activities via Programs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.286943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.286943Z digest=sha256:76a7173daf2c03a6885f885ca2580656a502c56b2859e0819cbac0640f532acd

Observation 71db451c-40d2-45f9-8b25-7e66452edb7b · outbound

This paper cites Habitat: A Platform for Embodied AI Research, 2019.

ViSTa Dataset: Do vision-language models understand sequential tasks? Habitat: A Platform for Embodied AI Research, 2019

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:48.002322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T16:45:47.292786Z digest=sha256:c5edf23075257e1815310bce6c7987840c12f8fac48f019ca8e46cf05d6a4a7f

Observation bb1bd04c-aecc-4cc9-89e0-8037a97f1019 · outbound

This paper cites Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots.

ViSTa Dataset: Do vision-language models understand sequential tasks? Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.297556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.297556Z digest=sha256:95ebae3464dacfc0f7e377811e0273a73f4b2df723fa305fac840b5075cb36ab

Observation 2279356c-1ac0-415b-9d63-798035e8fa5d · outbound

This paper cites ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks.

ViSTa Dataset: Do vision-language models understand sequential tasks? ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.303205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.303205Z digest=sha256:714d0151203de18ce024519f4b6ac457f2a4ffb76c5be6369804225658fe9a7a

Observation f0770c95-1ffe-4f66-bf58-fe6bfe25b17b · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

ViSTa Dataset: Do vision-language models understand sequential tasks? A Short Note on the Kinetics-700 Human Action Dataset

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.308527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.308527Z digest=sha256:491e2a14b0dbf0ec10c2ce90f0345292b0aa8ab9bd68a2328b84b09687b5241f

Observation c5d0ccb6-c9d5-4759-9588-a02d4e5e6e0b · outbound

This paper cites BEDD: the minerl BASALT evaluation and demonstrations dataset for training and benchmarking agents that solve fuzzy tasks.

ViSTa Dataset: Do vision-language models understand sequential tasks? BEDD: the minerl BASALT evaluation and demonstrations dataset for training and benchmarking agents that solve fuzzy tasks

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:47.986056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T16:45:47.314036Z digest=sha256:32712acb04f22c083052aa67b82af189d0da7ce1ba4e1fb56bdd2e2ba2dd8e85

Observation 7826661d-af97-436e-821c-bab9a859cffa · outbound

This paper cites Skill Reinforcement Learning and Planning for Open-World Long-Horizon Tasks.

ViSTa Dataset: Do vision-language models understand sequential tasks? Skill Reinforcement Learning and Planning for Open-World Long-Horizon Tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.319151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.319151Z digest=sha256:64b21f69847e1721fb0f43025086b7ac1248074f4d83905ac8b2bd4493451ef8

Observation d2683055-6a7f-4111-9dd9-1f7db5471581 · outbound

This paper cites Learning transferable visual models from natural language supervision.

ViSTa Dataset: Do vision-language models understand sequential tasks? Learning transferable visual models from natural language supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.324717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.324717Z digest=sha256:dd865b434a3dccb1660a55adeef5c54c3200e77b9192c1acaf712ab433810417

Observation e028cb0c-5dc6-43e0-a6a0-b2a019949f23 · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

ViSTa Dataset: Do vision-language models understand sequential tasks? Reproducible scaling laws for contrastive language-image learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.329565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.329565Z digest=sha256:ebbb2ed82bb3150628f6b36f20077c4d54bcaf6c0a184f89e48ddd412470a80c

Observation 7cb16373-0c7a-453d-b7b7-d9eb3609c871 · outbound

This paper cites LAION-5B: An open large-scale dataset for training next generation image-text models.

ViSTa Dataset: Do vision-language models understand sequential tasks? LAION-5B: An open large-scale dataset for training next generation image-text models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.334727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.334727Z digest=sha256:2fafd3daa8011bbeafdccfcf13eb59c57d334727c262bb99cb20ecfa524a2412

Observation b1c1dc7f-dd84-4c26-a1f5-56965d91caf5 · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

ViSTa Dataset: Do vision-language models understand sequential tasks? InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.339786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.339786Z digest=sha256:55e4b74a95a34bbdeb0ee71eed91493b5dfa55afd7ba0fbca25006886c55fa2f

Observation a5599288-1fb0-4b3a-a851-82adc3c08e5a · outbound

This paper cites GPT-4o System Card.

ViSTa Dataset: Do vision-language models understand sequential tasks? GPT-4o System Card

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.344793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.344793Z digest=sha256:2003407e643cff8a02507474de4a141dd590acf8381c707724f309a0df2bc7cb

Observation b56b9f6a-7f3e-411b-b0e0-2f5d1da7985b · outbound

This paper cites an unresolved cited work.

ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:47.959247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T16:45:47.350936Z digest=sha256:c9d89a9f6107506c5b3a160821843c2396f5bce219f8e668f0272935a4c25451

Observation 0d463477-81d5-4b03-80d1-0e4d632067d4 · outbound

This paper cites likely does not describe the video.

ViSTa Dataset: Do vision-language models understand sequential tasks? likely does not describe the video

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:47.941465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T16:45:47.356076Z digest=sha256:ef9e4742297ee07bc928958e64c0d6582d2dff05e246a464417f8260b1b8c71f

Observation 30a37c82-3666-4da5-88af-79c089e76955 · outbound

This paper cites Player holds an oak fence block in their hands.

ViSTa Dataset: Do vision-language models understand sequential tasks? Player holds an oak fence block in their hands

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:47.924572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T16:45:47.361398Z digest=sha256:882936c0749c3651294a7e4ff12835a1ee548774a36096cf8ffad4525eefa63f

Observation 70e1214b-42bd-4e42-9f27-ab6895ef9a60 · outbound

This paper cites an unresolved cited work.

ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:47.907865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T16:45:47.366264Z digest=sha256:72fe0d6f39e0c8579412f1d9498020ac782c099c4e68cd86a899b28e65d82aaa

Observation df5360ce-2569-4aa7-b5ad-0b5ad313d235 · outbound

This paper cites an unresolved cited work.

ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:47.891764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T16:45:47.371881Z digest=sha256:d53a6dafd35f0b9f8ac4f1dc7d734dc311b43454d8835239a20e45c3a511937d

Observation a87f0140-0eb4-4f33-af8e-bb5b875d2568 · outbound

This paper cites an unresolved cited work.

ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:47.876042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T16:45:47.377065Z digest=sha256:51b6ea6e329129e1adfc3897b9f56f4e124b4cd65e97809401f20aed08346198

Observation 57397734-49ee-4851-95fa-7ba559af7cbb · outbound

This paper cites an unresolved cited work.

ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:47.858713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T16:45:47.382130Z digest=sha256:b64ba8acaaeb0b60f5ecc6973ea5b5c849097e209399cc7b6e284357b65d120b

Observation 1514cef4-3eb7-405d-aa8d-230ac9932723 · outbound

This paper cites an unresolved cited work.

ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:47.840209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T16:45:47.387037Z digest=sha256:7cd8f538a1644b601b55f0e6b74c17220e38f3c4588847606bd7411c6500ca96

Observation 17260624-6b44-4002-beb6-dc5def08f35c · outbound

This paper cites an unresolved cited work.

ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:47.823881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T16:45:47.392457Z digest=sha256:125fa50371ffa4a25f5b65dcf3ddfea8c2787d7f96f78723fd80d1ae891cad09

Observation 6b87587f-48bf-46ac-b72a-5ea90ab7105e · outbound

This paper cites an unresolved cited work.

ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:47.807962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T16:45:47.397470Z digest=sha256:1e50f3cdc3d2a3ca4e5140cacc288e0bacc6014fff6e4a8605b2bb8c2ea4abae

Observation 23868eb5-550a-4cdc-aaaf-bc4eb6984541 · outbound

This paper cites an unresolved cited work.

ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:47.792703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T16:45:47.402130Z digest=sha256:ece90ec716f442430d10fead1866004e7a41b30eafcbb4790c28095532c73d21

Observation 89dd84ec-b538-41e0-aba9-af172bd47aad · outbound

This paper cites Figure 10: The prompt used to obtain frame-by-frame descriptions for Minecraft videos.

ViSTa Dataset: Do vision-language models understand sequential tasks? Figure 10: The prompt used to obtain frame-by-frame descriptions for Minecraft videos

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:47.776841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T16:45:47.406893Z digest=sha256:9d214f6f79ed27a7095623170ac4104d34994d01569c75d40c51bd32c465891e

Observation 70ef0f7b-7277-46a0-89bb-879759f4bd46 · outbound

This paper cites Paxion: Patching Action Knowledge in Video-Language Foundation Models.

ViSTa Dataset: Do vision-language models understand sequential tasks? Paxion: Patching Action Knowledge in Video-Language Foundation Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.276058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.276058Z digest=sha256:af48550b3c7d95c2ff8916479fe1a09e5f3d4d1b5daa3e7b9b5861bb9efc69d8

Pith citing papers

No inbound Pith citation observations are available.