Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T01:34:33.493545Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2607.14635.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T01:34:33.493545Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 168c138b-5b28-48a2-9e18-51b61348cba2 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models RT-1: Robotics transformer for real-world control at scale,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d41ccb1-28fa-4344-8721-eb0f3c1270f8 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models OpenVLA: An open-source vision-language-action model,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4340d76-8cb1-40bd-9926-c82e604cc9ac · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15d08ad3-b1d1-4906-bbc5-4ace3c238c75 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models RT-2: Vision-language-action models transfer web knowledge to robotic control,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b41b2c2-a51e-4c51-959b-28a852d27df9 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models π 0.5: A vision-language-action model with open-world generalization,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05b9df07-540b-435f-9d44-3ad02a2d2a90 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77c6d70d-6cb4-45f1-ba0e-04c1d9600072 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models VQ-VLA: Improving vision-language-action models via scaling vector-quantized action tokenizers,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1533588-0a9b-4320-a324-d7079b8c57bb · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Latent action pretraining from videos,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 476587b7-83dd-4d1d-9db3-723908f040ae · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Learning to act anywhere with task-centric latent actions,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3014be4-9765-487d-8d81-2ee225112ea3 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models RotVLA: Rotational Latent Action for Vision-Language-Action Model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef5d9e8a-f341-4bfa-91ca-2e9dbd3a5b62 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Learning transferable visual models from natural language supervision,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a7b6a1c-7f0c-4f82-805e-08098bf12975 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Align before fuse: Vision and language representation learning with momentum distillation,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e83d065-144b-446d-aecb-f7e79013b687 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9de61cc-a106-44b7-9630-e6e93b987c5e · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Visual instruction tuning,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed05aa5d-8654-43c9-8fe2-d5c4f95ef435 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Perceiver: General perception with iterative attention,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7716786-d9b1-4637-94f3-fc2a37d4d9fe · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Flamingo: A visual language model for few-shot learning,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 923415a7-9bd2-49a4-88fe-a6d3bd47c635 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c49c74d-9986-4f96-b0a7-db0f3047c11a · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models InstructBLIP: Towards general-purpose vision-language models with instruction tuning,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9637b991-567d-4564-8175-23d732540a7d · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ead814d6-4973-4ae1-9688-f1eda41eb078 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 074e1e22-2d11-4cd4-b40b-49822b0b9199 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bcd7638-d956-4d7a-9b0c-18e40d3d4ef2 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models LARA: Latent Action Representation Alignment for Vision-Language-Action Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56aa65d7-e7da-4bf5-a083-24c74d54eef3 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Making Foresight Actionable: Repurposing Representation Alignment in World Action Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a87a132-008f-48dd-9d21-a3f5767226a7 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models A formal basis for the heuristic determination of minimum cost paths,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cec96d1-84a1-4269-8b49-73017f76d0ec · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models The dynamic window approach to collision avoidance,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7acdcb17-7df6-4517-bd60-2c2202c9e3e3 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models ViKiNG: Vision-based kilometer-scale navigation with geographic hints,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45b6e0e3-aa74-46fc-a565-6864dce056a6 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models GNM: A general navigation model to drive any robot,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60b2c79c-6854-4342-9f7e-59017e491a0f · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models ViNT: A foundation model for visual navigation,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a20c86d-0c4d-4318-afe6-1498b8d62cc5 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models NoMaD: Goal Masked Diffusion Policies for Navigation and Exploration
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0e19cde-c6eb-4aca-8ce8-022e72b98cc1 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models LM-Nav: Robotic navigation with large pre-trained models of language, vision, and action,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3517109-29f0-42f2-8a4e-e7f4e012a460 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models NaVid: Video-based vlm plans the next step for vision- and-language navigation,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 641a3b0e-511b-44fb-a13e-ee9f7b980fa6 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Uni-NaVid: A video-based vision-language-action model for unifying embodied navigation tasks,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8c0ee6e-9933-4b35-9cf4-7dee1f335325 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Embodied navigation foundation model,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d593800d-480f-4f4d-b9fd-8912113eed2b · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Navigation world models,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42860d4a-8a13-45c9-adb6-3eccb20c7d28 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Habitat: A platform for embodied ai research,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ed974af-8670-469b-bac0-d40b8af23c2b · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models ObjectNav revisited: On evaluation of embodied agents navigating to objects,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d62678d-e558-41a8-b12a-c0ee58e35ac4 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10b45c3a-54d0-4c01-9a52-5529cd98d10e · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models In the main training setup, this parsing uses thespecial-token branchexclusively
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 308a49be-7698-4501-bf4d-2d46507a8aa4 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd1d99b8-d009-4994-9894-fb5cacfddc32 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models make a left turn
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49c84480-0031-416f-b8f9-295d3eb852dc · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Thedirect-fusion baselineuses the minimal shared-context interface described in Sec
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c118e976-ecdd-4dfc-83a7-1c27e862da10 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models These set- tings differ in which parts of the inherited multimodal pathway are exposed to action-loss gradients
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7554865d-e876-417b-9474-1fda0ca23fbb · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 520758be-a9a1-48c8-bd17-3eb81162391c · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Token-wise rewriting and rewriting-subspace analyses examine how strongly the inherited pathway is rewritten and whether that rewriting is broad or targeted
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fc04309-d7d6-4a5f-ab9e-e3044dfd14de · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models These statistics are supporting measurements: they are not intended to replace the token-level visualizations in Sec
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f877ce6-2ddb-4ece-971e-ff9a85ac6098 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models We partition tokens into boundary, control, spatial, and other groups, and report each group’s share of the total rewriting budget in Table VIII
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a284d3e-0b81-4b44-bc9d-dee06831325f · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models We next ask whether this selectivity is also reflected in how rewritten dimensions are organized
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3a8b086-09c5-4be3-ad73-9e8c228cfc88 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Unresolved cited work
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71385eb9-5416-4727-a888-a2b4a0614e02 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models These statistics provide additional views of how closely an action-loss-exposed atten- tion map remains aligned with its action-update-blocked refer- ence
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6701a273-65d1-432f-8b5c-3b6f98e49d3a · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models 18 provides complementary distributional views of attention stability across the full set of phrase-level comparisons
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6967e63-079c-47b5-b8e4-eb3ea5aecc2d · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Unresolved cited work
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4f04fbf-6567-4331-8616-3e12c1f68dd2 · outbound
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models 20 shows that the all-head average can hide substantial head-level heterogeneity
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.