Pith. sign in

Paper Citation Record · LEDGER

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models

As of 13 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2607.14635.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.14635 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T01:34:33.493545Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 168c138b-5b28-48a2-9e18-51b61348cba2 · outbound

This paper cites RT-1: Robotics transformer for real-world control at scale,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models RT-1: Robotics transformer for real-world control at scale,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:29.433041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:29.433041Z digest=sha256:d038163bfa32a596c044ec9a57c67ebd419f9938c618485979e47f2645b72513

Observation 9d41ccb1-28fa-4344-8721-eb0f3c1270f8 · outbound

This paper cites OpenVLA: An open-source vision-language-action model,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models OpenVLA: An open-source vision-language-action model,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:29.498298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:29.498298Z digest=sha256:c22afd6b6e491e0a738fd6a4c1be7adc0c39a36f69fe54d8a4bcdb0620d540ff

Observation f4340d76-8cb1-40bd-9926-c82e604cc9ac · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:29.563364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:29.563364Z digest=sha256:b979e88f50d40c7df28d6ade31d5cd391e8f652edad228ce901e0eaf5f642426

Observation 15d08ad3-b1d1-4906-bbc5-4ace3c238c75 · outbound

This paper cites RT-2: Vision-language-action models transfer web knowledge to robotic control,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models RT-2: Vision-language-action models transfer web knowledge to robotic control,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:29.631505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:29.631505Z digest=sha256:c82332da252d7afe50c96111dc89d4606974ac73bccf6111e6febb1db60dbca2

Observation 7b41b2c2-a51e-4c51-959b-28a852d27df9 · outbound

This paper cites π 0.5: A vision-language-action model with open-world generalization,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models π 0.5: A vision-language-action model with open-world generalization,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:29.689972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:29.689972Z digest=sha256:1eb74f479ebdedb8406732dd6a1ae0a9653860f2b31487e451a2efb7223f8799

Observation 05b9df07-540b-435f-9d44-3ad02a2d2a90 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:29.776432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:29.776432Z digest=sha256:a3a3dab91affe63bb7b46436abd80395bf15631ef20d7267977e686d1988d42e

Observation 77c6d70d-6cb4-45f1-ba0e-04c1d9600072 · outbound

This paper cites VQ-VLA: Improving vision-language-action models via scaling vector-quantized action tokenizers,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models VQ-VLA: Improving vision-language-action models via scaling vector-quantized action tokenizers,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:29.861393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:29.861393Z digest=sha256:cf19627b8bee1ddee1c26d07c93cf77784fd8049b620c0f26cdceb64c1f0032a

Observation b1533588-0a9b-4320-a324-d7079b8c57bb · outbound

This paper cites Latent action pretraining from videos,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Latent action pretraining from videos,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:29.942388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:29.942388Z digest=sha256:22aa8bb16adb88d630088408dfd3293723edf791132823644ef2f1e4e8d261c6

Observation 476587b7-83dd-4d1d-9db3-723908f040ae · outbound

This paper cites Learning to act anywhere with task-centric latent actions,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Learning to act anywhere with task-centric latent actions,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:30.060679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:30.060679Z digest=sha256:12564ede441c8cfb2cf70fa393c1770ea3a55a5f60acfbe8ab7150968979993e

Observation b3014be4-9765-487d-8d81-2ee225112ea3 · outbound

This paper cites RotVLA: Rotational Latent Action for Vision-Language-Action Model.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models RotVLA: Rotational Latent Action for Vision-Language-Action Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:30.181563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:30.181563Z digest=sha256:9d685ceba9e45a300046c9ff8c5b76a67d4047cf8dc6a34a9a671b3f56e5fb69

Observation ef5d9e8a-f341-4bfa-91ca-2e9dbd3a5b62 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Learning transferable visual models from natural language supervision,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:30.319349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:30.319349Z digest=sha256:ba3873e4c2f23c31b04e8c868e922508b2d5ae2bceb13a7747ab5121b1584b6b

Observation 2a7b6a1c-7f0c-4f82-805e-08098bf12975 · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distillation,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Align before fuse: Vision and language representation learning with momentum distillation,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:30.461696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:30.461696Z digest=sha256:2619213b5034e895f2814802a18d1f2c221d5a02d133f798a3e0a331991269a0

Observation 0e83d065-144b-446d-aecb-f7e79013b687 · outbound

This paper cites BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:30.599740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:30.599740Z digest=sha256:5762b1c83c1bd661eed281468d32f664b6ddcb41d1cd93081d1e7c395788dea7

Observation e9de61cc-a106-44b7-9630-e6e93b987c5e · outbound

This paper cites Visual instruction tuning,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Visual instruction tuning,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:30.685381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:30.685381Z digest=sha256:ecb7937637b3aae2838edea8301d15f1ae2c2c3999d9f95afdb877f61b886e7c

Observation ed05aa5d-8654-43c9-8fe2-d5c4f95ef435 · outbound

This paper cites Perceiver: General perception with iterative attention,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Perceiver: General perception with iterative attention,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:30.777303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:30.777303Z digest=sha256:7b75f1a46ac83022ab33c193d8c78a9fca7f1e57be64dc7c480efe9e0f582d19

Observation d7716786-d9b1-4637-94f3-fc2a37d4d9fe · outbound

This paper cites Flamingo: A visual language model for few-shot learning,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Flamingo: A visual language model for few-shot learning,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:30.870154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:30.870154Z digest=sha256:d46d729cb755dc1e57bcdf509f5ca282662b03614f9d38eb9123388c6c315d14

Observation 923415a7-9bd2-49a4-88fe-a6d3bd47c635 · outbound

This paper cites BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:30.960688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:30.960688Z digest=sha256:29b1b7f89c6788f6faf9de15a02c58e1bdac438c4c103902c8fef26eb408311a

Observation 1c49c74d-9986-4f96-b0a7-db0f3047c11a · outbound

This paper cites InstructBLIP: Towards general-purpose vision-language models with instruction tuning,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models InstructBLIP: Towards general-purpose vision-language models with instruction tuning,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.053667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.053667Z digest=sha256:c09c1c1c8f1b7dacbd4cc07d9b37e4bcc693c09194f5cc0f1f24827d0edd9e32

Observation 9637b991-567d-4564-8175-23d732540a7d · outbound

This paper cites VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.166730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.166730Z digest=sha256:b7631b2d9da5e46b075dd8a4509cf10fce3a384d2e9f57d301910be416cf170e

Observation ead814d6-4973-4ae1-9688-f1eda41eb078 · outbound

This paper cites DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.269443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.269443Z digest=sha256:9c66345205bb32e5935ee01e9cb30061db0a8da1dd4fc07454a8a36de56e4d1e

Observation 074e1e22-2d11-4cd4-b40b-49822b0b9199 · outbound

This paper cites HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.334140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.334140Z digest=sha256:49c150d9401b9332a2d2946ec77df903af96f0239a5a3b32d98dd7d67d15fea4

Observation 0bcd7638-d956-4d7a-9b0c-18e40d3d4ef2 · outbound

This paper cites LARA: Latent Action Representation Alignment for Vision-Language-Action Models.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models LARA: Latent Action Representation Alignment for Vision-Language-Action Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.411913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.411913Z digest=sha256:3f032fcc34a47a7e6a481fc78cd9422d54a5e9aedc54ba4be0f9bc9d3eddb960

Observation 56aa65d7-e7da-4bf5-a083-24c74d54eef3 · outbound

This paper cites Making Foresight Actionable: Repurposing Representation Alignment in World Action Models.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Making Foresight Actionable: Repurposing Representation Alignment in World Action Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.476784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.476784Z digest=sha256:0e06ce07eb7c3e217c2336abb9c7a6fc9cce7fc8250bb2da4ec60fbb97bd987b

Observation 8a87a132-008f-48dd-9d21-a3f5767226a7 · outbound

This paper cites A formal basis for the heuristic determination of minimum cost paths,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models A formal basis for the heuristic determination of minimum cost paths,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.541114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.541114Z digest=sha256:887336c882aa93c02661ef756bfbcb7bdfdc48bf69f7e8b0a588f15503809fd4

Observation 9cec96d1-84a1-4269-8b49-73017f76d0ec · outbound

This paper cites The dynamic window approach to collision avoidance,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models The dynamic window approach to collision avoidance,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.602905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.602905Z digest=sha256:66b73939712a04a71603dfd881f5b9f6844c407c8e3a6a999a4a01743271ecb3

Observation 7acdcb17-7df6-4517-bd60-2c2202c9e3e3 · outbound

This paper cites ViKiNG: Vision-based kilometer-scale navigation with geographic hints,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models ViKiNG: Vision-based kilometer-scale navigation with geographic hints,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.688754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.688754Z digest=sha256:3ca44cd8c07b59d1114a5750e1a3755207aa72840252907c701347048746aa40

Observation 45b6e0e3-aa74-46fc-a565-6864dce056a6 · outbound

This paper cites GNM: A general navigation model to drive any robot,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models GNM: A general navigation model to drive any robot,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.745987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.745987Z digest=sha256:eb7d9710fbcbb9b693d2bb339f74c2ac94a1dffd7a8cecc7e7b2b37c9f872dcd

Observation 60b2c79c-6854-4342-9f7e-59017e491a0f · outbound

This paper cites ViNT: A foundation model for visual navigation,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models ViNT: A foundation model for visual navigation,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.799338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.799338Z digest=sha256:42c6a9a75c4ce0231d4964c6fbc62cbac5879b7805f76154dfb6c7616453a333

Observation 3a20c86d-0c4d-4318-afe6-1498b8d62cc5 · outbound

This paper cites NoMaD: Goal Masked Diffusion Policies for Navigation and Exploration.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models NoMaD: Goal Masked Diffusion Policies for Navigation and Exploration

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.849984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.849984Z digest=sha256:5a2c6c97628612a32a4763a6f9ade1a1f52975fbd8e3e70e9837d4eb8d2cbfde

Observation c0e19cde-c6eb-4aca-8ce8-022e72b98cc1 · outbound

This paper cites LM-Nav: Robotic navigation with large pre-trained models of language, vision, and action,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models LM-Nav: Robotic navigation with large pre-trained models of language, vision, and action,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.928705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.928705Z digest=sha256:c4e145176e3c2f943ea00874b0968d911acfe3331b7a703072259fab9a1c3087

Observation c3517109-29f0-42f2-8a4e-e7f4e012a460 · outbound

This paper cites NaVid: Video-based vlm plans the next step for vision- and-language navigation,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models NaVid: Video-based vlm plans the next step for vision- and-language navigation,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.989633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.989633Z digest=sha256:09e4d07420f911491a0c4e58321d9a318f49a4f351e50a3d338a234771aa708a

Observation 641a3b0e-511b-44fb-a13e-ee9f7b980fa6 · outbound

This paper cites Uni-NaVid: A video-based vision-language-action model for unifying embodied navigation tasks,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Uni-NaVid: A video-based vision-language-action model for unifying embodied navigation tasks,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:32.050463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:32.050463Z digest=sha256:2824bc77581410da86cca5bcda55812a8352cd7a31be7635c3b49013d2c9e6fa

Observation a8c0ee6e-9933-4b35-9cf4-7dee1f335325 · outbound

This paper cites Embodied navigation foundation model,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Embodied navigation foundation model,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:32.102790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:32.102790Z digest=sha256:8809c6f5e1153fa53b9e901259ab55a0cdb7c2b0fa416abd973fa4b6364912d2

Observation d593800d-480f-4f4d-b9fd-8912113eed2b · outbound

This paper cites Navigation world models,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Navigation world models,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:32.173003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:32.173003Z digest=sha256:d1bdb97c0e6744851abaea4a39317cc4ab764046dd97dee52f3d813f1b4f2257

Observation 42860d4a-8a13-45c9-adb6-3eccb20c7d28 · outbound

This paper cites Habitat: A platform for embodied ai research,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Habitat: A platform for embodied ai research,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:32.225527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:32.225527Z digest=sha256:3ba6803b47e2f24dc53d118466ed336d6c0fd8f32f5001655bab9ac31e313ddd

Observation 5ed974af-8670-469b-bac0-d40b8af23c2b · outbound

This paper cites ObjectNav revisited: On evaluation of embodied agents navigating to objects,.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models ObjectNav revisited: On evaluation of embodied agents navigating to objects,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:32.327882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:32.327882Z digest=sha256:ddd71987b707baf45100e31760f079650fc4d2f748123fcb97c45549ba413853

Observation 4d62678d-e558-41a8-b12a-c0ee58e35ac4 · outbound

This paper cites an unresolved cited work.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:32.403060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:32.403060Z digest=sha256:e5a24a3d1b23efc9fcfc31f97329e245e2dcecc84f5e64608a8e56a4e547f451

Observation 10b45c3a-54d0-4c01-9a52-5529cd98d10e · outbound

This paper cites In the main training setup, this parsing uses thespecial-token branchexclusively.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models In the main training setup, this parsing uses thespecial-token branchexclusively

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:32.482082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:32.482082Z digest=sha256:db70d3c758c0c6d69f3f904172d51b9c93bff64be06b1d7a1bbb30f2f750cc37

Observation 308a49be-7698-4501-bf4d-2d46507a8aa4 · outbound

This paper cites an unresolved cited work.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:32.565185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:32.565185Z digest=sha256:d399aa094bcb4ab2766f8565420607ca87c8db58b25e99fb449d7d6837a55e79

Observation cd1d99b8-d009-4994-9894-fb5cacfddc32 · outbound

This paper cites make a left turn.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models make a left turn

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:32.642066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:32.642066Z digest=sha256:5f249b4d7611175e9fc70f984a109967050876eb4ca65b82bf352a8f8ab65baa

Observation 49c84480-0031-416f-b8f9-295d3eb852dc · outbound

This paper cites Thedirect-fusion baselineuses the minimal shared-context interface described in Sec.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Thedirect-fusion baselineuses the minimal shared-context interface described in Sec

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:32.746572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:32.746572Z digest=sha256:e34587bdee34a39c3ad1ac97bde2f394105459f4dbb44dfcfe64e253dbbb222c

Observation c118e976-ecdd-4dfc-83a7-1c27e862da10 · outbound

This paper cites These set- tings differ in which parts of the inherited multimodal pathway are exposed to action-loss gradients.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models These set- tings differ in which parts of the inherited multimodal pathway are exposed to action-loss gradients

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:32.820960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:32.820960Z digest=sha256:12482b2e330e66fa523a547182fc8dc8c27ef79923ac58051709238face5155f

Observation 7554865d-e876-417b-9474-1fda0ca23fbb · outbound

This paper cites an unresolved cited work.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:32.876278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:32.876278Z digest=sha256:f9f6869fd1fa4328f9996fc20db4ee3532920f3b594deb344c7d37a718ce239e

Observation 520758be-a9a1-48c8-bd17-3eb81162391c · outbound

This paper cites Token-wise rewriting and rewriting-subspace analyses examine how strongly the inherited pathway is rewritten and whether that rewriting is broad or targeted.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Token-wise rewriting and rewriting-subspace analyses examine how strongly the inherited pathway is rewritten and whether that rewriting is broad or targeted

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:32.984898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:32.984898Z digest=sha256:831fffce66b3f3c5967f86b4b2973a12ed9ea1a138e0674e2921f2b7323892e7

Observation 1fc04309-d7d6-4a5f-ab9e-e3044dfd14de · outbound

This paper cites These statistics are supporting measurements: they are not intended to replace the token-level visualizations in Sec.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models These statistics are supporting measurements: they are not intended to replace the token-level visualizations in Sec

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:33.064517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:33.064517Z digest=sha256:7cce53401f1eda4bf048a0590ac81edf48d32cc6380cd287a69dbcb8eda239c0

Observation 1f877ce6-2ddb-4ece-971e-ff9a85ac6098 · outbound

This paper cites We partition tokens into boundary, control, spatial, and other groups, and report each group’s share of the total rewriting budget in Table VIII.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models We partition tokens into boundary, control, spatial, and other groups, and report each group’s share of the total rewriting budget in Table VIII

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:33.108432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:33.108432Z digest=sha256:91b35ce4e832aedcd1cfeb04dc6246f2f6d30373b20a882c4e2849007ff667a2

Observation 5a284d3e-0b81-4b44-bc9d-dee06831325f · outbound

This paper cites We next ask whether this selectivity is also reflected in how rewritten dimensions are organized.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models We next ask whether this selectivity is also reflected in how rewritten dimensions are organized

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:33.159920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:33.159920Z digest=sha256:8154f772e2f81bf29d3df6b38183d1ca9347c703cfba8b753ff43bddd9995533

Observation e3a8b086-09c5-4be3-ad73-9e8c228cfc88 · outbound

This paper cites an unresolved cited work.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:33.236218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:33.236218Z digest=sha256:dc333e94515c07ca9a004499135fd42455ca8e1f359b76c9346503e3665fa1d2

Observation 71385eb9-5416-4727-a888-a2b4a0614e02 · outbound

This paper cites These statistics provide additional views of how closely an action-loss-exposed atten- tion map remains aligned with its action-update-blocked refer- ence.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models These statistics provide additional views of how closely an action-loss-exposed atten- tion map remains aligned with its action-update-blocked refer- ence

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:33.281574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:33.281574Z digest=sha256:e213d351de4f3c8e21825f0e8b835323225944664492958d7289b2f2404e1ba7

Observation 6701a273-65d1-432f-8b5c-3b6f98e49d3a · outbound

This paper cites 18 provides complementary distributional views of attention stability across the full set of phrase-level comparisons.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models 18 provides complementary distributional views of attention stability across the full set of phrase-level comparisons

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:33.348531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:33.348531Z digest=sha256:fe644dce292e35e98941dbcd372da4c13ba6eb92ede7b52e83ffc292a02940a4

Observation c6967e63-079c-47b5-b8e4-eb3ea5aecc2d · outbound

This paper cites an unresolved cited work.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:33.434174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:33.434174Z digest=sha256:33a44c1c47e47e76b833864476015132d3b081c96eb6cb0383eca7adbc9ae584

Observation d4f04fbf-6567-4331-8616-3e12c1f68dd2 · outbound

This paper cites 20 shows that the all-head average can hide substantial head-level heterogeneity.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models 20 shows that the all-head average can hide substantial head-level heterogeneity

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:33.493545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:33.493545Z digest=sha256:4e7853047ae930dd40efcd6401921e06bf879285f90a8e99891f992bb2ae8c1b

Pith citing papers

No inbound Pith citation observations are available.