Pith. sign in

Paper Citation Record · LEDGER

CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2412.20451.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.20451 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:26:03.787816Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T13:16:59.192241Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e40c096e-f28b-4109-bf45-bfaee6bd7b19 · inbound

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge cites this paper.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:03.787816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:03.787816Z digest=sha256:ffcd2e8c752e1bea0c6a0a7966b03b83bb9fa97ec47a9cfb3d4f35c9325512a7

Observation 1e4cf007-3d7a-4a75-b37c-c233baff2257 · inbound

FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies cites this paper.

FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T05:48:48.013404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:48:48.013404Z digest=sha256:776e3e6593907ffd2673d4ba026556b1e8ceffb51c299be13c71afa3053c637f

Observation 2cb37b32-e59c-4e28-904d-aa6761045a37 · inbound

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies cites this paper.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:04.939069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:04.939069Z digest=sha256:4a7ae2cb5c950abda2306600f4dc63b2408b20176b677179bea35bcc4f7db779

Observation 46fd2d0a-a3d7-4d3c-8c44-198cfb7a3dba · inbound

VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models cites this paper.

VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:58:26.715669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T00:54:36.125258Z digest=sha256:8a1abc6358e3126027f90db8111efbdd7140f4cc7497b0a6e41cebcc445e5e25

Observation e0ea0d9d-76c4-4400-9481-c83e3acaa1d2 · inbound

TriRelVLA: Triadic Relational Structure for Generalizable Embodied Manipulation cites this paper.

TriRelVLA: Triadic Relational Structure for Generalizable Embodied Manipulation CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:36:09.184585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T14:54:48.058400Z digest=sha256:160b9c9e0e4d2ca765373d3231e47bb0b4f9c1515f2c11b355670b73444e3e13

Observation 97c78094-a882-4347-91c7-b03004757dd9 · inbound

General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling cites this paper.

General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance

Reference 144

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:33:27.807771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-29T13:33:03.368006Z digest=sha256:97ea6fc9b54321f2e4633958529603efc9baaf25f5c3a7f5397ed25de825e304

Observation 9aaf89e3-443e-4bd7-882d-8f4c0dd255ae · inbound

GraspGen-X: Cross-Embodiment 6-DOF Diffusion-based Grasping cites this paper.

GraspGen-X: Cross-Embodiment 6-DOF Diffusion-based Grasping CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T21:16:13.226599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T17:26:12.044711Z digest=sha256:46b6642041a235c31eb7d439b6ddec8d88e38f39dc3460fcf553d524855e9269

Observation af54694b-a523-42e1-8949-d905413d30aa · inbound

AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding cites this paper.

AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:16:59.193734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T01:23:02.576098Z digest=sha256:5054eaea4de210f382513245c349e8d2f57b3178c11e658496c5fb22ef5dc59f

Observation 3b8c218e-6a88-40b9-90ab-70045d15195f · inbound

Data Pyramid for Embodied Manipulation cites this paper.

Data Pyramid for Embodied Manipulation CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance

Reference 195

Resolution
unresolved
no resolver link, observed 2026-07-31T06:18:55.737801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:18:55.737801Z digest=sha256:98c6f7acee31c70976ab3f86f4295799110a27e5a429be800804b5927c11987f