Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T11:25:15.746328Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 1 inbound Pith citation observation for arXiv:2411.18270.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T11:25:15.746328Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T16:51:32.663520Z
A source-named dated measurement, never combined with another source.
Source: cited_works
15 of 15 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7172ec51-9ca8-457b-bdc2-b266f47c8556 · outbound
Grid-augmented vision: A simple yet effective approach for enhanced spatial understanding in multi-modal agents Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 006ce885-1fac-4058-b8b3-7d56b990c00f · outbound
Grid-augmented vision: A simple yet effective approach for enhanced spatial understanding in multi-modal agents Lisa: Reasoning segmentation via large language model
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dffe833-07fd-466c-8d61-9fffd69e08cb · outbound
Grid-augmented vision: A simple yet effective approach for enhanced spatial understanding in multi-modal agents Visual instruction tuning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5e76e83-acfe-42e3-914e-4f1aebd4b197 · outbound
Grid-augmented vision: A simple yet effective approach for enhanced spatial understanding in multi-modal agents Learning transferable visual models from natural language supervision
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2ae4e0b-0674-40b0-97ca-0ba6288ef644 · outbound
Grid-augmented vision: A simple yet effective approach for enhanced spatial understanding in multi-modal agents Learning synergies between pushing and grasping with self-supervised deep reinforcement learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3b38ed5b-05c9-4da1-b902-42a404e89e10 · outbound
Grid-augmented vision: A simple yet effective approach for enhanced spatial understanding in multi-modal agents U-net: Convolutional networks for biomedical image segmentation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96edc13b-0f37-4ef3-9543-41ab151f17b9 · outbound
Grid-augmented vision: A simple yet effective approach for enhanced spatial understanding in multi-modal agents End-to-end learning for structured prediction energy networks
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation af7b19c7-8707-4730-aaa9-6bfc03b1b569 · outbound
Grid-augmented vision: A simple yet effective approach for enhanced spatial understanding in multi-modal agents Deep residual learning for image recognition
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00af768d-4baa-4701-990f-5eacdfa60330 · outbound
Grid-augmented vision: A simple yet effective approach for enhanced spatial understanding in multi-modal agents Densely connected convolutional networks
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f9e1fc0-3382-498a-9189-9f464aabfe7d · outbound
Grid-augmented vision: A simple yet effective approach for enhanced spatial understanding in multi-modal agents Rethinking the value of labels for improving class-imbalanced learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a5e783f4-1e2e-4dfa-9f80-cd12ee10891c · outbound
Grid-augmented vision: A simple yet effective approach for enhanced spatial understanding in multi-modal agents An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e78e2ec3-3344-493d-9f5b-7e84464cee28 · outbound
Grid-augmented vision: A simple yet effective approach for enhanced spatial understanding in multi-modal agents Attention is all you need
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d3ee5ee-b516-43a6-a896-1e1efa0277d4 · outbound
Grid-augmented vision: A simple yet effective approach for enhanced spatial understanding in multi-modal agents End-to-end object detection with transformers
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2b3b00de-7705-49c2-b9dc-23f1f552aabe · outbound
Grid-augmented vision: A simple yet effective approach for enhanced spatial understanding in multi-modal agents End-to-end learning for lane keeping of self-driving cars
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5e63ce16-e934-42bd-b573-0dd430c3d9e7 · outbound
Grid-augmented vision: A simple yet effective approach for enhanced spatial understanding in multi-modal agents Segment anything
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 958a67cc-ee55-4b13-9a38-8bf77d5d972b · inbound
How Auxiliary Reasoning Unleashes GUI Grounding in VLMs Grid-augmented vision: A simple yet effective approach for enhanced spatial understanding in multi-modal agents
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.