Pith. sign in

Paper Citation Record · LEDGER

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning

As of 15 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2412.16599.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.16599 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T10:29:05.612175Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 86650132-cfd7-4c3e-ba1c-24f31881fd7f · outbound

This paper cites World Models.

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning World Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T10:29:05.487011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:29:05.487011Z digest=sha256:3d9f77bbf7d89c4360a6e6d721b6a1c9b4abe0fd97afa8849231411cf33081a9

Observation e1de751e-d817-4cbe-8b07-792ac818aa0e · outbound

This paper cites Reasoning with Language Model is Planning with World Model.

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning Reasoning with Language Model is Planning with World Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T10:29:05.493027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:29:05.493027Z digest=sha256:752e552d9e14d1b208ebff409b1ec0dc0d54e8daee464f549c0386753f882985

Observation 06652ee1-ac18-40bc-b991-4db7b0a3fdb3 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T10:29:05.498538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:29:05.498538Z digest=sha256:248d770d49a583754da96652d9ecc9638c46d9ea0883c1c93ded93f9a6fd3b7b

Observation b9d84a79-efa2-4d21-9d46-5e40b1c062fc · outbound

This paper cites Penetrative ai: Making llms comprehend the physical world,.

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning Penetrative ai: Making llms comprehend the physical world,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:29:06.076764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T10:29:05.504412Z digest=sha256:37977f7762d03f1eb85018cba7bcb1fbc2a8454b62a93a7cde63307a025ba150

Observation 263c16ad-82e0-4eb2-b2ba-d0f7479bb9a1 · outbound

This paper cites Stepgame: A new benchmark for robust multi-hop spatial reasoning in texts,.

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning Stepgame: A new benchmark for robust multi-hop spatial reasoning in texts,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:29:06.060982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T10:29:05.509538Z digest=sha256:6eae556d4b4e5326a3162f5a4106815ad4cbfca1a455a56d085a529499745525

Observation 4496aaef-c2ed-474e-acb7-37ec9f096352 · outbound

This paper cites Eval- uating spatial understanding of large language models,.

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning Eval- uating spatial understanding of large language models,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:29:06.045388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T10:29:05.514722Z digest=sha256:e74c01c95a6b6ca4180fbd95de48264e30976e028f878f1ee0377ad339cf3dbc

Observation 6b9fc6ba-4703-4fa6-a88d-dbcf55aad462 · outbound

This paper cites SpartQA: : A Textual Question Answering Benchmark for Spatial Reasoning.

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning SpartQA: : A Textual Question Answering Benchmark for Spatial Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T10:29:05.521033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:29:05.521033Z digest=sha256:8ad45c1cd707c179c96355cbd8c5fba678389b5c2067ec8a1c3bff6c10a8fdee

Observation dee31865-f735-4b78-9d25-15fd47421547 · outbound

This paper cites Can Large Language Models be Good Path Planners? A Benchmark and Investigation on Spatial-temporal Reasoning.

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning Can Large Language Models be Good Path Planners? A Benchmark and Investigation on Spatial-temporal Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T10:29:05.526241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:29:05.526241Z digest=sha256:62fa021f3eddaeb543f0be3338149ce6beda31b214267ab1d5fd0ba525d3558a

Observation b6b1b0c3-c5a0-44a2-8877-cb0afababbe2 · outbound

This paper cites Advancing spatial reasoning in large language models: An in-depth evaluation and enhancement using the stepgame benchmark,.

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning Advancing spatial reasoning in large language models: An in-depth evaluation and enhancement using the stepgame benchmark,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:29:06.028494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T10:29:05.531882Z digest=sha256:a05babe0d4ad74e5a4b621d3de25f198d6bcf653b8845f4ddbca96deff18f5b7

Observation b76c1a05-afa7-40b1-9724-21d0df1cc0e3 · outbound

This paper cites GPT-4 Technical Report.

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning GPT-4 Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T10:29:05.536601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:29:05.536601Z digest=sha256:5177c7d801f31ac62b154c17a8da09b027c17414e2413f891332206cd5864b86

Observation 634969cb-f05d-4736-ac92-1799865ddd12 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning Learning transferable visual models from natural language supervision,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T10:29:05.541338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:29:05.541338Z digest=sha256:8aa47d75cc78cbb1cee9c03fd05a543361ecd3ff7b9910453797a685cb889ef5

Observation 4fb1afa3-aee2-40c9-b75a-972bd31ae5b9 · outbound

This paper cites Scaling up vision-language pre-training for image captioning,.

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning Scaling up vision-language pre-training for image captioning,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:29:05.997392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T10:29:05.546151Z digest=sha256:de469086a11f5ee572cb7a33473badc1105dad25a79dc8277add50ce118755ea

Observation ee1d367d-2c0c-49b0-8369-e01b9a11d65c · outbound

This paper cites Visual instruction tuning,.

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning Visual instruction tuning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:29:05.980749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T10:29:05.550294Z digest=sha256:0d9de4055cb7f22594bb0370c75e5404e85e3ed10183fb1183d3c2a0b018b459

Observation 4da937dd-f47f-441b-b5dd-a729538d505d · outbound

This paper cites MODE: a multimodal open-domain dialogue dataset with explanation,.

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning MODE: a multimodal open-domain dialogue dataset with explanation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:29:05.963609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T10:29:05.554412Z digest=sha256:c0a3517fa4e7e36e878665547749f2588a6242ecf51c19395196b3367ec1638c

Observation cdceea7b-3961-4914-9e6a-af78bac8385a · outbound

This paper cites Canny, The complexity of robot motion planning.

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning Canny, The complexity of robot motion planning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:29:05.946630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T10:29:05.558686Z digest=sha256:ce8a5a0c589d533560d524d5b58f8be049970ad87490f921f416ee566ca03ce1

Observation d9e5c415-fc7b-4cc8-9539-f16119af22f1 · outbound

This paper cites an unresolved cited work.

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T10:29:05.562695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:29:05.562695Z digest=sha256:5f98f3f13d6ffe8c705bf6c8c50014250cccae005f4afb9e3c8cf227c6341820

Observation 930a76a6-6306-4e9c-a06c-9835a40853af · outbound

This paper cites Intelligent control and decision- making demonstrated on a simple compass-guided robot,.

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning Intelligent control and decision- making demonstrated on a simple compass-guided robot,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:29:05.916989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T10:29:05.566617Z digest=sha256:00e18a10b9356de61c07f9ebac778553694209312941ddef6b3c2e01ecdbac8d

Observation b1884237-a6bb-45f9-850a-40c0cbe2aa08 · outbound

This paper cites Plasticity of human spatial cognition: Spatial language and cognition covary across cultures,.

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning Plasticity of human spatial cognition: Spatial language and cognition covary across cultures,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:29:05.899360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T10:29:05.570617Z digest=sha256:1418e76d3bdf76efe5ca398e6618899c6c1c5fdf20f31604dec503b173f64c5a

Observation abb1f37a-3880-4ab1-8888-6c32844914e1 · outbound

This paper cites an unresolved cited work.

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:29:05.883313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T10:29:05.574525Z digest=sha256:bf106d1c80e9c7e8e6934196ddc986c6f6e4e9c2ff642ac7ed8f68914dfd8adc

Observation 9e875c7d-0722-46b8-aaf1-4a8b6bb602b9 · outbound

This paper cites Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models.

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T10:29:05.578679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:29:05.578679Z digest=sha256:1e9e0f8b82d72722fcc44157086ea68ee8fce7075afc1cf7454f84cbe58b8393

Observation e5a606e8-74fe-4f05-bc13-4f53f5c4652a · outbound

This paper cites Visual spatial reasoning,.

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning Visual spatial reasoning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:29:05.866361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T10:29:05.583011Z digest=sha256:ced9d2ffc0d7535207aae79cde840343e38845153d8e75f451e93fe1ad4c919c

Observation 557c7b6f-1b2d-44cc-86b2-e53dbf62dc5f · outbound

This paper cites What's "up" with vision-language models? Investigating their struggle with spatial reasoning.

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T10:29:05.587498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:29:05.587498Z digest=sha256:9f6f6d9b8011a2686481f217f171f8e9cfc243437f9d5c38747e52432ce2c107

Observation e8b71fba-e78b-4a87-828b-981acb1d03ae · outbound

This paper cites Improved baselines with visual instruction tuning,.

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning Improved baselines with visual instruction tuning,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:29:05.850085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T10:29:05.592077Z digest=sha256:aa9134ee124c03a7eb69122db0dd8863eeabc9f611edb8784b84bd365db24337

Observation 12ee3148-c720-496e-b595-1eeb9305831c · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku,.

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning The claude 3 model family: Opus, sonnet, haiku,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T10:29:05.597085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:29:05.597085Z digest=sha256:deed6a49b43f0c5259f143ef97e3585c29e018c8670d559642e647d00a9b518f

Observation 3b5a6f9e-fe49-470e-99c8-cac06d98752d · outbound

This paper cites Gpt-4o mini: A smaller, cheaper ai model,.

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning Gpt-4o mini: A smaller, cheaper ai model,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:29:05.822641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T10:29:05.601779Z digest=sha256:a10ff8d0f437968badb8fb2b77a8b175065d4c9d767ff05bf8d8071efb456372

Observation 447d4899-bc7a-4edc-9c37-ffd815ad3125 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T10:29:05.606668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:29:05.606668Z digest=sha256:9e3f2332a2140373e69bfd952230365420f750e4cd2f859166e3eda09e51a134

Observation 683d4dd7-22e4-4a3b-8873-1ad1e09eeae8 · outbound

This paper cites Ali iconfont,.

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning Ali iconfont,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:29:05.802894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T10:29:05.612175Z digest=sha256:26fc9a19304c08750104a357a17d4898aaf77644587eee1d1095abfa6dbf5a47

Pith citing papers

No inbound Pith citation observations are available.