Pith. sign in

Paper Citation Record · LEDGER

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation

As of 15 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2506.11380.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11380 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:14:29.482855Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8bf112e4-dc4c-453b-8352-2d8b5b85dced · outbound

This paper cites an unresolved cited work.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:14:30.500766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T04:14:29.011158Z digest=sha256:a4465ee526e8968442f6621406b1b8daa7ed540a17b435600d9885972f8851a6

Observation e609a620-eb69-4781-b6fc-ba4c1f745fb6 · outbound

This paper cites LLM+P: Empowering Large Language Models with Optimal Planning Proficiency.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation LLM+P: Empowering Large Language Models with Optimal Planning Proficiency

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:14:28.867443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:14:28.867443Z digest=sha256:24a487b84b4924601084dd8fba9b08caa4d1335d063c8d9c3c4d203bd027864d

Observation 66bc283c-1e62-4a2d-920b-c5d976a44efc · outbound

This paper cites an unresolved cited work.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:14:30.332972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T04:14:29.106652Z digest=sha256:ebfd904a8654f6115d0bcc98265187e83171edcc6d2a05cff13b90ef5929ebae

Observation 321b4cd2-42c1-439b-a103-54ddd6c14e09 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:14:28.955153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:14:28.955153Z digest=sha256:afb382a6a82e41ed71a4c83f5948cf3b077d92c0e6a25f82744c9c96dfbf199e

Observation 06fdeab4-66da-4dd4-b4dd-9b766fa1cee8 · outbound

This paper cites an unresolved cited work.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:14:30.408760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T04:14:29.058811Z digest=sha256:fd27ba95dd8b1266e080b24fb83a46d74cd1c3825d3d7ceef6b51096ff803394

Observation 712feb65-3b2e-4312-9fd6-b0966e665e53 · outbound

This paper cites Let’s improve a step toGwith visual information.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Let’s improve a step toGwith visual information

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:14:30.250877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T04:14:29.150192Z digest=sha256:5e2d59275d4b833d25afdac6ad9ee69c603934df48222b32e69f588b670de773

Observation a05a9ca8-2151-4fd8-9381-5f379e8da042 · outbound

This paper cites an unresolved cited work.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:14:30.153594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T04:14:29.200602Z digest=sha256:f6ec4f4489f18b8dcd596c722cf963a11468af4d9fa2feec399e2029fdcef1e4

Observation 77a8981f-afc0-449a-9d31-6e5ea9f32dbd · outbound

This paper cites an unresolved cited work.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:14:30.057830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T04:14:29.243794Z digest=sha256:50d3140177cfc32a5bae545ba58012a96a9e5e880ff09787e6b34cc609aa97dd

Observation de4054d2-4b13-4090-a2bb-7f310d10c171 · outbound

This paper cites Tie” for “Textual Quality.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Tie” for “Textual Quality

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:14:29.957433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T04:14:29.290892Z digest=sha256:289feca49f28badd16f6a2eff41ab11f9fd91f0e6769cc31e87565cf4eb5c840

Observation 97f39116-2264-4adb-9b17-92084d95a3c1 · outbound

This paper cites an unresolved cited work.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:14:29.879480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T04:14:29.344193Z digest=sha256:016a04be261218da6b7c088b36c6051079ef654f5d1b79089d0f93d45be415d5

Observation 52783f23-e09a-4689-be77-670abb294039 · outbound

This paper cites an unresolved cited work.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:14:29.798378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T04:14:29.396630Z digest=sha256:d5d4bb9cadaf8e3aadb5b332bb84aacd0b933a61597e9a2cca493ff3171ea873

Observation 06bc4e8b-cb5f-4fe8-a427-3d13067a0b17 · outbound

This paper cites an unresolved cited work.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:14:29.711568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T04:14:29.437039Z digest=sha256:9f8349de785c494b8ad67478f28d100b4392a63dc10547eb67b5bf84e8823fca

Observation 6f0d2a26-9721-4c60-a1a0-4ab35859dfc3 · outbound

This paper cites an unresolved cited work.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:14:29.630085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T04:14:29.482855Z digest=sha256:3d9a61a5d242deb4d375c2bf298d43924d3d6642ec8b0e9c4cea5e17fdb9b743

Observation baed8f9a-6e54-423d-8192-56728cce2a6f · outbound

This paper cites InstructPix2Pix: Learning to Follow Image Editing Instructions.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T04:14:28.819047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:14:28.819047Z digest=sha256:ea3a0bbfb0c835903d82d7461e8637cec26fd2bfefee29f53a580e1376d40a4d

Observation b11bbddd-215b-449a-951e-b61d403315af · outbound

This paper cites LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language Models.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:14:28.908967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:14:28.908967Z digest=sha256:f0369e02801b1f6ae446d6e56ba1b3d3f43b296ee39b54c90de03ae912a62974

Pith citing papers

No inbound Pith citation observations are available.