Pith. sign in

Paper Citation Record · LEDGER

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation

As of 8 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2506.11380.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11380 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:14:29.482855Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8bf112e4-dc4c-453b-8352-2d8b5b85dced · outbound

This paper cites an unresolved cited work.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:14:30.500766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:14:29.011158Z digest=sha256:bd8dc52847463859e22ff86d6c8fd4ed1adeb974c75a33694461a7fe2df3a795

Observation e609a620-eb69-4781-b6fc-ba4c1f745fb6 · outbound

This paper cites LLM+P: Empowering Large Language Models with Optimal Planning Proficiency.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation LLM+P: Empowering Large Language Models with Optimal Planning Proficiency

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:14:28.867443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:14:28.867443Z digest=sha256:b0be5e1048fd926f0fa26145680b296696303e3233e8db31e4eb879313197dba

Observation 66bc283c-1e62-4a2d-920b-c5d976a44efc · outbound

This paper cites an unresolved cited work.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:14:30.332972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:14:29.106652Z digest=sha256:6d4a8063e982493604d70378aac8fcf98d8b3dd2e4756dcc5d93e0adc96fae30

Observation 321b4cd2-42c1-439b-a103-54ddd6c14e09 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:14:28.955153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:14:28.955153Z digest=sha256:4e7b7f8d7a6b63d9e3afe5d85aa485efa7a26f511caf4e3db29e723ac771f342

Observation 06fdeab4-66da-4dd4-b4dd-9b766fa1cee8 · outbound

This paper cites an unresolved cited work.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:14:30.408760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:14:29.058811Z digest=sha256:36b67446ee5bbd30615a91c50c96e80768f16e20f786f7ebed5aabd7c5dcf52b

Observation 712feb65-3b2e-4312-9fd6-b0966e665e53 · outbound

This paper cites Let’s improve a step toGwith visual information.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Let’s improve a step toGwith visual information

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:14:30.250877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:14:29.150192Z digest=sha256:5e76121248377af096c81dde8b928726fa50c7a06c0a2dd439b1663f7bb8b29f

Observation a05a9ca8-2151-4fd8-9381-5f379e8da042 · outbound

This paper cites an unresolved cited work.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:14:30.153594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:14:29.200602Z digest=sha256:96e4caa90c7505cc5acac01f4f4aa895952836c4fa2d3ced989513c2e719cdf4

Observation 77a8981f-afc0-449a-9d31-6e5ea9f32dbd · outbound

This paper cites an unresolved cited work.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:14:30.057830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:14:29.243794Z digest=sha256:e73aeced51c5eda259fa5f4c23b6037d2bd122eda65385c1a2874397f36317a1

Observation de4054d2-4b13-4090-a2bb-7f310d10c171 · outbound

This paper cites Tie” for “Textual Quality.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Tie” for “Textual Quality

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:14:29.957433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:14:29.290892Z digest=sha256:7a4428b28de664af4e75cef8ff8e4e27171bed06d8c6c8d43ca7fa207de735a6

Observation 97f39116-2264-4adb-9b17-92084d95a3c1 · outbound

This paper cites an unresolved cited work.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:14:29.879480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:14:29.344193Z digest=sha256:0e027e09d8ce254dc3852ea333b093349500687ea88c1cc0f159021f4ccd519e

Observation 52783f23-e09a-4689-be77-670abb294039 · outbound

This paper cites an unresolved cited work.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:14:29.798378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:14:29.396630Z digest=sha256:c10b7cbd1f5834ec108a5dfaa60848ce5c14c6a0ec54f2bfff1a33ae055fc4b3

Observation 06bc4e8b-cb5f-4fe8-a427-3d13067a0b17 · outbound

This paper cites an unresolved cited work.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:14:29.711568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:14:29.437039Z digest=sha256:cfb3a2fb7161269df467e326f9e6c91e3f9fc72f187ef2c10561ff1ceb34983f

Observation 6f0d2a26-9721-4c60-a1a0-4ab35859dfc3 · outbound

This paper cites an unresolved cited work.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:14:29.630085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:14:29.482855Z digest=sha256:ada82e504f5f1f587a8169226d464478cb862b3d0d0e97a79fa5c68e389ce1ab

Observation baed8f9a-6e54-423d-8192-56728cce2a6f · outbound

This paper cites InstructPix2Pix: Learning to Follow Image Editing Instructions.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T04:14:28.819047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:14:28.819047Z digest=sha256:ccbff6aaf5ab22dff7ea355ca65c0e01f392e69a6334593217e486c328198614

Observation b11bbddd-215b-449a-951e-b61d403315af · outbound

This paper cites LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language Models.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:14:28.908967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:14:28.908967Z digest=sha256:1db70db671133b6d1246870f6f03a2aa68b5f9a277468adda33bfbe9b2f92daa

Pith citing papers

No inbound Pith citation observations are available.