Pith. sign in

Paper Citation Record · LEDGER

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?

As of 16 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 8 inbound Pith citation observations for arXiv:2504.19267.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.19267 v3

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T06:01:23.332049Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:44:02.721295Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T20:17:33.741732Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact4
  • verified fuzzy22
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6760388d-8789-4527-9afe-589bd0a2d0a4 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T06:01:23.124363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:01:23.124363Z digest=sha256:cd9ac02fa0ebbbc140929c8fc140326883222126275b73ceac9a50b9a756a489

Observation 5d826cc0-354f-4f4c-b9df-a0d0f8bb1d0e · outbound

This paper cites AugmentGest: Can Random Data Cropping Augmentation Boost Gesture Recognition Performance?.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? AugmentGest: Can Random Data Cropping Augmentation Boost Gesture Recognition Performance?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T06:01:23.129181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:01:23.129181Z digest=sha256:ab9e03b63ed3a0033d068d5d12de6e57ae3cb117b3a1dae6eea5a572662e4170

Observation dfd2f52f-2930-4980-bb98-2023403a8dc0 · outbound

This paper cites Commonsense knowledge aware concept selec- tion for diverse and informative visual storytelling.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Commonsense knowledge aware concept selec- tion for diverse and informative visual storytelling

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:01:24.043134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:01:23.133659Z digest=sha256:96d821dbfbf9446d68cbae3f61aaa3afb0245389b4f83206e1e34c22d115b874

Observation 8535f83d-afbf-4c11-9220-9e127ba9bdf9 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:01:24.031500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:01:23.138038Z digest=sha256:a0b95326a1acadc82849c5d9c14742a4ebb502e025e993e131b603e6f54af187

Observation a50d37f2-08dd-4a4f-ba75-e3bff22fc8a6 · outbound

This paper cites Content Planning for Neural Story Generation with Aristotelian Rescoring.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Content Planning for Neural Story Generation with Aristotelian Rescoring

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-16T06:01:23.702521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:01:23.142181Z digest=sha256:e9146d9b9d156291c876dce7fd92bdbd96920d697b4b01a4c30380661778f8e1

Observation c8c37ef6-9f0f-4c5d-9d23-2b7b1aa90832 · outbound

This paper cites Contextualize, Show and Tell: A Neural Visual Storyteller.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Contextualize, Show and Tell: A Neural Visual Storyteller

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T06:01:23.146325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:01:23.146325Z digest=sha256:625b861b72756e5db321a5cc9dec9056f5bbb3abe6c6db625e440f626e0ff120

Observation c9d55bb7-81e8-4bb7-9ddc-644d5df70121 · outbound

This paper cites LEMUR Neural Network Dataset: Towards Seamless AutoML.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? LEMUR Neural Network Dataset: Towards Seamless AutoML

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T06:01:23.150619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:01:23.150619Z digest=sha256:ee46a99e3ffa2ac91158a0658857e9dc8e8957636be6d98b519ec55c35e4324f

Observation 085bb07d-7049-4840-b448-a2383c2e28fa · outbound

This paper cites Diverse and rel- evant visual storytelling with scene graph embeddings.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Diverse and rel- evant visual storytelling with scene graph embeddings

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:01:24.019570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:01:23.154531Z digest=sha256:8cf5830a8d499c563345d25d11bb4bca9fc735d715dbcfd8be7de25da676560a

Observation e0349bc7-5996-42b1-bed6-007de2b9a893 · outbound

This paper cites Knowledge-enriched visual storytelling.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Knowledge-enriched visual storytelling

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:01:24.006104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:01:23.158076Z digest=sha256:20dcba056878e37958301edf2b8d22d6684282ff51b94a52f60a16b28adca897

Observation 7226fa20-1556-436a-940c-aabfbe67aa07 · outbound

This paper cites Plot and Rework: Modeling Storylines for Visual Storytelling.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Plot and Rework: Modeling Storylines for Visual Storytelling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T06:01:23.162160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:01:23.162160Z digest=sha256:41aa266e8c2e4b1b8fa94f00e2eee8dd583524c2eb4ae76c8f14c635ed3b14df

Observation e1970102-7c0b-4aac-9eb6-b318903257dd · outbound

This paper cites Visual storytelling.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Visual storytelling

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:01:23.990040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:01:23.166999Z digest=sha256:a885c91e1b1caf030cecc3b956c6d15ad9f8ed2491ffee7e88d82581343ae966

Observation 80041430-73ed-4c4f-a890-0066ea7718b0 · outbound

This paper cites GLAC Net: GLocal Attention Cascading Networks for Multi-image Cued Story Generation.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? GLAC Net: GLocal Attention Cascading Networks for Multi-image Cued Story Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T06:01:23.172231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:01:23.172231Z digest=sha256:de57ed4c531958116a984ed1c3b9ead03e046053cd36d16e3b9832736b9ce39d

Observation 8af941f5-374c-4ba1-987f-13f9f840a108 · outbound

This paper cites Optuna vs Code Llama: Are LLMs a New Paradigm for Hyperparameter Tuning?.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Optuna vs Code Llama: Are LLMs a New Paradigm for Hyperparameter Tuning?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T06:01:23.176415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:01:23.176415Z digest=sha256:e04dc0acb00c1fb8cc090dc9d1ffa4862c97bfd358349b3f36c8fda0e9b701ec

Observation 98ce1c0a-4834-4e40-9e6d-74accd52accd · outbound

This paper cites Nngpt: Neural network model gener- ation.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Nngpt: Neural network model gener- ation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:01:23.977923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:01:23.180622Z digest=sha256:973c9c4395545b64b68760ce453848ef743854c6ddeb34c593e91034ef97086c

Observation 7e0425fc-f678-4fae-bcd2-60ca1909c238 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T06:01:23.185104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:01:23.185104Z digest=sha256:d2bfea519b72d398f1bf9fb9b351c44fe39ab44b9640b84d07a909b19a4dc779

Observation ccaa7531-48cf-48c6-8c84-f72c74f78554 · outbound

This paper cites A character- centric neural model for automated story generation.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? A character- centric neural model for automated story generation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:01:23.956941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:01:23.189005Z digest=sha256:af2e8bb11bc6fa3ec048764d058a8571df594d9efbd8a5578b5da6a35054b947

Observation aea480ca-f4f3-479d-aa55-a27b25dc95be · outbound

This paper cites Visual instruction tuning.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Visual instruction tuning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:01:23.944861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:01:23.193585Z digest=sha256:dd5d9c43c43a374db64081511b9d2863ce97905e84f85acd553f85516a066fe3

Observation f33c85e4-5d4d-4c1c-ab0f-5b95b9944e7b · outbound

This paper cites Visual relationship detection with language priors.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Visual relationship detection with language priors

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:01:23.933243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:01:23.197458Z digest=sha256:0769b9b76243ed9f28d3806e6b653444f93691bb297d42b931d9b58741a8e5c6

Observation 5af4da34-a208-4bb5-b6e1-b0d39ac5dea9 · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T06:01:23.201259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:01:23.201259Z digest=sha256:effbe9905f29ecc9d48ca21a39befe8e151edec80b0befc34c1e25dd87c34aac

Observation d30d19fe-72d8-4f89-ab49-b4a270c02dab · outbound

This paper cites Step-by-Step: Separating Planning from Realization in Neural Data-to-Text Generation.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Step-by-Step: Separating Planning from Realization in Neural Data-to-Text Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T06:01:23.205297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:01:23.205297Z digest=sha256:5be93f1e0f6db0b90df1d7f2a08396623b1ff76c8564545dafa3d759e4b2f6ef

Observation 80074fde-1985-48e3-8180-1a6853fcf6b1 · outbound

This paper cites Planning with learned entity prompts for abstractive summarization.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Planning with learned entity prompts for abstractive summarization

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:01:23.921060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:01:23.209327Z digest=sha256:cc898b19d43c21f87041b449e6f5d933a085635741795cb44a11b4829a96e334

Observation caac109f-fd31-45e1-8269-e7040447eac9 · outbound

This paper cites Data-to - text generation with variational sequential planning.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Data-to - text generation with variational sequential planning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:01:23.908657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:01:23.214444Z digest=sha256:e1a029af738b00f6fb8fbaea4e669e0eb7406c78575466341cdd907a4d71a993

Observation 93617cae-c03e-4cef-a190-97f97c8862f5 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Learning transferable visual models from natural language supervi- sion

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T06:01:23.218507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:01:23.218507Z digest=sha256:5c3b5914ec08f6cb53b1ea651c8bddbcd78bb179e7b3b93f4f6b325bba5da45a

Observation e8b801c0-7a0a-434c-a615-558d6e4d1dfe · outbound

This paper cites PlotMachines: Outline-Conditioned Generation with Dynamic Plot State Tracking.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? PlotMachines: Outline-Conditioned Generation with Dynamic Plot State Tracking

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T06:01:23.222500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:01:23.222500Z digest=sha256:4e364767d1650365d671fb092655ebf800a179b4059b2a21845b8a08b0fb9cc6

Observation 1744fcef-258b-41cd-9c1c-5113ebd97892 · outbound

This paper cites Explo r- ing the collaboration between vision models and llms for en- hanced image classification.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Explo r- ing the collaboration between vision models and llms for en- hanced image classification

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:01:23.887922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:01:23.226364Z digest=sha256:de0dbb3d7a1c4befb4ca8ff24d738f5b7d1723d9edb4a1e9598f9205d02066d9

Observation 1e794097-1dcf-4006-bb43-7f98d7be9a30 · outbound

This paper cites Storygpt-v: Larg e language models as consistent story visualizers.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Storygpt-v: Larg e language models as consistent story visualizers

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:01:23.872990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:01:23.230179Z digest=sha256:0fdb704cd098f7ec518fc8d0fb5b2646a18d8997c043e2040aae63a70319eb63

Observation 9ce6d469-26ef-430d-a29f-60cc7c259f6c · outbound

This paper cites GROOViST: A Metric for Grounding Objects in Visual Storytelling.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? GROOViST: A Metric for Grounding Objects in Visual Storytelling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T06:01:23.234251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:01:23.234251Z digest=sha256:f05e4cf91794826744c4e3d58466e333f47b79f1879c25e7a096431d7adeefec

Observation 1bf0ffcf-37df-4592-8d4e-3da5001bc2a9 · outbound

This paper cites Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and Repetition.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and Repetition

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-16T06:01:23.585212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:01:23.238386Z digest=sha256:e27da9f91b675d79fd9d2203a39f22f829a469a596da2bd54e0826ec3e22732f

Observation 818e9292-15ac-4b03-8570-65409f501950 · outbound

This paper cites Vokenization: Improving Language Understanding with Contextualized, Visual-Grounded Supervision.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Vokenization: Improving Language Understanding with Contextualized, Visual-Grounded Supervision

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-16T06:01:23.567815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:01:23.242223Z digest=sha256:2f81d77ab5f72f9d0c0d3b3a6da2eb690f56be4645542f38a273d4b98d2517d0

Observation 3265e071-60ac-4744-9e67-c528d6571b8e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? LLaMA: Open and Efficient Foundation Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T06:01:23.246251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:01:23.246251Z digest=sha256:0969fc4642c4e0c3e5db5b668a076033f17643c577a8418b9e09e95eebaa4fda

Observation 8b678470-e1ee-4d09-820d-04a7f6294c33 · outbound

This paper cites Multimodal few-shot learning with frozen language models.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Multimodal few-shot learning with frozen language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:01:23.861352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:01:23.249987Z digest=sha256:e1a37de7d8291c15dd60f22f98a114aab43bb2959b25f8f62d1ed56aae79f855

Observation ed4bc41c-829b-4201-b0eb-309f1de6bef7 · outbound

This paper cites RoViST:Learning Robust Metrics for Visual Storytelling.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? RoViST:Learning Robust Metrics for Visual Storytelling

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T06:01:23.253660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:01:23.253660Z digest=sha256:b86877c37b2068bef3aef205d0148564739e5d2c646f6e6337687e5dab6fb7f0

Observation 90319ff0-983c-43ce-a91d-f84a9dd7496d · outbound

This paper cites Storytelling from an image stream using scene graphs.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Storytelling from an image stream using scene graphs

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:01:23.849200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:01:23.258234Z digest=sha256:1d6379d9f3a02616dfb0d97845b8b4af9f1b4fb7f8f002a76c6e8c907e28aace

Observation 7e9d90ce-c6a9-4208-8eb3-e48c1bb83a32 · outbound

This paper cites No Metrics Are Perfect: Adversarial Reward Learning for Visual Storytelling.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? No Metrics Are Perfect: Adversarial Reward Learning for Visual Storytelling

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T06:01:23.261733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:01:23.261733Z digest=sha256:1ad68c8539837cb806f426a1e3978f2ffc8ca721c7bcdc77b2364f3334413cbe

Observation 4a460b36-13f7-4994-892e-2be8b8539735 · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T06:01:23.265728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:01:23.265728Z digest=sha256:ea5233082485aca1d62746451c901ea7e7a9de8a5d49a33265ac6304b110e69e

Observation 2a82f8d6-699b-4ecb-a986-a04434b420e5 · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? NExT-GPT: Any-to-Any Multimodal LLM

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T06:01:23.269598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:01:23.269598Z digest=sha256:cb7976dc3b76c1e15b4fb10fcca06a6ac597517733db6eb24c07111a09398305

Observation 584e468c-d11a-4434-9a59-f8198e516b7e · outbound

This paper cites Imagine, reason and write: Visual storytelling with graph knowledge and relational reasonin g.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Imagine, reason and write: Visual storytelling with graph knowledge and relational reasonin g

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:01:23.837361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:01:23.273622Z digest=sha256:b13dae2cafa18e36ab0fed008dc6613644c73fa26a48b1ed62e7b776be6fb0b6

Observation 4d6658c8-87ef-4ac5-9b47-afb7a9b566c5 · outbound

This paper cites A Skeleton-Based Model for Promoting Coherence Among Sentences in Narrative Story Generation.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? A Skeleton-Based Model for Promoting Coherence Among Sentences in Narrative Story Generation

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-16T06:01:23.483658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:01:23.276960Z digest=sha256:b09b9dfe59623260ec7e69fb03f2b6aa4289df85567739fa1b72e06daebcdcac

Observation 81abbb5d-2571-481a-b227-705f781cf8c3 · outbound

This paper cites Re3: Generating Longer Stories With Recursive Reprompting and Revision.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Re3: Generating Longer Stories With Recursive Reprompting and Revision

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T06:01:23.281033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:01:23.281033Z digest=sha256:ca928e052eb65a4b497ddf1fd2cb6046ce1de5a2178c1b486ef85c13e4c25834

Observation fd7ce6b2-fbbd-4e5a-ac2f-12a7f3a3e7d1 · outbound

This paper cites StoryLLaV A: Enhancing visual storytelling with multi- modal large language models.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? StoryLLaV A: Enhancing visual storytelling with multi- modal large language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:01:23.825998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:01:23.288909Z digest=sha256:489a0b3739ebd55824f3c5ed899b8c1ff9dceb680d9f40e8e456799e5e1fb2ce

Observation 5fc1bee9-fdf8-46e6-a486-4cf1310fd674 · outbound

This paper cites Knowledgeable storyteller: A commonsense-driven generative model for visual story- telling.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Knowledgeable storyteller: A commonsense-driven generative model for visual story- telling

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:01:23.814059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:01:23.292765Z digest=sha256:68eab466f8781a826b11147b6aa69216ad5846bb0cc943b807f94d6c2f83e9ac

Observation b3baa85f-5e35-4e1f-abb4-1468ea767f85 · outbound

This paper cites Plan-and-write: Towards bet- ter automatic storytelling.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Plan-and-write: Towards bet- ter automatic storytelling

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:01:23.801754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:01:23.297024Z digest=sha256:72dc012b455ba68bdb5a08f28c66e25a3a99bb35cfb5cf76eb29478f14554bc6

Observation b17d90b2-ac7a-4d08-b6d0-ce6b3971cc2a · outbound

This paper cites Minigpt- 5: Interleaved vision-and-language generation via genera tive vokens.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Minigpt- 5: Interleaved vision-and-language generation via genera tive vokens

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T06:01:23.301577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:01:23.301577Z digest=sha256:491b8a107e559a6db1fb73262c8469c1275d3f9dba1ff12325e29a8756588307

Observation 58f04644-0dcd-40fe-9b6b-61c53672ccd9 · outbound

This paper cites Towards a Unified Multi-Dimensional Evaluator for Text Generation.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Towards a Unified Multi-Dimensional Evaluator for Text Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T06:01:23.305677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:01:23.305677Z digest=sha256:88ebf04522ef15ccbc6f57bcea7d63247d0a4ad3b3310d535f3819d9faf50803

Observation eb85e769-c408-448e-9b4c-7b16dc0142fd · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T06:01:23.309985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:01:23.309985Z digest=sha256:66c0266aaa582caac2e7b9c2f2cd1944697672839995c6b82986a5d321d6014d

Observation 27241bec-5521-4c9d-be65-0144dd6df5c0 · outbound

This paper cites an unresolved cited work.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-16T06:01:23.789566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:01:23.314440Z digest=sha256:7617ce1ac57a0bb04445abba9cd5cfa0395e62c7a1ee7824f2e5d5d6bb376181

Observation e06c6b87-40b2-4378-9430-c2fe299c4093 · outbound

This paper cites an unresolved cited work.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-16T06:01:23.777656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:01:23.318477Z digest=sha256:b4b7d64a9f69cff7236e7d5d7c023fe30b8cae3ce324c34f376f098a3a4d02d0

Observation 4a8981f1-f0d1-4987-8332-aa1b8bcf7d92 · outbound

This paper cites ollie.” The phrase “nailed it!.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? ollie.” The phrase “nailed it!

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:01:23.766114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:01:23.322954Z digest=sha256:506b209de50b96d948b9ab012435c2a670dcd3e417855a0226874b7bac1e288e

Observation 62ff5fcf-ab47-4b20-a9ca-9d090a25b9be · outbound

This paper cites The human story pro- vides a rich description of the family’s shared experience, including interactions and conversation, accurately capt ur- ing the sense of warmth and connection.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? The human story pro- vides a rich description of the family’s shared experience, including interactions and conversation, accurately capt ur- ing the sense of warmth and connection

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:01:23.754291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:01:23.327982Z digest=sha256:509d499e4983cb5acb32b52c527311aa7f6baad92622b5bc06486a9ace65f60f

Observation 0d924b66-180b-479d-bd4c-8587c8a8185e · outbound

This paper cites the daughter wouldn’t leave them alone,.

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? the daughter wouldn’t leave them alone,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:01:23.742235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T06:01:23.332049Z digest=sha256:62dc668dfc1a83e29ba243ee5e3ecd0d16a30097fb04b404b3f1863df872c08d

Pith citing papers

Observation f80cf10c-37b6-49f8-bbb3-814a4ef833e4 · inbound

AugmentGest: Can Random Data Cropping Augmentation Boost Gesture Recognition Performance? cites this paper.

AugmentGest: Can Random Data Cropping Augmentation Boost Gesture Recognition Performance? VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:44:02.721295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:44:02.721295Z digest=sha256:b88eb2c4b4849df93fa5bd59f75a52fe7d14d690f56130347cb5c97599292545

Observation 23c4dfaf-a800-4f42-b141-baab6dfd2d6b · inbound

Preparation of Fractal-Inspired Computational Architectures for Advanced Large Language Model Analysis cites this paper.

Preparation of Fractal-Inspired Computational Architectures for Advanced Large Language Model Analysis VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:25:28.649236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-17T23:22:59.323897Z digest=sha256:35d773708ea4b6dcdf8f70d49ca657963a8e27f011b03cd7fa6fa28f4dbdb744

Observation 279727a7-8e76-4b32-9465-35ef91b13615 · inbound

Preparation of Fractal-Inspired Computational Architectures for Advanced Large Language Model Analysis cites this paper.

Preparation of Fractal-Inspired Computational Architectures for Advanced Large Language Model Analysis VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:25:31.127664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-21T19:25:25.465712Z digest=sha256:d41524c6ab8ee5ea036e0da215f3307af1e56ca2ed8968748b8a954d39a2b577

Observation a6664b61-924b-4acf-a077-4eb2a6544ca4 · inbound

Enhancing LLM-Based Neural Network Generation: Few-Shot Prompting and Efficient Validation for Automated Architecture Design cites this paper.

Enhancing LLM-Based Neural Network Generation: Few-Shot Prompting and Efficient Validation for Automated Architecture Design VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:11:11.778912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T19:08:18.706760Z digest=sha256:81062d87994e4cd5ad75b824be1081a81b421e3b9aa45aa5e34d731a2eed2152

Observation 44ff5b58-02d1-4003-b081-9728fe09c8bd · inbound

Closed-Loop LLM Discovery of Non-Standard Channel Priors in Vision Models cites this paper.

Closed-Loop LLM Discovery of Non-Standard Channel Priors in Vision Models VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:38:00.527331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T14:34:24.849584Z digest=sha256:720dd1b0ce5a3d7b0ecc5fe560c26f3501db6e528d4345b797767b204d8955f3

Observation e7106e4e-214a-4f3d-83bc-d6882f0a6488 · inbound

Delta-Based Neural Architecture Search: LLM Fine-Tuning via Code Diffs cites this paper.

Delta-Based Neural Architecture Search: LLM Fine-Tuning via Code Diffs VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:11:06.529683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T16:28:40.262681Z digest=sha256:ce093b8d290dca5e26563ce0f83e831c348e5dc88b0f55fd118e814edce38daf

Observation b3594f60-72c7-43c8-b0fc-ea4a8a6cafbe · inbound

Controllable Narrative Rendering for Enhanced Assisted Writing cites this paper.

Controllable Narrative Rendering for Enhanced Assisted Writing VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:28.453392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-03T00:15:45.193859Z digest=sha256:05619f35434ec4c30edf5963d7b0ead3f8696ba430a169647c2e2d75d2f02394

Observation 7eaa31c1-b8bb-450a-9d45-ba4591770b32 · inbound

LEMUR 2: Unlocking Neural Network Diversity for AI cites this paper.

LEMUR 2: Unlocking Neural Network Diversity for AI VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:17:33.742935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-10T20:15:59.064352Z digest=sha256:667c5ee2cec7ed89ff83a23e0ba34c0a15a9da602df8fff27d00de8eab6690ad