Pith. sign in

Paper Citation Record · LEDGER

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation

As of 20 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 1 inbound Pith citation observation for arXiv:2504.14011.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.14011 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:02:00.476424Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:35:56.638407Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T23:35:56.901371Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact1
  • verified fuzzy45
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6358db80-c949-4c9f-9fef-971d10b46a5c · outbound

This paper cites Generative adversarial nets,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Generative adversarial nets,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.695392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:01:59.835477Z digest=sha256:07a8e8b9773acfeb54e19392f36208a1cc6244f3d8bc9bc76dc280c8451c345b

Observation 54e12cc5-7824-4ff5-8ebc-739b4c4e5a2f · outbound

This paper cites VITON: An Image- Based Virtual Try-On Network,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation VITON: An Image- Based Virtual Try-On Network,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.669688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:01:59.863259Z digest=sha256:203524890624311cbb21f20d01e042a54c2d5c8d04d0883db7b39b4ad7194bda

Observation 310887ff-a728-417c-926b-b7fbf62355a2 · outbound

This paper cites VITON-HD: High-Resolution Virtual Try-On via Misalignment-Aware Normalization,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation VITON-HD: High-Resolution Virtual Try-On via Misalignment-Aware Normalization,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.625200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:01:59.893015Z digest=sha256:62fa4cd07ba43cedc4e49799f864e1bee32e0a465f906a4f763a45a9689ae43f

Observation 898a85d7-061c-497b-abb8-2ebdec373e2d · outbound

This paper cites VITON- GT: An Image-based Virtual Try-On Model with Geometric Transfor- mations,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation VITON- GT: An Image-based Virtual Try-On Model with Geometric Transfor- mations,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.610530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:01:59.898231Z digest=sha256:688a17671b806faed7acf0e570bcc2cce1635073f4c80ed376f2ae75520697e4

Observation ad1a6c13-be98-4c0d-9e06-939a1189e432 · outbound

This paper cites Dress Code: High-Resolution Multi-Category Virtual Try-On,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Dress Code: High-Resolution Multi-Category Virtual Try-On,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.594784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:01:59.903569Z digest=sha256:d1680d720bb6602680ecb697536be613aee96d88043152d3537c2e4d46c897ec

Observation 2a6b6daf-b74b-41f6-a696-d44a90a805c9 · outbound

This paper cites Deep Unsupervised Learning using Nonequilibrium Thermodynamics,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Deep Unsupervised Learning using Nonequilibrium Thermodynamics,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.579480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:01:59.908523Z digest=sha256:433ccd555df40e951e623283d291cc0ef8fc3bf59047d08dbca2b04b4ee47693

Observation 96f3a5c4-ba04-46a1-b9ff-3dabe488fe48 · outbound

This paper cites Denoising Diffusion Probabilistic Models,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Denoising Diffusion Probabilistic Models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.564828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:01:59.913678Z digest=sha256:2765c9e26a5680c98639b36e3bb655a6a8e7e7bf5be0eabc747aa38036fa03e5

Observation 8e65ccd1-4c0b-4cfb-94a7-38046502fc18 · outbound

This paper cites High- Resolution Image Synthesis With Latent Diffusion Models,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation High- Resolution Image Synthesis With Latent Diffusion Models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.549706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:01:59.923988Z digest=sha256:3300ebb1ba59eb7f764f84386a699dd9e607e2a285e3f5d2a7456789f1064f34

Observation 802c9a52-fe79-41b2-aea1-7a52d553b849 · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Scaling Rectified Flow Transformers for High-Resolution Image Synthesis,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.534136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:01:59.961768Z digest=sha256:d8d77119bcd7a43b2961cad89ede911388aa4da8ddbb78b4c75da03d5d723ff3

Observation 21ca5d6f-95f6-4df5-b016-a9c96512e10c · outbound

This paper cites LaDI-VTON: Latent Diffusion Textual-Inversion Enhanced Vir- tual Try-On,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation LaDI-VTON: Latent Diffusion Textual-Inversion Enhanced Vir- tual Try-On,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.518828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:01:59.989358Z digest=sha256:4f9c0fb6a0b4ea8a5192d1b4f7a3c3df110a6f9f9cfde574cffed618c7e685cb

Observation 9998c28e-4ba0-4d05-931e-dfda49e1123f · outbound

This paper cites StableVITON: Learning Semantic Correspondence with Latent Diffusion Model for Virtual Try- On,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation StableVITON: Learning Semantic Correspondence with Latent Diffusion Model for Virtual Try- On,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.503781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:01:59.994262Z digest=sha256:24ee4d3f80b175aebf53ec219438f8407775864c950986e7d82d197360d12b34

Observation b1183fd8-0c81-4ec3-8bd0-b22e19c10b79 · outbound

This paper cites StableGarment: Garment-Centric Generation via Stable Diffusion.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation StableGarment: Garment-Centric Generation via Stable Diffusion

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T12:01:59.999194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:01:59.999194Z digest=sha256:c78d3d1a3577b69a4972598eb6854d09e40e8014aecc3010fc7bc96cb7933cf7

Observation e607e973-c717-4960-9869-fab0a254e56c · outbound

This paper cites Improving Diffusion Models for Authentic Virtual Try-on in the Wild,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Improving Diffusion Models for Authentic Virtual Try-on in the Wild,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.489635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.005587Z digest=sha256:0796dd2872acea36c812ca18593e519db4d89fb756affe04cc1e7fd2b71a0db2

Observation 21d2b234-b596-4153-bf72-d5b0690d5734 · outbound

This paper cites Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image Editing,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image Editing,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.474954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.010925Z digest=sha256:43961fc92b2c81f265c9690552743dca55e8e0f83554fa01bc6e7dda901aa00d

Observation b9d35268-193a-45aa-ad78-514736ab3815 · outbound

This paper cites Multimodal-Conditioned Latent Diffusion Models for Fashion Image Editing.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Multimodal-Conditioned Latent Diffusion Models for Fashion Image Editing

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-16T12:02:00.632201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.016480Z digest=sha256:b6cb162e63e5f9798fdf8984ec3ba2f067f468d7bba7892cbc3683f7ea5904a8

Observation 19a71e40-02d8-4b1d-a17d-630daa9decdf · outbound

This paper cites TexFit: Text-Driven Fashion Image Editing with Diffusion Models,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation TexFit: Text-Driven Fashion Image Editing with Diffusion Models,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.434562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.022430Z digest=sha256:c1312a4bb43cdf961f43a066b0bcf2ef3fac3c7423eb328ffd7fd87bfa72373f

Observation 1e7b2a24-6778-4550-a0b6-653e1060fd1c · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T12:02:00.028238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:02:00.028238Z digest=sha256:4479a3ec05ff2a9ef5886584b4fd082bf6504079767e9e330db2dc82cc10460c

Observation 30e6f005-7ab9-4828-bf9b-d2d0ca860e8a · outbound

This paper cites Self-RAG: Learning to retrieve, generate, and critique through self-reflection,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Self-RAG: Learning to retrieve, generate, and critique through self-reflection,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.390143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.036897Z digest=sha256:a0f35b2ff63511f6ce665cfd19c92620232b75cf1060071219d391696c8ee642

Observation c409e811-f24c-4d05-a6bb-8ba259550739 · outbound

This paper cites RAVEN: Multitask Retrieval Augmented Vision-Language Learning.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation RAVEN: Multitask Retrieval Augmented Vision-Language Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T12:02:00.083810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:02:00.083810Z digest=sha256:536893d1b2a010001f684323f464f2a358db6f58ab791c3dd97a1fc34f2d8a85

Observation 3cf3c6a3-1bca-4dc6-8381-63ff05223254 · outbound

This paper cites An Image is Worth One Word: Personalizing Text- to-Image Generation using Textual Inversion,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation An Image is Worth One Word: Personalizing Text- to-Image Generation using Textual Inversion,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.367512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.109969Z digest=sha256:3f91b58d32d354056bf0074949f220b9c968a1b971db6ca2ef02ab54c1cef7e5

Observation 074f9807-20f7-4a99-a96e-2c49a96d1918 · outbound

This paper cites Zero-Shot Composed Image Retrieval with Textual Inversion,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Zero-Shot Composed Image Retrieval with Textual Inversion,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.352195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.115161Z digest=sha256:299fd84a7b08296733303233249024a61900c3aa317dc0e191b0f362aa8bd4a9

Observation 4eeaf0a1-eaae-402c-9d62-173e7db7b071 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Learning Transferable Visual Models From Natural Language Supervision,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.336875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.120421Z digest=sha256:0a41e8dd0dc1dc816a9598c7c440f138bbebebc018246dde4c1682f44dd2a1d8

Observation 3fa8aa8d-34fe-4b57-8ec3-e518583d0eb2 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T12:02:00.125021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:02:00.125021Z digest=sha256:2a90a99e686e9fb97f80fcef2db4d18a0c27386198b3d40c7aac4c49a8b441e8

Observation 2cdde2b3-2a22-465a-8283-c758025fda52 · outbound

This paper cites Diffusion Models Beat GANs on Image Synthesis,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Diffusion Models Beat GANs on Image Synthesis,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.321042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.130274Z digest=sha256:06e32c5c105960b87dcc9c7fbd1a6b21595c1e99129343b97a2038151ec49ce9

Observation f9908f3d-4ce2-4cc3-afcc-128685f441b3 · outbound

This paper cites Taming the Power of Diffusion Models for High-Quality Virtual Try-On with Appearance Flow,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Taming the Power of Diffusion Models for High-Quality Virtual Try-On with Appearance Flow,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.305788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.135051Z digest=sha256:ab1a53ca1e9257ce77ee265ad5c54c1a13b481faac63fc69061c2f619e379f28

Observation b7facbc8-fac3-4f19-add0-2a6dbfa8209b · outbound

This paper cites Size Does Matter: Size-aware Virtual Try-on via Clothing-oriented Transformation Try-on Network,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Size Does Matter: Size-aware Virtual Try-on via Clothing-oriented Transformation Try-on Network,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.290133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.139665Z digest=sha256:f53c51a5cdaa3921395e72faf920f77ef4bc6aab2b4d7785b4e2aa40f7492f32

Observation 2c6a5a4d-abd5-4802-9446-593ba3009471 · outbound

This paper cites TryOnDiffusion: A Tale of Two UNets,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation TryOnDiffusion: A Tale of Two UNets,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.274479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.144543Z digest=sha256:ac34a0f2f08123607d70a3e6a2b93218974deb8d14b357c9e6f8ac606f30de53

Observation e7d871c4-1ca2-4019-b1d5-3dcdb256b5a8 · outbound

This paper cites CatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion Models,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation CatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion Models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.259559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.149190Z digest=sha256:eab26146647b4437ff1425a35d9d67eeb82bbc8be479ce12867d8cdac4e29202

Observation 91942acb-7201-46b3-8e39-1ea26869eb4a · outbound

This paper cites U-Net: Convolutional Net- works for Biomedical Image Segmentation,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation U-Net: Convolutional Net- works for Biomedical Image Segmentation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.244414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.173553Z digest=sha256:6b02e70930ebbacf4cc91de39579b3ada088f80a7d9c2923fcb77e3001a7532b

Observation 4730167d-d512-4da7-8d67-8eacd5dca136 · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T12:02:00.187479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:02:00.187479Z digest=sha256:cc631dba4821c21d5e9f08af8d347c221260ab8297de3f8c297e6c97621fc9d4

Observation a1b802bc-0380-4518-bb8f-d34e60fecaec · outbound

This paper cites UPRISE: Universal Prompt Retrieval for Improving Zero-Shot Evaluation,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation UPRISE: Universal Prompt Retrieval for Improving Zero-Shot Evaluation,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.225695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.221721Z digest=sha256:70d1affb83f9601d1e9c08a491ef25957b6148ee0cb275197b808efa721c8a2f

Observation 9718fb06-6477-4e66-a335-c71aa2098dd7 · outbound

This paper cites REPLUG: Retrieval-Augmented Black-Box Language Models.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation REPLUG: Retrieval-Augmented Black-Box Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T12:02:00.226471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:02:00.226471Z digest=sha256:80ae87a865b1cca595a3f98e9d5e8c4d819a95d961ac86977f7eea3f468c344d

Observation a64a5777-cf3c-4333-98f7-a9a8e60e975f · outbound

This paper cites Generate rather than Retrieve: Large Language Models are Strong Context Generators,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Generate rather than Retrieve: Large Language Models are Strong Context Generators,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.208672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.231742Z digest=sha256:1506dfdd221ff67f64e131389d00e0314d49ca27fa9c76a5ddeb22f8566259e0

Observation 23eac7b3-24c5-444a-9f77-26a001819641 · outbound

This paper cites RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T12:02:00.236299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:02:00.236299Z digest=sha256:b8ec8cd1075a640a7e9681061fcc055f8a569b6e09fa1e62702b747f70d9c3a9

Observation 1179b5c5-c751-4437-9c1f-1ccb6354652d · outbound

This paper cites In-Context Retrieval-Augmented Language Models,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation In-Context Retrieval-Augmented Language Models,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.190648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.241334Z digest=sha256:f84b22709e572876d0b6494ae188914e7e076e2d4d43fd9ca53673161a79ae65

Observation 0d484499-f0be-4efe-866e-8be41dce0568 · outbound

This paper cites To- wards Retrieval-Augmented Architectures for Image Captioning,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation To- wards Retrieval-Augmented Architectures for Image Captioning,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.174541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.246090Z digest=sha256:acbecd96b7361f4b9db6db4ce9ba923c0c77bd5b589680cb09237f73fa9ea780

Observation 5237c635-1fca-4ca4-9344-0de2126e8416 · outbound

This paper cites SmallCap: Lightweight Image Captioning Prompted With Retrieval Augmentation,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation SmallCap: Lightweight Image Captioning Prompted With Retrieval Augmentation,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.156947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.250941Z digest=sha256:f1f03c3a91f5c385f6307764da1893cb9b8b463f643577f0ab3e545d04e13e1b

Observation abeb0310-5682-49bf-a1d1-fdbbfdc54251 · outbound

This paper cites With a Little Help from your own Past: Prototypical Memory Networks for Image Captioning,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation With a Little Help from your own Past: Prototypical Memory Networks for Image Captioning,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.139520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.255722Z digest=sha256:476a5bd46bf85ddb2d96df5b71fdcdda290d7998e51df766390f95034f16f551

Observation 263561da-78b8-4137-9880-84addeb7a4dc · outbound

This paper cites REVEAL: Retrieval-Augmented Visual- Language Pre-Training With Multi-Source Multimodal Knowledge Memory,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation REVEAL: Retrieval-Augmented Visual- Language Pre-Training With Multi-Source Multimodal Knowledge Memory,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.121844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.261125Z digest=sha256:29c5d6b43cfdd2bcd7f2369806fda7df822bd09c00871a3d5aaeb4e4b59171b7

Observation ce8c26df-9c3f-477a-8fb9-c28c88305e56 · outbound

This paper cites Aug- menting Multimodal LLMs with Self-Reflective Tokens for Knowledge- based Visual Question Answering,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Aug- menting Multimodal LLMs with Self-Reflective Tokens for Knowledge- based Visual Question Answering,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.103386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.265807Z digest=sha256:0a97a52c0c35d01843dd5e912c2f074348e030e5fd075eba71c7eeae80702ba9

Observation c2f7266f-7647-4efb-8568-7e8b19c5d82c · outbound

This paper cites Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document Retrieval,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document Retrieval,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.086763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.312977Z digest=sha256:1ad789871ba4f2023b6600c7e68f74fb2f16ac897673b9ed319a5b83948c1502

Observation 52931a46-e0d0-4297-a4db-9e8ecfe5ff89 · outbound

This paper cites Re-Imagen: Retrieval-Augmented Text-to-Image Generator.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Re-Imagen: Retrieval-Augmented Text-to-Image Generator

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T12:02:00.351960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:02:00.351960Z digest=sha256:f795dc397b7c7fed02f8e27f8c6299484f71a509709e7d944b5c76bc909ee38b

Observation 1349a857-82bf-4c76-b535-c908643293e6 · outbound

This paper cites Retrieval-Augmented Diffusion Models,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Retrieval-Augmented Diffusion Models,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.067630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.362372Z digest=sha256:43a033011e5d31f60dcfedfe695dfa7ad6677e0b7bb932788a4ae2bfd7b74c0c

Observation c4cd907a-1766-484a-865b-c7b195480a27 · outbound

This paper cites Retrieval-Augmented Multi- modal Language Modeling,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Retrieval-Augmented Multi- modal Language Modeling,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.050418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.367791Z digest=sha256:6bf6eb64dfd0189cf7589e5ee0d2ba4fae72094af52f40593f7fd7a05070ba23

Observation 57b5a98a-dd79-4660-97c2-70dbcc4a4474 · outbound

This paper cites Multi- Concept Customization of Text-to-Image Diffusion,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Multi- Concept Customization of Text-to-Image Diffusion,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:01.032395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.372667Z digest=sha256:e769a8d099506d5dfb1fbee104b3768fde073dd7059f7167799bec5a6936c3be

Observation ce70fd5b-80da-4781-8105-4ee34322cd3c · outbound

This paper cites Style Aligned Image Generation via Shared Attention,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Style Aligned Image Generation via Shared Attention,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:00.998408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.377703Z digest=sha256:c5e8483e7b9cb8683e197482c89ecfb082130b7c3496d9495c9dff9b96d57a72

Observation e838c5d9-28fa-46fb-914c-2fb238f9ef4c · outbound

This paper cites Cross-Image Attention for Zero-Shot Appearance Transfer,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Cross-Image Attention for Zero-Shot Appearance Transfer,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:00.981433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.382608Z digest=sha256:543012b42cff992b6a358734a56f63df212b33222a510cb5997514080e16bc4d

Observation 60b12925-520a-41a1-b8ea-c4f2788a393b · outbound

This paper cites Adding Conditional Control to Text-to- Image Diffusion Models,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Adding Conditional Control to Text-to- Image Diffusion Models,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:00.944692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.387802Z digest=sha256:f5b580bc7c00fc0de19015800fcf9e9848c743db909207b0cf325efa7c8bdaee

Observation 624b32a6-9386-451f-bf10-f456c85a9879 · outbound

This paper cites Reproducible Scaling Laws for Contrastive Language-Image Learning,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Reproducible Scaling Laws for Contrastive Language-Image Learning,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:00.895056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.392250Z digest=sha256:cc78916f5b5e803d20e07dee39566c13ae7fe4f05f6b74d9b07fe891c07744be

Observation afe231f6-ca34-40d6-854e-f0637d9ed73f · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T12:02:00.397756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:02:00.397756Z digest=sha256:064939889d0ce28b52c947ac10974e37462502c249d4b9898b96e1d06547581c

Observation a2dc77d1-09e7-4742-af15-4b9f9f048ec5 · outbound

This paper cites SpaText: Spatio-Textual Repre- sentation for Controllable Image Generation,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation SpaText: Spatio-Textual Repre- sentation for Controllable Image Generation,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:00.869184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.438357Z digest=sha256:c19728670e04cfdbbb2b420a3454fa054bc3bf9cf86523ab553d3d23e568a9cd

Observation aa4bf8f6-67f9-4d0c-bb92-7d65205232e2 · outbound

This paper cites Decoupled Weight Decay Regularization,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Decoupled Weight Decay Regularization,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T12:02:00.459044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:02:00.459044Z digest=sha256:a364d9b66d0c6d81cdbf3e9a6dcba068b493ded4da4a8364083637a1e45d20bc

Observation 1e2b5bb1-4976-4359-91d3-e191d8597de1 · outbound

This paper cites The Unreasonable Effectiveness of Deep Features as a Perceptual Metric,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation The Unreasonable Effectiveness of Deep Features as a Perceptual Metric,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:00.810317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.463702Z digest=sha256:cf44cac2670505b0d270e68c93178984c0434a73744dc5bbb32f923cd0ed706c

Observation 79b8a6cc-cba4-4fff-9d12-3f894c5f3eca · outbound

This paper cites Image quality assessment: from error visibility to structural similarity,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Image quality assessment: from error visibility to structural similarity,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:00.745870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.467898Z digest=sha256:4b6ec54dcaf65e31c071bca0d5efa528587b876627aa6b6c10119573adf0a746

Observation 6b2d1085-4011-42a1-bfff-2991a6d90a2b · outbound

This paper cites GANs trained by a two time-scale update rule converge to a Nash equilibrium,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation GANs trained by a two time-scale update rule converge to a Nash equilibrium,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:00.730081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.472149Z digest=sha256:04258bf21988f21e065487553c4dd4e1233a84fb0cf9c5c914702ea8f284aa1f

Observation f9784e97-8bc0-48dc-92ad-5dd58a75b6f4 · outbound

This paper cites Demysti- fying MMD GANs,.

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation Demysti- fying MMD GANs,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:02:00.713365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:02:00.476424Z digest=sha256:1970abc6e0a56bb61ea2bd053090135b5d1a22d63cb1ed8d8b8f771227798ef1

Pith citing papers

Observation d86c8ca0-ce33-4a21-9e16-3b49d0ddde45 · inbound

Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation cites this paper.

Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T23:35:56.906988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T23:35:56.638407Z digest=sha256:3bb2072d996ddcdceed26ea8ed99b3d94ef8746af69f60c35a8ba282debe4a7e