Pith. sign in

Paper Citation Record · LEDGER

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents

As of 17 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 2 inbound Pith citation observations for arXiv:2507.17462.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.17462 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:53:38.496401Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T00:17:39.409452Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T00:25:09.923270Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ca5a907-6541-4edd-be98-86e26c94919c · outbound

This paper cites CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:31.294528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:31.294528Z digest=sha256:4c19ed350770b8f0a99caa4253492ffe08e86937a7aeea3af53fcd496262b152

Observation f0a1a531-f8d7-473f-a58b-97a808b6eff5 · outbound

This paper cites Scaling Robot Learning with Semantically Imagined Experience,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Scaling Robot Learning with Semantically Imagined Experience,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:46.735035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:53:31.361436Z digest=sha256:0ff626bf49aa81e34207606930e295053c59adf920af27728e4ea21a5d21eff6

Observation 8baeeaea-cb0f-4687-bc98-7a2727894b97 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:31.519794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:31.519794Z digest=sha256:b0ca84b2932fb95b6224699e932b4df28de0892584620efd0d46c19cf75d23ff

Observation 4be03cb1-ac15-4dd4-bb8c-6a463f158267 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents OpenVLA: An Open-Source Vision-Language-Action Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:31.611898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:31.611898Z digest=sha256:d0fb6a734b43f4a5a9a7dd208b1a9bd2170045d2dba586c8e248b41cb750835a

Observation 04a71937-6bfb-4ee6-a0e0-869bd17dd796 · outbound

This paper cites MagicDrive: Street view generation with diverse 3d geometry control,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents MagicDrive: Street view generation with diverse 3d geometry control,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:46.492290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:53:31.715847Z digest=sha256:df409125cf524e8207f9aab4945e41af89db2a83e0ec9aacdf93576e914a086b

Observation ad4d3104-934c-424c-a9f3-e97ebffce656 · outbound

This paper cites BEVControl: Accurately Controlling Street-view Elements with Multi-perspective Consistency via BEV Sketch Layout.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents BEVControl: Accurately Controlling Street-view Elements with Multi-perspective Consistency via BEV Sketch Layout

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:31.818402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:31.818402Z digest=sha256:d3055a1dc1c6823e5ffed62463899238cf50cb32b0c438e782a8d7a260a4530f

Observation 35592f22-efe7-47ff-92fd-c461a4bc71ed · outbound

This paper cites Street-view image generation from a bird’s-eye view layout,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Street-view image generation from a bird’s-eye view layout,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:46.224327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:53:31.939060Z digest=sha256:baef7e7fc19b32e20d87a767a264a9beda4e1847d5824bb57ba0f23c35653e47

Observation 5ab518d0-e587-4160-9616-96361742f5ca · outbound

This paper cites Dragvideo: Interactive drag-style video editing,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Dragvideo: Interactive drag-style video editing,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:45.962585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:53:32.047960Z digest=sha256:e72950471b2bb84d957c6cff82371c7784d847003e02ec0f810cbf7f4c140846

Observation 26ed3f74-1906-4583-8acb-be8ce7bf7b73 · outbound

This paper cites Video-p2p: Video editing with cross-attention control,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Video-p2p: Video editing with cross-attention control,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:45.700618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:53:32.210605Z digest=sha256:8b25c97a6d8c90ce61708282acedf509ff5d6198581cb13c5c99bd61a6a41982

Observation 6c75b870-d27c-432c-bd5f-a7f2bfce9af3 · outbound

This paper cites Visual commonsense-aware representation network for video captioning,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Visual commonsense-aware representation network for video captioning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:45.383791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:53:32.355751Z digest=sha256:58ff5d0c641f7d5c4960e88e84dd4f2c3d546e79c5a130e34129e2d4472cb5f8

Observation a17b54cc-3fe0-44e2-bea7-a313f0067b13 · outbound

This paper cites Diffusion model-based image editing: A survey,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Diffusion model-based image editing: A survey,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:45.135486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:53:32.484985Z digest=sha256:fece3de38efc441d2b491007ed0e904c953d724ff5837af54f863fdbc1603b5d

Observation 858904f2-5e64-41b3-bb0c-63489df87bf6 · outbound

This paper cites Feditnet++: Few- shot editing of latent semantics in gan spaces with correlated attribute disentanglement,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Feditnet++: Few- shot editing of latent semantics in gan spaces with correlated attribute disentanglement,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:44.811434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:53:32.605859Z digest=sha256:aa9852c9618c6685caadad2cd6a4d595b30afb7c4d5f8b2751b73a119b03aab2

Observation de2cc8c5-818d-4812-82d7-c105f1d3360d · outbound

This paper cites Gaussctrl: Multi-view consistent text-driven 3d gaussian splatting edit- ing,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Gaussctrl: Multi-view consistent text-driven 3d gaussian splatting edit- ing,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:32.743292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:32.743292Z digest=sha256:f48b726f221418974cd9e16eeaa57a1f477ab6505ac5a9fa95e4dffc55bcbefa

Observation f33ea19d-8b14-464f-98c4-669954140c4f · outbound

This paper cites Efficient dynamic scene editing via 4d gaussian-based static-dynamic separation,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Efficient dynamic scene editing via 4d gaussian-based static-dynamic separation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:44.429108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:53:32.872349Z digest=sha256:5de1adb244f047f7dc42aa2eccbf6b4cd0896e5685970d01fd81b8f5f5ce011a

Observation 240dac32-42fa-4a83-95bf-d3b4d47eab96 · outbound

This paper cites Generating long videos of dynamic scenes,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Generating long videos of dynamic scenes,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:44.173267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:53:33.031861Z digest=sha256:95555f12eab9ab12c1114126a84d8ebdaf4c5a937c19d5db090efa6c509ae25c

Observation b2344980-f9e9-4149-aac0-0f5ced496595 · outbound

This paper cites Storydiffusion: Consistent self-attention for long-range image and video generation,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Storydiffusion: Consistent self-attention for long-range image and video generation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:43.868743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:53:33.186072Z digest=sha256:0dec969a9b6e828e934ef0b83dc45cc684d71e569870a92f43506b7df216040b

Observation fef81db5-5e53-49eb-85f1-025222e676f7 · outbound

This paper cites Evalcrafter: Benchmarking and evaluating large video generation models,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Evalcrafter: Benchmarking and evaluating large video generation models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:43.588423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:53:33.288535Z digest=sha256:873edbafc58a9b979ba4f3ce34c6d61e14f9283dcb7a0fdaea4e3a50769724be

Observation d6105e95-80cd-4038-9ef7-ce23d0e3fe6c · outbound

This paper cites Towards long video understanding via fine-detailed video story generation,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Towards long video understanding via fine-detailed video story generation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:43.394123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:53:33.414870Z digest=sha256:37395d8c60296d7842adfe99d4cf38351ea7713d84f45a6deb2df3c611df50db

Observation b30c98a2-6965-4090-9e4b-9012ea9e36a0 · outbound

This paper cites Maskgwm: A generalizable driving world model with video mask reconstruction,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Maskgwm: A generalizable driving world model with video mask reconstruction,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:43.094305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:53:33.564446Z digest=sha256:ea9b6e5a4d608b72d43fe173de5aef0dca6402b2cb5d179baebaed664655633a

Observation 537a1492-6b2a-408f-880c-16d2b2a1bf8c · outbound

This paper cites A Survey on Long Video Generation: Challenges, Methods, and Prospects.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents A Survey on Long Video Generation: Challenges, Methods, and Prospects

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:33.702516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:33.702516Z digest=sha256:666f541f2dcc5d320323f780b581b829ae78bd1615d6a6a5bce09b8d00a1181d

Observation 64896346-6477-4c74-a9cf-a30b6a39df53 · outbound

This paper cites Dall-e-bot: Introducing web- scale diffusion models to robotics,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Dall-e-bot: Introducing web- scale diffusion models to robotics,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:42.734812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:53:33.900587Z digest=sha256:5750c7abeb21935a9bd133f39f52c4eb002d655a24c4b945816c9126f4787645

Observation ec05318f-4a89-4cac-a4b0-b51e7bc9083a · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:34.065833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:34.065833Z digest=sha256:eb6a8290384f3b6d1b0ef6a453cab49afd9d4d5c84c46384ac62ef2b6017b259

Observation 7fe0b49a-3265-49fe-a6aa-54abf355c307 · outbound

This paper cites GenAug: Retargeting behaviors to unseen situations via Generative Augmentation.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents GenAug: Retargeting behaviors to unseen situations via Generative Augmentation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:34.194140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:34.194140Z digest=sha256:8f9f96d6890886f0b5bbe8793aeec8c63aa9b773bb37f24ea2c0017b91ff490f

Observation e7569c55-ebc9-4a05-9667-47468ebe9be1 · outbound

This paper cites Semantically controllable augmentations for generalizable robot learning,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Semantically controllable augmentations for generalizable robot learning,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:34.333520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:34.333520Z digest=sha256:f66250136fd591debdf6573c617c80ecbadb6353f39ff8f0133e5079b5fdbfcb

Observation 3f9cb659-3d2d-4a89-9357-d4a00e51e36a · outbound

This paper cites Roboagent: Generalization and efficiency in robot manipu- lation via semantic augmentations and action chunking,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Roboagent: Generalization and efficiency in robot manipu- lation via semantic augmentations and action chunking,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:42.481792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:53:34.459950Z digest=sha256:c841808131ef348bfc5338b7e44cb5da1de5aa60916f68ebc7754009a8c8ccf4

Observation f48df0bf-9d3b-4967-bb78-03bd870c2454 · outbound

This paper cites Segment anything,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Segment anything,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:34.586243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:34.586243Z digest=sha256:bf9612341360b8104e37cd5853aea20f775994f7d1f60140fdda6b9aeb18b433

Observation e8b228d8-3a7f-428f-bc33-44f5c8b892d9 · outbound

This paper cites EnerVerse-AC: Envisioning Embodied Environments with Action Condition.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents EnerVerse-AC: Envisioning Embodied Environments with Action Condition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:34.708735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:34.708735Z digest=sha256:b542ad857597813c352e4b7346012092941ec651402e0c58b20b20f4dd9d0076

Observation 27c73efe-2b8f-4912-8279-17ef1c9a34df · outbound

This paper cites Zero-1-to-3: Zero-shot one image to 3d object,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Zero-1-to-3: Zero-shot one image to 3d object,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:42.247498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:53:34.849354Z digest=sha256:8914193c707abe08c16a22158aa04db033e11f48e638f93cd52a9891e1a57c0b

Observation 0bce7733-0683-4bd4-b8a0-e45affa1d359 · outbound

This paper cites InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:34.980027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:34.980027Z digest=sha256:c4e04e662d0af90d9f093fdec4988b13eb480a37302080e291087ff5ade9aea4

Observation 1c2321a9-0043-42e0-8703-03a03eeb6977 · outbound

This paper cites 3D-Adapter: Geometry-Consistent Multi-View Diffusion for High-Quality 3D Generation.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents 3D-Adapter: Geometry-Consistent Multi-View Diffusion for High-Quality 3D Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:35.168365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:35.168365Z digest=sha256:d33996fe56bbbba48700c5a37e11bf3d7af5f2e4ec6d802029e949fb543e829b

Observation 2d9265b2-33db-4031-a976-f83f19887450 · outbound

This paper cites Dge: Direct gaussian 3d editing by consistent multi-view editing,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Dge: Direct gaussian 3d editing by consistent multi-view editing,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:42.001632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:53:35.339800Z digest=sha256:f818b341367c989e70c446a669566b3e624b1162d3d98dbe890997c286bf9e48

Observation e8c810d4-e30d-4347-ae5c-2f106ca3aa20 · outbound

This paper cites Imfine: 3d inpainting via geometry-guided multi-view refinement,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Imfine: 3d inpainting via geometry-guided multi-view refinement,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:41.681720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:53:35.479138Z digest=sha256:2b776c15f2efeafaccabe7473715ccb1bc0da59553ed01eb79de0319c8ef090d

Observation 822d2eb3-7671-400a-b957-cd6d7ef32a5d · outbound

This paper cites High- resolution image synthesis with latent diffusion models,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents High- resolution image synthesis with latent diffusion models,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:35.647926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:35.647926Z digest=sha256:c24a052236470603edfeda10bd3a8b64c09e442a5ab361d07478dae956e60c75

Observation efaa2062-6f9e-416b-8b88-5f1e572d0c7a · outbound

This paper cites Towards language-driven video inpainting via multimodal large language models,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Towards language-driven video inpainting via multimodal large language models,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:41.373678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:53:35.820351Z digest=sha256:cce09ac98ea1c200f0db990bc58843a33f1c183fb762bb5bff4cc80b04c4101b

Observation 0d35a874-59eb-41ee-8366-83b33dab9d1c · outbound

This paper cites Brush2prompt: Contextual prompt generator for object inpainting,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Brush2prompt: Contextual prompt generator for object inpainting,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:41.065267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:53:35.926503Z digest=sha256:fd06a319cfb049bca9fcae20fc91c30d42d1230c2b7dde8a18c3f5febee63c51

Observation 9e760139-0a9d-4a6d-9ac5-19d643a221d6 · outbound

This paper cites Sketch- guided image inpainting with partial discrete diffusion process,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Sketch- guided image inpainting with partial discrete diffusion process,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:40.819891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:53:36.052743Z digest=sha256:f0e0a73b3f5ca13f2c6651768a0ddab7c8616d3461fe35ea42634ae3e7d5b99c

Observation 337a5171-982a-4350-842d-46761fed50a2 · outbound

This paper cites Few-shot image generation via style adaptation and content preservation,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Few-shot image generation via style adaptation and content preservation,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:40.501753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:53:36.256605Z digest=sha256:c7621d2fa9ba2c2d964b815249eecd82d1072203d3fe63fe11a1dac190dc1f43

Observation 2b5dfbe2-4686-4566-ac14-412938512fcb · outbound

This paper cites Learning transferable visual models from natural language supervision,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Learning transferable visual models from natural language supervision,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:36.417013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:36.417013Z digest=sha256:d8c7198910becb6bcc9c11c13fdf1237ed4dad5f6ec559092141320e869b0f1b

Observation c1f65cef-63aa-4d4d-b954-2bf34b686539 · outbound

This paper cites Video diffusion models,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Video diffusion models,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:36.559772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:36.559772Z digest=sha256:b0b9e944da2a9e236b105e749df44f1b7dd7af6356104f043e355d3b25b5c7cd

Observation e59b2166-dec2-49ad-b81d-6e467a3a7ffb · outbound

This paper cites Enerverse: Envisioning embodied future space for robotics manipulation,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Enerverse: Envisioning embodied future space for robotics manipulation,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:36.744744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:36.744744Z digest=sha256:83634c4bbef558e645ffd6d3808f41aeba58a5db9e6bd2268ff7926cd0b024d9

Observation 02991ebf-d39f-4ce1-b936-b379fb0b1826 · outbound

This paper cites Epidiff: Enhancing multi-view synthesis via localized epipolar-constrained diffusion,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Epidiff: Enhancing multi-view synthesis via localized epipolar-constrained diffusion,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:40.178023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:53:36.907756Z digest=sha256:af773bb7a11f3de7b2fb763e0481eed50702ea076c0be50da87be4c25d860e81

Observation 4e7cb55a-f64f-4b26-b156-4c97ebcb9b65 · outbound

This paper cites Ar-diffusion: Asynchronous video generation with auto-regressive diffusion,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Ar-diffusion: Asynchronous video generation with auto-regressive diffusion,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:39.812003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:53:37.094974Z digest=sha256:0aa806ab6fc5043ef4c2191455761e1d9dda397ce9730691c8acfe6ca5c536aa

Observation 475c8523-758a-435a-a4cc-b3631bffabf6 · outbound

This paper cites Progressive autoregressive video diffusion models,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Progressive autoregressive video diffusion models,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:39.489380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:53:37.295899Z digest=sha256:75dd2118f41d84f81a6962e8eef2c966710ce991e0644dfbc77e2cd8eeabff61

Observation 8a3c1759-7c0a-4c46-8052-9405e6541aa0 · outbound

This paper cites From slow bidirectional to fast autoregressive video diffusion models,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents From slow bidirectional to fast autoregressive video diffusion models,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:37.457637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:37.457637Z digest=sha256:1fd0f9070530cbfe25b02052b6913f7290d0e2f340bae084a3434aa6def1b368

Observation a207e586-b1bb-4667-bdb0-8ad1045082c9 · outbound

This paper cites Qwen2.5-VL Technical Report.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Qwen2.5-VL Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:37.599945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:37.599945Z digest=sha256:7839bcd3a4955df020813d85cf67112ea44fd1e7d6401013fdbcb2ab945c05cf

Observation 72315267-80ee-42a1-b6e5-0598d6458ecc · outbound

This paper cites Step1X-Edit: A Practical Framework for General Image Editing.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Step1X-Edit: A Practical Framework for General Image Editing

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:37.812216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:37.812216Z digest=sha256:a8f573c717a711ce1902997c3d9d20cbc2d21b36b64759430d74b3630f9100ab

Observation 80f9a6bd-903f-46d4-82d4-275ddf05a4f6 · outbound

This paper cites Robotwin: Dual-arm robot benchmark with generative digital twins (early version),.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Robotwin: Dual-arm robot benchmark with generative digital twins (early version),

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:39.242888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:53:37.961048Z digest=sha256:b8fe20cfda3068a2596df4637bb5b2bff19c9cb30b34d61483e9a3b29bf07ead

Observation 4b404dc2-4398-4132-8db1-459fd9fb345d · outbound

This paper cites Diffusion policy: Visuomotor policy learning via ac- tion diffusion,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Diffusion policy: Visuomotor policy learning via ac- tion diffusion,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:38.099075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:38.099075Z digest=sha256:74e43694d8f41904e3a2b37fd3e309583bf2816099dc59474f8973656b558e80

Observation cd782825-b3e2-4a52-a651-73a0e9b715f1 · outbound

This paper cites Schedule your edit: A simple yet effective diffu- sion noise schedule for image editing,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Schedule your edit: A simple yet effective diffu- sion noise schedule for image editing,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:38.933412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:53:38.294649Z digest=sha256:1370fc312dd6018fee1248df07e2b9a7295403992de36db91b6779418c566cff

Observation b0ffbd48-09d1-4a61-98ee-25730bfa6101 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:38.496401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:38.496401Z digest=sha256:7f444965280f104eed8e567923bb05a2595c0759cec531c1dd8964c41050cb90

Pith citing papers

Observation 6fb304ea-b10f-4f1d-8fde-8c04a951d874 · inbound

Embody4D: A Generalist Data Engine for Embodied 4D World Modeling cites this paper.

Embody4D: A Generalist Data Engine for Embodied 4D World Modeling ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:32.152811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T19:16:50.004359Z digest=sha256:a25bd564b4050541f3f919ca7b1092b3ba364c7cb5ee8dece9ead1b5f44f9a03

Observation 23da8182-92d7-4d43-a694-e877fb20b4a1 · inbound

Embody4D: A Generalist Data Engine for Embodied 4D World Modeling cites this paper.

Embody4D: A Generalist Data Engine for Embodied 4D World Modeling ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:25:09.924677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-01T00:17:39.409452Z digest=sha256:9968ed58161a0fb6521979dd261c37f663c73387f254697dd9685df5c7bb0ac4