Pith. sign in

Paper Citation Record · LEDGER

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents

As of 10 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 2 inbound Pith citation observations for arXiv:2507.17462.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.17462 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:53:38.496401Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T00:17:39.409452Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T00:25:09.923270Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ca5a907-6541-4edd-be98-86e26c94919c · outbound

This paper cites CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:31.294528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:31.294528Z digest=sha256:674ebe1b4c29106b8532dc48fff9164b4a28f0de36568ef3642843447be0b427

Observation f0a1a531-f8d7-473f-a58b-97a808b6eff5 · outbound

This paper cites Scaling Robot Learning with Semantically Imagined Experience,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Scaling Robot Learning with Semantically Imagined Experience,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:46.735035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:53:31.361436Z digest=sha256:01e13e83d35d68c517cf6d0247f51f2f776d23d44cfbbf6d03f71c49bad51f77

Observation 8baeeaea-cb0f-4687-bc98-7a2727894b97 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:31.519794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:31.519794Z digest=sha256:5e29435374612b8f171cb0be09237c7c79e7765efc1c6de196dd829d7e502106

Observation 4be03cb1-ac15-4dd4-bb8c-6a463f158267 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents OpenVLA: An Open-Source Vision-Language-Action Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:31.611898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:31.611898Z digest=sha256:f7c3428a0f7476fe04e578c48c01ea773ff38c61c0353f59cc3fcab450f98242

Observation 04a71937-6bfb-4ee6-a0e0-869bd17dd796 · outbound

This paper cites MagicDrive: Street view generation with diverse 3d geometry control,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents MagicDrive: Street view generation with diverse 3d geometry control,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:46.492290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:53:31.715847Z digest=sha256:e0b8edc75d378c5ef4e65bd5cbf5c2e8c9687a5341c4e08f8b80862f9a946fab

Observation ad4d3104-934c-424c-a9f3-e97ebffce656 · outbound

This paper cites BEVControl: Accurately Controlling Street-view Elements with Multi-perspective Consistency via BEV Sketch Layout.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents BEVControl: Accurately Controlling Street-view Elements with Multi-perspective Consistency via BEV Sketch Layout

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:31.818402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:31.818402Z digest=sha256:eed9657f813417b68f893bcc6c190ebb26abfaca54cf41068318ddbea4a7cde4

Observation 35592f22-efe7-47ff-92fd-c461a4bc71ed · outbound

This paper cites Street-view image generation from a bird’s-eye view layout,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Street-view image generation from a bird’s-eye view layout,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:46.224327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:53:31.939060Z digest=sha256:d054e740df51448062c1538a58b2d00fcc38118bbe7dea296228a971ebb1ace1

Observation 5ab518d0-e587-4160-9616-96361742f5ca · outbound

This paper cites Dragvideo: Interactive drag-style video editing,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Dragvideo: Interactive drag-style video editing,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:45.962585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:53:32.047960Z digest=sha256:c9190b49941534ebd0142a90a66ae1d3e5c38c2fe2741070e2e67e44189e41d5

Observation 26ed3f74-1906-4583-8acb-be8ce7bf7b73 · outbound

This paper cites Video-p2p: Video editing with cross-attention control,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Video-p2p: Video editing with cross-attention control,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:45.700618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:53:32.210605Z digest=sha256:8512c47e060d548be61f6e2e1e4d705fa9ba057ad261661799f96b406ce6d894

Observation 6c75b870-d27c-432c-bd5f-a7f2bfce9af3 · outbound

This paper cites Visual commonsense-aware representation network for video captioning,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Visual commonsense-aware representation network for video captioning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:45.383791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:53:32.355751Z digest=sha256:780c149da11b935daf56bed800f255e7aaac97dda7bfbd996b3270b86c379473

Observation a17b54cc-3fe0-44e2-bea7-a313f0067b13 · outbound

This paper cites Diffusion model-based image editing: A survey,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Diffusion model-based image editing: A survey,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:45.135486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:53:32.484985Z digest=sha256:9cb3da131b6aabcc79e124d90cded1290c0aecdf3ed33f216b6ead658c33404a

Observation 858904f2-5e64-41b3-bb0c-63489df87bf6 · outbound

This paper cites Feditnet++: Few- shot editing of latent semantics in gan spaces with correlated attribute disentanglement,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Feditnet++: Few- shot editing of latent semantics in gan spaces with correlated attribute disentanglement,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:44.811434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:53:32.605859Z digest=sha256:0eb1084f13e3efa3d6c05eb153785d2daab1a183be6b7ffbce4572f2847b5632

Observation de2cc8c5-818d-4812-82d7-c105f1d3360d · outbound

This paper cites Gaussctrl: Multi-view consistent text-driven 3d gaussian splatting edit- ing,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Gaussctrl: Multi-view consistent text-driven 3d gaussian splatting edit- ing,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:32.743292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:32.743292Z digest=sha256:49b94043d380ff8411aaae3927db035fb022da3332fcbb2be96c417a0eeba714

Observation f33ea19d-8b14-464f-98c4-669954140c4f · outbound

This paper cites Efficient dynamic scene editing via 4d gaussian-based static-dynamic separation,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Efficient dynamic scene editing via 4d gaussian-based static-dynamic separation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:44.429108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:53:32.872349Z digest=sha256:80316c23a96e8ddcb92d8a191fd08f59f5a30e6c98585920e486a56d5fee3a64

Observation 240dac32-42fa-4a83-95bf-d3b4d47eab96 · outbound

This paper cites Generating long videos of dynamic scenes,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Generating long videos of dynamic scenes,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:44.173267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:53:33.031861Z digest=sha256:028ef638b1fdc189571de020e830f3454d4fc6e87a0c0bb38de04972169246f9

Observation b2344980-f9e9-4149-aac0-0f5ced496595 · outbound

This paper cites Storydiffusion: Consistent self-attention for long-range image and video generation,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Storydiffusion: Consistent self-attention for long-range image and video generation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:43.868743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:53:33.186072Z digest=sha256:205f7aa38cc55e2e66704630ae5abc79c32439061d7809a4656cfdba6ffb63d0

Observation fef81db5-5e53-49eb-85f1-025222e676f7 · outbound

This paper cites Evalcrafter: Benchmarking and evaluating large video generation models,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Evalcrafter: Benchmarking and evaluating large video generation models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:43.588423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:53:33.288535Z digest=sha256:4b77980394865681463a997e9f2167af96366b22fab97e0ade0408baaee9479a

Observation d6105e95-80cd-4038-9ef7-ce23d0e3fe6c · outbound

This paper cites Towards long video understanding via fine-detailed video story generation,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Towards long video understanding via fine-detailed video story generation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:43.394123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:53:33.414870Z digest=sha256:e9ce0675a13074c5c659f9d89fefb21e8c0c193a88ca692618fbb52f9c11d600

Observation b30c98a2-6965-4090-9e4b-9012ea9e36a0 · outbound

This paper cites Maskgwm: A generalizable driving world model with video mask reconstruction,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Maskgwm: A generalizable driving world model with video mask reconstruction,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:43.094305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:53:33.564446Z digest=sha256:06f831c48adcb91bc43c91ad7d8504ee43b47009bd93b402ade4a15683b3970a

Observation 537a1492-6b2a-408f-880c-16d2b2a1bf8c · outbound

This paper cites A Survey on Long Video Generation: Challenges, Methods, and Prospects.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents A Survey on Long Video Generation: Challenges, Methods, and Prospects

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:33.702516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:33.702516Z digest=sha256:5d38eb9d03c154feb2a53efb50e183bcfd20ceb0ff8e1fefe30464d1d24904ae

Observation 64896346-6477-4c74-a9cf-a30b6a39df53 · outbound

This paper cites Dall-e-bot: Introducing web- scale diffusion models to robotics,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Dall-e-bot: Introducing web- scale diffusion models to robotics,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:42.734812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:53:33.900587Z digest=sha256:991fda814a3c7d000fb985311d0da8bf48d658cb20af225c33dc44f220f80a8a

Observation ec05318f-4a89-4cac-a4b0-b51e7bc9083a · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:34.065833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:34.065833Z digest=sha256:79e3be87567ac9e2e6e9f20fbdf8a325b7fcd4a7c25078c81cf4ccb78f67d88b

Observation 7fe0b49a-3265-49fe-a6aa-54abf355c307 · outbound

This paper cites GenAug: Retargeting behaviors to unseen situations via Generative Augmentation.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents GenAug: Retargeting behaviors to unseen situations via Generative Augmentation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:34.194140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:34.194140Z digest=sha256:5a7ded996fd6c7c8359df289180cbf596c3134c364d50eee6468bf3e42a35dff

Observation e7569c55-ebc9-4a05-9667-47468ebe9be1 · outbound

This paper cites Semantically controllable augmentations for generalizable robot learning,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Semantically controllable augmentations for generalizable robot learning,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:34.333520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:34.333520Z digest=sha256:aa81bf010e01c9bb4738698ed912b89b27b7f7dc0ee7a4ecf611241f751dd3fa

Observation 3f9cb659-3d2d-4a89-9357-d4a00e51e36a · outbound

This paper cites Roboagent: Generalization and efficiency in robot manipu- lation via semantic augmentations and action chunking,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Roboagent: Generalization and efficiency in robot manipu- lation via semantic augmentations and action chunking,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:42.481792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:53:34.459950Z digest=sha256:45de8c7412af5f3a36e1e3dbdbcc06a7ca496434290be3106f7e6a4c2800afd3

Observation f48df0bf-9d3b-4967-bb78-03bd870c2454 · outbound

This paper cites Segment anything,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Segment anything,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:34.586243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:34.586243Z digest=sha256:9a53116ffdf324d80a484f7d3cc7f25e2bda1eafae378b463972a5661077ada3

Observation e8b228d8-3a7f-428f-bc33-44f5c8b892d9 · outbound

This paper cites EnerVerse-AC: Envisioning Embodied Environments with Action Condition.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents EnerVerse-AC: Envisioning Embodied Environments with Action Condition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:34.708735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:34.708735Z digest=sha256:f0972df38803f8fc37c62dc4391cbbd1fd49c05612098843f6d2dad953bbd20f

Observation 27c73efe-2b8f-4912-8279-17ef1c9a34df · outbound

This paper cites Zero-1-to-3: Zero-shot one image to 3d object,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Zero-1-to-3: Zero-shot one image to 3d object,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:42.247498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:53:34.849354Z digest=sha256:0343cb792930bfaec4f9fca87539864ee90b0ee73cb278f6c52cc2aa9bb06fc5

Observation 0bce7733-0683-4bd4-b8a0-e45affa1d359 · outbound

This paper cites InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:34.980027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:34.980027Z digest=sha256:8fe1b94fb9ae300b7b53b420ad93b333b8cf64a73cc0e684af4b5bdb43c6d3a9

Observation 1c2321a9-0043-42e0-8703-03a03eeb6977 · outbound

This paper cites 3D-Adapter: Geometry-Consistent Multi-View Diffusion for High-Quality 3D Generation.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents 3D-Adapter: Geometry-Consistent Multi-View Diffusion for High-Quality 3D Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:35.168365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:35.168365Z digest=sha256:615e708f2af866191f5c1720f252ff90c11162e4de45e84c7bf6876effa707cd

Observation 2d9265b2-33db-4031-a976-f83f19887450 · outbound

This paper cites Dge: Direct gaussian 3d editing by consistent multi-view editing,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Dge: Direct gaussian 3d editing by consistent multi-view editing,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:42.001632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:53:35.339800Z digest=sha256:97a286bb9a2a16cad26590ced1f290dbb40dae1d2a3b85655d6d113d5101ad1c

Observation e8c810d4-e30d-4347-ae5c-2f106ca3aa20 · outbound

This paper cites Imfine: 3d inpainting via geometry-guided multi-view refinement,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Imfine: 3d inpainting via geometry-guided multi-view refinement,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:41.681720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:53:35.479138Z digest=sha256:d977958aa1f702bdb85fdcb4ec92305bac398ce3487e37816f942b0370ced6e7

Observation 822d2eb3-7671-400a-b957-cd6d7ef32a5d · outbound

This paper cites High- resolution image synthesis with latent diffusion models,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents High- resolution image synthesis with latent diffusion models,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:35.647926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:35.647926Z digest=sha256:a98438dd75a929f3b459e0a84e1847736bb445d3610de0b946759ad4fbbb0ecb

Observation efaa2062-6f9e-416b-8b88-5f1e572d0c7a · outbound

This paper cites Towards language-driven video inpainting via multimodal large language models,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Towards language-driven video inpainting via multimodal large language models,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:41.373678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:53:35.820351Z digest=sha256:8c1665a987bfd242e8b52a3f71e58dbd74c137b37cf39acd575b2feee339fc84

Observation 0d35a874-59eb-41ee-8366-83b33dab9d1c · outbound

This paper cites Brush2prompt: Contextual prompt generator for object inpainting,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Brush2prompt: Contextual prompt generator for object inpainting,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:41.065267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:53:35.926503Z digest=sha256:c4cc765c5c29422db5182a9e8c74aa236d6281c4a7d3e97ad57bd689e24e0705

Observation 9e760139-0a9d-4a6d-9ac5-19d643a221d6 · outbound

This paper cites Sketch- guided image inpainting with partial discrete diffusion process,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Sketch- guided image inpainting with partial discrete diffusion process,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:40.819891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:53:36.052743Z digest=sha256:43e1d959bc08f4cbb650a3be2b22bbcab23afaa27707a0eafad269636386c5a6

Observation 337a5171-982a-4350-842d-46761fed50a2 · outbound

This paper cites Few-shot image generation via style adaptation and content preservation,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Few-shot image generation via style adaptation and content preservation,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:40.501753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:53:36.256605Z digest=sha256:02e50dd16ebde142f3dba65f0ec5e1fdffc741d735f9bb8fbff1aef7d77fbd30

Observation 2b5dfbe2-4686-4566-ac14-412938512fcb · outbound

This paper cites Learning transferable visual models from natural language supervision,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Learning transferable visual models from natural language supervision,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:36.417013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:36.417013Z digest=sha256:24ce9b237f7e7a57b5fed12595220a6e299a8843226dfd90f3e67e175348264f

Observation c1f65cef-63aa-4d4d-b954-2bf34b686539 · outbound

This paper cites Video diffusion models,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Video diffusion models,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:36.559772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:36.559772Z digest=sha256:27d4767975af705035fc04977a7b29f2bb8ea83dd5ed21b343f2a64cf81d3c4b

Observation e59b2166-dec2-49ad-b81d-6e467a3a7ffb · outbound

This paper cites Enerverse: Envisioning embodied future space for robotics manipulation,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Enerverse: Envisioning embodied future space for robotics manipulation,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:36.744744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:36.744744Z digest=sha256:a2391fbfc4356cca84e0f5ecf2237cfc1385e90139b093732bbdced3f56a009e

Observation 02991ebf-d39f-4ce1-b936-b379fb0b1826 · outbound

This paper cites Epidiff: Enhancing multi-view synthesis via localized epipolar-constrained diffusion,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Epidiff: Enhancing multi-view synthesis via localized epipolar-constrained diffusion,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:40.178023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:53:36.907756Z digest=sha256:b88ff82f2b67182c32f9a7c2bbd3e578c261b39dfda30fee07b5aeff658d15d1

Observation 4e7cb55a-f64f-4b26-b156-4c97ebcb9b65 · outbound

This paper cites Ar-diffusion: Asynchronous video generation with auto-regressive diffusion,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Ar-diffusion: Asynchronous video generation with auto-regressive diffusion,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:39.812003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:53:37.094974Z digest=sha256:eb0927deaf43e4a15829736f1322ea53d01d0e1ef67042290574419baf937948

Observation 475c8523-758a-435a-a4cc-b3631bffabf6 · outbound

This paper cites Progressive autoregressive video diffusion models,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Progressive autoregressive video diffusion models,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:39.489380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:53:37.295899Z digest=sha256:f5e3023461df495a8b1ad2675f85f95c468e9dda6ef05bde6be1794a8f6c2b41

Observation 8a3c1759-7c0a-4c46-8052-9405e6541aa0 · outbound

This paper cites From slow bidirectional to fast autoregressive video diffusion models,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents From slow bidirectional to fast autoregressive video diffusion models,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:37.457637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:37.457637Z digest=sha256:9babc67a9625020629d0520f4ad2c9521479a2d2dabc13d204373be1b341c3c7

Observation a207e586-b1bb-4667-bdb0-8ad1045082c9 · outbound

This paper cites Qwen2.5-VL Technical Report.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Qwen2.5-VL Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:37.599945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:37.599945Z digest=sha256:b17ec4a043ddbd8e7f76700a59d9546a83c12414f65660dd5b75e23dd4ddcc07

Observation 72315267-80ee-42a1-b6e5-0598d6458ecc · outbound

This paper cites Step1X-Edit: A Practical Framework for General Image Editing.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Step1X-Edit: A Practical Framework for General Image Editing

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:37.812216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:37.812216Z digest=sha256:606d6e343692ed8efc71c34543880eb8aeef20ea4740beed9a0714d89339301b

Observation 80f9a6bd-903f-46d4-82d4-275ddf05a4f6 · outbound

This paper cites Robotwin: Dual-arm robot benchmark with generative digital twins (early version),.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Robotwin: Dual-arm robot benchmark with generative digital twins (early version),

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:39.242888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:53:37.961048Z digest=sha256:74f2e94398ffab7a57af0e55216f77f50e9e22225384304a5a2f9db72cc6fc8e

Observation 4b404dc2-4398-4132-8db1-459fd9fb345d · outbound

This paper cites Diffusion policy: Visuomotor policy learning via ac- tion diffusion,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Diffusion policy: Visuomotor policy learning via ac- tion diffusion,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:38.099075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:38.099075Z digest=sha256:81014367030d5e55552f0d257b0ef7d10c7e8ae0b7c36f1a74e4b931fa9943de

Observation cd782825-b3e2-4a52-a651-73a0e9b715f1 · outbound

This paper cites Schedule your edit: A simple yet effective diffu- sion noise schedule for image editing,.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Schedule your edit: A simple yet effective diffu- sion noise schedule for image editing,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:38.933412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:53:38.294649Z digest=sha256:91187fe2997b0159bdbcd3c04faaf66e1a23b5d0f25815bf7f41c34d605d2fde

Observation b0ffbd48-09d1-4a61-98ee-25730bfa6101 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:38.496401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:38.496401Z digest=sha256:c905ae241f467eeef05f400b31fb7cfc61ec5ffbcf574f0e59f1ef387039695b

Pith citing papers

Observation 6fb304ea-b10f-4f1d-8fde-8c04a951d874 · inbound

Embody4D: A Generalist Data Engine for Embodied 4D World Modeling cites this paper.

Embody4D: A Generalist Data Engine for Embodied 4D World Modeling ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:32.152811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T19:16:50.004359Z digest=sha256:3a624d753fb959f22bcb88d2816de6f9a568c82965c8e2eb3141dd8ee0888ebe

Observation 23da8182-92d7-4d43-a694-e877fb20b4a1 · inbound

Embody4D: A Generalist Data Engine for Embodied 4D World Modeling cites this paper.

Embody4D: A Generalist Data Engine for Embodied 4D World Modeling ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:25:09.924677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T00:17:39.409452Z digest=sha256:8123f1a17d2e34036dce4afc95ffa74b8d79de784477d7db72db87c9e52a3d20