REVIEW 3 cited by
Optical Flow Representation Alignment Mamba Diffusion Model for Medical Video Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Medical video generation models are expected to have a profound impact on the healthcare industry, including but not limited to medical education and training, surgical planning, and simulation. Current video diffusion models typically build on image diffusion architecture by incorporating temporal operations (such as 3D convolution and temporal attention). Although this approach is effective, its oversimplification limits spatio-temporal performance and consumes substantial computational resources. To counter this, we propose Medical Simulation Video Generator (MedSora), which incorporates three key elements: i) a video diffusion framework integrates the advantages of attention and Mamba, balancing low computational load with high-quality video generation, ii) an optical flow representation alignment method that implicitly enhances attention to inter-frame pixels, and iii) a video variational autoencoder (VAE) with frequency compensation addresses the information loss of medical features that occurs when transforming pixel space into latent features and then back to pixel frames. Extensive experiments and applications demonstrate that MedSora exhibits superior visual quality in generating medical videos, outperforming the most advanced baseline methods. Further results and code are available at https://wongzbb.github.io/MedSora
Forward citations
Cited by 3 Pith papers
-
MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos
MedVideoCap-55K, a 55,803-clip caption-rich medical video dataset, enables MedGen, a LoRA fine-tune of HunyuanVideo that reports top open-source scores and near-commercial quality on medical video benchmarks.
-
FEAT: Full-Dimensional Efficient Attention Transformer for Medical Video Generation
FEAT combines spatial, temporal, and channel attention in a diffusion transformer and reports improved medical video generation metrics with lower parameter counts than Endora.
-
SG2VID: Scene Graphs Enable Fine-Grained Control for Video Synthesis
SG2VID conditions a latent video diffusion model on scene graphs with temporal features to generate controllable surgical videos across cataract and cholecystectomy datasets.
Discussion (0). Sign in to comment.