Pith. sign in

REVIEW 4 major objections 5 minor 100 references

Deblur-Avatar: Animatable Avatars from Motion-Blurred Monocular Videos

T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Modeling the body's pose trajectory during exposure lets 3D Gaussian avatars be trained on motion-blurred monocular video and still render sharp, re-poseable humans.

desk verdict A plausible first solution to a real problem, but the sharp-avatar claim needs external validation before it fully lands. read the letter →

arxiv 2501.13335 v4 pith:P3S4PWVV submitted 2025-01-23 cs.CV

classification cs.CV
keywords 3DGaussianSplattinganimatablehumanavatarmotionblurmonocularvideoposetrajectorySMPLneuralrenderingpose-dependentfusion
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Existing human avatar methods assume video frames are sharp, so everyday capture with fast movement or long exposure produces blurry, low-detail reconstructions. This paper proposes to treat motion blur as a first-class signal instead of a nuisance: it models the body's pose trajectory during each frame's exposure, renders a sequence of sharp virtual frames, and averages them to reproduce the observed blur. The pose trajectories and the 3D Gaussians are optimized jointly, so the Gaussians converge to a sharp, re-poseable avatar. A pose-dependent fusion mask lets static body parts train as sharp while moving parts train as blurred, which the ablation shows is important. If correct, the method would let casual monocular video, blur included, produce high-fidelity animatable avatars without a separate deblurring step.

What carries the argument

The load-bearing mechanism is the Slerp-based human motion trajectory model. Given a blurred frame with an estimated body pose, the method learns two poses $\theta_{\text{start}}$ and $\theta_{\text{end}}$ at the exposure endpoints; intermediate poses follow spherical linear interpolation in quaternion space (Eqs. 8–9), so the body is assumed to move at constant angular velocity during exposure. This trajectory yields $n$ virtual sharp images via the standard deformable-Gaussian avatar pipeline (non-rigid cloth deformation, rigid skinning, rasterization), and Eq. (6) averages them into the blurred image. The second component is a pose-dependent fusion MLP that takes the pose latent code, view embedding, position encoding, and the rendered virtual image to output a per-pixel mask blending the blurred and sharp renders during training. At inference the mask and trajectory modules are dropped, leaving a standard sharp Gaussian avatar renderer.

What would settle it

Train on a synthetic sequence with a known ground-truth pose trajectory that includes strong acceleration (for example, a fast punch or a sudden stop) and compare the optimized start/end poses to the true ones. If the recovered trajectory deviates substantially from the ground truth while the rendered blur still matches the input, the trajectory model is absorbing error into the Gaussians rather than recovering true motion; alternatively, if residual ghosting remains along the accelerating limb, the constant-velocity Slerp prior is the bottleneck.

Watch

Extended reading notes

Core claim

The paper's central claim is that human-motion blur in monocular avatar capture can be inverted by explicitly modeling the exposure period with a human motion trajectory. For each blurred frame, the method optimizes two poses of the parametric body model SMPL, one at exposure start and one at end, interpolates between them with spherical linear interpolation to obtain $n$ virtual poses, deforms canonical 3D Gaussians to each virtual pose, rasterizes $n$ sharp images, and averages them to synthesize the blurred observation (Eqs. 5–9). A pose-dependent fusion network then predicts a per-pixel mask that blends the averaged blurred render with the sharp virtual render, so the loss supervises each pixel in the regime that produced it. After training, only the sharp branch is rendered, giving a crisp animatable avatar at real-time frame rates. The paper presents this as the first framework aimed at sharp animatable avatars from motion-blurred monocular videos, with synthetic and real-world experiments supporting the claim.

Load-bearing premise

The method rests on the assumption that every frame's blur is the average of sharp renders of the body moving at constant velocity between two endpoint poses during exposure, with camera shake, acceleration, and non-rigid cloth motion treated as negligible.

Editorial extensions

If this is right

  • Monocular casual capture with fast movements or low-light long exposures becomes usable for avatar reconstruction without requiring sharp video.
  • Pre-deblurring with 2D video deblurring networks becomes unnecessary; the paper's experiments show such preprocessing helps baselines only marginally, while the 3D trajectory model outperforms it.
  • The avatar stays animatable to out-of-distribution poses because the sharp representation is in canonical Gaussian space, reposed via SMPL during inference.
  • Real-time rendering is preserved: the added modules cost training time but no inference cost, keeping roughly 50 FPS.
  • The number of virtual frames can stay small ($n=5$); more frames do not consistently improve quality and increase training cost.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same 'average a trajectory of articulated renders' recipe should transfer to other articulated subjects (quadrupeds, robots, hands) whenever a parametric skeleton and skinning are available; the core is the Slerp trajectory, not the specific body model.
  • Because the only pose supervision comes from a single pose estimate on a blurred frame, the optimization may absorb trajectory error into Gaussian parameters; an interesting extension would be to add temporal consistency across frames or to train a blur-aware pose estimator so the trajectory is constrained before avatar fitting.
  • The pose-dependent fusion mask is essentially a learned per-pixel blur map; it could be reused as a free motion-magnitude signal or to weight a confidence-aware loss, which the paper does not explore.
  • The constant-velocity Slerp assumption is a strong prior; replacing it with per-joint quadratic or learned trajectories could address acceleration, at the cost of more unknowns, and the ablation's cubic B-spline result suggests modest headroom.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes Deblur-Avatar, a 3D Gaussian Splatting framework for reconstructing sharp, animatable human avatars from monocular videos containing motion blur caused by human movement. The method models the exposure interval of each frame by estimating two SMPL poses (θ_start, θ_end) at the exposure endpoints and interpolating them with spherical linear interpolation (Eq. 8) to produce n virtual poses. Each virtual pose is used to deform canonical Gaussians and render a sharp image; the n images are averaged (Eq. 6) to synthesize the blurred frame. A pose-dependent fusion MLP (Eqs. 16–17) blends the averaged blurred rendering with one virtual sharp image, supervised by the input blurred frame, and is discarded at inference. Experiments compare against six human-avatar baselines, with and without a video-deblurring front end, on a synthesized ZJU-MoCap-Blur dataset and on a self-captured real dataset with qualitative evaluation only.

Significance. If the central claim holds, the paper addresses a real gap: existing animatable avatar methods assume sharp inputs, while human-motion blur is common in casual monocular capture. The paper's strengths include a physically motivated blur formation model, a new synthetic dataset built from real high-frame-rate ZJU-MoCap footage, a self-captured real dataset, and an ablation study that examines the trajectory representation and the number of virtual frames. The authors also release code and models, and the method renders at 50 FPS, which is practically relevant. However, the quantitative evidence for the sharp-avatar claim is currently limited to the synthetic benchmark, and the fusion mechanism has a potential collapse mode that could allow the model to fit the blurred input without recovering sharp Gaussians. These gaps need to be addressed before the claim of recovering sharp avatars from real motion-blurred monocular videos is fully supported.

major comments (4)
  1. [§IV-A, Eq. (6)] The quantitative evaluation is carried out on ZJU-MoCap-Blur, where each blurred frame is created by averaging 17–49 high-fps sharp frames (Sec. IV-A). This is precisely the temporal-averaging formation model used in the paper's forward model (Eq. 6). As a result, high PSNR/SSIM/LPIPS on this benchmark demonstrate that the method can invert its own assumed blur formation process; they do not demonstrate generalization to real motion blur, which may include camera shake, acceleration, and complex non-rigid cloth motion. The Real-Human-Blur dataset has no sharp ground truth and is used only for qualitative comparison (Sec. IV-C), so the central claim of recovering sharp avatars from real motion-blurred monocular videos is not quantitatively supported. I recommend adding a quantitative real-world benchmark (e.g., a dataset with high-speed sharp ground truth and real optical blur, or a synthetic benchmark generated with a different, richer blur model including camera motion) to test the transfer.
  2. [§III-E, Eq. (17)] The pose-dependent fusion mask can, in principle, collapse to M(x,j) ≈ 1 everywhere: then Cout(x,j) = B(x,j), and the training loss is minimized by fitting the blurred rendering to the blurred input. Because the sharp virtual image C(x,j) is never supervised directly and the mask is discarded at inference, the optimization does not force the Gaussians to represent sharp content; it only forces the blended output to match Bgt. The paper reports no statistics or visualizations of the learned masks, so it is unclear whether the reported gains come from genuinely sharp Gaussians or from the mask copying the blurred image. I suggest adding a regularizer that discourages large mask values (e.g., a prior towards M=0 in regions with small motion), directly supervising C against a deblurred estimate, or reporting mask statistics and sharpness metrics on the reconstructed virtual image.
  3. [§III-C, Eq. (8)] The trajectory model is a strong assumption: it represents the entire human motion during exposure as a constant-velocity Slerp between two SMPL poses, with no camera motion, acceleration, or per-frame temporal consistency. The paper's Limitation (1) explicitly states that the method relies on a single input pose and does not explore inter-frame relationships. Since the synthetic benchmark is generated by averaging real high-fps frames, it does not test the Slerp assumption against camera shake or complex human accelerations; the real dataset is only qualitative. This is a correctness-risk concern, not an internal inconsistency. A concrete test would be to add synthetic camera shake or to use a real blurred sequence with a synchronized high-speed sharp camera to measure whether the Slerp model remains accurate enough to produce sharp avatars.
  4. [§IV-B, Table I] All quantitative results are single training runs with no error bars or significance tests. Several PSNR differences are small (e.g., sequence 377: 30.36 vs. 30.29 for GART; sequence 386: 33.75 vs. 33.68 for 3DGS-Avatar). The claim that the method 'significantly outperforms' baselines is therefore not statistically supported. Reporting multiple seeds or paired per-frame comparisons would strengthen the quantitative claims.
minor comments (5)
  1. [Header, Index Terms] The Index Terms line reads 'All-in-Focus synthesis, main/ultra-wide camera, occlusion-aware networks', which appears to be a copy-paste artifact from an unrelated paper and does not match the content of this manuscript; please correct it.
  2. [Figure 1 caption] The caption notation 'LPIPS ∗=LPIPS∗103' is confusing; it should read 'LPIPS × 10^3' and clearly define the axes, since the figure plots PSNR against LPIPS.
  3. [Figure 5] The panel labels cite [56]+[85], [56]+[62], [56]+[35], and [56]+[82] while the text and Table II refer to [84] as the video deblurring method; the figure references are inconsistent and need to be aligned.
  4. [§IV-A, Real-Human-Blur] The sentence 'We train and evaluate the Real-Human-Blur dataset at the resolution of 540 × 540, 960 × 540, and 360 × 640' is unclear because the Real-Human-Blur dataset is used only for qualitative evaluation; please clarify how these resolutions are used and whether any quantitative evaluation is performed.
  5. [Throughout] There are several typos and formatting inconsistencies, including 'Canoncial' in Figure 2, 'Arah' for ARAH in Table I, and 'T V' in the author affiliation line; a careful proofreading pass is needed.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the blur formation model is a physical forward model, the trajectory parameters are optimized rather than fitted to the target metric, and the only same-author citation is a non-load-bearing related-work contrast.

full rationale

The paper's derivation chain is self-contained. The claimed result is that jointly optimizing human-motion trajectories and 3D Gaussians, rendering virtual sharp frames, averaging them (Eq. 6), and blending with a pose-dependent mask recovers sharp avatars from blurred monocular video. The forward model in Eqs. (5)-(9) is the standard temporal integration of irradiance: B(x) is the exposure-time integral of sharp images C_t(x), discretized as an average of n virtual sharp renders. This is a physical image-formation equation, not a definition of the output in terms of the input loss. The optimized quantities are the SMPL trajectory endpoints theta_start and theta_end and the Gaussian parameters, all supervised by the blurred RGB input; the reported metrics are computed against held-out sharp ground-truth frames from ZJU-MoCap-Blur, so the sharp-view PSNR/SSIM/LPIPS numbers are not statistically forced by the training objective. The synthetic benchmark does use temporal averaging to create blur, matching the assumed physics, but it also uses external high-frame-rate frame interpolation and re-estimated SMPL poses, and it evaluates novel views against original sharp frames; any limitation here concerns generalization to real blur, not circularity. The Real-Human-Blur evaluation is qualitative only, which limits evidence strength but is not a circular step. The one self-citation, DyBlurRF (ref. [26], sharing author Zhiguo Cao), appears only as a related-work contrast stating that prior dynamic deblurring lacks specialized human motion representation; no derivation or numerical result depends on it, so it is not load-bearing. The pose-dependent fusion mask could in principle absorb trajectory errors during training, but it is ablated (Table III), optimized with a delayed schedule, and discarded at inference; this is a correctness consideration, not an equivalence between input and output. No equation in the paper reduces to its own input by construction, no fitted parameter is renamed as a prediction, and no uniqueness claim is imported from the authors' prior work.

Assumptions & free parameters 3 free parameters · 5 assumptions · 1 invented entities

The central method rests on four domain assumptions, blur-as-average, pose interpolation, reliable blurry-frame annotations, and human-only motion, plus standard LBS deformation. The free parameters are the per-frame trajectory endpoints, the hand-chosen virtual frame count, and the loss weights. No new physical entity is claimed; the virtual frame sequence is a computational device.

free parameters (3)
  • per-frame SMPL start/end poses (theta_start, theta_end) = learned per frame, 72-d pose vector each
    Central to Eqs. (5)-(9): every blurred frame's motion trajectory is represented by these two optimized endpoints, with no temporal sticking prior across frames.
  • virtual frame count n = 5
    Chosen by hand in Sec. III-C as a quality and efficiency trade-off; Table IV shows n=5 is near-optimal, confirming it is a tuned constant rather than derived.
  • loss weights lambda1..lambda5 = 0.01, 0.1, 10 (decayed), 1, 100
    Hand-set in Sec. III-F; no sensitivity analysis is reported, so robustness to these choices is unknown.
assumptions (5)
  • domain assumption The observed blurred image is the normalized integral of instantaneous sharp images over exposure (Eqs. 5-6).
    Physical model of motion blur, but assumes linear sensor response and no background or camera motion; used as the forward model for optimization.
  • ad hoc to paper Human motion during exposure is well approximated by Slerp between two SMPL poses at exposure endpoints (Eq. 8).
    Adopted to make trajectory modeling tractable; not derived from biomechanics or data, and the paper itself notes that acceleration and complex motion would violate it (Limitations).
  • domain assumption SMPL parameters, camera calibration, and foreground masks from blurry video are sufficiently accurate.
    Pose and masks are estimated by EasyMocap, SPIN, and SAM on blurry frames (Sec. IV-A); inaccuracies propagate into the deformation and trajectory.
  • domain assumption Human motion is the dominant blur source; camera motion is negligible.
    The framework has no camera-motion blur term, so its applicability is limited to the regime claimed in Secs. I and II.
  • domain assumption Linear blend skinning with learned skinning and non-rigid MLPs can represent clothed human deformation.
    Inherited from 3DGS-Avatar [20]; needed to deform Gaussians from canonical to observation space for each virtual pose.
invented entities (1)
  • Virtual sharp image sequence, one per interpolated pose during exposure
    purpose: Discretizes Eq. (5) into n renders whose average forms the blurred image; used only in training, not at inference
    A modeling construction derived from the blur integral; it has no observable counterpart outside the paper and is not claimed as a physical discovery, but it is introduced specifically to make the blur model differentiable.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Deblur-Avatar: Animatable Avatars from Motion-Blurred Monocular Videos." pith.science (2026). https://pith.science/paper/P3S4PWVV

@misc{pith2026250113335,
  author       = {Pith},
  title        = {Pith review of: Deblur-Avatar: Animatable Avatars from Motion-Blurred Monocular Videos},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/P3S4PWVV}},
  note         = {Machine review of arXiv:2501.13335}
}
read the original abstract

We introduce a novel framework for modeling high-fidelity, animatable 3D human avatars from motion-blurred monocular video inputs. Motion blur is prevalent in real-world dynamic video capture, especially due to human movements in 3D human avatar modeling. Existing methods either (1) assume sharp image inputs, failing to address the detail loss introduced by motion blur, or (2) mainly consider blur by camera movements, neglecting the human motion blur which is more common in animatable avatars. Our proposed approach integrates a human movement-based motion blur model into 3D Gaussian Splatting (3DGS). By explicitly modeling human motion trajectories during exposure time, we jointly optimize the trajectories and 3D Gaussians to reconstruct sharp, high-quality human avatars. We employ a pose-dependent fusion mechanism to distinguish moving body regions, optimizing both blurred and sharp areas effectively. Extensive experiments on synthetic and real-world datasets demonstrate that our method significantly outperforms existing methods in rendering quality and quantitative metrics, producing sharp avatar reconstructions and enabling real-time rendering under challenging motion blur conditions. Code and models are available at https://github.com/xianrui-luo/deblur_avatar.

Figures

Figures reproduced from arXiv: 2501.13335 by the authors.

Figure 1
Figure 1. We propose the first method to render sharp animatable avatars with a motion-blurred monocular video as input. (b) demonstrates that our method restores sharper details than baselines. (c) LPIPS∗=LPIPS∗103 . Smaller circles denote higher FPS. motion blur in these blurred areas. To address the absence of existing human avatar mo￾tion deblurring datasets, we collect synthetic and real-world datasets and conduct an ext… view at source ↗
Figure 2
Figure 2. Pipeline of our method. For each frame with its corresponding body pose as input, we introduce human motion trajectory modeling to predict human poses θstart and θend at the start Tstart and end Tend of the exposure period of the frame. Then we use interpolation to obtain a virtual human pose sequence. With the pose of each virtual timestamp, we perform non-rigid deformation and rigid transformation to the 3D Gaussi… view at source ↗
Figure 3
Figure 3. Qualitative results on ZJU-MoCap-Blur. Compared with existing human avatar methods, our method generates sharper novel views and preserves more details with original motion-blurred inputs. Real-Human-Blur. Given the absence of publicly available, real-world datasets that are specifically tailored for tackling the challenge of motion blur in human avatar modeling, we have curated a dataset consisting of monocular mot… view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Qualitative results on Real-Human-Blur dataset. Our method generates sharper novel poses that are more faithful and preserve more details. TABLE I. Quantitative results on ZJU-MoCap-Blur. The input sequence is the motion-blurred images. The best performance is in boldf…
Figure 5
Figure 5. Figure 5: Qualitative results on ZJU-MoCap-Blur (pre-deblurred). The input sequences are deblurred by state-of-the-art video deblurring method [84]. Compared with existing human avatar methods, our method generates sharper novel views and preserves more details with original mot…
Figure 6
Figure 6. Figure 6: The qualitative comparisons of out-of-distribution pose animation on ZJU-MoCap-Blur. Our method produces fewer artifacts than baselines, demonstrating good generalization to unseen poses. TABLE III. Ablation study on the effectiveness of our proposed mod￾ules. We evalu…
Figure 7
Figure 7. Figure 7: Ablation Study on non-rigid deformation, motion trajectory modeling, and pose-dependent fusion. Implementing these modules pre￾serves more details and alleviates avatars’ motion-related artifacts. TABLE V. Ablation study on the trajectories representations. We evaluate…
Figure 8
Figure 8. Figure 8: Pre-deblurring results of state-of-the-art video deblurring method [84]. On the first row we show the data where this method performs well. Current video deblurring paradigms are fitting for motion blur from camera shake. However, the motion blur in the human avatar mo…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

100 extracted references · 59 canonical work pages

  1. [1]

    H-nerf: Neural radiance fields for rendering and temporal reconstruction of humans in motion,

    H. Xu, T. Alldieck, and C. Sminchisescu, “H-nerf: Neural radiance fields for rendering and temporal reconstruction of humans in motion,” in NeurIPS, 2021

  2. [2]

    A-nerf: Articulated neural radiance fields for learning human shape, appearance, and pose,

    S.-Y . Su, F. Yu, M. Zollh ¨ofer, and H. Rhodin, “A-nerf: Articulated neural radiance fields for learning human shape, appearance, and pose,” in NeurIPS, 2021

  3. [3]

    Neural articulated radiance field,

    A. Noguchi, X. Sun, S. Lin, and T. Harada, “Neural articulated radiance field,” in ICCV, 2021

  4. [4]

    Efficient neural radiance fields for interactive free-viewpoint video,

    H. Lin, S. Peng, Z. Xu, Y . Yan, Q. Shuai, H. Bao, and X. Zhou, “Efficient neural radiance fields for interactive free-viewpoint video,” in SIGGRAPH Asia 2022 Conference Papers , 2022

  5. [5]

    Vid2avatar: 3d avatar reconstruction from videos in the wild via self-supervised scene decomposition,

    C. Guo, T. Jiang, X. Chen, J. Song, and O. Hilliges, “Vid2avatar: 3d avatar reconstruction from videos in the wild via self-supervised scene decomposition,” in CVPR, 2023

  6. [6]

    Neuman: Neural human radiance field from a single video,

    W. Jiang, K. M. Yi, G. Samei, O. Tuzel, and A. Ranjan, “Neuman: Neural human radiance field from a single video,” in ECCV, 2022

  7. [7]

    Tava: Template-free animatable volumetric actors,

    R. Li, J. Tanke, M. V o, M. Zollh ¨ofer, J. Gall, A. Kanazawa, and C. Lassner, “Tava: Template-free animatable volumetric actors,” in ECCV, 2022

  8. [8]

    Animatable neural implicit surfaces for creating avatars from videos,

    S. Peng, S. Zhang, Z. Xu, C. Geng, B. Jiang, H. Bao, and X. Zhou, “Animatable neural implicit surfaces for creating avatars from videos,” arXiv preprint arXiv:2203.08133 , vol. 4, no. 5, 2022

Show all 100 references
  1. [9]

    Neural actor: Neural free-view synthesis of human actors with pose control,

    L. Liu, M. Habermann, V . Rudnev, K. Sarkar, J. Gu, and C. Theobalt, “Neural actor: Neural free-view synthesis of human actors with pose control,” ACM TOG, vol. 40, no. 6, pp. 1–16, 2021

  2. [10]

    Arah: Animatable volume rendering of articulated human sdfs,

    S. Wang, K. Schwarz, A. Geiger, and S. Tang, “Arah: Animatable volume rendering of articulated human sdfs,” in ECCV, 2022

  3. [11]

    Hdhumans: A hybrid approach for high-fidelity digital humans,

    M. Habermann, L. Liu, W. Xu, G. Pons-Moll, M. Zollhoefer, and C. Theobalt, “Hdhumans: A hybrid approach for high-fidelity digital humans,” Proceedings of the ACM on Computer Graphics and Inter- active Techniques, vol. 6, no. 3, pp. 1–23, 2023

  4. [12]

    3d gaussian splatting for real-time radiance field rendering

    B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Drettakis, “3d gaussian splatting for real-time radiance field rendering.” ACM TOG, vol. 42, no. 4, pp. 139–1, 2023

  5. [13]

    Gaussianavatar: Towards realistic human avatar modeling from a single video via animatable 3d gaussians,

    L. Hu, H. Zhang, Y . Zhang, B. Zhou, B. Liu, S. Zhang, and L. Nie, “Gaussianavatar: Towards realistic human avatar modeling from a single video via animatable 3d gaussians,” in CVPR, 2024

  6. [14]

    Splatarmor: Articulated gaussian splatting for animatable humans from monocular rgb videos,

    R. Jena, G. S. Iyer, S. Choudhary, B. Smith, P. Chaudhari, and J. Gee, “Splatarmor: Articulated gaussian splatting for animatable humans from monocular rgb videos,” arXiv preprint arXiv:2311.10812 , 2023

  7. [15]

    Hugs: Human gaussian splats,

    M. Kocabas, J.-H. R. Chang, J. Gabriel, O. Tuzel, and A. Ranjan, “Hugs: Human gaussian splats,” in CVPR, 2024

  8. [16]

    Gart: Gaussian articulated template models,

    J. Lei, Y . Wang, G. Pavlakos, L. Liu, and K. Daniilidis, “Gart: Gaussian articulated template models,” in CVPR, 2024

  9. [17]

    Animatable gaussians: Learning pose-dependent gaussian maps for high-fidelity human avatar model- ing,

    Z. Li, Z. Zheng, L. Wang, and Y . Liu, “Animatable gaussians: Learning pose-dependent gaussian maps for high-fidelity human avatar model- ing,” in CVPR, 2024

  10. [18]

    Animatable 3d gaussian: Fast and high-quality reconstruction of multiple human avatars,

    Y . Liu, X. Huang, M. Qin, Q. Lin, and H. Wang, “Animatable 3d gaussian: Fast and high-quality reconstruction of multiple human avatars,” arXiv preprint arXiv:2311.16482 , 2023

  11. [19]

    Gva: Reconstructing vivid 3d gaussian avatars from monoc- ular videos,

    X. Liu, C. Wu, J. Liu, X. Liu, C. Zhao, H. Feng, E. Ding, and J. Wang, “Gva: Reconstructing vivid 3d gaussian avatars from monoc- ular videos,” CoRR, 2024

  12. [20]

    3dgs-avatar: Animatable avatars via deformable 3d gaussian splatting,

    Z. Qian, S. Wang, M. Mihajlovic, A. Geiger, and S. Tang, “3dgs-avatar: Animatable avatars via deformable 3d gaussian splatting,” in CVPR, 2024. DEBLUR-A V ATAR: ANIMATABLE A V ATARS FROM MOTION-BLURRED MONOCULAR VIDEOS 11

  13. [21]

    Deblur-nerf: Neural radiance fields from blurry images,

    L. Ma, X. Li, J. Liao, Q. Zhang, X. Wang, J. Wang, and P. V . Sander, “Deblur-nerf: Neural radiance fields from blurry images,” in CVPR, 2022

  14. [22]

    Dp-nerf: Deblurred neural radiance field with physical scene priors,

    D. Lee, M. Lee, C. Shin, and S. Lee, “Dp-nerf: Deblurred neural radiance field with physical scene priors,” in CVPR, 2023

  15. [23]

    Bags: Blur agnostic gaussian splatting through multi-scale kernel modeling,

    C. Peng, Y . Tang, Y . Zhou, N. Wang, X. Liu, D. Li, and R. Chellappa, “Bags: Blur agnostic gaussian splatting through multi-scale kernel modeling,” arXiv preprint arXiv:2403.04926 , 2024

  16. [24]

    Bad-nerf: Bundle adjusted deblur neural radiance fields,

    P. Wang, L. Zhao, R. Ma, and P. Liu, “Bad-nerf: Bundle adjusted deblur neural radiance fields,” in CVPR, 2023

  17. [25]

    Bad-gaussians: Bundle adjusted deblur gaussian splatting,

    L. Zhao, P. Wang, and P. Liu, “Bad-gaussians: Bundle adjusted deblur gaussian splatting,” arXiv preprint arXiv:2403.11831 , 2024

  18. [26]

    Dyblurf: Dynamic neural radiance fields from blurry monocular video,

    H. Sun, X. Li, L. Shen, X. Ye, K. Xian, and Z. Cao, “Dyblurf: Dynamic neural radiance fields from blurry monocular video,” in CVPR, 2024

  19. [27]

    Video-based characters: creating new human performances from a multi-view video database,

    F. Xu, Y . Liu, C. Stoll, J. Tompkin, G. Bharaj, Q. Dai, H.-P. Seidel, J. Kautz, and C. Theobalt, “Video-based characters: creating new human performances from a multi-view video database,” in ACM SIGGRAPH 2011 papers , 2011, pp. 1–10

  20. [28]

    The relightables: V olumetric performance capture of humans with realistic relighting,

    K. Guo, P. Lincoln, P. Davidson, J. Busch, X. Yu, M. Whalen, G. Harvey, S. Orts-Escolano, R. Pandey, J. Dourgarian et al. , “The relightables: V olumetric performance capture of humans with realistic relighting,” ACM TOG, vol. 38, no. 6, pp. 1–19, 2019

  21. [29]

    Photorealistic monocular 3d reconstruction of humans wearing clothing,

    T. Alldieck, M. Zanfir, and C. Sminchisescu, “Photorealistic monocular 3d reconstruction of humans wearing clothing,” in CVPR, 2022

  22. [30]

    Pifu: Pixel-aligned implicit function for high-resolution clothed human digitization,

    S. Saito, Z. Huang, R. Natsume, S. Morishima, A. Kanazawa, and H. Li, “Pifu: Pixel-aligned implicit function for high-resolution clothed human digitization,” in ICCV, 2019

  23. [31]

    Pifuhd: Multi-level pixel- aligned implicit function for high-resolution 3d human digitization,

    S. Saito, T. Simon, J. Saragih, and H. Joo, “Pifuhd: Multi-level pixel- aligned implicit function for high-resolution 3d human digitization,” in CVPR, 2020

  24. [32]

    High-quality streamable free- viewpoint video,

    A. Collet, M. Chuang, P. Sweeney, D. Gillett, D. Evseev, D. Calabrese, H. Hoppe, A. Kirk, and S. Sullivan, “High-quality streamable free- viewpoint video,” ACM TOG, vol. 34, no. 4, pp. 1–13, 2015

  25. [33]

    Dynamicfusion: Recon- struction and tracking of non-rigid scenes in real-time,

    R. A. Newcombe, D. Fox, and S. M. Seitz, “Dynamicfusion: Recon- struction and tracking of non-rigid scenes in real-time,” in CVPR, 2015

  26. [34]

    Rapid avatar capture and simulation using commodity depth sensors,

    A. Feng, A. Shapiro, W. Ruizhe, M. Bolas, G. Medioni, and E. Suma, “Rapid avatar capture and simulation using commodity depth sensors,” in ACM SIGGRAPH 2014 Talks , 2014, pp. 1–1

  27. [35]

    Flyfusion: Realtime dynamic scene reconstruction using a flying depth camera,

    L. Xu, W. Cheng, K. Guo, L. Han, Y . Liu, and L. Fang, “Flyfusion: Realtime dynamic scene reconstruction using a flying depth camera,” TVCG, vol. 27, no. 1, pp. 68–82, 2019

  28. [36]

    Monocular, one-stage, regression of multiple 3d people,

    Y . Sun, Q. Bao, W. Liu, Y . Fu, M. J. Black, and T. Mei, “Monocular, one-stage, regression of multiple 3d people,” in ICCV, 2021

  29. [37]

    Vibe: Video inference for human body pose and shape estimation,

    M. Kocabas, N. Athanasiou, and M. J. Black, “Vibe: Video inference for human body pose and shape estimation,” in CVPR, 2020

  30. [38]

    Expressive body capture: 3d hands, face, and body from a single image,

    G. Pavlakos, V . Choutas, N. Ghorbani, T. Bolkart, A. A. Osman, D. Tzionas, and M. J. Black, “Expressive body capture: 3d hands, face, and body from a single image,” in CVPR, 2019

  31. [39]

    Smpl: A skinned multi-person linear model,

    M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black, “Smpl: A skinned multi-person linear model,” in Seminal Graphics Papers: Pushing the Boundaries, Volume 2 , 2023, pp. 851–866

  32. [40]

    Nerf: Representing scenes as neural radiance fields for view synthesis,

    B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoor- thi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” Communications of the ACM , vol. 65, no. 1, pp. 99–106, 2021

  33. [41]

    Selfrecon: Self reconstruc- tion your digital avatar from monocular video,

    B. Jiang, Y . Hong, H. Bao, and J. Zhang, “Selfrecon: Self reconstruc- tion your digital avatar from monocular video,” in CVPR, 2022

  34. [42]

    Humannerf: Free-viewpoint rendering of moving people from monocular video,

    C.-Y . Weng, B. Curless, P. P. Srinivasan, J. T. Barron, and I. Kemelmacher-Shlizerman, “Humannerf: Free-viewpoint rendering of moving people from monocular video,” in CVPR, 2022

  35. [43]

    Monohuman: Animatable human neural field from monocular video,

    Z. Yu, W. Cheng, X. Liu, W. Wu, and K.-Y . Lin, “Monohuman: Animatable human neural field from monocular video,” in CVPR, 2023

  36. [44]

    Animatable neural radiance fields from monocular rgb videos,

    J. Chen, Y . Zhang, D. Kang, X. Zhe, L. Bao, X. Jia, and H. Lu, “Animatable neural radiance fields from monocular rgb videos,” arXiv preprint arXiv:2106.13629, 2021

  37. [45]

    High-fidelity clothed avatar reconstruction from a single image,

    T. Liao, X. Zhang, Y . Xiu, H. Yi, X. Liu, G.-J. Qi, Y . Zhang, X. Wang, X. Zhu, and Z. Lei, “High-fidelity clothed avatar reconstruction from a single image,” in CVPR, 2023

  38. [46]

    Tech: Text-guided reconstruction of lifelike clothed humans,

    Y . Huang, H. Yi, Y . Xiu, T. Liao, J. Tang, D. Cai, and J. Thies, “Tech: Text-guided reconstruction of lifelike clothed humans,” in 2024 International Conference on 3D Vision (3DV) . IEEE, 2024

  39. [47]

    Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans,

    S. Peng, Y . Zhang, Y . Xu, Q. Wang, Q. Shuai, H. Bao, and X. Zhou, “Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans,” in CVPR, 2021

  40. [48]

    Snarf: Differentiable forward skinning for animating non-rigid neural implicit shapes,

    X. Chen, Y . Zheng, M. J. Black, O. Hilliges, and A. Geiger, “Snarf: Differentiable forward skinning for animating non-rigid neural implicit shapes,” in ICCV, 2021

  41. [49]

    Leap: Learning articulated occupancy of people,

    M. Mihajlovic, Y . Zhang, M. J. Black, and S. Tang, “Leap: Learning articulated occupancy of people,” in CVPR, 2021

  42. [50]

    Arch: Animatable reconstruction of clothed humans,

    Z. Huang, Y . Xu, C. Lassner, H. Li, and T. Tung, “Arch: Animatable reconstruction of clothed humans,” in CVPR, 2020

  43. [51]

    Motion-aware 3d gaussian splatting for efficient dynamic scene reconstruction,

    Z. Guo, W. Zhou, L. Li, M. Wang, and H. Li, “Motion-aware 3d gaussian splatting for efficient dynamic scene reconstruction,” IEEE Transactions on Circuits and Systems for Video Technology , 2024

  44. [52]

    Drivable 3d gaussian avatars,

    W. Zielonka, T. Bagautdinov, S. Saito, M. Zollh ¨ofer, J. Thies, and J. Romero, “Drivable 3d gaussian avatars,” arXiv preprint arXiv:2311.08581, 2023

  45. [53]

    Human gaussian splatting: Real-time rendering of animat- able avatars,

    A. Moreau, J. Song, H. Dhamo, R. Shaw, Y . Zhou, and E. P ´erez- Pellitero, “Human gaussian splatting: Real-time rendering of animat- able avatars,” in CVPR, 2024

  46. [54]

    Haha: Highly articulated gaussian human avatars with textured mesh prior,

    D. Svitov, P. Morerio, L. Agapito, and A. Del Bue, “Haha: Highly articulated gaussian human avatars with textured mesh prior,” arXiv preprint arXiv:2404.01053, 2024

  47. [55]

    Tgavatar: Reconstructing 3d gaussian avatars with transformer-based tri-plane,

    R. Hu, X. Wang, Y . Yan, and C. Zhao, “Tgavatar: Reconstructing 3d gaussian avatars with transformer-based tri-plane,” IEEE Transactions on Circuits and Systems for Video Technology , 2025

  48. [56]

    Gauhuman: Articulated gaussian splatting from monocular human videos,

    S. Hu, T. Hu, and Z. Liu, “Gauhuman: Articulated gaussian splatting from monocular human videos,” in CVPR, 2024

  49. [57]

    Gaussianbody: Clothed human reconstruction via 3d gaussian splatting,

    M. Li, S. Yao, Z. Xie, K. Chen, and Y .-G. Jiang, “Gaussianbody: Clothed human reconstruction via 3d gaussian splatting,”arXiv preprint arXiv:2401.09720, 2024

  50. [58]

    Moss: Motion-based 3d clothed human synthesis from monocular video,

    H. Wang, X. Cai, X. Sun, J. Yue, S. Zhang, F. Lin, and F. Wu, “Moss: Motion-based 3d clothed human synthesis from monocular video,” arXiv preprint arXiv:2405.12806 , 2024

  51. [59]

    Gomavatar: Efficient animatable human modeling from monocular video using gaussians-on-mesh,

    J. Wen, X. Zhao, Z. Ren, A. G. Schwing, and S. Wang, “Gomavatar: Efficient animatable human modeling from monocular video using gaussians-on-mesh,” in CVPR, 2024

  52. [60]

    Splattingavatar: Realistic real-time human avatars with mesh-embedded gaussian splatting,

    Z. Shao, Z. Wang, Z. Li, D. Wang, X. Lin, Y . Zhang, M. Fan, and Z. Wang, “Splattingavatar: Realistic real-time human avatars with mesh-embedded gaussian splatting,” in CVPR, 2024

  53. [61]

    Chase: 3d-consistent human avatars with sparse inputs via gaussian splatting and contrastive learning,

    H. Zhao, H. Wang, C. Yang, and W. Shen, “Chase: 3d-consistent human avatars with sparse inputs via gaussian splatting and contrastive learning,” arXiv preprint arXiv:2408.09663 , 2024

  54. [62]

    Sg-gs: Photo- realistic animatable human avatars with semantically-guided gaussian splatting,

    H. Zhao, C. Yang, H. Wang, X. Zhao, and W. Shen, “Sg-gs: Photo- realistic animatable human avatars with semantically-guided gaussian splatting,” arXiv preprint arXiv:2408.09665 , 2024

  55. [63]

    Blind deconvolution using a normalized sparsity measure,

    D. Krishnan, T. Tay, and R. Fergus, “Blind deconvolution using a normalized sparsity measure,” in CVPR, 2011

  56. [64]

    Single- image blind deblurring using multi-scale latent structure prior,

    Y . Bai, H. Jia, M. Jiang, X. Liu, X. Xie, and W. Gao, “Single- image blind deblurring using multi-scale latent structure prior,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 30, no. 7, pp. 2033–2045, 2019

  57. [65]

    High-quality motion deblurring from a single image,

    Q. Shan, J. Jia, and A. Agarwala, “High-quality motion deblurring from a single image,” ACM TOG, vol. 27, no. 3, pp. 1–10, 2008

  58. [66]

    Non-uniform deblur- ring for shaken images,

    O. Whyte, J. Sivic, A. Zisserman, and J. Ponce, “Non-uniform deblur- ring for shaken images,” in CVPR, 2010

  59. [67]

    Deep multi-scale convolutional neural network for dynamic scene deblurring,

    S. Nah, T. Hyun Kim, and K. Mu Lee, “Deep multi-scale convolutional neural network for dynamic scene deblurring,” in CVPR, 2017

  60. [68]

    Deblurgan: Blind motion deblurring using conditional adversarial networks,

    O. Kupyn, V . Budzan, M. Mykhailych, D. Mishkin, and J. Matas, “Deblurgan: Blind motion deblurring using conditional adversarial networks,” in CVPR, 2018

  61. [69]

    Scale-recurrent network for deep image deblurring,

    X. Tao, H. Gao, X. Shen, J. Wang, and J. Jia, “Scale-recurrent network for deep image deblurring,” in CVPR, 2018

  62. [70]

    Deblurgan-v2: Deblur- ring (orders-of-magnitude) faster and better,

    O. Kupyn, T. Martyniuk, J. Wu, and Z. Wang, “Deblurgan-v2: Deblur- ring (orders-of-magnitude) faster and better,” in ICCV, 2019

  63. [71]

    Human-aware motion deblurring,

    Z. Shen, W. Wang, X. Lu, J. Shen, H. Ling, T. Xu, and L. Shao, “Human-aware motion deblurring,” in ICCV, 2019

  64. [72]

    Restormer: Efficient transformer for high-resolution image restoration,

    S. W. Zamir, A. Arora, S. Khan, M. Hayat, F. S. Khan, and M.- H. Yang, “Restormer: Efficient transformer for high-resolution image restoration,” in CVPR, 2022

  65. [73]

    Strip- former: Strip transformer for fast image deblurring,

    F.-J. Tsai, Y .-T. Peng, Y .-Y . Lin, C.-C. Tsai, and C.-W. Lin, “Strip- former: Strip transformer for fast image deblurring,” in ECCV, 2022

  66. [74]

    Multiscale structure guided diffusion for image deblurring,

    M. Ren, M. Delbracio, H. Talebi, G. Gerig, and P. Milanfar, “Multiscale structure guided diffusion for image deblurring,” in ICCV, 2023

  67. [75]

    Multi-scale residual low-pass filter network for image deblurring,

    J. Dong, J. Pan, Z. Yang, and J. Tang, “Multi-scale residual low-pass filter network for image deblurring,” in ICCV, 2023

  68. [76]

    Generalized video deblurring for dynamic scenes,

    T. Hyun Kim and K. Mu Lee, “Generalized video deblurring for dynamic scenes,” in CVPR, 2015

  69. [77]

    Cascaded deep video deblurring using temporal sharpness prior,

    J. Pan, H. Bai, and J. Tang, “Cascaded deep video deblurring using temporal sharpness prior,” in CVPR, 2020

  70. [78]

    Deep video deblurring for hand-held cameras,

    S. Su, M. Delbracio, J. Wang, G. Sapiro, W. Heidrich, and O. Wang, “Deep video deblurring for hand-held cameras,” in CVPR, 2017

  71. [79]

    Efficient spatio-temporal recurrent neural network for video deblurring,

    Z. Zhong, Y . Gao, Y . Zheng, and B. Zheng, “Efficient spatio-temporal recurrent neural network for video deblurring,” in ECCV, 2020. 12 MANUSCRIPT SUBMITTED TO IEEE TRANS. ON CIRCUIT SYST. VIDEO TECHNOL

  72. [80]

    Mc-blur: A comprehensive benchmark for image deblurring,

    K. Zhang, T. Wang, W. Luo, W. Ren, B. Stenger, W. Liu, H. Li, and M.- H. Yang, “Mc-blur: A comprehensive benchmark for image deblurring,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 34, no. 5, pp. 3755–3767, 2023

  73. [81]

    Spatio-temporal deformable attention network for video deblurring,

    H. Zhang, H. Xie, and H. Yao, “Spatio-temporal deformable attention network for video deblurring,” in ECCV, 2022

  74. [82]

    Multi-attention convolutional neural network for video deblurring,

    X. Zhang, T. Wang, R. Jiang, L. Zhao, and Y . Xu, “Multi-attention convolutional neural network for video deblurring,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 4, pp. 1986– 1997, 2021

  75. [83]

    Efficient video deblurring guided by motion magnitude,

    Y . Wang, Y . Lu, Y . Gao, L. Wang, Z. Zhong, Y . Zheng, and A. Ya- mashita, “Efficient video deblurring guided by motion magnitude,” in ECCV, 2022

  76. [84]

    Deep discriminative spatial and temporal network for efficient video deblurring,

    J. Pan, B. Xu, J. Dong, J. Ge, and J. Tang, “Deep discriminative spatial and temporal network for efficient video deblurring,” in CVPR, 2023

  77. [85]

    Exblurf: Efficient radiance fields for extreme motion blurred images,

    D. Lee, J. Oh, J. Rim, S. Cho, and K. M. Lee, “Exblurf: Efficient radiance fields for extreme motion blurred images,” in ICCV, 2023

  78. [86]

    Deblurring 3d gaussian splatting,

    B. Lee, H. Lee, X. Sun, U. Ali, and E. Park, “Deblurring 3d gaussian splatting,” arXiv preprint arXiv:2401.00834 , 2024

  79. [87]

    Robust gaussian splatting,

    F. Darmon, L. Porzi, S. Rota-Bul `o, and P. Kontschieder, “Robust gaussian splatting,” arXiv preprint arXiv:2404.04211 , 2024

  80. [88]

    Deblur-gs: 3d gaussian splatting from camera motion blurred images,

    W. Chen and L. Liu, “Deblur-gs: 3d gaussian splatting from camera motion blurred images,” Proceedings of the ACM on Computer Graph- ics and Interactive Techniques , vol. 7, no. 1, pp. 1–15, 2024

  81. [89]

    Crim-gs: Continuous rigid motion-aware gaussian splatting from motion blur images,

    J. Lee, D. Kim, D. Lee, S. Cho, and S. Lee, “Crim-gs: Continuous rigid motion-aware gaussian splatting from motion blur images,” arXiv preprint arXiv:2407.03923, 2024

  82. [90]

    Structure-from-motion revisited,

    J. L. Sch ¨onberger and J.-M. Frahm, “Structure-from-motion revisited,” in CVPR, 2016

  83. [91]

    Animating rotation with quaternion curves,

    K. Shoemake, “Animating rotation with quaternion curves,” in Pro- ceedings of the 12th annual conference on Computer graphics and interactive techniques, 1985

  84. [92]

    Adam: A method for stochastic optimization,

    D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” arXiv preprint arXiv:1412.6980 , 2014

  85. [93]

    Davanet: Stereo deblurring with view aggregation,

    S. Zhou, J. Zhang, W. Zuo, H. Xie, J. Pan, and J. S. Ren, “Davanet: Stereo deblurring with view aggregation,” in CVPR, 2019

  86. [94]

    Ntire 2019 challenge on video deblurring and super- resolution: Dataset and study,

    S. Nah, S. Baik, S. Hong, G. Moon, S. Son, R. Timofte, and K. Mu Lee, “Ntire 2019 challenge on video deblurring and super- resolution: Dataset and study,” in CVPRW, 2019

  87. [95]

    Extracting motion and appearance via inter-frame attention for efficient video frame interpolation,

    G. Zhang, Y . Zhu, H. Wang, Y . Chen, G. Wu, and L. Wang, “Extracting motion and appearance via inter-frame attention for efficient video frame interpolation,” in CVPR, 2023

  88. [96]

    Motion capture from internet videos,

    J. Dong, Q. Shuai, Y . Zhang, X. Liu, X. Zhou, and H. Bao, “Motion capture from internet videos,” in ECCV, 2020

  89. [97]

    Learning to reconstruct 3d human pose and shape via model-fitting in the loop,

    N. Kolotouros, G. Pavlakos, M. J. Black, and K. Daniilidis, “Learning to reconstruct 3d human pose and shape via model-fitting in the loop,” in ICCV, 2019

  90. [98]

    Sam 2: Segment anything in images and videos,

    N. Ravi, V . Gabeur, Y .-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. R ¨adle, C. Rolland, L. Gustafson, E. Mintun, J. Pan, K. V . Alwala, N. Carion, C.-Y . Wu, R. Girshick, P. Doll ´ar, and C. Feichtenhofer, “Sam 2: Segment anything in images and videos,” arXiv preprint arXiv:...

  91. [99]

    Amass: Archive of motion capture as surface shapes,

    N. Mahmood, N. Ghorbani, N. F. Troje, G. Pons-Moll, and M. J. Black, “Amass: Archive of motion capture as surface shapes,” in ICCV, 2019

  92. [100]

    Ai choreographer: Music conditioned 3d dance generation with aist++,

    R. Li, S. Yang, D. A. Ross, and A. Kanazawa, “Ai choreographer: Music conditioned 3d dance generation with aist++,” in ICCV, 2021

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.