Motion Projection Consistency Based 3D Human Pose Estimation with Virtual Bones from Monocular Videos

28 Jun 2021  ·  Guangming Wang, Honghao Zeng, Ziliang Wang, Zhe Liu, Hesheng Wang ·

Real-time 3D human pose estimation is crucial for human-computer interaction. It is cheap and practical to estimate 3D human pose only from monocular video. However, recent bone splicing based 3D human pose estimation method brings about the problem of cumulative error. In this paper, the concept of virtual bones is proposed to solve such a challenge. The virtual bones are imaginary bones between non-adjacent joints. They do not exist in reality, but they bring new loop constraints for the estimation of 3D human joints. The proposed network in this paper predicts real bones and virtual bones, simultaneously. The final length of real bones is constrained and learned by the loop constructed by the predicted real bones and virtual bones. Besides, the motion constraints of joints in consecutive frames are considered. The consistency between the 2D projected position displacement predicted by the network and the captured real 2D displacement by the camera is proposed as a new projection consistency loss for the learning of 3D human pose. The experiments on the Human3.6M dataset demonstrate the good performance of the proposed method. Ablation studies demonstrate the effectiveness of the proposed inter-frame projection consistency constraints and intra-frame loop constraints.

PDF Abstract

Datasets


Task Dataset Model Metric Name Metric Value Global Rank Result Benchmark
3D Human Pose Estimation Human3.6M Virtual Bones (T=243 CPN) Average MPJPE (mm) 44.8 # 86
Using 2D ground-truth joints No # 1
Multi-View or Monocular Monocular # 1
3D Human Pose Estimation Human3.6M Virtual Bones (T=243 GT) Average MPJPE (mm) 32.5 # 36
Using 2D ground-truth joints Yes # 1
Multi-View or Monocular Monocular # 1
3D Human Pose Estimation Human3.6M Virtual Bones (T=9 GT) Average MPJPE (mm) 35.4 # 45
Using 2D ground-truth joints Yes # 1
Multi-View or Monocular Monocular # 1
3D Human Pose Estimation Human3.6M Virtual Bones (T=9 CPN) Average MPJPE (mm) 47.4 # 106
Using 2D ground-truth joints No # 1
Multi-View or Monocular Monocular # 1

Methods


No methods listed for this paper. Add relevant methods here