Lifting Monocular Events to 3D Human Poses

21 Apr 2021  ·  Gianluca Scarpellini, Pietro Morerio, Alessio Del Bue ·

This paper presents a novel 3D human pose estimation approach using a single stream of asynchronous events as input. Most of the state-of-the-art approaches solve this task with RGB cameras, however struggling when subjects are moving fast. On the other hand, event-based 3D pose estimation benefits from the advantages of event-cameras, especially their efficiency and robustness to appearance changes. Yet, finding human poses in asynchronous events is in general more challenging than standard RGB pose estimation, since little or no events are triggered in static scenes. Here we propose the first learning-based method for 3D human pose from a single stream of events. Our method consists of two steps. First, we process the event-camera stream to predict three orthogonal heatmaps per joint; each heatmap is the projection of of the joint onto one orthogonal plane. Next, we fuse the sets of heatmaps to estimate 3D localisation of the body joints. As a further contribution, we make available a new, challenging dataset for event-based human pose estimation by simulating events from the RGB Human3.6m dataset. Experiments demonstrate that our method achieves solid accuracy, narrowing the performance gap between standard RGB and event-based vision. The code is freely available at

PDF Abstract


Introduced in the Paper:


Used in the Paper:

ImageNet Human3.6M DHP19

Results from the Paper

Task Dataset Model Metric Name Metric Value Global Rank Result Benchmark
3D Human Pose Estimation DHP19 Lifting Events MPJPE3D 92.09 # 2
3D Human Pose Estimation Human3.6M Events Average MPJPE (mm) 116.4 # 309