Search Results for author: Kunyu Peng

Found 39 papers, 32 papers with code

RoDLA: Benchmarking the Robustness of Document Layout Analysis Models

no code implementations • 21 Mar 2024 • Yufan Chen, Jiaming Zhang, Kunyu Peng, Junwei Zheng, Ruiping Liu, Philip Torr, Rainer Stiefelhagen

To address this, we are the first to introduce a robustness benchmark for DLA models, which includes 450K document images of three datasets.

Benchmarking Document Layout Analysis

Paper
Add Code

Skeleton-Based Human Action Recognition with Noisy Labels

1 code implementation • 15 Mar 2024 • Yi Xu, Kunyu Peng, Di Wen, Ruiping Liu, Junwei Zheng, Yufan Chen, Jiaming Zhang, Alina Roitberg, Kailun Yang, Rainer Stiefelhagen

In this study, we bridge this gap by implementing a framework that augments well-established skeleton-based human action recognition methods with label-denoising strategies from various research areas to serve as the initial benchmark.

Action Recognition Denoising +3

Paper
Code

EchoTrack: Auditory Referring Multi-Object Tracking for Autonomous Driving

1 code implementation • 28 Feb 2024 • Jiacheng Lin, Jiajun Chen, Kunyu Peng, Xuan He, Zhiyong Li, Rainer Stiefelhagen, Kailun Yang

This paper introduces the task of Auditory Referring Multi-Object Tracking (AR-MOT), which dynamically tracks specific objects in a video sequence based on audio expressions and appears as a challenging problem in autonomous driving.

Autonomous Driving Multi-Object Tracking +1

Paper
Code

LF Tracy: A Unified Single-Pipeline Approach for Salient Object Detection in Light Field Cameras

1 code implementation • 30 Jan 2024 • Fei Teng, Jiaming Zhang, Jiawei Liu, Kunyu Peng, Xina Cheng, Zhiyong Li, Kailun Yang

Previous approaches predominantly employ a custom two-stream design to discover the implicit angular feature within light field cameras, leading to significant information isolation between different LF representations.

Data Augmentation object-detection +2

Paper
Code

Fourier Prompt Tuning for Modality-Incomplete Scene Segmentation

1 code implementation • 30 Jan 2024 • Ruiping Liu, Jiaming Zhang, Kunyu Peng, Yufan Chen, Ke Cao, Junwei Zheng, M. Saquib Sarfraz, Kailun Yang, Rainer Stiefelhagen

Integrating information from multiple modalities enhances the robustness of scene perception systems in autonomous vehicles, providing a more comprehensive and reliable sensory framework.

Autonomous Vehicles Scene Segmentation

Paper
Code

Navigating Open Set Scenarios for Skeleton-based Action Recognition

1 code implementation • 11 Dec 2023 • Kunyu Peng, Cheng Yin, Junwei Zheng, Ruiping Liu, David Schneider, Jiaming Zhang, Kailun Yang, M. Saquib Sarfraz, Rainer Stiefelhagen, Alina Roitberg

In real-world scenarios, human actions often fall outside the distribution of training data, making it crucial for models to recognize known actions and reject unknown ones.

Novelty Detection Open Set Action Recognition +3

Paper
Code

Quantized Distillation: Optimizing Driver Activity Recognition Models for Resource-Constrained Environments

1 code implementation • 10 Nov 2023 • Calvin Tanama, Kunyu Peng, Zdravko Marinov, Rainer Stiefelhagen, Alina Roitberg

The framework enhances 3D MobileNet, a neural architecture optimized for speed in video classification, by incorporating knowledge distillation and model quantization to balance model accuracy and computational efficiency.

Activity Recognition Autonomous Driving +4

Paper
Code

Unveiling the Hidden Realm: Self-supervised Skeleton-based Action Recognition in Occluded Environments

1 code implementation • 21 Sep 2023 • Yifei Chen, Kunyu Peng, Alina Roitberg, David Schneider, Jiaming Zhang, Junwei Zheng, Ruiping Liu, Yufan Chen, Kailun Yang, Rainer Stiefelhagen

To integrate action recognition methods into autonomous robotic systems, it is crucial to consider adverse situations involving target occlusions.

Action Recognition Imputation +1

Paper
Code

Elevating Skeleton-Based Action Recognition with Efficient Multi-Modality Self-Supervision

1 code implementation • 21 Sep 2023 • Yiping Wei, Kunyu Peng, Alina Roitberg, Jiaming Zhang, Junwei Zheng, Ruiping Liu, Yufan Chen, Kailun Yang, Rainer Stiefelhagen

These works overlooked the differences in performance among modalities, which led to the propagation of erroneous knowledge between modalities while only three fundamental modalities, i. e., joints, bones, and motions are used, hence no additional modalities are explored.

Action Recognition Knowledge Distillation +3

Paper
Code

Towards Privacy-Supporting Fall Detection via Deep Unsupervised RGB2Depth Adaptation

1 code implementation • 23 Aug 2023 • Hejun Xiao, Kunyu Peng, Xiangsheng Huang, Alina Roitberg1, Hao Li, Zhaohui Wang, Rainer Stiefelhagen

In this paper, we introduce a privacy-supporting solution that makes the RGB-trained model applicable in depth domain and utilizes depth data at test time for fall detection.

Domain Adaptation

Paper
Code

OAFuser: Towards Omni-Aperture Fusion for Light Field Semantic Segmentation

2 code implementations • 28 Jul 2023 • Fei Teng, Jiaming Zhang, Kunyu Peng, Yaonan Wang, Rainer Stiefelhagen, Kailun Yang

To avoid feature loss during network propagation and simultaneously streamline the redundant information from the light field camera, we present a simple yet very effective Sub-Aperture Fusion Module (SAFM) to embed sub-aperture images into angular features without any additional memory cost.

Autonomous Driving Scene Understanding +1

Paper
Code

Tightly-Coupled LiDAR-Visual SLAM Based on Geometric Features for Mobile Agents

no code implementations • 15 Jul 2023 • Ke Cao, Ruiping Liu, Ze Wang, Kunyu Peng, Jiaming Zhang, Junwei Zheng, Zhifeng Teng, Kailun Yang, Rainer Stiefelhagen

On the other hand, the entire line segment detected by the visual subsystem overcomes the limitation of the LiDAR subsystem, which can only perform the local calculation for geometric features.

Autonomous Navigation Pose Estimation +2

Paper
Add Code

Open Scene Understanding: Grounded Situation Recognition Meets Segment Anything for Helping People with Visual Impairments

1 code implementation • 15 Jul 2023 • Ruiping Liu, Jiaming Zhang, Kunyu Peng, Junwei Zheng, Ke Cao, Yufan Chen, Kailun Yang, Rainer Stiefelhagen

Grounded Situation Recognition (GSR) is capable of recognizing and interpreting visual scenes in a contextually intuitive way, yielding salient activities (verbs) and the involved entities (roles) depicted in images.

Grounded Situation Recognition Navigate +1

Paper
Code

RelaMiX: Exploring Few-Shot Adaptation in Video-based Action Recognition

2 code implementations • 15 May 2023 • Kunyu Peng, Di Wen, David Schneider, Jiaming Zhang, Kailun Yang, M. Saquib Sarfraz, Rainer Stiefelhagen, Alina Roitberg

Domain adaptation is essential for activity recognition to ensure accurate and robust performance across diverse environments, sensor types, and data sources.

Action Recognition Unsupervised Domain Adaptation

Paper
Code

FishDreamer: Towards Fisheye Semantic Completion via Unified Image Outpainting and Segmentation

1 code implementation • 24 Mar 2023 • Hao Shi, Yu Li, Kailun Yang, Jiaming Zhang, Kunyu Peng, Alina Roitberg, Yaozu Ye, Huajian Ni, Kaiwei Wang, Rainer Stiefelhagen

This paper raises the new task of Fisheye Semantic Completion (FSC), where dense texture, structure, and semantics of a fisheye image are inferred even beyond the sensor field-of-view (FoV).

Image Outpainting Semantic Segmentation

Paper
Code

360BEV: Panoramic Semantic Mapping for Indoor Bird's-Eye View

1 code implementation • 21 Mar 2023 • Zhifeng Teng, Jiaming Zhang, Kailun Yang, Kunyu Peng, Hao Shi, Simon Reiß, Ke Cao, Rainer Stiefelhagen

Seeing only a tiny part of the whole is not knowing the full circumstance.

Semantic Segmentation

Paper
Code

MuscleMap: Towards Video-based Activated Muscle Group Estimation in the Wild

1 code implementation • 2 Mar 2023 • Kunyu Peng, David Schneider, Alina Roitberg, Kailun Yang, Jiaming Zhang, Chen Deng, Kaiyu Zhang, M. Saquib Sarfraz, Rainer Stiefelhagen

In this paper, we tackle the new task of video-based Activated Muscle Group Estimation (AMGE) aiming at identifying active muscle regions during physical activity in the wild.

Human Activity Recognition Knowledge Distillation +1

Paper
Code

Delivering Arbitrary-Modal Semantic Segmentation

1 code implementation • CVPR 2023 • Jiaming Zhang, Ruiping Liu, Hao Shi, Kailun Yang, Simon Reiß, Kunyu Peng, Haodong Fu, Kaiwei Wang, Rainer Stiefelhagen

To make this possible, we present the arbitrary cross-modal segmentation model CMNeXt.

Ranked #1 on Semantic Segmentation on DSEC

Segmentation Semantic Segmentation +1

123

Paper
Code

MateRobot: Material Recognition in Wearable Robotics for People with Visual Impairments

1 code implementation • 28 Feb 2023 • Junwei Zheng, Jiaming Zhang, Kailun Yang, Kunyu Peng, Rainer Stiefelhagen

People with Visual Impairments (PVI) typically recognize objects through haptic perception.

Material Recognition Semantic Segmentation

Paper
Code

Behind Every Domain There is a Shift: Adapting Distortion-aware Vision Transformers for Panoramic Semantic Segmentation

1 code implementation • 25 Jul 2022 • Jiaming Zhang, Kailun Yang, Hao Shi, Simon Reiß, Kunyu Peng, Chaoxiang Ma, Haodong Fu, Philip H. S. Torr, Kaiwei Wang, Rainer Stiefelhagen

In this paper, we address panoramic semantic segmentation which is under-explored due to two critical challenges: (1) image distortions and object deformations on panoramas; (2) lack of semantic annotations in the 360-degree imagery.

Ranked #1 on Semantic Segmentation on SynPASS

Pseudo Label Segmentation +2

Paper
Code

Multi-modal Depression Estimation based on Sub-attentional Fusion

1 code implementation • 13 Jul 2022 • Ping-Cheng Wei, Kunyu Peng, Alina Roitberg, Kailun Yang, Jiaming Zhang, Rainer Stiefelhagen

Failure to timely diagnose and effectively treat depression leads to over 280 million people suffering from this psychological disorder worldwide.

Paper
Code

Trans4Map: Revisiting Holistic Bird's-Eye-View Mapping from Egocentric Images to Allocentric Semantics with Vision Transformers

1 code implementation • 13 Jul 2022 • Chang Chen, Jiaming Zhang, Kailun Yang, Kunyu Peng, Rainer Stiefelhagen

Humans have an innate ability to sense their surroundings, as they can extract the spatial representation from the egocentric perception and form an allocentric semantic map via spatial transformation and memory updating.

Semantic Segmentation

Paper
Code

Is my Driver Observation Model Overconfident? Input-guided Calibration Networks for Reliable and Interpretable Confidence Estimates

no code implementations • 10 Apr 2022 • Alina Roitberg, Kunyu Peng, David Schneider, Kailun Yang, Marios Koulakis, Manuel Martinez, Rainer Stiefelhagen

In this work, we for the first time examine how well the confidence values of modern driver observation models indeed match the probability of the correct outcome and show that raw neural network-based approaches tend to significantly overestimate their prediction quality.

Action Recognition Image Classification

Paper
Add Code

A Comparative Analysis of Decision-Level Fusion for Multimodal Driver Behaviour Understanding

no code implementations • 10 Apr 2022 • Alina Roitberg, Kunyu Peng, Zdravko Marinov, Constantin Seibold, David Schneider, Rainer Stiefelhagen

Visual recognition inside the vehicle cabin leads to safer driving and more intuitive human-vehicle interaction but such systems face substantial obstacles as they need to capture different granularities of driver behaviour while dealing with highly limited body visibility and changing illumination.

Paper
Add Code

Indoor Navigation Assistance for Visually Impaired People via Dynamic SLAM and Panoptic Segmentation with an RGB-D Sensor

no code implementations • 3 Apr 2022 • Wenyan Ou, Jiaming Zhang, Kunyu Peng, Kailun Yang, Gerhard Jaworek, Karin Müller, Rainer Stiefelhagen

Then, poses and speed of tracked dynamic objects can be estimated, which are passed to the users through acoustic feedback.

Motion Estimation Object +1

Paper
Add Code

Towards Robust Semantic Segmentation of Accident Scenes via Multi-Source Mixed Sampling and Meta-Learning

1 code implementation • 19 Mar 2022 • Xinyu Luo, Jiaming Zhang, Kailun Yang, Alina Roitberg, Kunyu Peng, Rainer Stiefelhagen

Autonomous vehicles utilize urban scene segmentation to understand the real world like a human and react accordingly.

Ranked #1 on Semantic Segmentation on DADA-seg (using extra training data)

Autonomous Vehicles Meta-Learning +3

Paper
Code

MatchFormer: Interleaving Attention in Transformers for Feature Matching

1 code implementation • 17 Mar 2022 • Qing Wang, Jiaming Zhang, Kailun Yang, Kunyu Peng, Rainer Stiefelhagen

While detector-based methods coupled with feature descriptors struggle in low-texture scenes, CNN-based methods with a sequential extract-to-match pipeline, fail to make use of the matching capacity of the encoder and tend to overburden the decoder for matching.

Homography Estimation Pose Estimation +1

158

Paper
Code

Bending Reality: Distortion-aware Transformers for Adapting to Panoramic Semantic Segmentation

1 code implementation • CVPR 2022 • Jiaming Zhang, Kailun Yang, Chaoxiang Ma, Simon Reiß, Kunyu Peng, Rainer Stiefelhagen

To get around this domain difference and bring together semantic annotations from pinhole- and 360-degree surround-visuals, we propose to learn object deformations and panoramic image distortions in the Deformable Patch Embedding (DPE) and Deformable MLP (DMLP) components which blend into our Transformer for PAnoramic Semantic Segmentation (Trans4PASS) model.

Ranked #2 on Semantic Segmentation on SynPASS

Scene Understanding Semantic Segmentation +1

Paper
Code

TransDARC: Transformer-based Driver Activity Recognition with Latent Space Feature Calibration

1 code implementation • 2 Mar 2022 • Kunyu Peng, Alina Roitberg, Kailun Yang, Jiaming Zhang, Rainer Stiefelhagen

This module operates in the latent feature-space enriching and diversifying the training set at feature-level in order to improve generalization to novel data appearances, (e. g., sensor changes) and general feature quality.

Human Activity Recognition

Paper
Code

TransKD: Transformer Knowledge Distillation for Efficient Semantic Segmentation

2 code implementations • 27 Feb 2022 • Ruiping Liu, Kailun Yang, Alina Roitberg, Jiaming Zhang, Kunyu Peng, Huayao Liu, Yaonan Wang, Rainer Stiefelhagen

Semantic segmentation benchmarks in the realm of autonomous driving are dominated by large pre-trained transformers, yet their widespread adoption is impeded by substantial computational costs and prolonged training durations.

Autonomous Driving Knowledge Distillation +3

Paper
Code

Delving Deep into One-Shot Skeleton-based Action Recognition with Diverse Occlusions

2 code implementations • 23 Feb 2022 • Kunyu Peng, Alina Roitberg, Kailun Yang, Jiaming Zhang, Rainer Stiefelhagen

Yet, the research of data-scarce recognition from skeleton sequences, such as one-shot action recognition, does not explicitly consider occlusions despite their everyday pervasiveness.

Ranked #1 on Action Classification on Toyota Smarthome dataset (Accuracy metric)

Action Classification Action Recognition +2

Paper
Code

Should I take a walk? Estimating Energy Expenditure from Video Data

1 code implementation • 1 Feb 2022 • Kunyu Peng, Alina Roitberg, Kailun Yang, Jiaming Zhang, Rainer Stiefelhagen

To study this underresearched task, we introduce Vid2Burn -- an omni-source benchmark for estimating caloric expenditure from video data featuring both, high- and low-intensity activities for which we derive energy expenditure annotations based on models established in medical literature.

Video Recognition

Paper
Code

Affect-DML: Context-Aware One-Shot Recognition of Human Affect using Deep Metric Learning

1 code implementation • 30 Nov 2021 • Kunyu Peng, Alina Roitberg, David Schneider, Marios Koulakis, Kailun Yang, Rainer Stiefelhagen

Human affect recognition is a well-established research area with numerous applications, e. g., in psychological care, but existing methods assume that all emotions-of-interest are given a priori as annotated training examples.

Emotion Recognition Metric Learning +1

Paper
Code

Transfer beyond the Field of View: Dense Panoramic Semantic Segmentation via Unsupervised Domain Adaptation

1 code implementation • 21 Oct 2021 • Jiaming Zhang, Chaoxiang Ma, Kailun Yang, Alina Roitberg, Kunyu Peng, Rainer Stiefelhagen

We look at this problem from the perspective of domain adaptation and bring panoramic semantic segmentation to a setting, where labelled training data originates from a different distribution of conventional pinhole camera images.

Ranked #7 on Semantic Segmentation on DensePASS (using extra training data)

Autonomous Vehicles Segmentation +2

Paper
Code

Trans4Trans: Efficient Transformer for Transparent Object and Semantic Scene Segmentation in Real-World Navigation Assistance

1 code implementation • 20 Aug 2021 • Jiaming Zhang, Kailun Yang, Angela Constantinescu, Kunyu Peng, Karin Müller, Rainer Stiefelhagen

In this paper, we build a wearable system with a novel dual-head Transformer for Transparency (Trans4Trans) perception model, which can segment general- and transparent objects.

Ranked #2 on Semantic Segmentation on DADA-seg (using extra training data)

Navigate Scene Segmentation +1

Paper
Code

Trans4Trans: Efficient Transformer for Transparent Object Segmentation to Help Visually Impaired People Navigate in the Real World

1 code implementation • 7 Jul 2021 • Jiaming Zhang, Kailun Yang, Angela Constantinescu, Kunyu Peng, Karin Müller, Rainer Stiefelhagen

Common fully glazed facades and transparent objects present architectural barriers and impede the mobility of people with low vision or blindness, for instance, a path detected behind a glass door is inaccessible unless it is correctly perceived and reacted.

Ranked #1 on Semantic Segmentation on Trans10K

Navigate Semantic Segmentation +1

Paper
Code

HIDA: Towards Holistic Indoor Understanding for the Visually Impaired via Semantic Instance Segmentation with a Wearable Solid-State LiDAR Sensor

no code implementations • 7 Jul 2021 • Huayao Liu, Ruiping Liu, Kailun Yang, Jiaming Zhang, Kunyu Peng, Rainer Stiefelhagen

To tackle these issues, we propose HIDA, a lightweight assistive system based on 3D point cloud instance segmentation with a solid-state LiDAR sensor, for holistic indoor detection and avoidance.

Ranked #17 on 3D Instance Segmentation on ScanNet(v2)

3D Instance Segmentation Point Cloud Segmentation +2

Paper
Add Code

MASS: Multi-Attentional Semantic Segmentation of LiDAR Data for Dense Top-View Understanding

1 code implementation • 1 Jul 2021 • Kunyu Peng, Juncong Fei, Kailun Yang, Alina Roitberg, Jiaming Zhang, Frank Bieder, Philipp Heidenreich, Christoph Stiller, Rainer Stiefelhagen

At the heart of all automated driving systems is the ability to sense the surroundings, e. g., through semantic segmentation of LiDAR sequences, which experienced a remarkable progress due to the release of large datasets such as SemanticKITTI and nuScenes-LidarSeg.

3D Object Detection Graph Attention +4

Paper
Code

PillarSegNet: Pillar-based Semantic Grid Map Estimation using Sparse LiDAR Data

no code implementations • 10 May 2021 • Juncong Fei, Kunyu Peng, Philipp Heidenreich, Frank Bieder, Christoph Stiller

The recent publication of the SemanticKITTI dataset stimulates the research on semantic segmentation of LiDAR point clouds in urban scenarios.

2D Semantic Segmentation Segmentation +1

Paper
Add Code

Cannot find the paper you are looking for? You can Submit a new open access paper.