Search Results for author: Yansong Tang

Found 23 papers, 13 papers with code

Self-similarity-based super-resolution of photoacoustic angiography from hand-drawn doodles

1 code implementation2 May 2023 Yuanzheng Ma, Wangting Zhou, Rui Ma, Sihua Yang, Yansong Tang, Xun Guan

To address this challenge, we propose a novel approach that employs a super-resolution PAA method trained with forged PAA images.

Image Generation Super-Resolution +1

Global Knowledge Calibration for Fast Open-Vocabulary Segmentation

no code implementations16 Mar 2023 Kunyang Han, Yong liu, Jun Hao Liew, Henghui Ding, Yunchao Wei, Jiajun Liu, Yitong Wang, Yansong Tang, Yujiu Yang, Jiashi Feng, Yao Zhao

Recent advancements in pre-trained vision-language models, such as CLIP, have enabled the segmentation of arbitrary concepts solely from textual inputs, a process commonly referred to as open-vocabulary semantic segmentation (OVS).

Knowledge Distillation Open Vocabulary Semantic Segmentation +3

FLAG3D: A 3D Fitness Activity Dataset with Language Instruction

no code implementations CVPR 2023 Yansong Tang, Jinpeng Liu, Aoyang Liu, Bin Yang, Wenxun Dai, Yongming Rao, Jiwen Lu, Jie zhou, Xiu Li

With the continuously thriving popularity around the world, fitness activity analytic has become an emerging research topic in computer vision.

Action Generation Action Recognition +2

Global Spectral Filter Memory Network for Video Object Segmentation

1 code implementation11 Oct 2022 Yong liu, Ran Yu, Jiahao Wang, Xinyuan Zhao, Yitong Wang, Yansong Tang, Yujiu Yang

Besides, we empirically find low frequency feature should be enhanced in encoder (backbone) while high frequency for decoder (segmentation head).

Semantic Segmentation Semi-Supervised Video Object Segmentation +1

HorNet: Efficient High-Order Spatial Interactions with Recursive Gated Convolutions

4 code implementations28 Jul 2022 Yongming Rao, Wenliang Zhao, Yansong Tang, Jie zhou, Ser-Nam Lim, Jiwen Lu

In this paper, we show that the key ingredients behind the vision Transformers, namely input-adaptive, long-range and high-order spatial interactions, can also be efficiently implemented with a convolution-based framework.

Image Classification Object Detection +2

Learning from Temporal Spatial Cubism for Cross-Dataset Skeleton-based Action Recognition

1 code implementation17 Jul 2022 Yansong Tang, Xingyu Liu, Xumin Yu, Danyang Zhang, Jiwen Lu, Jie zhou

Different from the conventional adversarial learning-based approaches for UDA, we utilize a self-supervision scheme to reduce the domain shift between two skeleton-based action datasets.

Action Recognition Self-Supervised Learning +2

ScalableViT: Rethinking the Context-oriented Generalization of Vision Transformer

2 code implementations21 Mar 2022 Rui Yang, Hailong Ma, Jie Wu, Yansong Tang, Xuefeng Xiao, Min Zheng, Xiu Li

The vanilla self-attention mechanism inherently relies on pre-defined and steadfast computational dimensions.

Semantic-Aware Auto-Encoders for Self-Supervised Representation Learning

1 code implementation CVPR 2022 Guangrun Wang, Yansong Tang, Liang Lin, Philip H.S. Torr

Inspired by perceptual learning that could use cross-view learning to perceive concepts and semantics, we propose a novel AE that could learn semantic-aware representation via cross-view image reconstruction.

Image Reconstruction Representation Learning +1

LAVT: Language-Aware Vision Transformer for Referring Image Segmentation

1 code implementation CVPR 2022 Zhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen, Hengshuang Zhao, Philip H. S. Torr

Referring image segmentation is a fundamental vision-language task that aims to segment out an object referred to by a natural language expression from an image.

Image Segmentation Referring Expression +2

Unsupervised Embedding Learning from Uncertainty Momentum Modeling

no code implementations19 Jul 2021 Jiahuan Zhou, Yansong Tang, Bing Su, Ying Wu

We justify that the performance limitation is caused by the gradient vanishing on these sample outliers.

Comprehensive Instructional Video Analysis: The COIN Dataset and Performance Evaluation

no code implementations20 Mar 2020 Yansong Tang, Jiwen Lu, Jie zhou

We believe the introduction of the COIN dataset will promote the future in-depth research on instructional video analysis for the community.

Action Detection

COIN: A Large-scale Dataset for Comprehensive Instructional Video Analysis

no code implementations CVPR 2019 Yansong Tang, Dajun Ding, Yongming Rao, Yu Zheng, Danyang Zhang, Lili Zhao, Jiwen Lu, Jie zhou

There are substantial instructional videos on the Internet, which enables us to acquire knowledge for completing various tasks.

Action Detection

Deep Progressive Reinforcement Learning for Skeleton-Based Action Recognition

no code implementations CVPR 2018 Yansong Tang, Yi Tian, Jiwen Lu, Peiyang Li, Jie zhou

In this paper, we propose a deep progressive reinforcement learning (DPRL) method for action recognition in skeleton-based videos, which aims to distil the most informative frames and discard ambiguous frames in sequences for recognizing actions.

Action Recognition reinforcement-learning +3

Cannot find the paper you are looking for? You can Submit a new open access paper.