BioGPT: Generative Pre-trained Transformer for Biomedical Text Generation and Mining

microsoft/biogpt 19 Oct 2022

Pre-trained language models have attracted increasing attention in the biomedical domain, inspired by their great success in the general natural language domain.

Document Classification Language Modelling +3

1,261
3.64 stars / hour

Multimodal Chain-of-Thought Reasoning in Language Models

amazon-science/mm-cot 2 Feb 2023

By incorporating the vision features in both stages, the model is able to generate effective rationales that contribute to answer inference.

253
2.56 stars / hour

Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-Image Diffusion Models

AttendAndExcite/Attend-and-Excite 31 Jan 2023

Recent text-to-image generative models have demonstrated an unparalleled ability to generate diverse and creative imagery guided by a target text prompt.

Generative Semantic Nursing

216
1.43 stars / hour

BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

salesforce/lavis 30 Jan 2023

The cost of vision-and-language pre-training has become increasingly prohibitive due to end-to-end training of large-scale models.

Image Captioning Image Retrieval +5

1,875
0.79 stars / hour

Open Source Vizier: Distributed Infrastructure and API for Reliable and Flexible Blackbox Optimization

google/vizier 27 Jul 2022

Vizier is the de-facto blackbox and hyperparameter optimization service across Google, having optimized some of Google's largest products and research efforts.

Hyperparameter Optimization Transfer Learning

820
0.74 stars / hour

STEPS: Joint Self-supervised Nighttime Image Enhancement and Depth Estimation

ucaszyp/steps 2 Feb 2023

By fitting a bridge-shaped curve to the illumination map distribution, both regions are suppressed and two tasks are bridged naturally.

Depth Estimation Image Enhancement

80
0.74 stars / hour

Learning the Beauty in Songs: Neural Singing Voice Beautifier

MoonInTheRiver/DiffSinger ACL 2022

Furthermore, we propose a latent-mapping algorithm in the latent space to convert the amateur vocal tone to the professional one.

Dynamic Time Warping

1,900
0.58 stars / hour

NaturalSpeech: End-to-End Text to Speech Synthesis with Human-Level Quality

heatz123/naturalspeech 9 May 2022

In this paper, we answer these questions by first defining the human-level quality based on the statistical significance of subjective measure and introducing appropriate guidelines to judge it, and then developing a TTS system called NaturalSpeech that achieves human-level quality on a benchmark dataset.

Speech Synthesis Text-To-Speech Synthesis

74
0.57 stars / hour

DAMO-YOLO : A Report on Real-Time Object Detection Design

tinyvision/damo-yolo 23 Nov 2022

In this report, we present a fast and accurate object detection method dubbed DAMO-YOLO, which achieves higher performance than the state-of-the-art YOLO series.

Neural Architecture Search object-detection +1

1,321
0.55 stars / hour

InstructPix2Pix: Learning to Follow Image Editing Instructions

timothybrooks/instruct-pix2pix 17 Nov 2022

We propose a method for editing images from human instructions: given an input image and a written instruction that tells the model what to do, our model follows these instructions to edit the image.

Language Modelling Text-based Image Editing +1

3,182
0.53 stars / hour