AltDiffusion

Introduced by Chen et al. in AltCLIP: Altering the Language Encoder in CLIP for Extended Language Capabilities

In this work, we present a conceptually simple and effective method to train a strong bilingual multimodal representation model. Starting from the pretrained multimodal representation model CLIP released by OpenAI, we switched its text encoder with a pretrained multilingual text encoder XLM-R, and aligned both languages and image representations by a two-stage training schema consisting of teacher learning and contrastive learning. We validate our method through evaluations of a wide range of tasks. We set new state-of-the-art performances on a bunch of tasks including ImageNet-CN, Flicker30k- CN, and COCO-CN. Further, we obtain very close performances with CLIP on almost all tasks, suggesting that one can simply alter the text encoder in CLIP for extended capabilities such as multilingual understanding. Our models and code are available at https://github.com/FlagAI-Open/FlagAI.

Source: AltCLIP: Altering the Language Encoder in CLIP for Extended Language Capabilities

Read Paper See Code

Papers

Paper	Code	Results	Date	Stars

Tasks

Task	Papers	Share
Blocking	1	10.00%
Concept Alignment	1	10.00%
Cross-Modal Retrieval	1	10.00%
Image Classification	1	10.00%
Image Retrieval	1	10.00%
Image-to-Text Retrieval	1	10.00%
Text-to-Image Generation	1	10.00%
XLM-R	1	10.00%
Zero-Shot Cross-Modal Retrieval	1	10.00%

Usage Over Time

This feature is experimental; we are continuously improving our matching algorithm.

Components

Component	Type	Add Remove
AltCLIP	Vision and Language Pre-Trained Models
Diffusion	Image Generation Models

Categories

Add Remove

Image Generation Models