Search Results for author: Julian Schulz

Found 1 papers, 1 papers with code

Steering Llama 2 via Contrastive Activation Addition

3 code implementations9 Dec 2023 Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, Alexander Matt Turner

We introduce Contrastive Activation Addition (CAA), an innovative method for steering language models by modifying their activations during forward passes.

Multiple-choice

Cannot find the paper you are looking for? You can Submit a new open access paper.