Kevin Meng

Twitter· GitHub· Scholar· LinkedIn

About

Selected Work

Mass-Editing Memory in a Transformer

ICLR 2023 · with Arnab Sen Sharma, Alex Andonian, Yonatan Belinkov, and David Bau

MEMIT is an algorithm that updates a language model with thousands of new memories at once, scaling to GPT-J (6B) and GPT-NeoX (20B), orders of magnitude beyond prior work.

Locating and Editing Factual Associations in GPT

NeurIPS 2022 · with David Bau, Alex Andonian, and Yonatan Belinkov

We developed causal tracing, a technique to locate where GPT models store facts, and ROME, a method for editing facts directly into model weights.

Studying the Approximate Linearity of Apple's NeuralHash

ICML 2022 Workshop on ML for Cybersecurity · with Jagdeep Bhatia

We find that Apple's NeuralHash is approximately linear. We use that insight to design black-box attacks that evade detection, generate near-collisions, and leak information about hashed images.

Adversarial Training on Transformer Networks

ACM Transactions · with Damian Jimenez, Chengkai Li, and others

We adversarially trained a claim-spotting model and deployed it on ClaimBuster, a fact-checking platform used by thousands worldwide.