LexSym: Compositionality as Lexical Symmetry

Ekin Akyurek; Jacob Andreas

LexSym: Compositionality as Lexical Symmetry

Ekin Akyurek, Jacob Andreas

📝 Paper

Anthology

Underline 📺 Watch Video on Underline Add to Favorites

Main: Semantics: Lexical Main-oral Paper

Session 5: Semantics: Lexical (Oral)

Conference Room: Pier 2&3

Conference Time: July 11, 16:15-17:45 (EDT) (America/Toronto)

Global Time: July 11, Session 5 (20:15-21:45 UTC)

Keywords: compositionality

TLDR: In tasks like semantic parsing, instruction following, and question answering, standard deep networks fail to generalize compositionally from small datasets. Many existing approaches overcome this limitation with model architectures that enforce a compositional process of sentence interpretation. In...

You can open the #paper-P2367 channel in a separate window.

Abstract: In tasks like semantic parsing, instruction following, and question answering, standard deep networks fail to generalize compositionally from small datasets. Many existing approaches overcome this limitation with model architectures that enforce a compositional process of sentence interpretation. In this paper, we present a domain-general and model-agnostic formulation of compositionality as a constraint on symmetries of data distributions rather than models. Informally, we prove that whenever a task can be solved by a compositional model, there is a corresponding data augmentation scheme --- a procedure for transforming examples into other well-formed examples --- that imparts compositional inductive bias on any model trained to solve the same task. We describe a procedure called LexSym that discovers these transformations automatically, then applies them to training data for ordinary neural sequence models. Unlike existing compositional data augmentation procedures, LexSym can be deployed agnostically across text, structured data, and even images. It matches or surpasses state-of-the-art, task-specific models on COGS semantic parsing, SCAN and Alchemy instruction following, and CLEVR-CoGenT visual question answering datasets.