STREET: A Multi-Task Structured Reasoning and Explanation Benchmark

Danilo Neves Ribeiro; Shen Wang; Xiaofei Ma; Henghui Zhu; Rui Dong; Deguang Kong; Juliette Burger; Anjelica Ramos; William Yang Wang; Zhiheng Huang

STREET: A Multi-Task Structured Reasoning and Explanation Benchmark

Danilo Neves Ribeiro, Shen Wang, Xiaofei Ma, Henghui Zhu, Rui Dong, Deguang Kong, Juliette Burger, Anjelica Ramos, William Yang Wang, Zhiheng Huang

Add to Favorites

1st Workshop on Natural Language Reasoning and Structured Explanations (@ACL 2023) Long Paper

TLDR: We introduce STREET, a unified multi-task and multi-domain natural language reasoning and explanation benchmark. Unlike most existing question-answering (QA) datasets, we expect models to not only answer questions, but also produce step-by-step structured explanations describing how premises in the

RocketChat
Abstract

You can open the #paper-ACL_98 channel in a separate window.

Abstract: We introduce STREET, a unified multi-task and multi-domain natural language reasoning and explanation benchmark. Unlike most existing question-answering (QA) datasets, we expect models to not only answer questions, but also produce step-by-step structured explanations describing how premises in the question are used to produce intermediate conclusions that can prove the correctness of a certain answer. We perform extensive evaluation with popular language models such as few-shot prompting GPT-3 and fine-tuned T5. We find that these models still lag behind human performance when producing such structured reasoning steps. We believe this work will provide a way for the community to better train and test systems on multi-step reasoning and explanations in natural language.