FOLIO: Natural Language Reasoning with First-Order Logic

Han, Simeng; Schoelkopf, Hailey; Zhao, Yilun; Qi, Zhenting; Riddell, Martin; Zhou, Wenfei; Coady, James; Peng, David; Qiao, Yujie; Benson, Luke; Sun, Lucy; Wardle-Solano, Alex; Szabo, Hannah; Zubova, Ekaterina; Burtell, Matthew; Fan, Jonathan; Liu, Yixin; Wong, Brian; Sailor, Malcolm; Ni, Ansong; Nan, Linyong; Kasai, Jungo; Yu, Tao; Zhang, Rui; Fabbri, Alexander R.; Kryscinski, Wojciech; Yavuz, Semih; Liu, Ye; Lin, Xi Victoria; Joty, Shafiq; Zhou, Yingbo; Xiong, Caiming; Ying, Rex; Cohan, Arman; Radev, Dragomir

Computer Science > Computation and Language

arXiv:2209.00840 (cs)

[Submitted on 2 Sep 2022 (v1), last revised 17 May 2024 (this version, v2)]

Title:FOLIO: Natural Language Reasoning with First-Order Logic

Abstract:Large language models (LLMs) have achieved remarkable performance on a variety of natural language understanding tasks. However, existing benchmarks are inadequate in measuring the complex logical reasoning capabilities of a model. We present FOLIO, a human-annotated, logically complex and diverse dataset for reasoning in natural language (NL), equipped with first-order logic (FOL) annotations. FOLIO consists of 1,430 examples (unique conclusions), each paired with one of 487 sets of premises used to deductively reason for the validity of each conclusion. The logical correctness of the premises and conclusions is ensured by their FOL annotations, which are automatically verified by an FOL inference engine. In addition to the main NL reasoning task, NL-FOL pairs in FOLIO constitute a new NL-FOL translation dataset. Our experiments on FOLIO systematically evaluate the FOL reasoning ability of supervised fine-tuning on medium-sized language models. For both NL reasoning and NL-FOL translation, we benchmark multiple state-of-the-art language models. Our results show that a subset of FOLIO presents a challenge for one of the most capable {Large Language Model (LLM)} publicly available, GPT-4.

Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2209.00840 [cs.CL]
	(or arXiv:2209.00840v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2209.00840

Submission history

From: Simeng Han [view email]
[v1] Fri, 2 Sep 2022 06:50:11 UTC (7,677 KB)
[v2] Fri, 17 May 2024 15:06:25 UTC (7,375 KB)

Computer Science > Computation and Language

Title:FOLIO: Natural Language Reasoning with First-Order Logic

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:FOLIO: Natural Language Reasoning with First-Order Logic

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators