DERELAB: Probing Defeasible Reasoning and Confirmation Bias in LLMs with a Generative Benchmark

arXiv:2608.30413 · cs.AI · Submitted 2026-08-31 · Read on arXiv

cs.AI

Submitted: 2026-08-31

Updated: 2026-08-31

Code: https://github.com/Jayanta47/DeReLab

Terminology

Sources

Related papers