Inter-Rater Reliability of LLM and Rule-Based Annotation for Inferential Narrative Features: Three Studies on a Turkish Corpus
cs.CL
Submitted: 2026-09-12
Updated: 2026-09-12
Comments: v1.0. 10 pages, 3 tables. Also archived at Zenodo: https://doi.org/10.5281/zenodo.21740239
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Datasets that ship automatically generated feature annotations invite a question rarely asked of them: would a human agree with those labels? This report answers that for the Objective Projection
Terminology
Abstract
Datasets that ship automatically generated feature annotations invite a question rarely asked of them: would a human agree with those labels? This report answers that for the Objective Projection corpus, a Turkish narrative dataset whose scenes carry a per-scene applied rules field from a rule-based detector over six craft features -- two prohibitions (emotion labelling, simile) and four positive techniques (materialized metaphor, micro-focus, temporal anchor, atmosphere contradiction). Three studies are reported. Study 1 (n = 120) scores the detector against blind labels from the scheme's own author. Study 2 (n = 100, a disjoint scene set) scores the detector plus Gemini 2.5 Flash and Grok against an independent non-expert rater whose labels were locked before any machine ran. Study 2b re-runs the identical protocol with Claude Fable 5 (High) and ChatGPT 5.5. The central result concerns one rule. On materialized metaphor -- closest to the methodology's theoretical core -- the five machine labellers returned positive rates of 0, 1, 40, 72 and 78 out of 100 scenes, against a human count of 9. Cohen's κ was at or indistinguishable from chance for five of six labellers, across both human references and both scene sets: 0.004, 0.015, 0.000, 0.019, 0.027. Raw agreement ranged from 74.7% to 84.5%, an artefact of class imbalance rather than a sign of competence. We deliberately do not resolve this into a single story. Two readings survive: the feature is genuinely inferential and beyond current automatic detection, or the rule's definition is not yet operational enough for any rater to apply consistently -- including the human. Distinguishing them needs a second independent human rater, which this report does not have and therefore does not claim.
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering