STRIVE: Probing Reasoning Limits in Graded Plausibility Generation and Evaluation
For psycholinguists, STRIVE automates the labor-intensive construction of controlled event sets, but its limited accuracy on hard cases suggests it is not yet a complete replacement for human input.
STRIVE is an LLM-based framework for generating and evaluating controlled event sets for psycholinguistic plausibility studies. With a global reasoning scratchpad and evaluator-guided refinement, it improved high-quality set generation from 16.7% to 75.0%, though near-boundary events remain challenging with evaluator accuracy at 57% on the hardest condition.
Event knowledge concerns who does what to whom. Psycholinguists use event-plausibility judgments to examine how this knowledge supports human language processing. To isolate plausibility effects, these studies require controlled event sets in which one event slot varies across plausibility levels while all other event features remain fixed. Constructing such sets manually is labor-intensive. We therefore introduce STRIVE, an LLM-based framework for jointly generating and evaluating controlled event sets crossing plausibility class (plausible vs. implausible) with intended classification difficulty (easy vs. hard). Given a verb, STRIVE constructs a shared event frame, then produces one event per condition by varying one slot while holding all others fixed. In experiments with six models across 60 verbs, GPT-5.1 produced high-quality sets only 16.7% of the time using the baseline generation prompt. Adding a global reasoning scratchpad and evaluator-guided refinement raised this rate to 75.0%. Greater reasoning effort also improved evaluator--human agreement. Nevertheless, events near the plausibility boundary remain most difficult. They elicit the greatest human disagreement, and the best evaluator reaches only 57% accuracy on the implausible-hard condition, indicating a need for human input. Overall, STRIVE offers a scalable approach to reducing manual effort by automating initial event-set generation and evaluation for psycholinguistic studies.