François Roewer-Després

h-index6
2papers
156citations

2 Papers

15.4AIJun 4, 2024Code
ACCORD: Closing the Commonsense Measurability Gap

François Roewer-Després, Jinyue Feng, Zining Zhu et al.

We present ACCORD, a framework and benchmark suite for disentangling the commonsense grounding and reasoning abilities of large language models (LLMs) through controlled, multi-hop counterfactuals. ACCORD introduces formal elements to commonsense reasoning to explicitly control and quantify reasoning complexity beyond the typical 1 or 2 hops. Uniquely, ACCORD can automatically generate benchmarks of arbitrary reasoning complexity, and so it scales with future LLM improvements. Benchmarking state-of-the-art LLMs -- including GPT-4o (2024-05-13), Llama-3-70B-Instruct, and Mixtral-8x22B-Instruct-v0.1 -- shows performance degrading to random chance with only moderate scaling, leaving substantial headroom for improvement. We release a leaderboard of the benchmark suite tested in this work, as well as code for automatically generating more complex benchmarks.

1.2CYNov 30, 2020
Continuous Subject-in-the-Loop Integration: Centering AI on Marginalized Communities

Francois Roewer-Despres, Janelle Berscheid

Despite its utopian promises as a disruptive equalizer, AI - like most tools deployed under the guise of neutrality - has tended to simply reinforce existing social structures. To counter this trend, radical AI calls for centering on the marginalized. We argue that gaps in key infrastructure are preventing the widespread adoption of radical AI, and propose a guiding principle for both identifying these infrastructure gaps and evaluating whether proposals for new infrastructure effectively center marginalized voices.