Nicolas Fay

CL
h-index30
6papers
3,483citations
Novelty45%
AI Score41

6 Papers

21.8CLFeb 7, 2025
Evaluating Personality Traits in Large Language Models: Insights from Psychological Questionnaires

Pranav Bhandari, Usman Naseem, Amitava Datta et al.

Psychological assessment tools have long helped humans understand behavioural patterns. While Large Language Models (LLMs) can generate content comparable to that of humans, we explore whether they exhibit personality traits. To this end, this work applies psychological tools to LLMs in diverse scenarios to generate personality profiles. Using established trait-based questionnaires such as the Big Five Inventory and by addressing the possibility of training data contamination, we examine the dimensional variability and dominance of LLMs across five core personality dimensions: Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism. Our findings reveal that LLMs exhibit unique dominant traits, varying characteristics, and distinct personality profiles even within the same family of models.

14.6CLJun 18
Beyond Uniform Forgetting: A Study of Sequential Direct Preference Optimization Across Preference Settings

Pranav Bhandari, Nicolas Fay, Amitava Datta et al.

Aligning language models with human preferences often requires optimising multiple behavioural objectives. A practical approach is to apply these objectives sequentially using preference optimisation methods such as Direct Preference Optimisation (DPO), but it remains unclear whether later training uniformly degrades preferences learned earlier or whether the effect depends on the relationship between objectives. We study sequential DPO across four preference settings covering distributional conflict, multi-attribute interaction, strong safety signal, and compatible response-quality objectives. Using Llama-3.1-8B-Instruct with LoRA adapters, we evaluate all objectives after every stage with a fixed base-model reference. We find that sequential DPO does not produce a single forgetting pattern; preference change ranges from partial degradation to stability, pair-level redistribution, or positive transfer depending on objective relationship, signal strength, and training order. Pair-level analysis using length-normalised policy margins shows that aggregate metrics can mask heterogeneous changes across preference pairs, whereas quartile decomposition reveals that high-confidence pairs can either degrade or improve depending on the setting. Mechanistic diagnostics show that Stage~2 gradients and adapter updates are near-orthogonal to the previous objective across all settings, providing little evidence that direct gradient opposition is the primary driver. These findings suggest that future sequential alignment pipelines should account for objective compatibility and signal strength, rather than assuming that later objectives affect earlier preferences uniformly.

17.0CLFeb 17, 2025
Can LLM Agents Maintain a Persona in Discourse?

Pranav Bhandari, Nicolas Fay, Michael Wise et al.

Large Language Models (LLMs) are widely used as conversational agents, exploiting their capabilities in various sectors such as education, law, medicine, and more. However, LLMs are often subjected to context-shifting behaviour, resulting in a lack of consistent and interpretable personality-aligned interactions. Adherence to psychological traits lacks comprehensive analysis, especially in the case of dyadic (pairwise) conversations. We examine this challenge from two viewpoints, initially using two conversation agents to generate a discourse on a certain topic with an assigned personality from the OCEAN framework (Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism) as High/Low for each trait. This is followed by using multiple judge agents to infer the original traits assigned to explore prediction consistency, inter-model agreement, and alignment with the assigned personality. Our findings indicate that while LLMs can be guided toward personality-driven dialogue, their ability to maintain personality traits varies significantly depending on the combination of models and discourse settings. These inconsistencies emphasise the challenges in achieving stable and interpretable personality-aligned interactions in LLMs.

6.7CLOct 29, 2025
Activation-Space Personality Steering: Hybrid Layer Selection for Stable Trait Control in LLMs

Pranav Bhandari, Nicolas Fay, Sanjeevan Selvaganapathy et al.

Large Language Models exhibit implicit personalities in their generation, but reliably controlling or aligning these traits to meet specific needs remains an open challenge. The need for effective mechanisms for behavioural manipulation of the model during generation is a critical gap in the literature that needs to be fulfilled. Personality-aware LLMs hold a promising direction towards this objective. However, the relationship between these psychological constructs and their representations within LLMs remains underexplored and requires further investigation. Moreover, it is intriguing to understand and study the use of these representations to steer the models' behaviour. We propose a novel pipeline that extracts hidden state activations from transformer layers using the Big Five Personality Traits (Openness, Conscientiousness, Extraversion, Agreeableness and Neuroticism), which is a comprehensive and empirically validated framework to model human personality applies low-rank subspace discovery methods, and identifies trait-specific optimal layers across different model architectures for robust injection. The resulting personality-aligned directions are then operationalised through a flexible steering framework with dynamic layer selection, enabling precise control of trait expression in LLM outputs. Our findings reveal that personality traits occupy a low-rank shared subspace, and that these latent structures can be transformed into actionable mechanisms for effective steering through careful perturbations without impacting the fluency, variance and general capabilities, helping to bridge the gap between psychological theory and practical model alignment.

1.2SIFeb 9, 2019Code
Network connectivity dynamics affect the evolution of culturally transmitted variants

José Segovia Martín, Bradley Walker, Nicolas Fay et al.

The distribution of cultural variants in a population is shaped by both neutral evolutionary dynamics and by selection pressures, which include several individual cognitive biases, demographic factors and social network structures. The temporal dynamics of social network connectivity, i.e. the order in which individuals in a population interact with each other, has been largely unexplored. In this paper we investigate how, in a fully connected social network, connectivity dynamics, alone and in interaction with different cognitive biases, affect the evolution of cultural variants. Using agent-based computer simulations, we manipulate population connectivity dynamics (early, middle and late full-population connectivity); content bias, or a preference for high-quality variants; coordination bias, or whether agents tend to use self-produced variants (egocentric bias), or to switch to variants observed in others (allocentric bias); and memory size, or the number of items that agents can store in their memory. We show that connectivity dynamics affect the time-course of variant spread, with lower connectivity slowing down convergence of the population onto a single cultural variant. We also show that, compared to a neutral evolutionary model, content bias accelerates convergence and amplifies the effects of connectivity dynamics, whilst larger memory size and coordination bias, especially egocentric bias, slow down convergence. Furthermore, connectivity dynamics affect the frequency of high quality variants (adaptiveness), with late connectivity populations showing bursts of rapid change in adaptiveness followed by periods of relatively slower change, and early connectivity populations following a single-peak evolutionary dynamic. In this way, we provide for the first time a direct connection between the order of agents' interactions and punctuational evolution.

1.2SIJun 29, 2014
Human Communication Systems Evolve by Cultural Selection

Nicolas Fay, Monica Tamariz, T Mark Ellison et al.

Human communication systems, such as language, evolve culturally; their components undergo reproduction and variation. However, a role for selection in cultural evolutionary dynamics is less clear. Often neutral evolution (also known as 'drift') models, are used to explain the evolution of human communication systems, and cultural evolution more generally. Under this account, cultural change is unbiased: for instance, vocabulary, baby names and pottery designs have been found to spread through random copying. While drift is the null hypothesis for models of cultural evolution it does not always adequately explain empirical results. Alternative models include cultural selection, which assumes variant adoption is biased. Theoretical models of human communication argue that during conversation interlocutors are biased to adopt the same labels and other aspects of linguistic representation (including prosody and syntax). This basic alignment mechanism has been extended by computer simulation to account for the emergence of linguistic conventions. When agents are biased to match the linguistic behavior of their interlocutor, a single variant can propagate across an entire population of interacting computer agents. This behavior-matching account operates at the level of the individual. We call it the Conformity-biased model. Under a different selection account, called content-biased selection, functional selection or replicator selection, variant adoption depends upon the intrinsic value of the particular variant (e.g., ease of learning or use). This second alternative account operates at the level of the cultural variant. Following Boyd and Richerson we call it the Content-biased model. The present paper tests the drift model and the two biased selection models' ability to explain the spread of communicative signal variants in an experimental micro-society.