On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and PerspectiveYue Huang, Chujie Gao, Siyuan Wu et al.
For researchers and practitioners in AI safety and governance, this work provides a structured guideline and dynamic evaluation tool to assess trustworthiness of generative models, though it is largely a survey and framework proposal without novel empirical results.
25.3AIMay 11
Positive Alignment: Artificial Intelligence for Human FlourishingRuben Laukkonen, Seb Krier, Chloé Bakalar et al.
For AI alignment researchers, this paper introduces a complementary agenda to safety-focused alignment, aiming to broaden the scope of alignment to include proactive support for human flourishing.
19.5AIMar 16
Are Dilemmas and Conflicts in LLM Alignment Solvable? A View from Priority GraphZhenheng Tang, Xiang Liu, Qian Wang et al.
This addresses alignment issues for LLMs in autonomous scenarios, but it is incremental as it builds on existing alignment research without solving fundamental philosophical dilemmas.
22.5CYJun 1
Legal Alignment for Safe and Ethical AINoam Kolt, Nicholas Caputo, Jack Boeglin et al.
For AI researchers and legal scholars, this paper provides a taxonomy and research agenda for integrating law into AI alignment, but it is a survey without empirical results or concrete numbers.
NARRA-Gym for Evaluating Interactive Narrative AgentsYue Huang, Yuchen Ma, Jiayi Ye et al.
For researchers evaluating LLMs in interactive narrative settings, this provides a more comprehensive benchmark than existing static or single-turn evaluations.
11.7CLMar 16
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value CodebookJaehyeok Lee, Xiaoyuan Yi, Jing Yao et al.
This addresses the problem of aligning LLMs with diverse cultural values for safety and engagement, offering a more nuanced evaluation method than existing benchmarks.
SWE-chat: Coding Agent Interactions From Real Users in the WildJoachim Baumann, Vishakh Padmakumar, Xiang Li et al.
This provides an empirical foundation for understanding AI coding agent performance in real developer workflows, addressing a gap in current research.
Synergy: A Next-Generation General-Purpose Agent for Open Agentic WebXiaohang Nie, Zihan Guo, Kezhuo Yang et al.
This addresses the problem of agent interoperability and social integration for developers and users in decentralized digital ecosystems, representing a novel architectural approach rather than an incremental improvement.
38.9CLAug 28
CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast AsiaBryan Chen Zhengyu Tan, Weihua Zheng, Thong T. Doan et al.
This work addresses the limitation of current cultural evaluations for LLMs, which often oversimplify culture to factual recall, by providing a more realistic multi-turn simulation for users seeking practical, culturally grounded help in East and Southeast Asia.
InterveneBench: Benchmarking LLMs for Intervention Reasoning and Causal Study Design in Real Social SystemsShaojie Shi, Zhengyu Shi, Lingran Zheng et al.
This addresses the need for better benchmarks to assess LLMs' causal reasoning in social science, which is incremental as it builds on existing benchmarking efforts.
AcademiClaw: When Students Set Challenges for AI AgentsJunjie Yu, Pengrui Lu, Weiye Si et al.
Provides a challenging benchmark for evaluating AI agents on real-world academic workflows, addressing the gap in assistant-level benchmarks for the OpenClaw ecosystem.
30.5MAMar 29
Emergent Social Intelligence Risks in Generative Multi-Agent SystemsYue Huang, Yu Jiang, Wenjie Wang et al.
It identifies a new class of risks in generative multi-agent systems for researchers and practitioners deploying such systems, though the study is exploratory and lacks quantitative benchmarks.
26.5CYApr 28
Responsible Evaluation of AI for Mental HealthHiba Arnaout, Anmol Goel, H. Andrew Schwartz et al.
For researchers and practitioners developing AI for mental health, this work provides a structured evaluation framework to address current gaps in clinical validity and equity.
10.0CLApr 3
Verbalizing LLMs' assumptions to explain and control sycophancyMyra Cheng, Isabel Sieh, Humishka Zope et al. · stanford
This addresses safety issues in AI by explaining and controlling sycophancy in LLMs, offering a new mechanism for understanding model behavior, though it is incremental in building on existing work on model interpretability.
16.6CRMay 17
AI Agents May Always Fall for Prompt InjectionsSahar Abdelnabi, Eugene Bagdasarian
For developers and researchers of AI agents, this paper highlights a fundamental limitation of current prompt injection defenses and provides a theoretical framework for understanding context-sensitive failures.
16.8CYApr 8
The ATOM Report: Measuring the Open Language Model EcosystemNathan Lambert, Florian Brand
This provides a comprehensive snapshot for researchers, entrepreneurs, and policy advisors tracking the open language model ecosystem.
16.2CYMay 12
LLM Harms: A Taxonomy and DiscussionKevin Chen, Saleh Afroogh, Abhejay Murali et al.
For AI developers and policymakers, this work provides a structured framework to identify and address LLM-related harms, though it is primarily a conceptual taxonomy without empirical validation.
35.2CLJul 7
Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and ReliabilityAlicia Parrish, Rajat Shinde, Sanket Badhe et al.
This work addresses the lack of culturally-aware AI safety evaluation for global deployments, exposing blind spots in current Western-centric benchmarks.
Beyond Alignment: Value Diversity as a Collective Property in Multicultural Agent SystemsShaoyang Xu, Jingshen Zhang, Long P. Hoang et al.
For researchers building multicultural LLM-based agent systems, this work establishes a necessary evaluation axis beyond alignment, revealing a persistent homogenization problem.
13.4CVMar 12
OrthoEraser: Coupled-Neuron Orthogonal Projection for Concept ErasureChuancheng Shi, Wenhua Wu, Fei Shen et al.
This addresses safety risks in text-to-image models by enabling more precise removal of harmful content without damaging benign attributes, representing a strong specific gain in a domain-specific area.