Neural Value Alignment: Human-AI Collaboration Under Goal-Action Ambiguity.
Journal:
IEEE transactions on cybernetics
Published Date:
Aug 24, 2026
Abstract
Value alignment plays a crucial role in human-artificial intelligence (AI) collaboration. Traditional approaches attempt to infer human goals from actions to guide AI policies. However, this behavior-level alignment faces an inherent challenge: the ambiguous mapping between goals and actions, as a single action might serve multiple possible goals, while different actions could achieve the same goal. To overcome these limitations, we propose neural value alignment (NVA), a unifying perspective that leverages key variables in human reinforcement learning (RL): reward prediction error (RPE) and state prediction error (SPE). RPE captures outcome discrepancies, refining AI goal inference, while SPE reflects state transition misalignment, shaping AI actions. Using a novel task paradigm that dissociated RPE and SPE, combined with electroencephalography (EEG) recordings, we demonstrated cortical decodability of RPE, SPE, and their co-occurrence, robustly across contexts. Simulations showed that RPE-SPE synergy accelerated value alignment, even under imperfect decoding. This study bridges RL and human-AI interaction, showing that RPE-SPE synergy enables flexible and human-compatible artificial systems operating under goal-action ambiguity.
Authors
Keywords
No keywords available for this article.