VTA dopamine neuron activity produces spatially organized stimulus and action value representations through conditioned reinforcement

Journal: bioRxiv
Published Date:

Abstract

What is the neural architecture by which dopamine (DA) determines choice? Reinforcement learning (RL) has suggested an algorithmic chain: prediction errors based on reward train predicted values for stimuli and actions, and thereby determine choice. Elements of this chain have been identified, but have yet to be convincingly assembled. Here we used optogenetic stimulation of putative prediction error neurons at the top of this chain -- i.e., DA neurons in the ventral tegmental area (VTA) -- and recorded the activity of thousands of neurons in their target structure -- the striatum -- in mice performing a task that dissociates VTA DA signaling, stimulus value, action value, and choice. Conventional RL models captured some features of the data, including DA-driven choice reinforcement and value correlates, but failed on two key observations: i) action and state value correlates were relatively weak in the VTA's main target, the nucleus accumbens (NAc), and ii) VTA DA could not reinforce actions in the absence of a reward-predictive cue (CS+). Instead, we find that VTA DA produces stimulus value representations in the NAc, which in turn generates action value representations downstream (e.g. the CP). We formalize these observations in a Conditioned Reinforcer model, where VTA DA specifically acts on stimuli rather than actions, and these stimulus value representations are used to drive choice reinforcement.

Authors

  • Pan-Vazquez
  • A.; Zimmerman
  • C. A.; McMannon
  • B.; Fabre
  • J. M. J.; Louka
  • M.; Jia
  • T.; Sagiv
  • Y.; West
  • S. J.; Faulkner
  • M.; Brain Laboratory
  • I.; Dayan
  • P.; Witten
  • I. B.

Categories