research

Neural network theory, learning dynamics and continual learning, language models, and computational neuroscience.

Why networks learn what they do

The through-line of my research is the theory of learning in neural networks: what a network will learn from data, and why. I have shown that efficient neural codes — long treated as a design principle for the brain — emerge naturally in any network trained by gradient descent, studied how to measure and regularize networks in the space of functions they compute rather than their weights, and shown that a trained network can be understood as a Bayesian ensemble of its tangent functions. That last result gives a probabilistic account of catastrophic forgetting and points toward gradient-based algorithms for continual learning. I have also studied why generalization is sometimes delayed long after the training loss falls (“grokking”), tracking how representational geometry changes during learning. I co-organized a COSYNE workshop on this theme (“Why networks learn what they do,” 2023) and co-authored the published lecture notes of the Gatsby Unit’s Analytical Connectionism summer school with Andrew Saxe, Jay McClelland, and colleagues.

  1. NeurIPS
    Continual learning with the neural tangent ensemble
    Ari Benjamin, Christian-Gernot Pehle, and Kyle Daruwalla
    Advances in Neural Information Processing Systems, 2024
  2. Nat. Comm.
    Efficient neural codes naturally emerge through gradient descent learning
    Ari S Benjamin, Ling-Qi Zhang, Cheng Qiu, Alan A Stocker, and Konrad P Kording
    Nature Communications, 2022
  3. PMLR
    An Introduction to Connectionist Theories of Semantic Cognition
    Ari S Benjamin, Anna-Lea Beyer, Marianne De Heer Kloots, Jaedong Hwang, Hajer Karoui, and 6 more authors
    In Analytical Connectionism School, 2026
  4. ICLR
    Measuring and regularizing networks in function space
    Ari S Benjamin, David Rolnick, and Konrad Kording
    In ICLR, 2019
  5. Shared visual illusions between humans and artificial neural networks
    Ari Benjamin, Cheng Qiu, Ling-Qi Zhang, Konrad Kording, and Alan Stocker
    In 2019 Conference on Cognitive Computational Neuroscience, 2019
  6. PMLR
    Delays in generalization match delayed changes in representational geometry
    Xingyu Zheng, Kyle Daruwalla, Ari S Benjamin, and David Klindt
    In UniReps: 2nd Edition of the Workshop on Unifying Representations in Neural Models, 2024
    Workshop paper

Knowledge and learning in language models

Language models raise the questions above in a new and pressing form. With Tingkai Liu and Tony Zador, I have worked on uncertainty-aware objectives for post-training language models at the token level. I am curious about how knowledge is formed and kept self-consistent in these models: what is learned in context versus in weights, how beliefs formed from one context should propagate to others, and whether the complementary-learning-systems perspective from neuroscience is the right lens for continual learning in modern AI. I think many of these questions can be understood analytically.

  1. Token-Level Uncertainty-Aware Objective for Language Model Post-Training
    Tingkai Liu, Ari S Benjamin, and Anthony M Zador
    arXiv preprint arXiv:2503.16511, 2025

Neuromodulation and cellular diversity

Much of my ongoing research is about building a connectionist framework for understanding neuromodulation — for example, treating neuromodulated networks as walking a manifold of weight configurations. I am also interested in the roles that diverse cell types play in neural circuits.

  1. PLOS CB
    A role for cortical interneurons as adversarial discriminators
    Ari S Benjamin and Konrad P Kording
    PLOS Computational Biology, 2023
  2. Walking the Weight Manifold: a Topological Approach to Conditioning Inspired by Neuromodulation
    Ari S Benjamin, Kyle Daruwalla, Christian Pehle, and Anthony M Zador
    arXiv preprint arXiv:2505.22994, 2025

Transformers for biological data

I design and train transformers end-to-end for problems in biology. In TissueFormer (BMC Bioinformatics, 2026), I built a transformer that extends single-cell foundation models to predict population-level phenotypes from groups of single cells while retaining single-cell resolution — applying it to predict COVID-19 severity from blood scRNA-seq and to identify cortical areas from mouse spatial transcriptomics.

  1. BMC
    Tissueformer: extending single-cell foundation models to predict population-level phenotypes
    Ari S Benjamin and Anthony M Zador
    BMC Bioinformatics, 2026

Machine learning for neural recordings

During my PhD I worked on machine learning methods for neural data: benchmarking how well modern ML predicts neural responses, and establishing best practices for neural decoding.

  1. The roles of supervised machine learning in systems neuroscience
    Joshua I Glaser, Ari S Benjamin, Roozbeh Farhoodi, and Konrad P Kording
    Progress in neurobiology, 2019
  2. Modern machine learning as a benchmark for fitting neural responses
    Ari S Benjamin, Hugo L Fernandes, Tucker Tomlinson, Pavan Ramkumar, Chris VerSteeg, and 3 more authors
    Frontiers in computational neuroscience, 2018
  3. Machine learning for neural decoding
    Joshua I Glaser, Ari S Benjamin, Raeed H Chowdhury, Matthew G Perich, Lee E Miller, and 1 more author
    eneuro, 2020

Past work

Bio-inspired materials science and molecular dynamics

Before I transitioned to neuroscience, I worked on bio-inspired materials science and molecular dynamics simulations. I was interested in self-assembly and how molecular interactions can lead to complex structures. The common thread that connected these interests was looking at nature in terms of its function, as might an engineer, as well as an interest in complex systems.

  1. Regulating ion transport in peptide nanotubes by tailoring the nanotube lumen chemistry
    Luis Ruiz, Ari Benjamin, Matthew Sullivan, and Sinan Keten
    The journal of physical chemistry letters, 2015
  2. Polymer conjugation as a strategy for long-range order in supramolecular polymers
    Ari Benjamin and Sinan Keten
    The Journal of Physical Chemistry B, 2016