Nived Rajaraman

Postdoctoral Researcher at Microsoft Research NYC

nived-compact.jpg

I am a member of the Reinforcement Learning group.

I recently finished my PhD at UC Berkeley, under the guidance of Jiantao Jiao and Kannan Ramchandran, affiliated with the BLISS and BAIR labs. While at Berkeley, I organized the BLISS and CLIMB seminars. I am fortunate to have spent summers working with Nevena Lazic and Dong Yin at Deepmind and with Ravishankar Krishnaswamy at MSR. Previously, I was a dual degree student in the Department of Electrical Engineering at IIT Madras. I am fortunate to have had Andrew Thangaraj as my thesis advisor and to have worked closely with Rahul Vaze.

I will be on the job market for positions starting Fall 2027.

Recent News

Jul 2026   Presenting a tutorial on the Foundations of Learning Reasoning Models at COLT 2026.

Jul 2026   Organizing the Second Workshop on the Foundations of Post-training at COLT 2026. Submit by May 26, 2026!

May 2026   Talked about the Provable Benefits of Autocurriculum at Google Research NYC.

Dec 2025   Presented about the Interactive Learning of Single Index Models at the RL Theory Seminar.

Research

A full list of papers is on Google Scholar.

🤖  Foundations of Reasoning Agents

The capability of a reasoning agent is shaped by several components: what data it was pre-trained on, what post-training / inference-time methods were used, and how these pieces interact with each other. My work studies these questions, with the broader goal of understanding: how can agents learn to tackle much harder tasks than what we can explicitly supervise them to solve?

Selected Work

🧬  Architecture from First Principles

The capability of a generative model depends not only on how it is trained, but also on how computation is organized within its architecture. By studying models on controlled tasks, my work tries to demystify the role of design choices such as tokenization, attention, and depth. The broader goal is to build a prescriptive theory for: which architectural ingredients are essential and how they influence training?

Selected work

🧠  Foundations of Interactive Decision Making

In practice, we would like agents which react to the world they are deployed in, continually learning and improving from the experience they gather. Work from my PhD builds the foundations of imitation learning, and rigorously studies how much interaction with the environment can benefit an agent. My later work shows that simple algorithms like stochastic gradient descent can take advantage of interaction, even when more complex algorithms would fail to.

Selected work