A commutative square A problem on the left and an answer on the right, joined by two wires. The top wire is labelled chain of thought and carries a green spider. The bottom wire is labelled internal computation and carries a red spider. A dashed line between the wires asks whether the two paths commute. chain of thought internal computation commutes? problem answer A commutative square A problem at the top and an answer at the bottom, joined by two wires. The left wire is labelled chain of thought and carries a green spider. The right wire is labelled internal computation and carries a red spider. A dashed line between the wires asks whether the two paths commute. chain of thought internal computation commutes? problem answer
When does a model's reasoning agree with what it actually computes?

Arjun Balaji

Co-founder of ImpactAI Foundry. Researcher in AI safety and alignment at NYU and Columbia University.

I co-founded ImpactAI Foundry, where we work with social impact organizations to run AI cohorts, build products and policy, and do AI safety and alignment research.

My own research asks when we can trust what a language model's reasoning tells us. I build verifiers for chain of thought, ask what guarantees they can give, and map where they break under distribution shift and optimization pressure. I'm a researcher in the Machine Learning for Good Lab at NYU and in the Tian Zheng Lab at Columbia University, and a graduate research assistant at Columbia University Irving Medical Center.

Underneath both is one concern, reliable AI in places where errors are costly. At Irving Medical Center I build voice and vision pipelines for dementia research and for training dental students to read radiographs and question AI output. The same concern runs through the policy work and the cohorts at ImpactAI Foundry.

Now

Selected research

all research·atlas

News