Research
When does a model's reasoning agree with what it actually computes? Each thread below reads one part of that question. The Atlas draws them as one diagram.
Reasoning verification
chain of thought. Cheap verifiers that read the written chain of thought.
-
Predicting Chain-of-Thought Correctness from Trajectory Geometry
Interpretable features of a chain of thought's trajectory give a cheap verifier with ROC-AUC 0.869. Reranking with it reaches 0.888 against 0.835 for majority vote. Step order carries the signal. Shuffling the steps costs an order-aware model 0.081 balanced accuracy, while matched order-free capacity adds only 0.005.
Compute for this paper came from the SAAR x AIM Intelligence Compute Partnership.
paper·bibtex
@article{balaji2026trajectory, title = {Predicting Chain-of-Thought Correctness from Trajectory Geometry}, author = {Balaji, Arjun}, journal = {Transactions on Machine Learning Research}, issn = {2835-8856}, year = {2026}, url = {https://openreview.net/forum?id=H9cBkEqVeY} } -
Sparse Spectral Signatures of Reasoning: Model-Agnostic Verification via Sentence-Level Graph Signals
Spectral entropy of a sentence graph built from a chain of thought separates correct from incorrect reasoning with AUC up to 0.77. This workshop paper is the origin of the trajectory geometry line.
paper·bibtex
@inproceedings{balaji2026sparse, title = {Sparse Spectral Signatures of Reasoning: Model-Agnostic Verification via Sentence-Level Graph Signals}, author = {Balaji, Arjun}, booktitle = {ICLR 2026 Workshop on Logical Reasoning of Large Language Models}, year = {2026}, url = {https://iclr.cc/virtual/2026/10017393} } -
On the Geometry of Reasoning Traces: Spectral and Ollivier–Ricci Signatures of Chain-of-Thought
Spectral and Ollivier–Ricci signatures of chain of thought. A connectivity error signal from the Fiedler value survives, and the broader geometric claims do not.
Guarantees
does it commute. What a verifier can certify, and when the certificate is empty.
-
Certified by Abstention: Distribution-Free Guarantees for Chain-of-Thought Verifiers at Small Calibration Budgets
Valid certificates for chain-of-thought verifiers mostly abstain at realistic label budgets and can be wrong most of the times they fire. The paper introduces a certification floor, a lattice condition for conformal selection, and a floor-started certificate that triples coverage at non-vacuous targets. The guarantees do not survive benchmark shift or best-of-n optimization. Across 7 open models, 5 verifier signals and about 37k graded traces, a residual-stream probe beats every readable signal.
code·bibtex
@misc{balaji2026certified, title = {Certified by Abstention: Distribution-Free Guarantees for Chain-of-Thought Verifiers at Small Calibration Budgets}, author = {Balaji, Arjun}, year = {2026}, note = {Preprint, under review}, url = {https://github.com/ArjunBalaji79/certified-by-abstention-code} } -
Verifying the Erdős–Gyárfás Conjecture up to 31 Vertices with SAT Modulo Symmetries
An LLM coding agent built a SAT Modulo Symmetries pipeline that verifies every graph of minimum degree 3 on at most 31 vertices contains a cycle of length 4, 8 or 16. This raises the counterexample bound from 17 to 32 in under six CPU core-hours, with every result independently checked. Cited by Tranquilli (arXiv 2608.02675).
Values and alignment
where the paths disagree. What decides the answer when the stated reasoning does not.
-
The Unauthored Tiebreak: What Decides When Two Constitutions Conflict
With two constitutions in one prompt, an explicit priority line governs the verdict. Without one, listing order decides, invisibly to the model's own stated reasoning.
-
Encoding Values: Injecting Morality into Machines via Prompt-Conditioned Moral Frames
How a moral question is framed to a model shapes its answer in systematic, measurable ways.
paper·bibtex
@inproceedings{balaji2025encoding, title = {Encoding Values: Injecting Morality into Machines via Prompt-Conditioned Moral Frames}, author = {Balaji, Arjun and Nema, Aarushi and Kamath, Neha and Raju, Jyothika Raju}, booktitle = {NeurIPS 2025 Creative AI Track}, year = {2025}, url = {https://neurips.cc/virtual/2025/loc/san-diego/129264} }
Agents and action
many paths. Many defensible paths from one problem, and monitors that miss what the agent does.
-
ForkSCOPE: Charting the Agentic Garden of Forking Paths
ForkSCOPE charts the multiverse of 207 autonomous LLM-agent data analyses of one dataset and one question, induces a fork and option decision vocabulary bottom-up from the code and write-ups, and ships an evidence-linked viewer.
-
What the Guard Misses, the Robot Executes: Implied Harm in VLA Instructions
Vision-language-action policies act on instructions without being able to refuse, so screening falls to monitors. Holding the robot task fixed and varying only how explicitly the harmful intent is stated, text guards flag blunt requests and miss most implied ones. Monitoring activations does not close the gap.
arXiv·bibtex
@misc{karne2026guard, title = {What the Guard Misses, the Robot Executes: Implied Harm in VLA Instructions}, author = {Karne, Sripad and Balaji, Arjun}, year = {2026}, eprint = {2610.05818}, archivePrefix = {arXiv}, url = {https://arxiv.org/abs/2610.05818} }
Other work
-
SENTINEL: Supervisory Text versus Contagion Graphs for Emerging-Market Crisis Early Warning
48 countries, 1994 to 2024. Contagion graphs add nothing over domestic indicators. An LLM-scored index of 448 IMF Article IV appraisals separates pre-crisis from calm periods with AUROC 0.71.
-
BiasCheck: An Analytical Framework for Contextual Bias Detection in Text, Models and Databases
Open-source framework for detecting contextual bias in text, models, retrieval pipelines and databases.
-
ProteoDockNet: Novel GNN-based ligand binding affinities prediction architecture via SMILES to key liver, kidney and brain proteins using QSAR data
A graph neural network that predicts ligand binding affinities for human liver, kidney and brain proteins from SMILES strings.
paper·code·bibtex
@article{setlur2024proteodocknet, title = {ProteoDockNet: Novel GNN-based ligand binding affinities prediction architecture via SMILES to key liver, kidney and brain proteins using QSAR data}, author = {Setlur, Anagha S and Niranjan, Vidya and Balaji, Arjun and Karunakaran, Chandrashekar}, journal = {Computational and Structural Biotechnology Reports}, volume = {1}, pages = {100011}, year = {2024}, doi = {10.1016/j.csbr.2024.100011} } -
Brain MRI surface registration
Unsupervised mesh deformation for cortical surface registration.
-
3D cardiac segmentation
Volumetric cardiac segmentation from MRI and echocardiograms.