Atlas

The square, unfolded. Each thread is a wire from problem to answer, each paper a spider on its wire. Green spiders read the written path, red ones read the gap between the path and the computation. Dashed curves join papers that share a method or a question.

Atlas of the research The commutative square unfolded. Four wires run from a problem on the left to an answer on the right, one per research thread. Papers sit on the wires as green or red spiders. Dashed curves join papers that share a method or a question. The same information is in the list below. Reasoning verification · chain of thought Guarantees · does it commute? Values and alignment · where the paths disagree Agents and action · many paths Trajectory geometry Spectral signatures Reasoning trace geometry Certified by abstention Erdős–Gyárfás with SAT The unauthored tiebreak Encoding values ForkSCOPE What the guard misses problem answer

hover a spider or a wire. click a spider to open it.

The same map as a list

Reasoning verification

chain of thought. Cheap verifiers that read the written chain of thought.

  1. Sparse Spectral Signatures of Reasoning: Model-Agnostic Verification via Sentence-Level Graph Signals

    Arjun Balaji·ICLR 2026 Workshop on Logical Reasoning of Large Language Models·Presented in Rio.

    Spectral entropy of a sentence graph built from a chain of thought separates correct from incorrect reasoning with AUC up to 0.77. This workshop paper is the origin of the trajectory geometry line.

  2. On the Geometry of Reasoning Traces: Spectral and Ollivier–Ricci Signatures of Chain-of-Thought

    Spectral and Ollivier–Ricci signatures of chain of thought. A connectivity error signal from the Fiedler value survives, and the broader geometric claims do not.

  3. Predicting Chain-of-Thought Correctness from Trajectory Geometry

    Arjun Balaji·Transactions on Machine Learning Research, 2026

    Interpretable features of a chain of thought's trajectory give a cheap verifier with ROC-AUC 0.869. Reranking with it reaches 0.888 against 0.835 for majority vote. Step order carries the signal. Shuffling the steps costs an order-aware model 0.081 balanced accuracy, while matched order-free capacity adds only 0.005.

    Compute for this paper came from the SAAR x AIM Intelligence Compute Partnership.

Guarantees

does it commute. What a verifier can certify, and when the certificate is empty.

  1. Verifying the Erdős–Gyárfás Conjecture up to 31 Vertices with SAT Modulo Symmetries

    Arjun Balaji·Preprint, Zenodo, 2026

    An LLM coding agent built a SAT Modulo Symmetries pipeline that verifies every graph of minimum degree 3 on at most 31 vertices contains a cycle of length 4, 8 or 16. This raises the counterexample bound from 17 to 32 in under six CPU core-hours, with every result independently checked. Cited by Tranquilli (arXiv 2608.02675).

  2. Certified by Abstention: Distribution-Free Guarantees for Chain-of-Thought Verifiers at Small Calibration Budgets

    Arjun Balaji·Under review at AISTATS 2027

    Valid certificates for chain-of-thought verifiers mostly abstain at realistic label budgets and can be wrong most of the times they fire. The paper introduces a certification floor, a lattice condition for conformal selection, and a floor-started certificate that triples coverage at non-vacuous targets. The guarantees do not survive benchmark shift or best-of-n optimization. Across 7 open models, 5 verifier signals and about 37k graded traces, a residual-stream probe beats every readable signal.

Values and alignment

where the paths disagree. What decides the answer when the stated reasoning does not.

  1. Encoding Values: Injecting Morality into Machines via Prompt-Conditioned Moral Frames

    Arjun Balaji, Aarushi Nema, Neha Kamath, Jyothika Raju Raju·NeurIPS 2025 Creative AI Track

    How a moral question is framed to a model shapes its answer in systematic, measurable ways.

  2. The Unauthored Tiebreak: What Decides When Two Constitutions Conflict

    Working paper, 2026

    With two constitutions in one prompt, an explicit priority line governs the verdict. Without one, listing order decides, invisibly to the model's own stated reasoning.

Agents and action

many paths. Many defensible paths from one problem, and monitors that miss what the agent does.

  1. ForkSCOPE: Charting the Agentic Garden of Forking Paths

    Arjun Balaji, Batuhan Duru Yeltekin, Tian Zheng·Under review at ACM CHI 2027·arXiv 2609.12438

    ForkSCOPE charts the multiverse of 207 autonomous LLM-agent data analyses of one dataset and one question, induces a fork and option decision vocabulary bottom-up from the code and write-ups, and ships an evidence-linked viewer.

  2. What the Guard Misses, the Robot Executes: Implied Harm in VLA Instructions

    Sripad Karne, Arjun Balaji·Under review at the SPAIS workshop, CoRL 2026·arXiv 2610.05818

    Vision-language-action policies act on instructions without being able to refuse, so screening falls to monitors. Holding the robot task fixed and varying only how explicitly the harmful intent is stated, text guards flag blunt requests and miss most implied ones. Monitoring activations does not close the gap.