AI for ScienceTrustworthy AIActive · 2024-2027

AIMing: Interactive Conjecture Proving

Human-AI collaboration for research-level mathematics.

1.7M
proof-state / premise pairs in NaturalPRISM
793
conjectures generated in our context ablation
32
arXiv mathematics categories in NaturalPRISM

Mathematicians increasingly ask language models for help, but a model is only useful if it is given the right context and reasons consistently. AIMing, funded by the National Science Foundation and led by the HAIE Lab, studies interactive conjecture proving: how humans and AI can work together to propose, retrieve and verify mathematical results.

The project contributes benchmarks and empirical studies. NaturalPRISM is a premise-retrieval task with 1.7 million pairs of intermediate proof states and the premise used in the next step, extracted from arXiv papers across all 32 mathematics categories. A paired ablation over 61 theorems shows which kinds of mathematical context (definitions, objects, assumptions, constructions) let a model recover the result a paper actually proves. TREAT and related studies test whether models recognize known theorems and reconstruct procedural routes when the same mathematics is written in equivalent forms, and whether LLM judges can verify valid but non-canonical solution steps.

Output

Publications

NeurIPS 2026 MATH-AI WorkshopForthcoming

The Shape of Mathematical Creativity: Measuring Mathematical Exploration in Formal Proof Generation

Fateme Mazdarani, Carlos Toxtli-Hernández

Website
MathNLP @ EMNLP 2026 WorkshopForthcoming

NaturalPRISM: A Natural Language Premise Retrieval Task for Research-level Intermediate Proof States

Harris Proctor, Li An, Austin LaHue, Michael Burr, Carlos Toxtli-Hernández, Vinita Gangaram Jansari, Luis David Garcia Puente, Benjamin E. Nye

Website
MathNLP @ EMNLP 2026 WorkshopForthcoming

Which Mathematical Context Supports Target-Aligned Conjecture Generation? A Paired Ablation Study

Austin LaHue, Michael Burr, Luis David Garcia Puente, Vinita Gangaram Jansari, Benjamin E. Nye, Carlos Toxtli-Hernández

PDF Website
MathNLP @ EMNLP 2026 WorkshopForthcoming

From Recognition to Reconstruction: Towards Robust Procedural Mathematical Reasoning

Fateme Mazdarani, Carlos Toxtli-Hernández

PDF Website
IEEE ICMLA 2026 Forthcoming

Beyond the Answer Key: Robustness Evaluation of Large Language Models for Step-Level Mathematical Verification

Fateme Mazdarani, Carlos Toxtli-Hernández

Website
SYNASC 2026 Forthcoming

TREAT: Evaluating Access to Formal Knowledge across Equivalent Mathematical Representations

Fateme Mazdarani, Carlos Toxtli-Hernández

Website