Austin LaHue
Ph.D. Student, Human-Centered Computing
- Ph.D. in Human-Centered Computing, Clemson University
- Human-AI Empowerment Lab, Clemson University
- In the lab since Fall 2025
- Expected graduation: Spring 2029
- Advisor: Dr. Carlos Toxtli-Hernández
- Publications
- 3
- First-authored
- 2
- Projects
- 1
- Lab co-authors
- 1
- Venues
- 2
About
Austin LaHue is a Ph.D. student in Human-Centered Computing at Clemson University and a graduate research assistant on the NSF AIMing project, advised by Dr. Carlos Toxtli. Austin studies how language models can support research-level mathematics, including which forms of mathematical context let a model generate the conjecture a paper actually proves (MathNLP 2026 at EMNLP). Austin also designs AI-mediated journaling systems and created SHARE, a prototype of an AI-mediated peer journaling platform, published at ACM HAI 2026.
Before Clemson, Austin worked as a data scientist and business intelligence analyst at The Walt Disney Company, building predictive models and data pipelines, and as a research assistant at Brigham Young University's Marriott School of Business, where Austin co-authored a study of the mental-health benefits of a daily one-hour social media break (ECIS 2024). Austin holds an M.S. in Data Science from The University of Texas at Austin and a B.S. in Global Supply Chain Management from Brigham Young University.
Highlights
- First author of SHARE (ACM HAI 2026) and a context ablation for conjecture generation (MathNLP 2026 at EMNLP)
- 1st Place, Dataiku Everyday AI Data Visualization Challenge (2025)
- Treasurer, School of Computing Graduate Student Association
- Reviewer for PACIS 2024 and the HAI 2026 poster session
Publications 3
NaturalPRISM: A Natural Language Premise Retrieval Task for Research-level Intermediate Proof States
As AI use progresses in research-level mathematics, automated information retrieval systems for constructing relevant prompts are increasingly important. Existing benchmarks and datasets for identifying which prior results are necessary to prove a statement do not align well with current workflows in which a natural language proof is in progress, but it is unclear what the next step(s) should be. We introduce NaturalPRISM, a benchmark task and dataset containing 1.7 million pairs of intermediate proof states and the premise invoked in the next step of the proof. All pairs are extracted from natural language research-level mathematics papers published on arXiv, and span all 32 subject categories. We evaluate the effect of fine-tuning retrieval models for domain-specific vs. domain-agnostic use cases, and find that the combination of sufficient training data with alignment between the training and testing distributions allows smaller specialized models to dramatically outperform larger general models.
@inproceedings{Proctor2026NaturalPRISM,
title = {NaturalPRISM: A Natural Language Premise Retrieval Task for Research-level Intermediate Proof States},
author = {Proctor, Harris and An, Li and LaHue, Austin and Burr, Michael and Toxtli-Hernandez, Carlos and Jansari, Vinita Gangaram and Garcia Puente, Luis David and Nye, Benjamin E.},
booktitle = {Proceedings of the 4th Workshop on Mathematical Natural Language Processing (MathNLP 2026)},
address = {Budapest, Hungary},
publisher = {Association for Computational Linguistics},
year = {2026},
month = October,
note = {Forthcoming},
url = {https://sites.google.com/view/mathnlp2026}
}Which Mathematical Context Supports Target-Aligned Conjecture Generation? A Paired Ablation Study
Generating mathematical conjectures from natural-language context is challenging because theorem-relevant information is distributed across multiple forms of mathematical context, including definitions, objects, assumptions, and constructions. We study which of these contextual signals help a model recover the specific claim supported by the surrounding setup, using 61 main theorems from 12 papers across four mathematical domains. For each theorem, we redact the result and represent the remaining setup using six context components. A paired ablation design compares full context, six leave-one-out conditions, and six single-component conditions, yielding 793 generated conjectures. Context components were largely complementary: removing one component usually had little effect, whereas retaining only one substantially reduced expected-theorem alignment. Local Mathematical Objects provided the strongest standalone signal, while Structural Constraints and Conditions produced the largest alignment decrease when removed. These findings suggest that target-aligned conjecture generation depends on combining information that identifies relevant objects with information that constrains admissible claims.
@inproceedings{LaHue2026Which,
title = {Which Mathematical Context Supports Target-Aligned Conjecture Generation? A Paired Ablation Study},
author = {LaHue, Austin and Burr, Michael and Garcia Puente, Luis David and Jansari, Vinita Gangaram and Nye, Benjamin E. and Toxtli-Hernandez, Carlos},
booktitle = {Proceedings of the 4th Workshop on Mathematical Natural Language Processing (MathNLP 2026)},
address = {Budapest, Hungary},
publisher = {Association for Computational Linguistics},
year = {2026},
month = October,
note = {Forthcoming},
url = {https://sites.google.com/view/mathnlp2026}
}In the news
- Three papers from the NSF AIMing project will appear at MathNLP 2026, co-located with EMNLP 2026 in Budapest: the NaturalPRISM premise-retrieval benchmark, a context ablation for conjecture generation, and robust procedural mathematical reasoning.
- Three Ph.D. students join the lab in 2025: Austin LaHue (HCC), Fateme Mazdarani (CS) and Christopher Kalahiki (CS).