Future of Work & Worker-Centered AI
AI tools that increase workers' agency, earnings and well-being.
Millions of people work through digital platforms that fuel AI systems, often under opaque algorithms and unequal power dynamics. Building on the director's award-winning research on the invisible labor of crowd work, the lab designs worker-centered AI: tools that make platforms more transparent, adapt to workers' cultures and languages, and measurably improve outcomes.
Recent work includes CultureFit, a culturally aware tool that adapts to monochronic and polychronic work styles and, in a field experiment with 55 workers from 24 countries, improved the earnings of workers from cultural backgrounds often overlooked in design (ACM CSCW 2024); transparent, explained task recommendation that raised perceived fairness, trust and empowerment (IEEE CAI 2026); a multidimensional framework for measuring how AI tools empower crowdworkers; qualification tests for admitting language models as workers in human-computation pipelines (HCOMP 2026); and AI assistants and self-quantification tools for knowledge work.
Guiding questions
- How can AI make digital labor platforms fairer and more transparent for workers?
- How do we measure whether an AI tool truly empowers the people who use it?
- How should language models be screened before they join human-computation pipelines?
Projects
Projects in this thrust
AI Assistants and Self-Quantification at Work
Privacy-preserving activity sensing and LLM coaching for knowledge workers.
The Future of Aviation Inspection: AI and Mixed Reality
AI and mixed reality as agents of transformation for inspectors.
Publications
14 publications
Human-Centered Automation
Human-Centered Automation (HCA) is becoming indispensable for organizations seeking to implement AI-driven systems, robotic process automation, and other advanced tools while keeping human needs at the core. Despite the widespread adoption of automation, many initiatives fall short due to a lack of alignment between technical capabilities and real-world user and organizational needs. This book unites insights from cognitive science, software engineering, business strategy, and human factors that view automation as something that works for people and not just systems. This book provides the tools and strategies to ensure efforts with automation succeed with people at the center. It takes readers through the entire HCA lifecycle; from process discovery and planning to system design, testing, and long-term oversight. With easy-to-understand frameworks and real-world examples across sectors, including healthcare, finance, and manufacturing, readers will gain tools to assess risk, define measurable outcomes, involve stakeholders, and build automation that is trustworthy, explainable, and effective. It offers the methodologies needed to drive meaningful, lasting impact through automation. Human-Centered Automation is essential for ergonomics and human factors professionals, researchers, and organizational leaders involved in designing or managing automation initiatives. Its appeal extends to professionals in AI and machine learning, UX design, process engineering, business operations, and policy development.
@book{ToxtliHernandez2026HumanCentered,
title = {Human-Centered Automation},
isbn = {9781003666400},
url = {http://dx.doi.org/10.1201/9781003666400},
doi = {10.1201/9781003666400},
publisher = {CRC Press},
author = {Toxli-Hernandez, Carlos},
year = {2026},
month = May
}Qualification by Calibration: A Readable Benchmark for Admitting Language Models to Human-Computation Tasks
When several language models are wired together into a human-computation pipeline, the system routes, arbitrates, and escalates work according to how confident each model says it is, so a team builder needs a way to screen candidate models the way crowdsourcing has long screened human contributors: with a small, inspectable qualification test. We present such an instrument, a compact battery of yes-or-no questions across everyday domains on which a model reports an answer and a confidence, with every gold label independently audited. Our central contribution, however, is the audited harness and the interface properties it measures, not accuracy discrimination: qualification verdicts for model workers are only as valid as the harness that administers them. Evaluating a diverse panel of contemporary models under three administration protocols, we show that a naive harness, with a fixed token budget and an unaudited parser, manufactures failing workers out of competent ones, misreading a model that answers essentially every item correctly as badly inaccurate, answer-biased, and overconfident. Properly administered, every model in the panel proves admissible on accuracy for these common-knowledge items, and the differentiators that remain, and that transfer to disjoint downstream tasks, are interface properties: calibration of stated confidence, format discipline, and availability under budget. A controlled manipulation further shows the calibration axis is dissociable from accuracy. We release the items, the audited protocol, all per-trial responses under every protocol, and the scoring code, so the check, and the audit of the check, are each a single command.
@inproceedings{Toxtli2026Qualification,
title = {Qualification by Calibration: A Readable Benchmark for Admitting Language Models to Human-Computation Tasks},
author = {Carlos Toxtli and Manuel Delaflor},
booktitle = {2026 ACM Conference on Human-AI Complementarity and Alignment (HCOMP 2026)},
year = {2026},
doi = {10.1145/3834580.3838735}
}Worker-Centered AI: Transparent Explanations for Trustworthy Task Recommendation in Crowd Work
Task recommendation algorithms in crowd work platforms can erode worker autonomy through strategies such as amplifying the visibility of undesirable or low-paying jobs to meet client demand. We explore whether transparency can shift these systems toward worker-centered values. We built a simulated crowd work platform comparing an opaque Baseline interface against a Transparent variant with explained recommendations and navigation cues. In a within-subjects lab study capturing an initial thirty-minute interaction (n=40), participants rated the transparent system substantially higher on fairness, trust, empowerment, and usability. Behaviorally, transparency supported more strategic task evaluation and increased recommendation alignment. Our results suggest that transparency, even in the form of lightweight disclosure, can improve trust and experience without altering the underlying algorithm. Therefore, we discuss design implications for worker-centered systems, including explanation scope and deployment trade-offs, while acknowledging that durable empowerment requires pairing transparency with meaningful agency and field validation under real economic constraints.
@inproceedings{Mazdarani2026Worker,
title = {Worker-Centered AI: Transparent Explanations for Trustworthy Task Recommendation in Crowd Work},
url = {http://dx.doi.org/10.1109/CAI68641.2026.11536153},
doi = {10.1109/cai68641.2026.11536153},
booktitle = {2026 IEEE Conference on Artificial Intelligence (CAI)},
publisher = {IEEE},
author = {Mazdarani, Fateme and Hernandez, Alberto Campos and Toxtli, Carlos},
year = {2026},
month = May,
pages = {295--302}
}Empowering the Crowd: Measuring the Impact of AI-Driven Tools on Crowdwork
Crowdsourcing platforms employ millions of workers worldwide to complete microtasks that fuel machine-learning pipelines, moderate user-generated content, and accelerate scientific discovery. The rapid ascent of large language models and other foundation models has shifted research attention from automation instead of humans to augmentation alongside humans. This chapter addresses a central challenge in this new landscape: understanding how, how much, and for whom AI-driven tools truly empower crowdworkers. We synthesize insights from a diverse body of empirical studies, spanning randomized lab experiments, field A/B tests, trace-log meta-analyses, and qualitative user studies, to evaluate the impact of AI on crowdwork. To structure this analysis, we introduce a multidimensional empowerment framework grounded in psychological and economic theory, focusing on five key dimensions: productivity and earnings, quality and reliability, learning and skill development, agency and autonomy, and well-being and inclusion. Our review of the evidence reveals a dual-edged reality: while AI assistance offers significant productivity gains, error reductions, and improvements in worker self-efficacy, it also introduces substantial risks, including automation bias, cultural exclusion, wage depression, and unequal access to tools. By distilling these findings, we formulate eight concrete design principles for creating more equitable and effective human-AI augmentation systems. The chapter concludes by outlining a forward-looking research agenda that treats crowdworkers not as replaceable cogs but as essential partners in designing the future of socio-technical systems.
@inbook{Toxtli_2026,
title = {Empowering the Crowd: Measuring the Impact of AI-Driven Tools on Crowdwork},
isbn = {9781836349464},
issn = {2753-894X},
url = {http://dx.doi.org/10.5772/intechopen.1014000},
doi = {10.5772/intechopen.1014000},
booktitle = {Crowdsourcing - Innovations in Digital Collaboration},
publisher = {IntechOpen},
author = {Toxtli, Carlos and Delgado Solorzano, Cecilia},
year = {2026},
month = Feb
}Human-AI Empowerment: An Interdisciplinary Perspective
As artificial intelligence (AI) advances, the question of how AI can empower humans over the long term has become increasingly important. This book, Human-AI Empowerment (HAIE), provides a timely exploration of strategies for aligning AI with long-term human goals, ensuring that AI acts as an empowering force across multiple dimensions. Drawing on interdisciplinary research from fields such as AI, HCI, psychology, education, economics, and social science, the book develops comprehensive frameworks for studying and optimizing AI’s impact on human empowerment. HAIE investigates empowerment from a human-centered computing (HCC) perspective, examining how AI systems can track and adapt to progressively achieve long-term goals. The book explores techniques for fostering a mutually beneficial human-AI synergy, delving into AI Empowerment approaches, applicable Human-Computer Interaction methods for long-term engagement, and insights from various disciplines on long-term goal management. Through integrative frameworks, empirical evidence, and ongoing work in the field, this volume informs academics and practitioners seeking to harness AI as a transformative technology for concretely empowering humanity. This book highlights the need for comprehensive approaches to understanding and shaping the future of human-AI collaboration, maximizing its potential to expand human possibilities and support the pursuit of mid-term and long-term goals.
@book{Toxtli_Hern_ndez_2025,
title = {Human-AI Empowerment: An Interdisciplinary Perspective},
isbn = {9781003536628},
url = {http://dx.doi.org/10.1201/9781003536628},
doi = {10.1201/9781003536628},
publisher = {Chapman and Hall/CRC},
author = {Toxtli-Hernández, Carlos},
year = {2025},
month = Sept
}AI-Powered Comment Triage for Efficient Collaboration and Feedback Management
In today's digital landscape, collaborative tools are critical for virtual teamwork, with comments as a key mechanism for communication and feedback. Our project, within the Natural Language Processing (NLP) domain, focuses on improving comment handling in collaborative environments using advanced machine learning methods. We developed a triage system that categorizes and prioritizes comments to help us efficiently address the most critical feedback. Building on previous work, we employed transformer models like BERT and RoBERTa, which showed strong performance in classifying comments when fine-tuned on our dataset. To enhance the handling of hierarchical structures, we experimented with Hierarchical Capsule Networks (HcapsNet) and Hierarchical Attention Networks (HAN). Additionally, GEMMA-2B, a large language model, demonstrated strong results in F1-score and precision while providing zero-shot and few-shot learning capabilities. The framework, tested in domains such as project management, academic collaboration, and document review, classifies and prioritizes comments based on six dimensions: urgency, importance, sentiment, actionability, resolution status, and thematic relevance.It incorporates rule-based logic alongside pre-trained NLP models, including GEMMA-2B for intent classification, Hugging Face models for sentiment analysis, and Latent Dirichlet Allocation (LDA) for topic modeling. This approach supports the efficient management of comments by prioritizing those that require immediate attention and improving the collaborative process.
@inproceedings{Pasam_2025,
series = {SAC ’25},
title = {AI-Powered Comment Triage for Efficient Collaboration and Feedback Management},
url = {http://dx.doi.org/10.1145/3672608.3707835},
doi = {10.1145/3672608.3707835},
booktitle = {Proceedings of the 40th ACM/SIGAPP Symposium on Applied Computing},
publisher = {ACM},
author = {Pasam, Vamsi Krishna and Pati, Sravani and Toxtli Hernandez, Carlos},
year = {2025},
month = Mar,
pages = {971--979},
collection = {SAC ’25}
}Customer Journey Mapping with Multimodal Large Language Models
The expansion of e-commerce requires accurate alignment of diverse multimodal content-such as user reviews, product images, and videos-with customer expectations. This study introduces a framework for mapping customer feedback to stages of the customer journey: Awareness, Consideration, Decision, and Post-Purchase. By analyzing user-generated content using Gemini 1.5 Flash, key features like sentiment, emotional tone, and product attributes were extracted to build a unified dataset. A hybrid classification approach was adopted, leveraging machine learning and rule-based logic. Among the models evaluated, Meta-LLaMA 3.2 3B and Meta-LLaMA 3.2-1B, large language models with robust contextual understanding, demonstrated the highest accuracy at 0.929, surpassing other models such as GPT-Neo, BERT, and RoBERTa. These results show how well the model captures complex customer opinion and aligns multimodal insights. The findings highlight discrepancies between consumer expectations and marketing narratives, particularly about the deliberation stage-a crucial phase in building trust. This research offers scalable and interpretable approaches to enhance customer journey mapping and improve content alignment in e-commerce.
@inproceedings{Pati_2025,
series = {WWW ’25},
title = {Customer Journey Mapping with Multimodal Large Language Models},
url = {http://dx.doi.org/10.1145/3701716.3717863},
doi = {10.1145/3701716.3717863},
booktitle = {Companion Proceedings of the ACM on Web Conference 2025},
publisher = {ACM},
author = {Pati, Sravani and Pasam, Vamsi Krishna and Toxtli Hernandez, Carlos},
year = {2025},
month = May,
pages = {2744--2748},
collection = {WWW ’25}
}A Culturally-Aware AI Tool for Crowdworkers: Leveraging Chronemics to Support Diverse Work Styles
Crowdsourcing markets are expanding worldwide, but often feature standardized interfaces that ignore the cultural diversity of their workers, negatively impacting their well-being and productivity. To transform these workplace dynamics, this paper proposes creating culturally-aware workplace tools, specifically designed to adapt to the cultural dimensions of monochronic and polychronic work styles. We illustrate this approach with "CultureFit," a tool that we engineered based on extensive research in Chronemics and culture theories. To study and evaluate our tool in the real world, we conducted a field experiment with 55 workers from 24 different countries. Our field experiment revealed that CultureFit significantly improved the earnings of workers from cultural backgrounds often overlooked in design. Our study is among the pioneering efforts to examine culturally aware digital labor interventions. It also provides access to a dataset with over two million data points on culture and digital work, which can be leveraged for future research in this emerging field. The paper concludes by discussing the importance and future possibilities of incorporating cultural insights into the design of tools for digital labor.
@article{Toxtli2024Culturally,
title = {A Culturally-Aware AI Tool for Crowdworkers: Leveraging Chronemics to Support Diverse Work Styles},
volume = {8},
issn = {2573-0142},
url = {http://dx.doi.org/10.1145/3686899},
doi = {10.1145/3686899},
number = {CSCW2},
journal = {Proceedings of the ACM on Human-Computer Interaction},
publisher = {Association for Computing Machinery (ACM)},
author = {Toxtli, Carlos and Curtis, Christopher and Savage, Saiph},
year = {2024},
month = Nov,
pages = {1--34}
}Assessing the Task Management Capabilities of LLM-Powered Agents
This research delves into the capabilities of Large Language Model (LLM)-powered agents across five critical dimensions of task management: decomposition, scheduling, delegation, and execution. Leveraging both surveys and interaction experiments, this study aims to unearth practical insights into the applications and constraints of LLM-powered agents in real-world settings. We describe our experimental setup, data collection methodologies, and analytic techniques, offering a nuanced understanding of these agents' efficiency and efficacy. Our initial findings underscore the nuanced performance of LLM-powered agents in task management, revealing their strengths in understanding complex tasks and their limitations in execution without human intervention. This research contributes to the burgeoning field of human-AI collaboration by providing empirical evidence on the capabilities and limitations of LLM-powered agents in task management.
@misc{Perera2024Assessing,
doi = {10.13140/RG.2.2.11776.85768},
url = {https://www.researchgate.net/doi/10.13140/RG.2.2.11776.85768},
author = {{Ravindu Perera} and {Adithya Ravi} and Toxtli, Carlos},
language = {en},
title = {Assessing the Task Management Capabilities of LLM-Powered Agents},
publisher = {Unpublished},
year = {2024},
howpublished = {Preprint, ResearchGate},
note = {Preprint}
}AI Assistants in the Workplace: Goal-Oriented Recommendations Using LLM
Self-quantifying technology enables users to evaluate their performance and define strategies for improvement. Workplace technology assists users in identifying activities and patterns that facilitate task completion. Software that measures workplace signals requires access to information from the user interface and peripherals. Current software solutions that track computer activity lack goal orientation and do not share raw data that could aid users and researchers in analyzing behavioral patterns at work. This paper presents Wellbot, an intelligent AI assistant capable of tracking workers' activities and providing personalized insights. The solution employs machine learning models to detect goal-oriented and recreational time. Based on users' goals and recent activities, Wellbot generates recommendations aided by Large Language Models. This work aims to enable tools that assist workers through an improved understanding of their context and goals.
@article{Perera_2024,
title = {AI Assistants in the Workplace: Goal-Oriented Recommendations Using LLM},
volume = {9},
issn = {2594-2352},
url = {http://dx.doi.org/10.47756/aihc.y9i1.140},
doi = {10.47756/aihc.y9i1.140},
number = {1},
journal = {Avances en Interacción Humano-Computadora},
publisher = {Asociacion Mexicana de Interaccion humano-Computadora (AMexIHC)},
author = {Perera, Ravindu and Gendron, Claire and Delgado, Cecilia and Hernández, Alberto Campos and Muñoz, Victor Rios and Rogers, Mathew and Toxtli, Carlos},
year = {2024},
month = Nov,
pages = {16--20}
}The Use of AI-powered Language Tools in Crowdsourcing to reduce Language Barriers
Crowdsourcing platforms gather people from different backgrounds to work on completing microtasks. The workers on those platforms are presented with tasks defined by requesters from all over the world. The tasks are defined in multiple languages, with English being the predominant language. It is unclear how language barriers can affect workers whose primary language is not English while completing tasks. In this paper, we study their practices and the impact of AI-powered tools on the completion of microtasks. We conduct a pre-test test field experiment in which workers share their current experience with the platform and receive basic training in the use of AI-powered tools. After a week, workers provide feedback on their practices and experiences. We find that workers who complete tasks in their primary language report less confidence in completing English tasks, and in general, workers perceive the use of the tools as a language-learning mechanism. We aim that the understanding of language barriers in crowd markets can promote more inclusive work conditions.
@inproceedings{Delgado2024Use,
title = {The Use of AI-powered Language Tools in Crowdsourcing to reduce Language Barriers},
url = {http://dx.doi.org/10.1109/iThings-GreenCom-CPSCom-SmartData-Cybermatics62450.2024.00110},
doi = {10.1109/ithings-greencom-cpscom-smartdata-cybermatics62450.2024.00110},
booktitle = {2024 IEEE International Conferences on Internet of Things (iThings) and IEEE Green Computing & Communications (GreenCom) and IEEE Cyber, Physical & Social Computing (CPSCom) and IEEE Smart Data (SmartData) and IEEE Congress on Cybermatics},
publisher = {IEEE},
author = {Delgado-Solorzano, Cecilia and Toxtli, Carlos},
year = {2024},
month = Aug,
pages = {601--608}
}SmartMonitor: Edge-Based Activity Monitoring from Visual Input
In this paper, we propose SmartMonitor, a system that utilizes an edge hardware device to record, analyze and log user activity while not interfering with said activity. The analysis of the activity can help users identify the time they spend on different tasks and provide real-time feedback from a Large Language Model (LLM) to give better awareness of user activity. The SmartMonitor passes through the HDMI signal from the video card and analyzes the user’s activity on edge using two artificial intelligence models, logs user activity and sends the log for analysis to a Large Language Model for feedback. SmartMonitor enables non-intrusive self-quantifying technology that both records and analyzes user activity while protecting privacy and reducing the processing burden on the user’s device, which can serve as an excellent research framework for behavioral analysis in the workplace and a way to enhance user's work activity.
@misc{Li2024SmartMonitor,
doi = {10.13140/RG.2.2.20165.46561},
url = {https://www.researchgate.net/doi/10.13140/RG.2.2.20165.46561},
author = {{Wangfan Li} and {Ravindu Perara} and Gendron, Claire and Delgado-Solórzano, Cecilia and Toxtli, Carlos},
language = {en},
title = {SmartMonitor: Edge-Based Activity Monitoring from Visual Input},
publisher = {Unpublished},
year = {2024},
howpublished = {Preprint, ResearchGate},
note = {Preprint}
}Exploring AI-Enhanced Multi-Screen Interaction in Extended Reality Workspaces
This demo paper explores AI-enhanced multi-screen interaction within extended reality (XR) workspaces using called VirtuaScreens. The system facilitates the user-centric addition of customized virtual monitors and employs machine learning to understand user preferences for monitor arrangements and application arrangements. We employed Large language and multimedia models to offer context-sensitive feedback and enhance the user experience. We created this research framework called VirtuaScreens for researchers and practitioners to use to understand the interaction of the multiple screens in the XR environment.
@misc{Perera2024Exploring,
doi = {10.13140/RG.2.2.31604.36481},
url = {https://www.researchgate.net/doi/10.13140/RG.2.2.31604.36481},
author = {{Ravindu Perera} and Toxtli, Carlos},
language = {en},
title = {Exploring AI-Enhanced Multi-Screen Interaction in Extended Reality Workspaces},
publisher = {Unpublished},
year = {2024},
howpublished = {Preprint, ResearchGate},
note = {Preprint}
}Designing AI Tools to Address Power Imbalances in Digital Labor Platforms
The artificial intelligence industry has been essential in creating new jobs for the deployment of real-world solutions. As a result, the implementation of these new jobs involves the execution of multiple human intelligence micro-tasks, such as data labeling tasks for training machine learning models. The workers who perform those tasks, also known as crowd workers, usually are independent workers within crowdsourcing platforms. These platforms are subject to the free market, where the forces of supply and demand produce various power dynamics among stakeholders. As a result, disassociation between stakeholders often generates unbalanced power dynamics where workers are paid below minimum wage and are intimidated to keep their reputation or face termination. Within this chapter, we introduce computational techniques to audit the workplace conditions of crowd workers and design tools to address these power imbalances, as a positive and more efficient alternative for the labor conditions of crowd workers. This chapter develops these objectives through the design and evaluation of three tools in digital labor platforms: “Invisible Labor Tracker,” “Reputation Agent,” and “CultureFit,” which we describe below. We will demonstrate the sustainability of systems that point to a future where AI can be used to audit and address power imbalances in the workplace.
@inbook{Toxtli_2023,
title = {Designing AI Tools to Address Power Imbalances in Digital Labor Platforms},
isbn = {9783031316425},
issn = {2524-4477},
url = {http://dx.doi.org/10.1007/978-3-031-31642-5_9},
doi = {10.1007/978-3-031-31642-5_9},
booktitle = {Torn Many Ways},
publisher = {Springer International Publishing},
author = {Toxtli, Carlos and Savage, Saiph},
year = {2023},
pages = {121--137}
}