Assessing LLMs for High Stakes Applications
• SEI Report
Publisher
Association for Computing Machinery (ACM)
DOI (Digital Object Identifier)
10.1145/3639477.3639720Topic or Tag
Abstract
Large Language Models (LLMs) promise strategic benefit for numerous application domains. The current state-of-the-art in LLMs, however, lacks the trust, security, and reliability which prohibits their use in high stakes applications. To address this, our work investigated the challenges of developing, deploying, and assessing LLMs within a specific high stakes application, intelligence reporting workflows. We identified the following challenges that need to be addressed before LLMs can be used in high stakes applications: (1) challenges with unverified data and data leakage, (2) challenges with fine tuning and inference at scale, and (3) challenges in reproducibility and assessment of LLMs. We argue that researchers should prioritize test and assessment metrics, as better metrics will lead to insight to further improve these LLMs.
Part of a Collection
AI Division Publications
Cite This SEI Report
Gallagher, S., Ratchford, J., Brooks, T., Brown, B., Heim, E., Nichols, B., McMillan, S., Rallapalli, S., Smith, C., VanHoudnos, N., Winski, N., & Mellinger, A. (2024, May 31). Assessing LLMs for High Stakes Applications. Retrieved August 12, 2026, from https://doi.org/10.1145/3639477.3639720.
@techreport{gallagher_2024,
author={Gallagher, Shannon and Ratchford, Jasmine and Brooks, Tyler and Brown, Bryan and Heim, Eric and Nichols, Bill and McMillan, Scott and Rallapalli, Swati and Smith, Carol and VanHoudnos, Nathan and Winski, Nick and Mellinger, Andrew},
title={Assessing LLMs for High Stakes Applications},
month={May},
year={2024},
institution={Software Engineering Institute, Carnegie Mellon University},
doi={10.1145/3639477.3639720},
url={https://doi.org/10.1145/3639477.3639720},
note={Accessed: 2026-Aug-12}
}
Gallagher, Shannon, Jasmine Ratchford, Tyler Brooks, Bryan Brown, Eric Heim, Bill Nichols, Scott McMillan, Swati Rallapalli, Carol Smith, Nathan VanHoudnos, Nick Winski, and Andrew Mellinger. "Assessing LLMs for High Stakes Applications." Software Engineering Institute, Carnegie Mellon University. Association for Computing Machinery (ACM), May 31, 2024. https://doi.org/10.1145/3639477.3639720.
S. Gallagher, J. Ratchford, T. Brooks, B. Brown, E. Heim, B. Nichols, S. McMillan, S. Rallapalli, C. Smith, N. VanHoudnos, N. Winski, and A. Mellinger, "Assessing LLMs for High Stakes Applications," Software Engineering Institute, Carnegie Mellon University. Association for Computing Machinery (ACM), 31-May-2024 [Online]. Available: https://doi.org/10.1145/3639477.3639720. [Accessed: 12-Aug-2026].
Gallagher, Shannon, Jasmine Ratchford, Tyler Brooks, Bryan Brown, Eric Heim, Bill Nichols, Scott McMillan, Swati Rallapalli, Carol Smith, Nathan VanHoudnos, Nick Winski, and Andrew Mellinger. "Assessing LLMs for High Stakes Applications." Software Engineering Institute, Carnegie Mellon University, Association for Computing Machinery (ACM), 31 May. 2024. https://doi.org/10.1145/3639477.3639720. Accessed 12 Aug. 2026.
Gallagher, Shannon; Ratchford, Jasmine; Brooks, Tyler; Brown, Bryan; Heim, Eric; Nichols, Bill; McMillan, Scott; Rallapalli, Swati; Smith, Carol; VanHoudnos, Nathan; Winski, Nick; & Mellinger, Andrew. Assessing LLMs for High Stakes Applications. Association for Computing Machinery (ACM). 2024. DOI: 10.1145/3639477.3639720. https://doi.org/10.1145/3639477.3639720