Virtual Workshop 2026 - AI Meets CI - Day 3 - Talk I: Considerations for Testing and Evaluation of AI Systems by Emily Castleton, Los Alamos National Laboratory

Talk I: Considerations for Testing and Evaluation of AI Systems, Emily Castleton, statistician, Los Alamos National Laboratory

Abstract from Emily Castleton: In this talk I will discuss some best practices and frameworks for testing and evaluating AI systems; in particular, best practices for benchmark creation, bespoke metrics for bespoke models, uncertainty quantification of metric values, and evaluation that goes beyond a standard leaderboard. These topics will be demonstrated using examples from large language models, vision-language models, seismic AI models, and multi-physics computer model emulation.

View Video


The NSF CI Compass virtual workshop - AI Meets CI: Intelligent Infrastructure for Major & Midscale Facilities featured themes to host speakers, presentations, and discussions surrounding artificial intelligence (AI) and cyberinfrastructure (CI) within the data lifecycle in facilities, cloud computing, and about societal and human intersections with AI.

Originally hosted January 12, 13, and 14, 2026

Learn more about the workshop here