
Accelerating the Development and Use of Generative AI for Science and Engineering: The Trillion Parameter Consortium (TPC)
- 22 November 2024 (8:30am-12pm)
- Location: Georgia World Congress Center, Atlanta, Georgia, USA)
Large-scale AI foundation models show great potential for scientific discovery, with promising results being obtained in areas ranging from self-driving laboratories to hypothesis generation. But realizing this potential at scale will require intellectual advances as well as unprecedented quantities of both computation to train models and multidisciplinary human effort to prepare diverse scientific data for use in model training and to construct evaluation suites to guide development. Only a small number of organizations have the resources to build models at state-of-the-art scales (e.g., trillions of parameters, trained using tens of trillions of tokens).
This reality is already catalyzing multi-institutional teams working together on open projects ranging from model architecture, evaluation, and training to collaboratively building and sharing high-quality, open, training data sets. This workshop features twelve lightning talks highlighting such collaborations, which motivated the formation in 2023 of the international Trillion Parameter Consortium (TPC). The lightning talks will highlight progress in various aspects of generative AI for science and engineering with presentations from academics, national laboratories, HPC centers, industry, institutes, and leaders from funding agencies.
Roughly 150 people attended the workshop, which included 12 talks selected from over 30 submissions.
Program
08:30 – 10:00 Session 1: Training, Data, Safety, and Evaluation
Moderator: Valerie Taylor (Argonne and University of Chicago, USA)
Welcome and Introduction
Rick Stevens, Charlie Catlett (Argonne National Laboratory and University of Chicago)
- Establishing a Methodology to Evaluate Large Language Models for Science
- Combined Talk:
- Skills, Safety, and Trust Evaluation of Large Language Models for Science
Cappello, Madireddy (Argonne, USA), Aula-Blasco (BSC, Spain), Bhattacharya (HPE Labs, India), Wahib (RIKEN, Japan) - Comprehensive Multi-Stage Evaluation of Language Models for Scientific Skill and Safety Red-Teaming
Madireddy, Cappello (Argonne, USA), Xu (HPE Labs, Singapore), Bhattacharya, Kumar (HPE Labs, India), Kailkhura (LLNL, USA), Foltin (HPE Labs, USA), Li (University of Chicago, USA)
- Skills, Safety, and Trust Evaluation of Large Language Models for Science
- CheckEmbed: Effective Verification of LLM Solutions to Open-Ended Tasks
(Hoefler, ETHZ, Switzerland) - Enabling cross-Facility LLM pre-Training
Mittone, Mulone, Colonnelli, Birke, Aldinucci (University of Turin, Italy) - Preparing Data at Scale: The Data Pipeline for AuroraGPT
Chia, Cucinell, Foster, Mallick, Underwood, Babuji, Brace, Gokdemir, Hippe, Khan, Siebenschuh (Argonne, University of Chicago, USA) - llm-recipes: A Framework for Seamless Integration and Efficient Continual Pre-Training of Large Language Models
Fujii, Nakamura, Yokota (Institute of Science Tokyo, Japan)
10:00 – 10:30 Break
10:30 – 12:00 Session 2: Scientific Applications
Moderator: Neeraj Kumar (PNNL, USA)
- Training Large-Scale Vision Transformer Foundation Models for Science and Engineering Applications
Wahib, Igarashi (RIKEN, Japan), Chen (AIST, Japan), Lyngaas, Wang (ORNL, USA) - Attribution in Large Language Models
Batista, Wahib (RIKEN, Japan) - Distributed document deduplication over slurm-based HPC environments
Rubio Pintado (BSC, Spain) - Agents for Climate Change Mitigation and Adaptation in Cities
Elhacham (BSC, Spain), Foster (Argonne, University of Chicago, USA) - “Fusion GPT” — A 1.5 B Parameter Foundation Model
Tang, Rodriguez (Princeton University, USA), Felker (Argonne, USA), Gomez (Harvard University, USA) - G-LED: Generative AI for Learning the Effective Dynamics of High-dimensional, Complex System
Gao, Kaltenbach, Koumoutsakos (Harvard University, USA) - Driving Autonomous Experiments and Molecular Explorations aided by Virtual Foundation Model OS
Foltin, Bhattacharya, Xu, Saranathan, Justine, Kumar, Shah, Tripathy, Faraboschi (HPE, USA), Ziatdinov (PNNL, USA), Ghosh, Roccapriore (ORNL, USA), Foster (Argonne, University of Chicago, USA) - Advancing Foundation Models in Earthquake Nowcasting
Fox, Jafari (University of Virginia, USA), Rundle (University of California-Davis, USA), Donnellan (Jet Propulsion Laboratory, USA), Grant Ludwig (University of California-Irvine, USA)
Program Committee
- Suparna Bhattacharya, HPE Labs (India)
- Jérôme Bobin, CEA (France)
- Charlie Catlett Argonne National Laboratory and University of Chicago (USA)
- Ian Foster, Argonne National Laboratory and University of Chicago (USA)
- Fabrizio Gagliardi, Barcelona Supercomputing Center (Spain)
- Neeraj Kumar, Pacific Northwest National Laboratory (USA)
- Satoshi Matsuoka, RIKEN Center for Computational Sciences (Japan)
- Paul Messina, Argonne National Laboratory (USA)
- Laura Morselli, CINECA (Italy)
- Irina Rish, Université de Montréal and Mila (Canada)
- Noah Smith, Allen Institute for Artificial Intelligence and University of Washington (USA)
- Rick Stevens, Argonne National Laboratory and University of Chicago (USA)
- Valerie Taylor, Argonne National Laboratory and University of Chicago (USA)
- Cong Xu, HPE Labs (USA)
- Rio Yokota, Institute of Science Tokyo (formerly Tokyo Tech) (Japan)

