Workshop: Harnessing the Power of Generative AI for Science and Engineering – The Trillion Parameter Consortium (TPC)

  • 10-March-2025
  • Venue: SCA 2025 in Singapore

Abstract

 Large-scale ’foundation’ AI models show great promise for scientific discovery, with promising results being obtained in areas ranging from self-driving laboratories to hypothesis generation. But realizing this promise at scale will require unprecedented quantities of both computation to train models and multidisciplinary human effort to prepare diverse scientific data for use in model training and to construct evaluation suites to guide development. Only a small number of organizations have the resources to build models at state-of-the-art scales (e.g., trillions of parameters, trained using tens of trillions of tokens). This reality is already motivating the formation of multi-institutional teams to work together on model architecture, evaluation, and training as well as on the collaborative building and sharing of high-quality training data sets. This workshop will highlight such collaborations, which are being catalyzed by the international Trillion Parameter Consortium (TPC). The workshop will highlight progress in various aspects of generative AI for science and engineering with presentations from academics, national laboratories, HPC centers, industry, institutes, and leaders from funding agencies. The workshop will also introduce the structure and strategies of the TPC, with an overview of high-priority areas in which new collaborators can contribute and benefit from joining the consortium.

Realizing the promise of large-scale, ‘foundation’ models for scientific discovery, supporting activities ranging from self-driving laboratories to hypothesis generation, will require unprecedented scale of computation to train large-scale AI models, coupled with an enormous, inherently multidisciplinary task of preparing diverse scientific data for use in model training. Simply put, only a relatively small number of organizations have the resources to build models at state-of-the-art scales (e.g., trillions of parameters, trained using tens of trillions of tokens).

These trends are motivating the formation of multi-institutional teams whose efforts can be accelerated through sharing strategies such as for model architecture, evaluation, and training as well as through collaboratively building and sharing high-quality training data sets. This workshop will highlight such collaborations, which are being catalyzed by the international Trillion Parameter Consortium (TPC), along with progress in various aspects of generative AI for science and engineering with presentations from academics, national laboratories, HPC centers, industry, institutes, and leaders from funding agencies.

This workshop will also provide a forum for continuing to develop a shared vision and goals for accelerating the use of generative AI for science and engineering, particularly providing the opportunity for participation by the European AI, HPC, and disciplinary science research communities. The workshop will introduce the structure and strategies of the TPC and an overview of high-priority areas in which new collaborators can contribute and benefit from joining the consortium.

Provisional Program

09:00-12:00 Opening Session: Large-Scale AI Model Development Around the World

  • Opening Remarks (Jens Domke, RIKEN CCS, Japan)
  • Keynote (Satoshi Matsuoka, RIKEN CCS, Japan)
  • Invited Talk (Arvind Ramanathan, Argonne National Laboratory and The University of Chicago, USA)
  • Invited Talk (Li Xiaoli, Agency for Science, Technology and Research (A*STAR), Singapore)
  • Invited Talk (Kimmo Koski, CSC – IT Center for Science, Finland)

Lunch Break

13:30-14:50 Selected Talks from Open Submissions Call

  • TPC Status and Roadmap (Charlie Catlett, Argonne National Laboratory and The University of Chicago, USA)
  • Algebraic Approaches to Combining Multiple Large Language Models (J. de Curtò, BSC-CNS, Spain)
  • MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models (Shuo Sun, A*STAR, Singapore)
  • Automated Detection of AI Training Jobs to Enhance Security In HPC Systems (Francesco Antici, University of Bologna, Italy)
  • Scientific Data Compression for Large Language Models (Maximilian Sander, Dresden University of Technology, Germany)
  • Advancing Autonomous Microscopy Agents with domain guided dynamic retrieval in a Virtual Foundation Model OS (Gayathri Saranathan, HP Labs, USA)
  • Faithful Reasoning over Scientific Claims (Mark Gahegan, University Aukland, New Zealand)
  • Concluding Remarks (Mohamed Wahib, RIKEN CCS, Japan)

Workshop Aims and Expected Outcomes

Today the science community has a growing number of projects aiming to harness generative AI, ranging from models trained for individual disciplines (e.g., UniverseTBD for astronomy or MAIRA-1 for radiological image analysis) to models aspiring to multi-disciplinary use (e.g., Olmo, AuroraGPT). There are also new data sources being released, such as DOLMA, and many efforts to identify scientific data (literature, databases, time series data, etc.) and transform those data for use training AI models.  There has been an explosion of new activities, each developing methods and tools to grapple with common challenges, from evaluation (for trustworthiness, safety, bias, performance, etc.) to training and data preparation workflows.  Concurrently, the enormous cost of computation for model training limits the number of groups that can realistically build and train large-scale models.  These trends all offer incredible opportunities to collaborate, both to accelerate progress toward optimal tools and methods and to enable groups to strategically collaborate to reduce duplication of effort, such as in data preparation of new scientific data sources. Additionally, many new challenges are coming to the fore with generative AI, including new ways of thinking about data sharing and attribution, about licensing artifacts, and about the importance of responsible development safe, ethical AI models.  An overarching need today in AI is openness – which includes open data, open source code for tools and workflows, and careful thought as to how, when, and whether to open the models themselves. Such openness is critical to progress in every area related to generative AI.

The enormity of these challenges, and of the resources needed for data preparation, pre-training new models, and responsibly preparing them for downstream applications has meant that progress is largely concentrated in industry, where there is limited, or in some cases, no visibility into the artifacts (models, data sets) or the processes used to create them.  This underscores the need for collaboration in the open science community—central to the motivation behind creating the international Trillion Parameter Consortium (TPC). The workshop aspires to stimulate new thinking, attract scientists to some of the emerging challenges associated with generative AI, and potentially catalyze the formation of new topical collaborative working groups that can be supported by TPC.

Impact

The trends discussed above put HPC at the center of the community’s quest to use generative AI for science, bringing together HPC and AI experts to discuss current trends and breakthroughs and explore potential collaborations.  The international Trillion Parameter Consortium (TPC) was formed in 2023 for this very purpose—to provide the community with a venue for identifying and pursuing collaborations that will accelerate progress and that will enable the broadest possible community to benefit from the limited HPC resources available to train large AI models.

A fundamental goal of this workshop is to attract early- and mid-career scientists, inclusive of the breadth of the SC Asia community from HPC providers to computer and computational scientists to disciplinary scientists and encompassing those involved in career development and training programs.   

For more information about the Trillion Parameter Consortium (TPC) please visit TPC.dev. If you are interested in participating in the TPC community, please contact us here.

Co-Chairs

  • Jens Domke, RIKEN-CCS (co-chair)
  • Mohamed Wahib, RIKEN-CCS (co-chair)

International Steering Group

  • Rick Stevens, Charlie Catlett, Ian Foster, and Valerie Taylor, Argonne National Laboratory
  • Fabrizio Gagliardi, Barcelona Supercomputing Center
  • Neeraj Kumar, Pacific Northwest National Laboratory
  • Satoshi Matsuoka, RIKEN Center for Computational Sciences
  • Irina Rish, Université de Montréal and Mila
  • Noah Smith, Allen Institute for Artificial Intelligence and University of Washington
  • Sean Smith, Australian National University

Speaker Biographies

Co-Chair: Jens Domke (Team Leader, RIKEN Center for Computational Science, Japan)

Jens Domke is the Team Leader of the Supercomputing Performance Research Team at the RIKEN Center for Computational Science (R-CCS), Japan. He received his doctoral degree from the TU Dresden, Germany, in 2017 for his work on HPC routing algorithms and interconnects. Jens contributed the DFSSSP and Nue routing algorithms to the subnet manager of InfiniBand, and built the first large-scale HyperX prototype at the Tokyo Institute of Technology. His research interests include system co-design, performance evaluation, extrapolation, and modelling, interconnect networks, and optimization of AI frameworks, parallel applications and architectures.

Co-Chair: Mohamed Wahib (Team Leader, RIKEN Center for Computational Science, Japan)

Mohamed Wahib received the PhD degree in computer science from Hokkaido University, Japan, in 2012. He is currently the team leader of the High Performance Artificial Intelligence Systems Research Team at the RIKEN Center for Computational Science (R-CCS), Japan. Prior he was a senior scientist with AIST/TokyoTech Open Innovation Laboratory, Tokyo, Japan, and a researcher at R-CCS. Before his graduate studies, Mohamed researched with Texas Instruments (TI) R&D Labs, Dallas, TX, USA, for four years. His research interests include the central topic of large-scale AI deployments in HPC infrastructures and performance-centric software development in the context of HPC.

Satoshi Matsuoka (Director, RIKEN Center for Computational Science, Japan)

Professor Satoshi Matsuoka from April 2018 has been the director of Riken Center for Computational Science (R-CCS), the Tier-1 national HPC center for Japan, developing and hosting Japan’s flagship ‘Fugaku’ supercomputer which has become the fastest supercomputer in the world in 2020 and 2021, supporting cutting edge HPC research, including investigating Post-Moore era computing, especially the future FugakuNEXT supercomputer. He led the TSUBAME series of supercomputers that received many international acclaims, at the Tokyo Institute of Technology, where he holds a professor position pursuing research in HPC, scalable Big Data, and AI. His longtime contribution was commended with the Medal of Honor with Purple ribbon by his Majesty Emperor Naruhito of Japan in 2022. He is a Fellow in ACM, ISC, IPSJ and the JSSST and has won numerous awards including ACM Gordon Bell Prizes, the IEEE-CS Sidney Fernbach Award, and the IEEE-CS Computer Society Seymour Cray Computer Engineering Award.

Arvind Ramanathan (Computational Biologist, Argonne National Laboratory and The University of Chicago, USA)

Arvind Ramanathan is a computational biologist in the Data Science and Learning Division at Argonne National Laboratory and a senior scientist at the University of Chicago Consortium for Advanced Science and Engineering (CASE). His research interests are at the intersection of data science, high performance computing and biological/biomedical sciences. His research focuses on three areas focusing on scalable statistical inference techniques: (1) for analysis and development of adaptive multi-scale molecular simulations for studying complex biological phenomena (such as how intrinsically disordered proteins self assemble, or how small molecules modulate disordered protein ensembles), (2) to integrate complex data for public health dynamics, and (3) for guiding design of CRISPR-Cas9 probes to modify microbial function(s). He has published over 30 papers, and his work has been highlighted in the popular media, including NPR and NBC News. He obtained his Ph.D. in computational biology from Carnegie Mellon University, and was the team lead for integrative systems biology team within the Computational Science, Engineering and Division at Oak Ridge National Laboratory.  More information about his group and research interests can be found at http://​ramanathanlab​.org.

Charlie Catlett (Senior Computer Scientist, Argonne National Laboratory and The University of Chicago, USA)

Charlie Catlett is a Senior Computer Scientist at the U.S. Department of Energy’s Argonne National Laboratory, and a Visiting Scientist at the University of Chicago. His research focuses on building cyberinfrastructure to embed edge-AI in urban, environmental, and emergency sensing and response settings. At UChicago he created the Urban Center for Computation and Data (UrbanCCD), an interdisciplinary urban sciences research center. From 1998-2005 he was founding chair of Grid Forum / Global Grid Forum and director of NSF’s TeraGrid initiative from 2004-2007. Charlie was part of the team that established the National Center for Supercomputing Applications (NCSA) in 1985, leading efforts there including the deployment and operation of the NSFNET backbone network, an early component of the Internet, and serving as Chief Technology Officer prior to joining Argonne and UChicago in 2000. He was one of GovTech magazine’s “25 Doers, Dreamers & Drivers” of 2016 and in 2019 received the Argonne Board of Governors Distinguished Performer award. Charlie is a Computer Engineering graduate of the University of Illinois at Urbana-Champaign.

J. de Curtò (BSC-CNS, Spain)

Shuo Sun (A*STAR, Singapore)

Francesco Antici (University of Bologna, Italy)

Maximilian Sander (Dresden University of Technology, Germany)

Gayathri Saranathan (HP Labs, Singapore)

Mark Gahegan (University of Aukland, New Zealand)

TBD (TBD)