Part of the TPC Seminar Series

Speaker: Dexter Pratt , Director of Software Development
Date: Wednesday, April 10, 2024
Time: 10:00 A.M. to 11:15 A.M. (Central Time)
Location: Virtual
Abstract:
Generative AI agents that perform multifaceted tasks, such as data analysis, hypothesis generation, experiment design, and medical diagnosis, are arriving and evolving rapidly. Even now, large language models (LLMs) can generate apparently credible complex outputs, such as detailed analyses of proteomic datasets or medical advice in response to user queries.
But how much trust should we place in these agents? How do they compare to humans performing the same tasks? We also need to compare outputs between agents, so that we can measure progress or select the best agent for a task.
Existing practices to evaluate complex behavior are expensive and slow: it takes significant time and effort by humans to review scientific manuscripts, qualify physicians to practice, and approve research proposals. How can we automate agent evaluation to match the pace of AI development, where new LLMs and improved methods arrive every week?
In this talk, Mr. Pratt will discuss challenges in automating agent evaluation along with ideas for approaches.
Biography:
Dexter Pratt is the Director of Software Development for the Ideker laboratory at UCSD, where he leads the Cytoscape (cytoscape.org) and NDEx (Network Data Exchange, ndexbio.org) projects. Prior to joining UCSD, he worked in industry, developing systems biology tools such as the BEL (Biological Expression Language) language for representing molecular and phenotypic interactions to serve as a substrate for causal reasoning. His early career was in “old” AI, notably at the Cyc project, an attempt to build a rule-based commonsense reasoning agent at a massive scale. His work has now turned towards the development of LLM-based agents to play “robot scientist” roles, starting with the critical problem of assessing the behavior of these agents.

