Patrick Emami, Sameera Horawalavithana, Truc Nguyen +11cs.AI cs.HC
Large language model-based agents are increasingly deployed as collaborators in scientific discovery yet most current work focuses on the autonomous capabilities of "AI Scientists". We argue that this overlooks the social aspects of scientific teamwork, and that studying AI Scientists as human-agent systems (HAS)--where the unit of analysis is the human-agent pair--is both underexplored and undervalued. We establish these points through literature and empirical analysis, and highlight recent incidences and studies which show that deploying agents in science without accounting for human-agent dynamics introduces near-term risks, including reduced diversity of scientific inquiry. Through analysis of real-world case studies, we show that scientists and agents can augment each other's capabilities. We call for new research that adopts the HAS lens to develop mathematical frameworks for understanding and fostering human-AI synergy in scientific discovery.
Raul Jimenez, Boris Bolliet, Francisco Villaescusa-Navarro +7cs.AI astro-ph.IM physics.soc-ph
Agentic artificial intelligence (AI) systems are beginning to assist, accelerate, and partially automate scientific discovery, performing tasks that span literature synthesis, code generation, data analysis, hypothesis proposal, and model criticism. We argue that this transition is qualitative rather than incremental, and that suitably designed multi-agent systems may evolve from passive computational tools into ``AI scientists'' that can expand the hypothesis-generating and verification capacity of science. Such systems must be developed and deployed within a scientific ecosystem fit for purpose: institutions must be redesigned for verification, accountability, interpretability, and dual-use safety. We sketch how multi-agent architectures, illustrated by the prototype framework \textit{Denario}, accelerate the discovery cycle and traverse model spaces beyond human reach; examine what this implies for authorship, peer review, and the enduring role of human scientists; and close with recommendations for governing AI as an epistemic actor rather than a mere instrument.
Bold claims that AI will accelerate scientific discovery have raced ahead of evidence from working scientists, yet large-scale, scientist-in-the-loop evidence is scarce. Here we mount the largest evaluation to date, inviting authors of 121,640 recent preprints in biology, medicine, chemistry, and social science to judge large language model (LLM)-generated ideas derived from their own papers. 6,749 representative scientists returned 25,139 rating sets on novelty, feasibility, probability of being true, and favorability of adoption. Three patterns emerge. First, non-reasoning LLMs collapse into a narrow "hivemind" of similar ideas while reasoning models explore a wider hypothesis space, but no model spontaneously proposes null hypotheses, a move humans make more freely. Second, scientists reward ideas resembling their own and prize probability over novelty, though social scientists tolerate risk more than life scientists; senior social scientists are the harshest critics, and their skepticism is earned, as LLMs falter most in pluralistic fields demanding context-aware interpretation and evolving theories. Third, automated evaluators, including LLM-as-a-judge and state-of-the-art (SOTA) models, agree weakly with expert judgment. Retrieval augmentation and scientist persona prompting yield marginal gains. A Qwen3-14B reward model we post-trained on human ratings captures nuances of taste, beats SOTA models by up to 27%, and closes the gap to the consistency of human peer reviewers. An analysis of 39 million papers from 2010 to 2025 links survey findings to macro-level patterns: following ChatGPT's release, null claims are sharply suppressed and ideas contract. Agent-based simulations further suggest that saturated fields should especially prize human uniqueness. For all the hype, today's AI for science remains a collaborator whose imagination and judgment benefit from human grounding.