AI for Science-领域全景与核心共识-关于 AI for Science,现有研究形成了哪些较可信且容易理解的核心结论-主要有哪些研究分支和代表性证/rapid-brief.md

AI for Science:领域全景与核心共识

证据级别:摘要级快速简报。 仅依据公开摘要,未核验全文方法、实验设置与结论边界。

基于公开摘要的证据扫描,AI for Science 已形成关于自主闭环、数据瓶颈、生成式设计、知识嵌入、科学 LLM 与开放扩散机制等若干可辨识的核心认识;这些认识多为综述层的一致报告,单篇量化结果仍待全文核验与复现。

当前摘要支持的核心认识

自主科研系统已从设想走向端到端闭环,是当前 AI for Science 最突出的演进方向

截至 2025 年的综述将“智能体科学(Agentic Science)”定位为 AI for Science 的演进阶段,即 AI 从部分协助走向完全科学能动性;自驱动实验室综述则报告,最先进的 SDL 已能自动化假设生成、实验设计、实验执行、数据分析、结论更新乃至下一轮假设迭代的几乎完整科学闭环。两篇综述在“AI 正在从任务特定工具变成自主研究伙伴”这一点上相互印证。

为什么重要: 这改变了读者评估相关研究时的基本预期:不应只看单项模型精度,而应关注自主闭环、数据与实验基础设施的整合程度。 证据边界: 仅基于两篇公开摘要的作者主张,未核验 SDL 的完整方法细节与真实自动化程度;不同实验室的自动化水平差异很大。 摘要证据: From AI for Science to Agentic Science: A Survey on Autonomous Scientific DiscoveryAutonomous ‘self-driving’ laboratories: a review of technology and policy implications

数据可用性是当前摘要扫描下最一致的 AI for Science 瓶颈

电池与能源存储综述明确列出数据短缺、网络基础设施不足、数据隐私、知识产权和伦理问题;材料发现综述将数据稀缺列为关键挑战;另有评论以“是否持续供应大规模高可用数据”作为释放 AI 科学潜力的首要问题;SciHorizon 还提出了包含质量、FAIR 性、可解释性和合规性的 AI 就绪数据评估框架。多篇来自不同领域的综述或评论独立指向同一卡点。

为什么重要: 帮助读者识别 AI4Science 的真正卡点:数据治理不是边缘问题,而是决定模型能否落地的首要条件;科研资助和项目设计应优先投资数据基础设施。 证据边界: 各论文领域和视角不同,未对具体瓶颈优先级做统一排序;AI 就绪数据评估框架属于单篇工作提出的框架。 摘要证据: AI for science in electrochemical energy storage: A multiscale systems perspective on transportation electrificationArtificial Intelligence and Generative Models for Materials Discovery -- A ReviewUnleashing the power of AI in science-key considerations for materials data preparationSciHorizon: Benchmarking AI-for-Science Readiness from Scientific Data to Large Language Models

生成式与逆设计是药物和材料两大核心分支的共同主线,但后期验证仍是共同短板

药物发现综述报告生成化学、机器学习和多属性优化已使若干化合物进入临床试验;材料科学综述报告生成模型可按目标性能从头设计催化剂、半导体、聚合物和晶体。两个分支的综述均把数据稀缺、可解释性以及后期可合成/临床验证列为未解瓶颈,说明生成式范式的进展与短板是跨分支共有的。

为什么重要: 有助于读者用“从生成到验证是否闭环”作为统一标尺,评估不同材料/药物方向的成熟度。 证据边界: 进入临床试验、可合成材料的具体数量与标准未在摘要中给出;部分进展来自单篇综述的作者报告;未核验任何具体药物管线或材料合成结果。 摘要证据: A survey of generative AI for <i>de novo</i> drug design: new frontiers in molecule and protein generationArtificial Intelligence for Drug Discovery: Are We There Yet?Artificial intelligence in drug discovery: A comprehensive review with a case study on hyperuricemia, gout arthritis, and hyperuricemic nephropathyArtificial Intelligence and Generative Models for Materials Discovery -- A ReviewA Survey on Graph Diffusion Models: Generative AI in Science for Molecule, Protein and Material

将领域知识嵌入神经网络是一条标志性的 AI for Science 方法学路线

物理信息神经网络(PINNs)通过把物理定律嵌入网络结构或损失函数来求解偏微分方程等复杂物理系统,被综述视为科学计算与深度学习中的变革性框架;另一篇综述则系统梳理了通过修改输入、损失函数和架构来纳入领域知识的做法,并报告这些技术能够显著改变深度神经网络性能。

为什么重要: 它揭示 AI4Science 并非只会“数据拟合”,而是存在把物理第一性原理与数据结合的一整套方法学分支,直接影响建模路线的选择。 证据边界: 两篇综述均为方法分类而非系统比较;“显著改变性能”不等于“一定提升”,具体增益因任务而异;未核验任何 PINNs 基准结果。 摘要证据: Physics-informed neural networks for PDE problems: a comprehensive reviewA review of some techniques for inclusion of domain-knowledge into deep neural networks

科学大语言模型已从零散建模走向有分类体系和基准评估的阶段

一项生物/化学领域综述按模型架构、能力、数据集和评估维度,系统梳理了面向文本知识、小分子、蛋白质、基因组序列的科学 LLM;SciHorizon 则从 AI 就绪数据与 LLM 科学能力(知识、理解、推理、多模态、价值观)两个角度提出评估框架,并已测评 50 多个开源与闭源 LLM。

为什么重要: 说明科学 LLM 领域已有分类与基准意识,评估科学 LLM 强弱需要使用专门基准,而不是沿用通用 NLP 评测。 证据边界: 科学 LLM 综述仅覆盖生物/化学;SciHorizon 的具体评估结果未在摘要中呈现;两者均属单篇工作,尚未成为统一共识。 摘要证据: Scientific Large Language Models: A Survey on Biological &amp; Chemical DomainsSciHorizon: Benchmarking AI-for-Science Readiness from Scientific Data to Large Language Models

AI 要从示范性创新扩散为普遍采用的科研方法,关键在开放数据科学这一“扩散引擎”而非算法本身

有论文提出需要构建跨学科思想供应链、通过开放研究快速转移技术能力、开发赋权研究者的 AI 工具并嵌入有效数据管理;Dagstuhl 研讨会报告也将跨学科共同体和人类—机器协作视为下一波进展的重要来源。两类来源都把开放基础设施与社区建设放在算法突破同等重要的位置。

为什么重要: 这改变了推进策略判断:仅增加算力或模型投入不够,开放基础设施、数据管理和跨学科社区是同等重要的条件。 证据边界: 这些来自观点性/综述性文献和研讨会报告,属于作者与参会社区的主张,缺乏对这些干预措施效果的实证检验。 摘要证据: Accelerating AI for science: open data science for scienceAI for Science: An Emerging AgendaAdvanced Research Directions on AI for Science, Energy, and Security: Report on Summer 2022 Workshops

自主 AI 科研系统已引发超出学术界的治理与认识论挑战

综述报告“云实验室”与自驱动实验室引出 AI 发明的可专利性问题(现行专利法只承认人类发明人)、安全与网络安全风险以及对技术劳动力的替代与创造效应;哲学讨论则指出,科学家对本质上不透明的 AI 应用形成认识论依赖,但 AI 不是可问责的行动者,现有科学信任理论难以覆盖这种关系。

为什么重要: 提醒读者,AI4Science 的成熟度不仅取决于性能指标,还取决于专利、安全、劳动力和科学信任等制度框架能否跟上。 证据边界: 劳动力影响的估计来自单篇综述的自身分析;哲学论文提出的是概念性批判而非实证结论;未核验任何专利判例或安全事件。 摘要证据: Autonomous ‘self-driving’ laboratories: a review of technology and policy implicationsWe Have No Satisfactory Social Epistemology of AI-Based ScienceA Survey of the Potential Long-term Impacts of AI

在控制方程发现上,符号—进化混合 AI 报告了可解释性与外推性的量化增益

一篇论文报告,结合符号主义与元启发式的“机器集体智能”方法可在确定性、随机和未表征动力学系统中自主恢复控制方程,将外推误差相对深度神经网络最多降低 6 个数量级,并把模型参数从约 50 万至 100 万压缩到 5 至 40 个可解释参数。

为什么重要: 这是一个具体、可理解的量化证据,说明符号—进化混合 AI 可能在可解释性和外推性上弥补纯深度学习的短板,值得关注和复现。 证据边界: 仅基于单篇公开摘要的作者报告,未核验具体实验设置、数据集和基线选择;结果不一定泛化到其他科学发现任务。 摘要证据: Machine Collective Intelligence for Explainable Scientific Discovery

当前研究版图

把当前摘要扫描中的文献按问题入口重排,可以看到 AI for Science 的研究版图主要由五类问题汇聚而成。第一类是“如何让机器自主完成科研闭环”:截至 2025 年的综述将 Agentic Science 定位为 AI 从部分协助走向完全科学能动性的阶段,自驱动实验室综述则报告最先进系统已经覆盖从假设生成到下一轮假设更新的几乎完整闭环(C1)。第二类是“如何为 AI 准备科学数据”:电池与能源存储综述、材料发现综述以及数据就绪度评论从各自领域把数据短缺与数据可用性列为最一致的瓶颈(C2)。第三类是“如何从目标性质生成候选分子与材料”:药物发现综述报告生成化学、机器学习与多属性优化已推动若干化合物进入临床试验,材料发现综述报告生成模型可按目标性能从头设计催化剂、半导体、聚合物和晶体(C3、C9)。第四类是“如何把科学知识嵌入神经网络”:PINNs 综述把物理定律嵌入网络结构与损失函数视为科学计算与深度学习中的变革性框架,另一篇综述系统梳理了输入、损失函数和架构三类知识注入途径(C4)。第五类是“如何用并评估科学 LLM”:生物/化学领域的科学 LLM 综述按架构、能力、数据集与评估维度做了系统梳理,SciHorizon 则从 AI 就绪数据和 LLM 科学能力两个角度提出评估框架并测评了 50 多个模型(C5)。在这五类主线的外围,还有文献综述自动化、大科学工作流嵌入和危机响应等应用分支(C10、C11、C12)。

证据类型:跨摘要综合 · 摘要证据:From AI for Science to Agentic Science: A Survey on Autonomous Scientific DiscoveryAutonomous ‘self-driving’ laboratories: a review of technology and policy implicationsAI for science in electrochemical energy storage: A multiscale systems perspective on transportation electrificationArtificial Intelligence and Generative Models for Materials Discovery -- A ReviewUnleashing the power of AI in science-key considerations for materials data preparationSciHorizon: Benchmarking AI-for-Science Readiness from Scientific Data to Large Language ModelsA survey of generative AI for <i>de novo</i> drug design: new frontiers in molecule and protein generationArtificial Intelligence for Drug Discovery: Are We There Yet?Artificial intelligence in drug discovery: A comprehensive review with a case study on hyperuricemia, gout arthritis, and hyperuricemic nephropathyA Survey on Graph Diffusion Models: Generative AI in Science for Molecule, Protein and MaterialPhysics-informed neural networks for PDE problems: a comprehensive reviewA review of some techniques for inclusion of domain-knowledge into deep neural networksScientific Large Language Models: A Survey on Biological &amp; Chemical DomainsExpert-Guided LLM Reasoning for Battery Discovery: From AI-Driven Hypothesis to Synthesis and CharacterizationThe emergence of large language models as tools in literature reviews: a large language model-assisted systematic reviewOpportunities in AI/ML for the Rubin LSST Dark Energy Science CollaborationMapping the Landscape of Artificial Intelligence Applications against COVID-19Artificial Intelligence for COVID-19: Rapid Review

这些入口并非并列的独立方向,而是围绕同一条科研流水线分布的:数据准备在前,知识注入与生成设计居中,自主闭环试图把这些环节串起来;科学 LLM 既是新的建模工具,也催生了新的评估问题(C1、C2、C3、C4、C5)。材料生成的综述还把数据稀缺、可解释性与可合成性列为挑战,并指出多模态模型、物理信息架构和闭环发现系统是克服这些限制的新兴路线,说明 C2、C3、C4 的边界在具体分支中正在互相渗透。

证据类型:跨摘要综合 · 摘要证据:From AI for Science to Agentic Science: A Survey on Autonomous Scientific DiscoveryAutonomous ‘self-driving’ laboratories: a review of technology and policy implicationsAI for science in electrochemical energy storage: A multiscale systems perspective on transportation electrificationArtificial Intelligence and Generative Models for Materials Discovery -- A ReviewUnleashing the power of AI in science-key considerations for materials data preparationSciHorizon: Benchmarking AI-for-Science Readiness from Scientific Data to Large Language ModelsA survey of generative AI for <i>de novo</i> drug design: new frontiers in molecule and protein generationArtificial Intelligence for Drug Discovery: Are We There Yet?Physics-informed neural networks for PDE problems: a comprehensive reviewA review of some techniques for inclusion of domain-knowledge into deep neural networksScientific Large Language Models: A Survey on Biological &amp; Chemical Domains

对比两分支:以自驱动实验室为代表的自主闭环路线,与以生成模型加逆设计为代表的离线设计路线,回答的是不同问题。自主闭环路线把问题定义成“让机器在同一个实验周期内完成假设—实验—分析—再假设”,其证据由 SDL 综述、Agentic Science 综述和 ChatBattery 的完整闭环案例共同支撑(C1、C9);离线生成路线把问题定义成“给定目标性质,从头生成候选分子/材料”,其证据来自药物与材料两个分支的综述(C3)。两者共享 C2 与 C3 中反复出现的数据稀缺、可解释性和后期可合成/临床试验验证瓶颈;差别主要在前者强调实验基础设施与硬件集成,后者强调分子/材料表示与生成模型架构。

证据类型:跨摘要综合 · 摘要证据:From AI for Science to Agentic Science: A Survey on Autonomous Scientific DiscoveryAutonomous ‘self-driving’ laboratories: a review of technology and policy implicationsAI for science in electrochemical energy storage: A multiscale systems perspective on transportation electrificationArtificial Intelligence and Generative Models for Materials Discovery -- A ReviewA survey of generative AI for <i>de novo</i> drug design: new frontiers in molecule and protein generationArtificial Intelligence for Drug Discovery: Are We There Yet?Expert-Guided LLM Reasoning for Battery Discovery: From AI-Driven Hypothesis to Synthesis and Characterization

这张版图是在摘要层面对研究问题入口的整理,不等同于各分支的成熟度排序;C8、C9、C10、C11 这类单篇作者主张只能作为分支内的单个证据,不能单独支撑“某分支已经成熟”的结论。

证据类型:有边界的编辑判断 · 摘要证据:Machine Collective Intelligence for Explainable Scientific DiscoveryExpert-Guided LLM Reasoning for Battery Discovery: From AI-Driven Hypothesis to Synthesis and CharacterizationThe emergence of large language models as tools in literature reviews: a large language model-assisted systematic reviewOpportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration

这些结论能相信到什么程度

要判断上述结论可以相信到什么程度,最直接的区分是证据性质而非结论数量。C1、C2、C3、C4、C5、C6、C12 属于跨论文/综述层面的综合,当前摘要扫描显示多个领域对同一判断有重复报告,较适合用来建立领域地图;C8、C9、C10、C11 属于单篇作者主张,即使给出量化指标(如外推误差降低 6 个数量级、容量提高 28.8%、GPT 模型精确度 83%),也只应视为候选证据。

证据类型:跨摘要综合 · 摘要证据:From AI for Science to Agentic Science: A Survey on Autonomous Scientific DiscoveryAutonomous ‘self-driving’ laboratories: a review of technology and policy implicationsAI for science in electrochemical energy storage: A multiscale systems perspective on transportation electrificationArtificial Intelligence and Generative Models for Materials Discovery -- A ReviewUnleashing the power of AI in science-key considerations for materials data preparationSciHorizon: Benchmarking AI-for-Science Readiness from Scientific Data to Large Language ModelsPhysics-informed neural networks for PDE problems: a comprehensive reviewA review of some techniques for inclusion of domain-knowledge into deep neural networksScientific Large Language Models: A Survey on Biological &amp; Chemical DomainsAccelerating AI for science: open data science for scienceAI for Science: An Emerging AgendaAdvanced Research Directions on AI for Science, Energy, and Security: Report on Summer 2022 WorkshopsMachine Collective Intelligence for Explainable Scientific DiscoveryExpert-Guided LLM Reasoning for Battery Discovery: From AI-Driven Hypothesis to Synthesis and CharacterizationThe emergence of large language models as tools in literature reviews: a large language model-assisted systematic reviewOpportunities in AI/ML for the Rubin LSST Dark Energy Science CollaborationMapping the Landscape of Artificial Intelligence Applications against COVID-19Artificial Intelligence for COVID-19: Rapid Review

即便在综述级综合内部,也有强度差别。例如 C3 说生成化学等已使若干化合物进入临床试验、生成模型可按目标性能设计材料,但没有在摘要中提供数量与验证标准;C4 的综述把 PINNs 称为变革性框架,却只做方法分类而非系统基准比较;C12 的疫情早期综述仅纳入 11 项研究并列出数据不足与验证不足。因此,“多篇综述报告同样方向”和“该方向已被实验充分验证”是两件不同的事。

证据类型:跨摘要综合 · 摘要证据:A survey of generative AI for <i>de novo</i> drug design: new frontiers in molecule and protein generationArtificial Intelligence for Drug Discovery: Are We There Yet?Artificial intelligence in drug discovery: A comprehensive review with a case study on hyperuricemia, gout arthritis, and hyperuricemic nephropathyArtificial Intelligence and Generative Models for Materials Discovery -- A ReviewPhysics-informed neural networks for PDE problems: a comprehensive reviewA review of some techniques for inclusion of domain-knowledge into deep neural networksMapping the Landscape of Artificial Intelligence Applications against COVID-19Artificial Intelligence for COVID-19: Rapid Review

当前摘要证据没有回答的问题包括:各综述中提及的模型在统一基准上的相对表现、生成出的候选物进入临床或实际合成的具体标准、ChatBattery 三种材料的长期稳定性与独立复现情况、以及 LLM 在数值型数据提取上的低准确性如何影响科学文献任务(C3、C9、C10)。这些问题需要回到全文核对方法、实验设置与基线选择,或由新的独立研究提供。

证据类型:有边界的编辑判断 · 摘要证据:A survey of generative AI for <i>de novo</i> drug design: new frontiers in molecule and protein generationArtificial Intelligence for Drug Discovery: Are We There Yet?Artificial intelligence in drug discovery: A comprehensive review with a case study on hyperuricemia, gout arthritis, and hyperuricemic nephropathyArtificial Intelligence and Generative Models for Materials Discovery -- A ReviewExpert-Guided LLM Reasoning for Battery Discovery: From AI-Driven Hypothesis to Synthesis and CharacterizationThe emergence of large language models as tools in literature reviews: a large language model-assisted systematic review

因此,最稳妥的读法是把本简报的所有判断都标为“摘要报告/跨摘要综合”,把单篇量化结果当作需要复现的假设。对读者而言,C1-C6 这类综述级认识可以用于快速建立认知地图;C8-C11 这类数字则应在引用前核对原始实验设计与全文结论,避免把单篇作者主张误当成领域共识。

证据类型:有边界的编辑判断 · 摘要证据:From AI for Science to Agentic Science: A Survey on Autonomous Scientific DiscoveryAutonomous ‘self-driving’ laboratories: a review of technology and policy implicationsAI for science in electrochemical energy storage: A multiscale systems perspective on transportation electrificationArtificial Intelligence and Generative Models for Materials Discovery -- A ReviewUnleashing the power of AI in science-key considerations for materials data preparationSciHorizon: Benchmarking AI-for-Science Readiness from Scientific Data to Large Language ModelsPhysics-informed neural networks for PDE problems: a comprehensive reviewA review of some techniques for inclusion of domain-knowledge into deep neural networksScientific Large Language Models: A Survey on Biological &amp; Chemical DomainsAccelerating AI for science: open data science for scienceAI for Science: An Emerging AgendaAdvanced Research Directions on AI for Science, Energy, and Security: Report on Summer 2022 WorkshopsMachine Collective Intelligence for Explainable Scientific DiscoveryExpert-Guided LLM Reasoning for Battery Discovery: From AI-Driven Hypothesis to Synthesis and CharacterizationThe emergence of large language models as tools in literature reviews: a large language model-assisted systematic reviewOpportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration

摘要来源