Periodic Labs 的 Liam Fedus 与 Ekin Dogus Cubuk 阐述合成超级智能
Synthesis Superintelligence: from Semiconductors to Superconductors — Periodic Labs’ Liam Fedus and Ekin Dogus Cubuk
Periodic Labs 的 Liam Fedus 与 Ekin Dogus Cubuk 在访谈中表示,合成超级智能要靠模型在真实物理实验中学习,而不能只增加互联网训练数据。
译文尚不完整,完整内容请切换到原文。
一年之后,它已被视为最杰出的 AI 科学家实验室之一,拥有令人眩晕的 人才密度,并在自主实验室建设方面取得了 惊人进展:
大多数人熟悉 Liam 和 Dogus 的常规资历,但在进一步了解 Periodic 时,我们发现了一条不可思议的“人才斜率”:后加入的每一位员工似乎都比前一位更出色:
从构建能对充满噪声的物理实验进行推理的 AI 系统,到打造每台仪器都能变得智能的实验室,Periodic Labs 押注 AI 的下一个前沿不会仅仅来自在更多互联网数据上训练,而会来自让模型与真实世界进行实验。在本期节目中,Periodic Labs 的 Liam Fedus 和 Ekin Dogus Cubuk 与 swyx 及 Brandon 一起,解释科学发现为何从根本上不同于数学和编程,以及打造能够真正发现新材料的 AI 科学家需要什么。
我们深入探讨 Periodic 对“合成超级智能”的愿景:以物理实验为基础的强化学习、AI 驱动的材料表征、模拟与密度泛函理论、高通量实验室,以及从做科学的全过程而非仅从已发表结果中学习的系统。Liam 和 Dogus 还解释了为何前沿模型仍需要实验、为何失败的实验可能是最有价值的训练数据之一、让每一件实验室设备拥有“140 IQ”意味着什么,以及自主实验如何能把数十年的科学试错压缩到数月。
我们讨论:
为何仅有智能还不够实现科学发现
当环境是物理世界时,强化学习会如何改变
为何科学需要在不确定性、噪声和信息缺失下进行推理
材料发现循环中的预测、合成与表征
为何物理学和材料科学仍远未“被解决”
“物质编译器”与 Periodic 的合成超级智能目标
相变、X 射线衍射与AI 驱动的材料表征
DFT、模拟,以及为何实验仍是最终的地面真值
室温超导体、新型磁体、电池,以及更高效的计算
为何量子计算未必能自动解决材料发现
让每一件实验室设备拥有“140 IQ”意味着什么
为何数据质量和阴性结果比单纯向科学投入更多算力更重要
在做科学的过程上训练模型,而非最终答案
为何即便未来的前沿模型仍需要进行物理实验
在AI、化学、物理和定制硬件领域扩展自主实验室
自动化实验如何能在发现新材料时大幅增加“运气表面积”
Liam Fedus
Ekin Dogus Cubuk
时间戳
00:00:00 引言
00:02:49 物理世界中的人工智能与强化学习
00:09:17 端到端材料发现闭环
00:13:28 为什么物理学尚未“被解决”
00:19:04 物质编译器与人工智能表征
00:30:22 DFT、模拟与实验真值
00:42:27 梦想材料、算力与量子计算
00:45:57 赋予每台实验室仪器“140 IQ”
00:50:02 实验室自动化
00:53:16 数据、模型与阴性结果
00:58:13 用科学过程训练人工智能
01:00:20 为什么前沿人工智能仍然需要实验
01:06:06 规模化自主实验室
01:10:45 组建团队并部署到产业
01:17:01 从人工智能副驾驶到科学成果
01:20:57 自动化搜寻超导体
文字稿
引言:智能必须与现实相遇
Swyx [00:00:00]: 好的。我们现在在 Periodic。今天我们请来了 Periodic 的 Liam 和 Doğuş。欢迎,也感谢邀请我们到你们这里来。
Liam Fedus [00:00:11]: 是啊,很高兴来到这里,也谢谢你们过来。
Swyx [00:00:13]: 我想先从你们网站上我很喜欢的一句话说起。它说:“智能是必要的,但并不充分。当想法被发现与现实一致时,新知识才会被创造出来。”这看起来如此直白,但为什么它并不显而易见?
Liam Fedus [00:00:29]: 我认为,这大致就是 Doğuş 和我在构建这一切时为 Periodic 想到的论题和前提,也就是你不能只靠思考找到解决方案。宇宙如此复杂,要想真正推动知识前沿并取得进展,你需要提出这些猜想,然后实际看看它们是否成立。无论把那本教科书或论文重读多少遍、怎样思考,或是把自己关进一个房间,都无法让你想通所有可能的实验结果,包括预期中的和意料之外的,你真正需要的是这个迭代过程。这就是我们把这些整合起来的核心理念,也是为什么从一开始我们就觉得:“我们需要拥有人工智能系统、物理世界的模拟,同时还要建起实体的高通量实验。”我们认为,一种不同类型的智能会从中涌现。
Swyx [00:01:21]: 我认为同样值得注意的是,像 OpenAI 和 Google 那样,你们在那些大型实验室里并没有那些资源。你们必须出来自己做,因为你们基本上是在边走边发明自己的打法。
打造面向人工智能科学的跨学科实验室
Ekin Doğuş Çubuk [00:01:34]: 是的,而且我觉得这样的团队从未存在过。我们试图把固体化学家、固体物理学家、实验学家、理论家、硬件工程师、LLM 专家和计算机科学家聚集到一起,而这些技术中有一些非常新近。高通量实验、机械臂大约只在过去三年左右才被尝试过。力场方面的专长,其中一些力场也非常新近。但没错,我们觉得拥有一个实验室非常重要。有这样的聚焦,并把所有这些人聚在一起共同工作,非常重要。我想,历史上最接近的例子是像贝尔实验室这样的地方,杰出的理论家、实验学家和化学家在那里共同工作,并取得了了不起的成就。所以我们正试图在这里做同样的事,是的。
物理世界强化学习与不确定性
Brandon [00:02:17]: 我想稍后再谈谈力场,但是
Swyx [00:02:19]: 是的
Brandon [00:02:19]: 在我们谈到那之前。对,它们是什么?但在我们谈到那之前,先说一下背景,我们的一些听众是 AI 工程师,另一些是科学家。对工程师而言:在像 OpenAI 这样的大型实验室工作,和你在这里所做的事情有什么区别?我想你已经暗示过这一点,但也许可以明确地说,你必须如何重新框定自己的思考过程?
Liam Fedus [00:02:49]: 嗯,我觉得一件有趣的事情是,现在我们的强化学习环境确实源自环境本身,源自我们的物理实验室。我们的数据来自我们的物理实验室,而这在某种意义上是我们的终极真相。仅仅针对一些论文或教科书中已知的答案做优化是不够的,因为我们正在超越那一步,我认为这是最大的差异之一。但与此同时,还存在比如说在不确定性下进行决策的问题。所以当你针对数学进行优化时,它具有很高的精度。你并不是真的在处理方差、不确定性或异常测量之类的东西,而这对我们的流程非常关键。例如,当我们在进行材料发现循环时,东西从炉子里出来时并没有被贴上标签,对吧?即便是标注过程也可能是随机且有噪声的。有时不同机器之间还会出现偏差。也许我们的遥测数据会不足。所以我们以为是在这个温度下运行的,但实际上温度偏离了某个差值。而一种能够接收这些噪声数据并像科学家那样做出智能决策的智能,则是一套不同的推理策略。所以我认为有大量共同点,比如标准的中期训练、强化学习、构建使用工具的智能体、方差缩减、确保基础设施在训练与推理之间没有不匹配。但我们必须超越这些,真正思考:在高度不确定的情况下,如何做准确的工作?如何真正极其高效地利用有限的一组数据?所以我们正在大幅扩大规模,但尽管如此,这与数字环境仍然非常不同——在数字环境中你可以任意增加更多环境或更多 rollout。我们没有同样的能力。所以样本效率是另一个关键点。
为什么实验是有噪声且不完整的
Brandon [00:04:48]: 这对我来说似乎有点抽象,而且这个主题在科学播客上已经出现了一段时间,但我认为,当你谈到实验不确定性时,如果能解释一下你在实验中具体做的某个过程会是什么样子,以及在进入人工智能这一面之前,人类会如何处理这个问题,可能会有帮助。你有没有一个具体例子,说明你制造的某种材料,以及在这个过程中你会遇到什么样的测量不确定性?
Ekin Doğuş Çubuk [00:05:17]: 我认为这有很多维度。其中之一是,我认为每一项科学实验都必须进行某种降维,而这与编程或数学非常不同。所以,当你做数学或编程时,所有上下文都可以提供给人类或 LLM。那是一堆公理,以及人们从中推导出的一些推论。那是一堆函数和 API 调用。因此,你推理所需的一切都摆在你面前,然后你只需要非常聪明地把它弄明白。在物理学中,如你所知,我们一开始面对的原子数量比我们能在计算机上存储的还要多。所以显然,我们必须从代表一个系统的原始维度数和比特数,缩减到我们能放进计算机的比特数。我想,这就是热力学最初的前提。他们最初对蒸汽机如何工作感到非常困惑,但后来发现,你可以找到五六个热力学变量,就能相当好地解释正在发生的事情,这令人难以置信。这就是物理学之美。
Brandon [00:06:13]: 你有一屋子的原子,数量是 10²³ 或 10²⁷
Ekin Doğuş Çubuk [00:06:19]: 正是如此
Brandon [00:06:19]: 或者这个房间里大概这么多原子,但你只需要温度、压强和一些简单变量,这就告诉你一切。不,它告诉你的主要是你所需要的
Ekin Doğuş Çubuk [00:06:28]: 很多
Brandon [00:06:29]: 在大多数情况下需要知道的东西,对吧?
Ekin Doğuş Çubuk [00:06:30]: 是的。
Brandon [00:06:30]: 那么假设你制造一种材料,你可以做的事情比如查看 X 射线结构,而这东西是一种信息损失很大的投影,对吧?它实际上并不是一个结构,或者说它并不能告诉你真实结构。它告诉你的是其中一些主成分。那么,人类会如何利用这些简单的主成分,然后对它们进行推理?接着你又如何把这一点扩展到人工智能?
Ekin Doğuş Çubuk [00:06:58]: 所以在理论物理这一边,对吧,我们已经认定能量、能量的平均值以及能量的涨落非常重要。因此很多热力学都是由此推导出来的。然后我们意识到反应势垒也非常重要。动力学非常重要。所以一个人可以尝试通过提出一些降维描述符,来对这个极其复杂的系统进行推理。人们也非常喜欢原子尺度的图景。你会去想局域邻域长什么样,它们是来自八面体还是四面体,而这一点尤其让化学家受益。好,这是理论这一边。而在实验这一边,你会尽量收集尽可能多的数据。所以你会有一本实验记录本,把你看到的一切都写下来,等等,但当然你也明白,你会漏掉很多东西。就像 Liam 提到的,会出现各种各样的问题。例如,你的炉子会随时间退化,因为在使用炉子的过程中,你烘烤的一些东西会蒸发并覆盖加热元件。所以你每用一次炉子,它就会变得越来越差。另一个问题是,你有一台炉子,但温度并不是完全均匀的,所以你把样品放在炉子的哪个位置会影响结果。我们现在还没有这个问题,但也许有一天会有。光学仪器受振动影响很大。所以每当有人走动。我记得我读博士的时候,有一个实验室遇到了很大的问题,因为有时候他们的结果和其他时候差别很大,他们最终追查到,在夜里某个时间楼上有人会走动,那种振动影响了激光装置。所以是的,科学非常难,但这也正是它特别的地方,对吧?Liam 和我一直在谈的另一件事是,目前很多改进都集中在数学、编程和理论计算机科学上,因为这对 LLM 来说更容易。但在现实生活中,大多数需要智能的事情实际上更像科学。有大量的不确定性,有大量的噪声,有大量缺失的上下文,但你必须成为那个有智能的存在,并弄清楚下一步该做什么。
Liam Fedus [00:08:53]: 我认为在此基础上还有一个非常有意思的点:做数学、做理论计算机科学、做这类事情的最优推理策略,未必是做科学的最优推理策略。
Ekin Doğuş Çubuk [00:09:04]: 是的。
材料发现循环:合成与表征
Brandon [00:09:05]: 当你的输出是一个有噪声的晶体结构,或者若干个——我不知道——集体变量之类的东西时,你如何定义一个 RL 奖励函数?比如。
Liam Fedus [00:09:17]: 也许,我来讲讲这个循环的一部分,并稍微深入一下。所以,作为材料发现循环的一部分,我们首先得弄清楚要制造什么。也就是原子组合在一起的稳定性,以及我们是否预期它具有我们想要的性质。所以我们不只是想要一种新颖的原子构型。我们希望它们组合在一起时,具有我们感兴趣的性质。接下来是如何合成它?你究竟如何把这个东西做出来?这也并非易事。所以即便你知道某种东西能够结合在一起,实际要弄清用什么样的工艺条件把它实现出来,也非常不容易。但再说回来这件事。一旦你完成了这两步并做出了某种东西,它并不会自带标签,所以你实际上必须对它进行表征,弄清楚你做出了什么。因此在这种情况下,与其把强化学习环境设想成我们只是启动一个实验、等上几天,然后尝试更新一下你是否找到了室温超导体,那是完全不可行的。
Brandon [00:10:15]: 是的。我正想问这个。是啊,你不能让 GPU 闲置一周之类的。
Liam Fedus [00:10:19]: 这。是的,这完全不可行
Brandon [00:10:20]: 是的
Liam Fedus [00:10:20]: 因为你没有足够的智能体来充分降低方差。rollout 时间是在不同仪器之间流转。你在等待。实验本身就存在一些物理限制,涉及实际合成或炉时。所以就是太慢、噪声太大。因此,我们思考这个 AI 项目的方式,基本上就是围绕我们已产生的这组数据来构建智能体。所以就表征而言,你要寻找的是对实际存在的物相之识别给予奖励的强化学习环境。给定原始实验数据,你能否有效地拟合它?我们的做法是向材料发射 X 射线。X 射线的频率或波长大致可测,或与原子间距相当,因此你会得到漂亮的衍射图案,它们某种程度上给出晶体结构的指纹。而要做到这一点,识别实际存在的是什么非常不容易。所以我们早期的一些 AI 工作,基本上就是在打造用来做这件事的系统。这也让我们能推进得快得多,因为如果你在进行这么多实验,你很快就会在理解产出了什么的能力上遇到瓶颈。但话说回来,现在这个强化学习环境相当直接,也就是你要对所存在物相的识别给予奖励。你要惩罚虚假物相,或没有任何真正化学合理性的东西。所以这有点像一个简短的例子,你可以针对不同环节这样做。而当你把这些拼接起来,那就是端到端的发现循环。
在新颖实验数据上训练,而非记忆答案
Brandon [00:12:02]: 好的。你
Liam Fedus [00:12:02]: 也许另一个
Brandon [00:12:03]: 是的。哦,抱歉
Liam Fedus [00:12:03]: 另一部分是,随着我们积累起一大篮子实验数据,你可以把它想成另一件事:基本上就是给世界的状态打上时间戳。所以你可以说:“在这个日期,这就是我们截至那时的实验证据。”然后你可以构建强化学习环境,在其中说:“好,给定这批实验证据,以及给定这些选择,科学家做出的下一个选择是什么,或者那次实验的结果是什么?”而且。这也非常有意思,因为它让我们能够做一些在外部会更难开展的项目类型。因为如果预训练模型已经记住了其中一些数据,然后你试图在那上面做强化学习,如果它已经知道答案,而你构建的强化学习任务又依赖于那个答案,它就可以假装在工作。所以,然后,因为这就像是,哦,好吧,它已经知道答案了,所以它实际上不必做困难的物理推理。它不必真正去做计算模拟。然后它得到了正确答案,接着你强化那条策略,并上调那些推理策略的权重。那些推理策略不会泛化到新的系统,所以对我们没有用。但随着我们积累起一篮子一篮子的实验数据并把事物拼接起来,让整个过程都保有谱系,这就赋予我们构建在其他地方根本不可行的新型环境的能力。
为什么已知的物理并不意味着完美的模拟
Swyx [00:13:27]: 我想问我的问题
Brandon [00:13:28]: 说真的
Swyx [00:13:28]: 但我担心我把话题带偏了。我还是直接问吧,因为我,再说一次,我来自非科学家背景,但算是工程这边的。我熟悉机器学习这一侧,但不熟悉、不熟悉物理这一侧。好,这个问题有几个版本,但我认为基本问题是:我们不是已经把大部分物理定律搞清楚了吗?那为什么我们还没有完美的模拟器?
Brandon [00:13:50]: 哦,这是个很好的问题。
Swyx [00:13:52]: 就是,这非常基础,就是。然后他就说,“哈,真可爱。”但就是
Brandon [00:13:56]: 不,这是个很好的问题
Swyx [00:13:57]: 老兄,我有这么多物理书。你什么意思,你什么意思我们还没做完?
Ekin Doğuş Çubuk [00:14:03]: 是的,所以这里有几件事在发生,对吧?所以
Swyx [00:14:07]: 是的
Ekin Doğuş Çubuk [00:14:07]: 我认为首先,我们绝对还没有把任何物理定律做完。即便是
Swyx [00:14:12]: 在量子尺度上,好吧,但是。也许还有非常大的尺度,但在人类尺度上,在材料尺度上,我们,我们。还剩下什么?
Ekin Doğuş Çubuk [00:14:22]: 我们还没有。所以这是个、这是个非常好的问题。而我们为什么还没做完,这一点非常有意思,对吧?所以有一个非常、常见的、人们常讲的故事。当 Dirac 在弄清楚量子力学时,大概是1920年代末,他写了一本关于量子力学的教科书,并且他大致把它表述为一切都完成了。然后有一句人们喜欢引用的话。我其实不确定它有多准确,但据说 Dirac 说:“剩下的就是化学。”其想法是他能够求解氢原子
Swyx [00:14:50]: 把所有东西都组合起来就行。
Ekin Doğuş Çubuk [00:14:51]: 是的。他也许能解出一条1D氢原子链,但目前还解不了,比如说,氮与氧的相互作用,不过没关系,那只是化学而已。
Swyx [00:15:00]: 是的。
Brandon [00:15:00]: 物理学家真的很喜欢单原子体系。真的是非常简单的系统。
Ekin Doğuş Çubuk [00:15:05]: 简单系统。
Brandon [00:15:05]: 而那就是你所能求解的全部了。
Ekin Doğuş Çubuk [00:15:07]: 球形奶牛。
Brandon [00:15:07]: 是啊,球形奶牛。
Ekin Doğuş Çubuk [00:15:08]: 但事实证明,我认为我们在过去一百年里学到的是,首先,那并不是真的。它并不是仅仅处理氢原子那种工作的简单延伸。而且也不仅仅是。其余的也不仅仅是化学。那里实际上有很多物理。我想我们仍未弄清楚的事情之一是高温超导。但还有很多其他的
Swyx [00:15:25]: 那对你来说,高温是175、200?
Ekin Doğuş Çubuk [00:15:28]: 哦,当人们说高温超导时,他们指的是非常规超导。所以存在常规超导,它主要只是由电子-声子耦合驱动的。在那种情况下,你多少可以看到同位素效应。如果你取同一个体系,只是让其中一种元素的质量不同,你就能看到超导温度以你所预期的速率下降。所以那种是常规的。但随后像铜氧化物这样温度更高的超导体,比如那些高于77 kelvin的,比如93 kelvin的,结果并不遵循那套物理。但我们不知道它们遵循什么物理。它们只是令人难以置信的超导体。但这还不止于此。所以那绝对是这些我们尚未知道其理论的非常热门的课题之一。但还有这么多其他我们不理解的事情。例如,即便是简单量子力学体系中的强关联,我们目前还无法模拟。密度泛函理论是一个非常强大的工具。它在某些事情上准确得令人难以置信,但只要。一旦出现某种强电子关联,它实际上就无法捕捉其中一些效应。有太多东西需要弄清楚。这就是为什么我真的很期待把AI应用到这个领域,因为在理论上、计算上和实验上,都有太多东西需要弄清楚。而我们在这里做出的每一项改进都应当改善人类生活,因为,我们越能理解材料、固体物理,就越能为它们制造出更好的器件。
Brandon [00:16:41]: 有一句Phil Anderson的名言,叫做“多即不同”,其基本意思是,你也许理解了某事物的所有基本定律,但当你把许多东西加在一起时,它们的行为在性质上截然不同,不同于简单物理所应告诉你的那样。所以理解大量事物如何共同表现,看起来似乎应该很简单。起始的定律是简单的,但集体行为如此复杂,以至于真的很难建模。
Ekin Doğuş Çubuk [00:17:07]: 它是涌现的、普适的,这太不可思议了。我们在深度学习模型中也看到了这一点,对吧?我们在深度学习中到处都能看到幂律。我们并不理解,但它非常让人联想到物理学中存在的涌现与普适性。
Brandon [00:17:20]: 物理学界已经有许多关于神经网络中涌现现象的理论论文。
Ekin Doğuş Çubuk [00:17:24]: 是的,完全正确,没错。
Brandon [00:17:25]: 是的。
合成预测与物质编译器
Swyx [00:17:26]: 然后我的另一个追问也只是关于这个过程,我们就称之为,对,过程和最终结果。这么说吧,你希望材料具备这些你所瞄准的性质,然后你必须弄清楚到达那里的工艺。把这些事情分开是否值得,这样你就可以。你只需训练一个模型,给定你想要的任何理论上的最终状态,它都能完美地逆向工程出任何过程?这,有意义吗?
Ekin Doğuş Çubuk [00:17:54]: 也许回到我们之前的谈话:既然原子数量超过了阿伏伽德罗数,我们永远无法知道所发生事情的确切背景。我认为这可能就是我们永远无法完美逆向工程一切的原因。但我们只是试图进行足够程度的逆向工程,以便能够提升材料的性能,基本上就是这样。
Swyx [00:18:13]: 是的。
Brandon [00:18:13]: 你问的是,是否存在一个正向模型,你放入一个输入就能预测输出,而你想要 under
Swyx [00:18:19]: 是的。
Brandon [00:18:20]: 你想要创建一种类似代理模型的东西,它能够理解
Liam Fedus [00:18:22]: 输入是那个的最终状态,这个原子构型,然后预测是
Swyx [00:18:27]: 因为这样你就可以,你几乎可以把工作拆开。你可以让一个团队做过程这一侧,然后另一个团队做性质这一侧,然后
Liam Fedus [00:18:33]: 对
Swyx [00:18:33]: 你知道,就让他们比拼。
Liam Fedus [00:18:36]: 是的,我认为在这些方面我们可以取得很多进展。把这些智能体工作流拆分到这些不同领域,我认为在合成预测方面我们也可以做很多实证工作。
Swyx [00:18:49]: 是的。
Liam Fedus [00:18:49]: 而给定那个原子构型,你能使用什么工具?如何,鉴于这些实验证据,鉴于先前的文献、先前的论文,你实际上如何得到这些东西?而你最终是在检验,你是否真的能够复现它?
Swyx [00:19:04]: 我这么说的原因也许也只是从一家公司的角度来思考。你知道,就像台积电某种程度上是半导体的晶圆厂,但他们并不设计半导体,你可以成为材料领域的台积电,人们来找你说,“嗯,这就是我想要的”,然后你就能弄清楚工艺,并且
Liam Fedus [00:19:21]: 是的,就像一个物质编译器。
Swyx [00:19:23]: 是的。
Liam Fedus [00:19:24]: 所以我想,给定这些要求
Swyx [00:19:26]: 非常《星际迷航》。
Liam Fedus [00:19:27]: 是的。但就是说,给定这些要求,这真的是一组有效的配置吗?它能、它能存在吗?
Swyx [00:19:33]: 这是一个好的目标函数吗?我不知道。它太大了。
Liam Fedus [00:19:37]: 是的。我们有点把我们的使命之一称为合成超级智能,所以我认为这与此相符。
相变、X 射线衍射与 RL 奖励
Swyx [00:19:44]: 是的。
Brandon [00:19:44]: 我想深入展开一下你刚才提到的一点,也就是你谈到的相。我认为这会回到 Doğuş 关于什么是涌现的评论。相变这个概念对很多听众来说可能比较陌生。所以你能解释一下什么是相变,以及为什么这对 RL 环境这类东西来说是一个有用的信号吗?
Ekin Doğuş Çubuk [00:20:08]: 相变非常迷人。我想人们在生活中最容易产生共鸣的也许是冰融化,或者水沸腾。你提高冰的温度,它看起来仍然像冰,行为也像冰,但到了某个温度,它就不再升温,然后变成液体。于是它突然从这种固相变成液相。另一个当然与我们非常相关的是 Ising 模型。我认为计算机科学和数学家也许也在用不同的名称研究它,但基本上,你可以有向上和向下的自旋,如果两者都向上,与一个向上、一个向下相比,它们会有不同的相互作用项。在足够高的温度下,它们通常只是随机地向上或向下。熵占了上风。你降低温度,看起来还是一样。但到了某个时刻,它们突然开始完美地对齐。是的,相变。相变非常好,因为它让物理学家在研究某些事物时更轻松。它们往往还具有某些空间相关性,这对我们也很有帮助。好,所以在我们的实验室里,相变通常是因为我们把不同的前驱体混合在一起,基本上是晶体,然后我们提高温度或做些别的事情来促使它们反应,接着原子开始反应并形成这种新晶体。而这会体现在 XRD 图谱中,因为通常如果原子的几何结构发生剧烈变化,XRD 图谱就会变化很大。但在一次典型的实际材料发现活动中,真正发生的情况是,当你第一次尝试某件事时,它并不成功,结果也不是简单的是或否,而通常是一个非常混合的相。它通常会含有一些前驱体。它可能含有一些非晶相,然后还会有一堆你也许并没打算制备的相。这就成了一个挑战,这也是为什么我们发现很难把模拟和 AI 结合起来做表征。所以模拟总是可以说:“哦,这个相看起来不像我们预测的那样,但这是它的一个轻微变体。让我现在对它做力场计算,看看它是否仍然稳定。”或者 AI 可以说:“根据之前的实验和这次实验,这大概不是我们想要制备的相,所以让我们改变条件。”
Liam Fedus [00:22:18]: 是的,而且在某些情况下
Swyx [00:22:19]: 听众朋友们,XRD 是 X 射线衍射,这一点你已经介绍过了。
Liam Fedus [00:22:22]: 是的。在某些情况下,有一种物相此前从未被记录过。它不存在于任何论文或数据库中。因此 AI 实际上必须去使用这些工具来判断:“那么,什么样的原子构型才能真正解释这些物相?”但这对于指导科研过程变得极其重要,因为,比方说我们心中有一个目标物相。我们确实需要逐步攀升,以便对该物相获得更好的测量。所以如果它是百分之一,而我们想,实际上,我们想提高这种物相纯度,我们就希望对这份实验数据有非常好的测量,从而判断什么存在、什么不存在。
Brandon [00:22:57]: 所以其中一个优势是,你有一个可以直接观测的明确变量。回到你如何处理不确定性的说法,你的回答中似乎有一部分是,你让可观测特征在某种程度上没有歧义。那里不存在解释空间。对吗?这就是背后的逻辑吗?还是说,这实际上只是你需要找到一种新物相,而这就是你关心的事情?
Liam Fedus [00:23:20]: 我认为在很多这类情况下仍然存在歧义。一个很笨的办法是重复实验。
Brandon [00:23:25]: 是的。
Liam Fedus [00:23:26]: 我们当然会做重复实验。
Brandon [00:23:27]: 是的。是的。
Liam Fedus [00:23:28]: 但不是,我认为——你知道,这是一项具有挑战性的任务,因为它并非完全确定性的,而且确实存在歧义以及其他问题。两种不同的物相实际上可能与同一种图谱相一致。
Brandon [00:23:42]: 是的。
Liam Fedus [00:23:42]: 所以我认为,这就是为什么系统运用一些化学直觉来判断“好吧,鉴于合成条件,鉴于这一点以及先验知识,这种情况极不可能”非常重要。而这可以用来帮助消歧。
化学先验、多模态测量与歧义
Swyx [00:23:58]: 所以你是在那里根据你的预期注入一些先验。
Ekin Doğuş Çubuk [00:24:03]: 是的。我认为热力学是最大的先验,对吧?
Swyx [00:24:06]: 确实如此。
Ekin Doğuş Çubuk [00:24:07]: 然后物理学也是一个很大的先验。我们在这里真正受益的另一点是多模态或材料表征。所以我们可以做 XRD,是的,有些物相在 XRD 图谱上看起来可能相似,但随后我们还可以测量它们的电学性质、磁学性质。我们可以测量它们的形貌,比如某种电子显微镜。然后,一些在 XRD 上看起来相似的物相,在这些维度中的某些维度上会看起来不同。这种多模态确实很有帮助。而且,AI 在这里也非常有用,因为人类的上下文和计算能力同样有限。所以如果你同时给人类 10 种不同的模态,并说“把这些全部一致地分析出来”,这有点困难。但对 AI 来说,做到这一点其实并不需要超级智能。这很棒,是的。
Liam Fedus [00:24:46]: 是的,超级智能就像是,它能做那么多计算。它能审视一切。而这,某种程度上就落到了那句关于在不确定性下做出更好决策的说法上,现在它正在把信号拼接起来,纵向跨越许多不同仪器、跨越不同实验、跨越重复实验,然后逐渐得到这个更底层的更好模型
Swyx [00:25:08]: 是的
Liam Fedus [00:25:08]: 而不是仅仅过度依赖来自一台仪器的一次测量。
信息极限、遥测与麦克斯韦妖
Swyx [00:25:12]: 然后让我,就是。所以你知道,我想你不久前也说过,数据多到任何合理的计算机、流程或任何东西都装不下。所以到某个时候,你不得不丢掉数据,即便你有所有这些重复实验,即便你大概有十种不同方法来测量同一件事。就是。因为万一你的某个测温的东西出了问题呢?那你什么时候把它丢掉?你怎么决定该怎么做?作为一个渴求数据的机器学习人,我就是想要一切,对吧?然后我就,你知道
Ekin Doğuş Çubuk [00:25:37]: 所以就是,是的。
Swyx [00:25:37]: 蹒跚着,学会一切。我不知道我不知道什么,就是扔、扔,把机器扔上去。
Ekin Doğuş Çubuk [00:25:42]: 是的。所以我们不丢弃任何数据。我想我也许想说的是,有可能。如果上帝在观察这个实验
Swyx [00:25:50]: 是的。
Ekin Doğuş Çubuk [00:25:50]: 实验中的数据比人类能装进计算机的还要多。不幸的是,我们这些凡人无法观测所有原子。所以即便数据多到我们存不下,我们实际上也无法获取那些数据。例如,我无法追踪所有原子的位置。所以不幸的是,我们仍处于这样一个阶段:我们收集到的任何数据都非常珍贵,我们不会删除其中任何一份。是的。
Brandon [00:26:11]: 也许一种把它和 AI 这边联系起来的思考方式,是想想探测,比如对一个基础模型做线性探测,对吧?这个模型拥有所有这些数据,所以上帝拥有材料正在做什么的完整图景,然后实验就是这个小小的线性探针,输出大概只有八比特之类的。所以你是在把这个庞然大物粗粒化成一点点信息,然后现在,要么传统上是人的工作,要么现在是你的智能体的工作,要从那八比特信息中重建
Liam Fedus [00:26:43]: 这是一个类比,是的。
Brandon [00:26:44]: 里面实际在发生什么。
Liam Fedus [00:26:45]: 这是一个类比:你无法对每一个参数权重、对该输入的每一次激活都有完全的可见性。是的,你得到的是它的某种子集,或者是贯穿整体的一些粗粒化特征。
Ekin Doğuş Çubuk [00:26:58]: 是的。但所以你大概会发现。你大概,确实会觉得,麦克斯韦妖的论证非常有趣,因为那种
Swyx [00:27:05]: 你能、你能再、复述一下吗?因为我不。他之前提过。我老是忘记。
Ekin Doğuş Çubuk [00:27:10]: 所以,在1800年代,人们意识到一个系统的熵必须增加,或者对于一个封闭系统
Swyx [00:27:18]: 是的。
Ekin Doğuş Çubuk [00:27:18]: 或者保持不变,但它不能下降。
Swyx [00:27:20]: 是的。
Ekin Doğuş Çubuk [00:27:21]: 这是热力学第二定律。曾经有一个思想实验在问:如果有一个非常微小、无所不知的恶魔,能够观察原子,然后每当某个原子能量较低时,就打开一扇门让它进入另一个房间,并不断这样做以降低系统的熵,从而打破热力学第二定律,那会发生什么?很长一段时间里,我认为没有一个非常令人满意的答案来解释为什么会这样。但后来在1900年代有一篇对我们来说更近的论文,Landauer提出了这样一个论点:删除信息必须消耗能量,而麦克斯韦妖通过做他所做的这件事,基本上会接触到如此多的信息,以至于他必须删除一些。而在。当他删除信息时,他必须消耗能量,这随后会确保熵增加。所以即使我们还没有达到能成为麦克斯韦妖的水平,我认为你的问题非常有远见,指向一个我们存储了如此多数据、不得不删除一些的未来。
Swyx [00:28:17]: 是的,不,我并没有,我显然没有想到那个、那个程度。这确实让我想起空调。空调差不多就是这样运作的。但同时,我认为这只是一个问题:为什么你不干脆买下世界上每一个传感器,然后过度。荒谬地过度配备仪器、过度复制一切?
Liam Fedus [00:28:33]: 哦,我们就是这么做的。
Swyx [00:28:34]: 对吧?是的,好,是的,好。就是这样。
Liam Fedus [00:28:35]: 是的,没错。
Swyx [00:28:35]: 好的,这样就说得通了。对。好吧。
Liam Fedus [00:28:37]: 是的,我认为,我认为这是一个关于
Swyx [00:28:39]: 已确认。
Liam Fedus [00:28:39]: 是的,正是。你能观测到单个原子吗,那是不可行的?但不,我们添加了大量遥测,这基本上减少了隐变量的数量,如果你愿意这么说的话。
Swyx [00:28:49]: 是的。是的。只是出于好奇,你们会担心自己只是在加利福尼亚吗?你们不想在尼泊尔和澳大利亚做这件事吗?
Ekin Doğuş Çubuk [00:28:59]: 是的,我们可以。我们有开设更多实验室的计划。
Swyx [00:29:02]: 因为显然重力和你面向哪里以及你。你在哪里很重要。
Ekin Doğuş Çubuk [00:29:07]: 专业能力也会有所不同,对吧?我们这里有某些类型的专业能力。例如,Stanford物理系、工程系拥有某些真正帮助我们的专业能力,因为我们可以从那里招聘。我们的研究人员可以合作。正如你所说,地理位置可能很重要,但不同城市在物理、化学方面的专业能力也会有一些差异。
Swyx [00:29:27]: 是的。哦,我们有Zoom
Ekin Doğuş Çubuk [00:29:28]: 是的。
Swyx [00:29:28]: 用来做那个。
Ekin Doğuş Çubuk [00:29:29]: 是的。但仍然
Swyx [00:29:31]: 并不是说你们的实验会受到你们所在位置的影响。
Liam Fedus [00:29:34]: 是的。是的,Periodic的未来不是单一实验室。
Swyx [00:29:38]: 是啊。你得登上私人飞机飞向太空,你懂的。
Liam Fedus [00:29:42]: 对。是的,我认为有必要建立这种实验室网络。而且,是的,我觉得正如 Doğuş 刚才指出的那样,基于专业能力或材料约束,这些各自都有一个最优地点。还有许多不同的约束,比如许可审批。
将模拟与物理实验相连接
Brandon [00:29:56]: 好的,所以其中之一是。所以有一件我一直很好奇的事是,你们如何处理计算工具?或者说,你们如何把计算工具整合进决策工作流,尤其是当计算和实验并不一定对得上的时候?从高层次理解来看,这也涉及理论。那么,这些是如何协同运作的?
Ekin Doğuş Çubuk [00:30:22]: 是的。所以,这里真正有帮助的一点是,有些事情在模拟中更容易做,而且在模拟中更准确。有些事情在实验中更容易做,而且在实验中更准确。一般来说,实验当然更准确,但某些测量用实验很难完成。于是你只能做出一个较差的版本,从而得到不准确的数据。我举几个例子。其中一个是,我们使用密度泛函理论来估算材料的形成焓,以及
Brandon [00:30:51]: 抱歉,稍停一下。你能解释一下 densital 密度泛函理论和形成熵吗?
密度泛函理论与形成焓
Ekin Doğuş Çubuk [00:30:55]: 是的,当然。
Brandon [00:30:56]: 抱歉,是焓,对。
Ekin Doğuş Çubuk [00:30:57]: 是的。所以,密度泛函理论大概是材料最常用的模拟方法。它大致来自 Hohenberg–Kohn 定理,Kohn 也因此获得了诺贝尔奖。他。当他意识到,没错,量子力学非常昂贵,因为你基本上对每个波函数都有这个指数级的希尔伯特空间,所以你必须求解这个非常困难的问题,尤其是随着体系尺寸增大。Kohn 和 Hohenberg 意识到,原来你并不需要在指数空间里去做。你所要做的只是考虑电荷密度,至少对于量子力学体系的基态是这样。所以这个,实际上这个定理相当简单。所以我觉得即使是物理学本科生也能理解。但核心思想是,量子力学体系基态的所有性质都是电荷密度的泛函,这很惊人,因为电荷密度只是这个三维对象,而希尔伯特空间和波函数则是指数级的。所以那是一个非常有趣的观察。然后 Kohn 又发表了另一篇论文,我想是和他的博士后 Sham 一起,Kohn–Sham 波函数,算是有了一种求解材料量子力学性质某种近似的实用方法。而今天,我们大量使用它,无论是在我们的实验室还是在外部。然后形成焓是我们可以做的测量之一,用来理解材料的能量。基本上,我们想问的是,这种材料具有的能量是多少?因为如果这个能量高于这些原子所能进入的其他材料,那么这种材料就不太可能被制备出来。所以稳定就意味着它位于其他材料及其能量的凸包上。这样说够清楚吗,还是我们应该。
Brandon [00:32:42]: 也许一个简单的高层次要点。所以密度泛函理论是一种方法,把某种在电子数量上呈指数级困难的东西
Ekin Doğuş Çubuk [00:32:51]: 对
Brandon [00:32:52]: 或者说在数量上,就你的体系尺寸而言,然后通过引入一些误差,近似地把它降到也许通常是一个 n 立方近似。所以现在你可以计算那些量,你知道的。你现在可以计算那些本来根本无法计算的东西,但要付出一定代价,对吧?如果你能完美地做到,你就会——我们就不需要实验室了。但我们做不到。这也许可以很好地呼应我们的,实际上是我们有史以来的第二集。Heather Kulik 曾说过,她有一句名言:“材料领域没有 AlphaFold,不只是在计算上,而是真值根本不存在。”比如,我们并不真正知道材料晶体结构长什么样,而 DFT 通常是我们最好的途径。
Ekin Doğuş Çubuk [00:33:34]: 是的,我也这么认为。也许我想补充两点是,DFT 按定义并不一定是一种近似。Hohenberg–Kohn 定理表明了这一点
Brandon [00:33:42]: 它不必
Ekin Doğuş Çubuk [00:33:42]: 它可以是精确的。但我认为,正如你所指出的,交换关联泛函我们尚未能够获得。然后对于基于电荷密度的模型,我们不知道动能泛函是什么。然后即便我们拥有完美的 DFT,我仍然认为我们需要实验室,因为我们。即便我们拥有完美的 DFT,我们也无法把十的二十三次方个原子放进一台计算机里。
Brandon [00:34:03]: 是的。
Swyx [00:34:04]: 你对那个数字非常执着。
Brandon [00:34:06]: 但是。不。嗯,好吧。是的。但随着算力呈指数级增长,你可以认为,在 n 的三次方缩放规模下,原则上你最终是可以算出结果的。你只要把停顿等得够久就行。
Ekin Doğuş Çubuk [00:34:17]: 哦,是的。我们需要巨型计算机,是的。
Brandon [00:34:19]: 但没错。但只要
Ekin Doğuş Çubuk [00:34:20]: 这得花上一段时间。
Brandon [00:34:20]: 是的。但只要 n 的三次方是一种不错的缩放方式,你就可以,你就可以。
Swyx [00:34:24]: 作为非科学家,我想做一个观察,不知道你们是否愿意把它到处抛出来之类的。我出身金融。我们有高斯 copula,那是一种给信用违约互换定价的方法,做法是计算相关性,并把一切归结为一个单一核。而且它非常。它在精神上也感觉和 VAE 很像,就是把所有这些东西浓缩成某种单一参数。我在想,是不是到处都是同一个技巧。你说的听起来比我的维度更高一些,VAE 就只是一个 sigma 之类的东西。但听起来是一回事?
Liam Fedus [00:35:01]: 是的,电荷密度文件仍然是一个相当大的文件。
Swyx [00:35:04]: 是的。对。你稍微挥挥手带过,但已经够好了。这是,这是最高有效位。是的。
Ekin Doğuş Çubuk [00:35:10]: 我个人不知道有什么更好的方法来预测一种新材料的稳定性。它绝对不完美,但绝对比我能想到的其他方法更好。是的,所以回到你的问题,我们使用它们的原因是,对某些事情来说,模拟比实验更容易,而且它们彼此互补得很好。然后我们可以把这个过程做成一个循环。我们觉得模拟本身永远不够,但在模拟、AI 和实验的循环中,我认为我们可以比以前快得多地取得进展。
当模拟失效时:微观结构与校准
Brandon [00:35:41]: 模拟让你能做的一件事,就是快得多地扩大规模。因为现在算力要上线容易得多
Ekin Doğuş Çubuk [00:35:48]: 是的
Brandon [00:35:48]: 比引进、新建新实验室更容易。你如何避免。或者说,假设你把大量 DFT 扩大规模,你如何避免让模型过度偏向于此,同时仍然立足于真实世界才是你最终关心的真相?如果你。你知道,你要怎样平衡这一点?
Ekin Doğuş Çubuk [00:36:05]: 嗯,我们也会扩大实验室的规模。
Brandon [00:36:06]: 是的,你们会扩大实验室规模。
Ekin Doğuş Çubuk [00:36:07]: 对,是的。
Brandon [00:36:07]: 但你可以。以现代算力,我可以想象把 DFT 扩大到数百万、数亿次,而你们的实验室,我猜,每天几百、几千次?我不知道。仍然相当。嗯,即便在真正的高通量情况下,相比之下你们仍然相当受限,对吧?
Ekin Doğuş Çubuk [00:36:27]: 是的。有一点是,人类,还有 LLM,都相当清楚 DFT 的局限性。例如,即便我们能像你说的那样,用 DFT 做一亿次试验
Brandon [00:36:39]: 是的
Ekin Doğuş Çubuk [00:36:39]: 我们实际上永远无法尝试它们的微观结构。微观结构指的是这样一种想法:DFT 通常模拟的是完美晶体,但晶体实际上通常并不完美,而它们的微观结构——也就是中等尺度上的结构——会真正影响性质。当然还有其他问题,比如超导温度这类某些性质无法轻易用 DFT 模拟。再说,即便你用 DFT 做一亿次试验,也仍然不够。所以我们必须依赖启发式方法、实验等等。
Liam Fedus [00:37:08]: 因此,从所有这些计算到实验室里真正执行的内容之间,存在一个巨大的筛选环节。
Swyx [00:37:13]: 那是由人介导的,还是 LLM,还是两者结合?
Ekin Doğuş Çubuk [00:37:17]: 可以是两者结合。
Swyx [00:37:17]: 是的。
Ekin Doğuş Çubuk [00:37:18]: 例如,在很长一段时间里,Materials Project——它曾是最好的开源 DFT 数据库——使用实验校准来校正 DFT。所以他们实际上会用实验数据来校准 DFT。我们在这里内部也会这样做。每当我们对某个化学体系感兴趣时,你会得到实验,也会得到模拟,然后就可以对它们进行校准。
无机材料为何与生物学不同
Brandon [00:37:37]: 这是另一件事。我和很多生物学家朋友聊过,当你说材料时,你不能只做一个 XRD 结构就真正知道结构是什么,他们总是非常困惑。因为如果你在生物学领域,你可以看蛋白质的晶体结构,通常或多或少能以亚埃、良好有序的埃级精度重建出来。材料的 XRD、X 射线 disfor——X 射线 diffu——不对
Ekin Doğuş Çubuk [00:38:05]: 衍射。
Brandon [00:38:06]: 衍射,抱歉。材料的实验 XRD 与,比方说,生物学领域相比。你会丢失哪些信息,以及为什么这有点像一种有损投影?
Ekin Doğuş Çubuk [00:38:16]: 回答这个问题时很难不让人觉得疯了。但你不觉得,有机化学和生物学几乎源于更低的 VC 维,就像更低的 Kolmogorov 复杂度模型?在那里,你几乎可以把它表示为一维序列,而且它几乎可以被编译。所以从某种意义上说,它在 Kolmogorov 复杂度的意义上更简单。
Swyx [00:38:38]: 一维?
Ekin Doğuş Çubuk [00:38:40]: 就像 DNA 就是
Brandon [00:38:43]: 是的。
Ekin Doğuş Çubuk [00:38:44]: 或者 RNA。
Brandon [00:38:44]: 我觉得大多数人会认为它更像是二维的,但是
Ekin Doğuş Çubuk [00:38:47]: 好吧,嗯,是的。我不是生物学家,所以
Brandon [00:38:49]: 是的。
Ekin Doğuş Çubuk [00:38:50]: 我确信他们是对的。但就无机化学而言,简直是疯狂的。它显然不是来自如此低的柯尔莫哥洛夫复杂度模型,因为三维无机晶体可以容纳金属、绝缘体、超导体、金刚石。差异真的非常疯狂。人们尝试过,但我们始终找不到一维表示。很难为三维无机晶体找到 SMILES 字符串。而且原子彼此相互作用的方式相当奇特,比如共价键、离子键、金属键。不过是的,这非常有意思。我有时会想,你们有没有研究过固态电解质的问题?为什么人们无法用固体取代电池中的液态电解质?会形成各种各样的枝晶。基本上锂会形成这些结构,伸入固态电解质并将其撑裂。你看着这些,会觉得与生物学相比这太原始了。在生物学中,这些系统能够实现如此复杂的纳米技术——技术进展,而我们甚至无法让两个界面彼此协同工作,因为锂会把它破坏掉。是的,但另一方面我想无机材料可以非常耐高温。你可以制造航天飞机。你可以制造硅,计算摩尔定律。所以是的,这是一种权衡。
Brandon [00:40:05]: 是的,我们和其他几位嘉宾谈过这一点。我认为在我看来一个很大的区别是,生物学方面,我们有一套工具包,基本上可以从数百万年的进化中借用
Ekin Doğuş Çubuk [00:40:19]: 是的
Brandon [00:40:19]: 那已经给了我们所有工具,我们只需复制它们。但进化也限制了,是的,蛋白质等的 VC 维相当低。像细胞这样的大规模系统涌现出的超越——行为可能会复杂得多。
Ekin Doğuş Çubuk [00:40:36]: 是的,还有意识。
Brandon [00:40:37]: 是的。是的,但也许甚至只是更具体地针对。如果你只看 X 射线,你有一块完美晶体,用 XRD 照射它,你得到、你得到光谱。那实际上仍然不能唯一地告诉你晶体结构是什么,对吧?
Swyx [00:40:54]: 不。
Brandon [00:40:55]: 所以,而这是某种
Swyx [00:40:56]: 什么?
Ekin Doğuş Çubuk [00:40:57]: 那只是平均。
Brandon [00:40:58]: 是的。
Ekin Doğuş Çubuk [00:40:58]: 那只是平均行为。但比如说,歧化是一件非常复杂的事,对吧,有可能你拥有完美晶体,平均来看它看起来像这个完美的 XRD。但实际上还有另一种模式在发生,原子略微向右或向左偏移,但平均而言它们处于中间。是的,这日子不好过,是的。
Brandon [00:41:16]: 而对于蛋白质,你基本上知道存在某种底层蛋白质遵循某些规则,所以你可以、你可以把正向模型与之结合起来。
Liam Fedus [00:41:23]: 对。然后你可以逐步走向、再走向单晶 XRD,那会有帮助。
Swyx [00:41:29]: 当我们开始做 Science Pod 的时候,里面有各种各样的元素:科学中的 AI、数学中的 AI、物理学中的 AI,诸如此类。我们之间还有一种比谁更难的较劲。我认为你可以提出科学上的论点,说材料科学是最难的。
Liam Fedus [00:41:43]: 嗯,我们选择它并不一定是因为难度。实际上有很多领域它反而更容易。我们有很好的模拟器,可以覆盖大类材料,而在生物学中,要模拟一个细胞、一个器官或一个完整生物体则极其困难。所以这是一个巨大的优势。
Swyx [00:42:01]: 是啊。这难道不有趣吗:在最小尺度上,它是一维或二维的,但随后它又多出许多数量级的尺度和结构,而这实际上引入了复杂性。这,这很奇怪。而材料中的复杂性也许更多在于微观结构。
Liam Fedus [00:42:20]: 是的,我仍然——在材料的不同长度尺度上,你仍然会遇到不同的复杂性。所以是的,微观结构以及更进一步。
梦想材料:超导体、磁体与电池
Brandon [00:42:26]: 是的。
Swyx [00:42:27]: 是的。就当是短暂地换换口味。从科幻的角度来说,假设你可以发明任何你想要的材料。什么更有价值?你知道,我这里有一份清单。显然有室温超导体。但你也提到了电池。我一直觉得是电池。只要比锂离子电池做得更好,就够了。那东西已经存在大约一百年了。最后一个是碳纳米管,主要是为了太空电梯。我不知道,就是。这就是我想到的三个。在你们的领域里,人们谈论的梦想材料还有哪些?
Ekin Doğuş Çubuk [00:43:03]: 绝对是超导体
Brandon [00:43:05]: 是的
Ekin Doğuş Çubuk [00:43:05]: 磁体。
Brandon [00:43:06]: 对。
Ekin Doğuş Çubuk [00:43:06]: Especially if we can lower the dependence on rare earth or transition metals that are hard to source, like cobalt, reducing cobalt in batteries. One of the things I think is very fundamentally important is, if we can get close to Landauer limit for compute energy efficiency.
The Landauer Limit, Reversible Computing, and Energy
Brandon [00:43:24]: Can you, explain Landauer limit?
Ekin Doğuş Çubuk [00:43:26]: Yeah. So if you look at how much energy we spend per, like FLOP or compute, it has followed the Moore’s Law type behavior, but it’s gotten exponentially more efficient. So we’ve been spending exponentially less energy per FLOP or compute. I haven’t checked recently. This was the case ten years ago when I last
Brandon [00:43:44]: Yeah, that’s great. I remember when I first learned about it, looked at that, and we were, what ten, twenty orders of magnitude away from it or something.
Liam Fedus [00:43:50]: Beyond the limit, yes, right.
Brandon [00:43:50]: And the last time I looked, it’s. We’re weirdly close. We’re within a factor of I don’t know three or four orders of magnitude away.
Ekin Doğuş Çubuk [00:43:59]: Yeah.
Brandon [00:43:59]: It’s actually like
Liam Fedus [00:44:00]: I just exponentially rolled out over many years.
Brandon [00:44:01]: Yeah, exponentially over many years are
Ekin Doğuş Çubuk [00:44:03]: Four or five, yeah.
Brandon [00:44:04]: Yeah, pretty impressive.
Swyx [00:44:05]: But in your lifetime, you
Brandon [00:44:06]: Yeah, no I. Yeah, since, just since I discovered this, what, ten or fifteen years ago? Yeah, no.
Ekin Doğuş Çubuk [00:44:11]: And one of the things we have to do is dissipate heat, because as these, computations are happening, it’s producing heat, which has to happen, but then we have to dissipate it. So if we can get close to the Landauer limit for how efficiently we do computation, that’s amazing for humanity, right? Because that means now we’re doing computation, which is probably one of the most fundamental things we do as humanity, but as efficiently as possible from energy perspective, which is a real occurrence in the universe.
Swyx [00:44:34]: That get affected by quantum computing? I’m just going to throw it out there. Theoretically, massively, embarrassingly parallel compute.
Ekin Doğuş Çubuk [00:44:41]: So it definitely gets affected by reversible computing. But honestly, I’m not an expert. I don’t understand. So if you can do reversible computing now, you’re not. You don’t actually have to spend energy to do computation because you’re not actually deleting any information. It’s reversible. But I don’t know if that works. I don’t know. Quantum computing, I assume, has to obey these laws somehow because they are still, like. They have to obey thermodynamic
Brandon [00:45:06]: It’s unitary, so
Ekin Doğuş Çubuk [00:45:07]: Yeah, exactly
Brandon [00:45:07]: It’s reversible, so for now.
Swyx [00:45:10]: One of my most memorable conversations with Elad Gil was he actually was a very skeptical person about quantum computing.
Ekin Doğuş Çubuk [00:45:15]: I see.
Swyx [00:45:15]: He’s like, he’s like, “Even if you had it today, there’s no applications.” I’m like, “Whoa.”
Ekin Doğuş Çubuk [00:45:19]: I
Swyx [00:45:19]: It breaks, it breaks the it breaks the security RSA. That’s it.
Ekin Doğuş Çubuk [00:45:22]: I agree with that. Like for them, people talk about how if you had a quantum computer today, you could simulate things so much better. And I asked them, “Okay, let’s accept that. What would you simulate?” Like, you’re still simulating perfect crystal. You still have the issues that DFT has if DFT was perfect in its prediction ability. So yeah, I do think people are kind of glossing over some things because quantum computing is so exciting. If we can compute in a quantum logic space instead of classical logic, it’s just so exciting that people are glossing over what it would actually do when it’s made. Maybe that’s fine because maybe once it’s made, it will do amazing things we can’t even imagine.
Putting Intelligence Into Every Laboratory Instrument
Swyx [00:45:57]: Good. Okay. I was going to move to the automating the lab side, where you’ve closed the loop on your experimentation. You can talk about robotic arms, talk about. But it’s basically just everything you’ve done here. My favorite quote from you was that every piece of equipment in your lab is going to have a hundred and forty IQ. Okay, what does that mean?
Liam Fedus [00:46:17]: In the early days, as we were scaling things up, we realized that we had huge bottlenecks imposed just by operating machinery. So, for one set of machinery, we would have technicians and scientists looking for particular morpho- morphology and trying to see, okay, what actually were we making, see if this is consistent with our intentions. And we really quickly, as we scaled up the lab, came into these bottlenecks where it just wasn’t keeping up. So we started doing some programmatic approaches to capturing the data on SEM. And the data that was being surfaced from this very simple program where it sort of takes a field of view, captures, zooms in captures, was just not all that useful. It was a little too dumb for what we really wanted to get. And so at that point, it was then, really pertinent to build AI systems directly onto the machines to start controlling these things. And now they have the full context as to what we were trying to achieve. So what was the intent of the experiment? What were we trying to synthesize? What were some other experimental evidence? And it’s actually looking, in the machine and capturing that data. And this is really valuable now because we’ve been talking about these hidden variables not being able to capture everything, but if you can do more intelligent data capture at the time of that experiment, your data for future AI systems and future computational predictions is that much better. You’ll have just a richer set of data. And so that’s sort of what we mean by just, everyth- everything on the lab has to be incredibly intelligent to just make the data as useful as possible. So that was, that was sort of the motivation and inspiration behind that.
Swyx [00:47:57]: What’s the state of the art there in terms of putting intelligence on every device, right?
Liam Fedus [00:48:02]: Well, I think one interesting aspect, is what’s the latency
Swyx [00:48:08]: Exactly
Liam Fedus [00:48:09]: Of controlling that instrument?
Swyx [00:48:10]: Because there’s cloud latency, but then there’s also device compute, which
Liam Fedus [00:48:13]: That’s right
Swyx [00:48:13]: Also.
Liam Fedus [00:48:14]: And then there’s sort of a time scale associated with different physical processes. And if calling out to an API or something is simply too slow, so the amount of reasoning or tokens or tool calls, is just not matched to the latency of that actual process, then that’s infeasible.
Swyx [00:48:33]: Okay. That is right.
Liam Fedus [00:48:34]: It’s almost like, self-driving cars, right? So.
Brandon [00:48:37]: I. Real quick. I just am actually surprised to hear that because the latency I’d imagine for reasoning seems small compared to a lot of these things take hours to run, right? Or maybe, you know. Or maybe I. Maybe your experiments are much faster or high-throughput or something. I’m just surprised to hear that, actually.
Liam Fedus [00:48:52]: So ultimately, the experiments do take, many hours, days, but there could be particular steps where latency would matter.
Ekin Doğuş Çubuk [00:49:00]: Yeah. You might want to have finer control. Another thing is, how expensive it is because the same reason some of these models can be really slow in analyzing data also makes them very expensive. And then finally, the human patience. If a human wants the analysis of a certain XRD pattern and they have to wait two hours before they get a good result, that’s very different, I think, than if they can get it in two minutes.
Brandon [00:49:25]: Oh, okay.
Liam Fedus [00:49:26]: Yeah. And then the reasoning, it’s. You know, you’re not doing a single call to that model, right? You’re doing, potentially many tokens, to kind of get to these patterns.
Ekin Doğuş Çubuk [00:49:35]: Run simulations in the loop.
Liam Fedus [00:49:37]: Exactly.
Ekin Doğuş Çubuk [00:49:38]: Deep research in the
Liam Fedus [00:49:39]: It’s not a single tool call. It’s not a single inference thread. So then the latency can blow up quite a bit.
Brandon [00:49:44]: I see what you mean.
Liam Fedus [00:49:45]: Yeah.
Brandon [00:49:46]: How much of your latency is tool calls? And I assume tool calls is largely DFT or computational- … and therefore. Yeah. So how much of your latency is derived from those versus actual reasoning?.
Liam Fedus [00:49:59]: It’s really process dependent.
Pragmatic Robotics, Automation, and Lab Reliability
Brandon [00:50:01]: Okay. Yeah.
Liam Fedus [00:50:01]: Yeah.
Swyx [00:50:02]: And then the other thing I think about is also, I guess, building up from small things to bigger things, where I assume that the general temptation or the typical development is incremental, where everything is human, operated, and then you find ways in which to automate it, and then you sort of build up from there. I worry that sometimes that is the way that people evolve things, but that’s a local minima, optima.
Ekin Doğuş Çubuk [00:50:25]: Like short-horizon optimization.
Liam Fedus [00:50:26]: Right.
Swyx [00:50:26]: Exactly. When actually you should get a humanoid and just put them in there.
Liam Fedus [00:50:31]: We. I think we’re of the opinion that solving humanoids would actually be slower to kind of getting to some of our goals.
Swyx [00:50:37]: Just checking. Again a lot of this is just like you do this every day. We, like
Liam Fedus [00:50:41]: Yes
Swyx [00:50:41]: See this, but we. I don’t know the reality of the situation. A lot of people having humanoids.
Liam Fedus [00:50:46]: Yeah. I think maybe one process is, by having a mix of humans and automation, you can identify really quickly what are some of the bottlenecks in the experimental process, and you start alleviating those bottlenecks one at a time. And so for example, if there’s some really tricky dexterity task that humans are excellent at but you’d have to spend months automating machine, maybe don’t spend a ton of time there. And maybe a more promising thing is something very routine that takes up a huge amount of scientists’ and technicians’ time, and it’s easy to automate. So I think it’there’s this pragmatism to it. But ultimately what we want from the lab is a huge quantity of data, high-quality data, diverse data, and those are our goals. And full autonomy is a non-goal. It’s sort of in the service that we use automation in the service of achieving the goals on the data.
Brandon [00:51:39]: Yeah.
Liam Fedus [00:51:40]: But also, I think another aspect to automation is, robots will just be. Make fewer mistakes potentially in the lab. And it’s been really helpful having AI systems with a full view over all of our data, because sometimes if there’s a permutation in data, then it can identify it. So for example, one of our steps at one point had, a cyclic error because one of the machines was loaded incorrectly, and the patterns were inconsistent. So the AI was reading through these things and says, “Well, given what was run, this is not expected.” And it’s looking at this basket of data longitudinally, and then it realized, “If I do this cyclic permutation and reverse it, everything is consistent.” And then we were able to go back to the physical infrastructure and understand that a mistake had been made in the loading. And I think managing data quality is so foundational to doing AI in the physical world. And these are the types of things that we’re building in. And so you can improve the operating process, but at that point you’re like, “Okay, this is another great opportunity for automation. How do we make this just so reliable, so durable that we never have those types of mistakes again?”
Ekin Doğuş Çubuk [00:52:51]: Yeah, the reducing the noise floor of experimental data I think is very important. Most experimental data in the literature have such a high noise floor that it’s actually usually worse than DFT accuracy. Because, when different people do experiments, different labs, different parts of the world, different times, it really introduces so much fluctuation to the result. So we are hoping that by standardizing these workflows, one of the biggest benefits will be that the noise floor will be lowered.
Open Models, Proprietary Data, and Compute Efficiency
Swyx [00:53:16]: You’ve been pretty public about how you use open source models and fine-tune them for stuff like this. Is it better. Is it basically. I can imagine a situation where, you have maybe the dumber models do one task, and then you optimize for that task, and then sort of like a generalist frontier model that supervises everything. Is that a good mental model to have? Are there more stages to this that I can think about?
Liam Fedus [00:53:37]: We basically use a mix of open source models and closed source models. You know, we’re huge beneficiaries of this from our simulation side, building up the ML stack, the simulation stack. We don’t need to push on that axis. But there’s a lot of areas where, the latency is too high, the cost is too high. But in some cases we can actually. You know, by having access to this data, we can actually push beyond the frontier of what these systems can do under even some the highest reasoning efforts. And so it’s not necessarily just about cost or Pareto efficiency, but in some cases you can push beyond it when you have access to data that no one else has. One, way to characterize this is in terms of compute efficiencies. So by having access to this data, you can be that much more compute efficient compared to some of the frontier models. And that’s been able. Been a really instrumental thing in our program.
Swyx [00:54:33]: Yeah. I’m so. Data quality or data access as a trade-off of compute I think there’s some amount of exchange rate of dollars for compute, for dollars for data that people. I feel like the pendulum might be swinging now towards data. Obviously you would, you would agree with that, but like.
Liam Fedus [00:54:50]: Yeah, especially like, talking about the noise floor of experimental data. We need to have high-quality data. If you throw a huge amount of compute against a bunch of noise, you’re not going to have a good thing emerge.
Swyx [00:55:00]: And it’s harder for you because you actively try to be sparse. You actually try to get null results and learn from that and
Liam Fedus [00:55:06]: That’s right.
Swyx [00:55:06]: Yeah.
Negative Results and Learning From Failed Experiments
Brandon [00:55:07]: Yeah, you’ve Yeah, you’ve talked about null results in several other venues. What does a null result look like for materials, and how do you use that effectively? Especially when I think your overall signal is probably quite sparse usually in terms of success.
Liam Fedus [00:55:19]: A null result could be we intended to produce some structure, and then all of the evidence points to us not producing that structure.
Brandon [00:55:28]: Do you ever intend to not produce a structure? Of course.
Liam Fedus [00:55:30]: Yeah.
Brandon [00:55:30]: Okay.
Liam Fedus [00:55:31]: Absolutely.
Brandon [00:55:31]: Okay. Cool. You do the negative controls
Liam Fedus [00:55:34]: Absolutely, yes.
Brandon [00:55:34]: Regularly. Okay.
Ekin Doğuş Çubuk [00:55:35]: There are impurity cases that will kill your property. Yeah.
Brandon [00:55:39]: Okay.
Ekin Doğuş Çubuk [00:55:40]: Can be toxic.
Liam Fedus [00:55:41]: Yeah, exactly. 100%.
Brandon [00:55:43]: Okay.
Liam Fedus [00:55:43]: And a negative result, in some cases is actually. You know, we intended to make this, the following thing, and then we’re actually able to identify a new structure that hadn’t previously been identified. So it’s. Negative in some respect. We intended to do something else, but something else emerged from the data.
Ekin Doğuş Çubuk [00:56:02]: Yeah. And also in general when you’re training a machine learning algorithm, especially like at a very basic level, if it’s a classification algorithm, if you don’t have negative samples, you can’t really train if everything is positive.
Brandon [00:56:13]: Yeah. Yeah.
Ekin Doğuş Çubuk [00:56:13]: And this is a particularly bad problem in material science because people usually publish crystals they could synthesize, but they usually don’t publish if they fail to synthesize a crystal. Sometimes they might. And also we never know really if a crystal is not synthesizable ever, right?
Brandon [00:56:30]: It could just be a skill issue.
Ekin Doğuş Çubuk [00:56:31]: Exactly.
Brandon [00:56:31]: Yeah. Yeah.
Ekin Doğuş Çubuk [00:56:31]: It could be a skill issue.
Brandon [00:56:32]: Especially with the literature.
Ekin Doğuş Çubuk [00:56:33]: Synthesis method issue, technology. Yeah. So, I think it really helps us when we do our own experiments and get negative results in the context of what we tried, so then we can even train a classification algorithm.
Liam Fedus [00:56:46]: I think there’s also another valid thing too
Brandon [00:56:48]: Yeah
Liam Fedus [00:56:49]: Of when you have this sort of string of negative results and then finally through process iteration, you’re able to get to that positive result, it’s a really interesting set of I’ll call it process engineering-type data. So it’s like through iteration, how did you actually get to that correct result? Because so much of material science has this ambiguity in actually how something was made. And so, yeah. And so there’s actually. Even if there’s a known material, it can be highly non-trivial to replicate that. Some things are, in like, a high school textbook, but other things are really at the frontier and we’re building up that know-how as well. And then building up a system that, given this string of negative results, how do you actually get to that
Brandon [00:57:35]: Yeah
Liam Fedus [00:57:35]: Positive case?
Brandon [00:57:36]: And going to the classifier result, I would almost assume that if you are doing everything in-house, most of your results will actually be negative rather than positive, which is kind of ironic because the literature only gives you positive results. So if you are sort of cold starting this, it seems like the problem is actually the reverse, that you
Ekin Doğuş Çubuk [00:57:52]: Yeah
Brandon [00:57:52]: Have an abundance of negative results and not enough positive.
Ekin Doğuş Çubuk [00:57:56]: And as you said if we were doing per experiment labels, it would be mostly negative.
Brandon [00:57:59]: Yeah.
Ekin Doğuş Çubuk [00:57:59]: But if we were doing per campaign and assuming campaigns end when we succeed, then it could be a bit more balanced.
Training on the Process of Science, Not Just Its Results
Brandon [00:58:08]: I see. So your reasoning traces could almost be over entire campaigns
Liam Fedus [00:58:12]: Absolutely.
Brandon [00:58:12]: And not. Okay.
Liam Fedus [00:58:13]: This is, this is the key thing. This type of data basically doesn’t exist anywhere else, and we spend so much of our time getting the full lineage of the scientific process into the model. And so tracking all this data, like the conversations, the intuitions, what was executed in the lab, where the computations run, where was the code written, stitching all this together is so valuable. And I think the overall goal is rather than training on the final output of science, you’re training on the process of doing science.
Brandon [00:58:44]: This feels very much like you. What you would try to do with RSI, where you are having a model trained to train better models. You’re now having a model trained to make better experiments. But
Liam Fedus [00:58:56]: Absolutely
Brandon [00:58:56]: The difference is, it’s not going back into the core model for. So it’s
Ekin Doğuş Çubuk [00:59:00]: Unless the physics models you’re developing are improving chips, which then improve AI. So this is the wider
Brandon [00:59:07]: The bigger. Yeah, the higher level RSI.
Liam Fedus [00:59:09]: That’s the big data flywheel. Yeah.
Brandon [00:59:10]: But does that mean that. Is it plausible that in Fable 6 or something, or GPT-8 could just have a. Which has been tuned on these sort of higher order, reasoning traces, could actually just do this without any
Liam Fedus [00:59:30]: Mm
Brandon [00:59:30]: Logic? Because it’s sort of the same thinking process, right?
Liam Fedus [00:59:33]: Yeah. I think there’s, decision-making under uncertainty. Obviously in machine learning, there’s noise, in running the AI loops. But we do think there’s a different set of challenges when you’re actually interfacing with the physical world. But also I think there’s another piece too, which is getting the compression of everything into weights is still really valuable. If inference time reasoning was sufficient, all of the frontier labs would have stopped training at GPT-4 and were like, “Okay, every. From now on out, we’re going to get really good at inference time improvements.” and so we think that by. You know, because of the differences between, physical sciences and machine learning, getting that compressed into our own weights will lead to different types of systems and different types of capabilities.
Ekin Doğuş Çubuk [01:00:20]: But also, we feel like even if Fable 7 gets really good at. Even better than what it is today, it will still have to run experiments to get results. And the reason for it is, right machine learning is really good at what it’s been trained on but scientific discovery is almost by definition what you haven’t been trained on. And that’s why we’re building these labs, so that, whether open models or closed models can use these labs to tinker with the universe, because we don’t feel like you can make a big discovery without trying things.
Liam Fedus [01:00:50]: Yeah, there’s
Brandon [01:00:51]: Yeah
Liam Fedus [01:00:51]: There’s not going to be like. No one’s going to zero shot the room-temperature superconductor.
Ekin Doğuş Çubuk [01:00:54]: Think theoretically
Brandon [01:00:55]: That would be pretty cool. Yeah. As someone who has worked closely with wet labs before, certainly I am more skeptical about zero shotting scientific results than some.
Liam Fedus [01:01:05]: Yeah.
Brandon [01:01:05]: But it is something that I think, some people might ask, so.
Liam Fedus [01:01:08]: Yeah. I think it’s really important to distinguish the results we see in math and theoretical physics from the physical world. Right? Like
Brandon [01:01:18]: Yeah
Liam Fedus [01:01:18]: These are two different things.
Brandon [01:01:19]: Two very different things.
Liam Fedus [01:01:20]: Yes.
Ekin Doğuş Çubuk [01:01:20]: I’m not even theoretical physics. So far it’s been theoretical computer science
Liam Fedus [01:01:22]: Yes
Ekin Doğuş Çubuk [01:01:22]: Coding and math.
Liam Fedus [01:01:23]: Right.
Brandon [01:01:24]: Yeah.
Ekin Doğuş Çubuk [01:01:25]: Maybe theoretical physics is next.
Liam Fedus [01:01:26]: Don’t come. Yeah.
Brandon [01:01:28]: Yeah.
Levels of Scientific Abstraction and the Human Role
Ekin Doğuş Çubuk [01:01:29]: Oh, yeah.
Swyx [01:01:29]: Can I get a sort of mental model of the layers, levels of extraction that you can go? So for example, in terms of automation, and I’m just on this theme again where we talked about the campaign, talked about individual essays and running tests and how they’re. Humans are bad at it, so we should stop humans from doing it. Are there others? So for example, one level higher than campaign could be, a physics theory that you’re testing? Or one level lower than campaign is what? And like, I just. I like to think about it from that point of view and Think about it from like, okay, well, this is an API call now. You never have to touch this again. And this one is. Yeah, no, this is still ninety percent human. And maybe we can sort of draw the map of the territory that way. Is there — are there other levels?
Ekin Doğuş Çubuk [01:02:17]: I can tell you some levels. I feel like you asked a good question. I don’t have a very systematic answer, but let’s talk about some levels. So one level is the atomistic structure. There’s an abstraction, right? There’s no perfect atomistic structure in anything we do, but it’s one approximation. Another one is the continuum model. So this one is not atomistic anymore, but it’s, a continuous mesh representing a material.
Swyx [01:02:39]: Okay.
Ekin Doğuş Çubuk [01:02:39]: And then in a different dimension, there’s the thermodynamics. Assuming, you do this experiment forever, what would be the final state? And then there’s another layer of abstraction, which is kinetics, acknowledging that we’re not doing this experiment forever, so time matters. So how quickly will a reaction happen or not, even if it’s lower energy or not? So thermodynamics, kinetics, atomistic, continuum. What else? Are there other levels of abstraction?
Liam Fedus [01:03:04]: Well, I think your point about okay, there could be an overarching new theoretical advance that informs multiple campaigns right?
Swyx [01:03:14]: It makes a difference. Kind of like as an investor, I want to go here’s the bottleneck, guys. And when we get this, we get everything else. I don’t know.
Ekin Doğuş Çubuk [01:03:22]: To me, from the beginning, our hypothesis has been that one of the big bottlenecks is automated characterization, because it’s not that hard to mix powders to get to try stuff. But if you can’t characterize and analyze it and then decide what the next step should be intelligently, you don’t really benefit much from mixing powders randomly. So we focused a lot on automated characterization, closing the loop with simulations, because that’s how you try a lot of things, make an informed decision, and then decide what to do the next day.
Brandon [01:03:51]: Maybe as a follow-up question, what is how do humans live in this loop? Where at what points do you pull drop the human in? I’m guessing you probably
Swyx [01:04:02]: You could, you could randomly drop in at any part
Ekin Doğuş Çubuk [01:04:04]: Of course.
Swyx [01:04:04]: And add value.
Liam Fedus [01:04:05]: Yes.
Swyx [01:04:05]: But where do you decide to spend your time?
Brandon [01:04:08]: It seems like it might be all levels. You have you have people in the lab who are physically moving materials, and then you also have scientists who are guiding decisions about what campaign, and then you have humans who are. What’s that? I’m just asking a question for you. Yeah.
Liam Fedus [01:04:21]: Yeah. One hundred percent. I think it’s always shifting, too.
Brandon [01:04:23]: Yeah.
Liam Fedus [01:04:24]: So, I think even just using the guidance of campaigns, in the early days, that was fully human-driven. Now, increasingly, it’s a mix of AI-driven and human-driven. And then Doğuş was saying early on we identified that characterization was a big bottleneck. In the early days, it was human-driven, and the scientists were very much overwhelmed, and the balance has shifted significantly towards AI-driven. And this has allowed the scientists who were previously spending all of their time doing these refinements to now elevate their work to something else. But I think we’re. It’s really sort of a function of time. There’it’s always changing at each of these levels.
Search, Optimization, and Design of Experiments
Swyx [01:05:04]: Yeah. In terms of just general search, are there things that perform well in let’s say, the physical world that don’t perform well in the other worlds or vice versa? Like so evolutionary search is pretty popular in LLMs. I don’t imagine it works well here.
Ekin Doğuş Çubuk [01:05:21]: So especially in simulations, people have been using evolutionary search, for a long time. It’s pretty good at structure prediction, for example, finding low energy structures.
Swyx [01:05:31]: Okay, so it works. Yeah.
Ekin Doğuş Çubuk [01:05:33]: And then process engineers in semiconductor industry will do DOE, design of experiments, and they’ll often use some kind of zeroth order whether it’s Bayesian optimization or evolutionary search or. I don’t think they use RL, but similar idea, yeah.
Swyx [01:05:47]: You have to have some kind of algorithm
Ekin Doğuş Çubuk [01:05:48]: Yeah
Swyx [01:05:49]: To just figure it out.
Ekin Doğuş Çubuk [01:05:50]: And all zeroth order optimization behaves the same anyway, right? So.
Scaling Labs and Reusing Data for Reinforcement Learning
Brandon [01:05:54]: Yeah. Okay, I want to go back to the question I had earlier, which was one of scaling. And so I-I’m curious, what is your vision for scaling the lab? So there’s different ways. Coming from the biology world there’s a lot of cool tricks you can do in biology. You can tag things, you can use DNA sequencing for all sorts of interesting readouts. And if you can map your complicated assay onto sequencing, you can now scale to millions or hundreds of millions or something. Can you do tricks like this? Ultimately, every assay has its own runtime. I’ll, I’I have a favorite blog post I would want to. I’ll quote in the show notes. But what does the runtimes look like for these labs? Is it. Is your scaling just for mater-materials just fundamentally linear in terms of you do you just need more resources to do more things, or can you paralyze things in clever ways and combine things, and are there tricks there?
Liam Fedus [01:06:52]: Well, I think one way to think about it is, the RL environment is not quite trivially run the experiment, wait for that agent to fully roll out, doing the calculations, going through all the instruments to the final outcome. We basically start producing data across all the different instruments and all these different campaigns. And one way you can begin to expand that for machine learning is now reinforcement learning environments can be constructed based on again dates for the experimental campaigns or computational campaigns at that point. Or you can also, start taking, subsets of instruments, and you’re like, “Okay, we’re going to create reinforcement learning environments on the basis of this instrument alone.” Or maybe it’s, across these different instruments to kind of get to some reward state. So that’s a way where. This finite basket of data, when viewed in different ways, can then be expanded, for training purposes. So it doesn’t get at the actual experiment run, but from an ML perspective.
Ekin Doğuş Çubuk [01:07:55]: Maybe I can give you some analogs for the bio examples you gave. So, one thing people have tried is combinatorial sputtering. I don’t know if you’ve ever heard of this. You have these sputtering targets, and you create this compositional gradient. So when you look at the final product, because there is gradient from each precursor, it creates different crystals at different spatial locations. So in one go, you maybe try hundreds, thousands of crystals. That’s an that’s similar, right, to your example.
Liam Fedus [01:08:23]: Oh, yeah, that makes sense.
Ekin Doğuş Çubuk [01:08:24]: And as far as I know, this hasn’t worked very well so far because it turns out there is diffusivity, so things move around. Another one that we thought about, but we haven’t done yet, but a lot of people talk about this is superconductivity measurements are a bottleneck because unlike XRD, there’s no high-throughput superconductivity measurement, and it can take one hour per measurement. So people have thought about taking the different candidates, mixing them all into one sample, and putting it in. And if, it has superconductivity, you know that one of the sources that. And then you can do a bit like how people are doing COVID testing to. So there are ideas like this exist. Some of them work well, some of them don’t. Yeah.
Custom Hardware, New Labs, and Scaling Ambition
Brandon [01:09:00]: So you, I think, are about to announce a big fundraise. How are you thinking about, scaling?
Liam Fedus [01:09:06]: We’ve been doing our design of our labs for a while, and so I think the resources allow us to continue to scale up the labs significantly, but also the compute, both from the AI side and the computational side. But I think really importantly too, it’s what are the new types of labs we’re able to build? So I think we’ve been proving out this loop, in our initial labs here in Menlo Park, but we’ll be continuing to expand to new types of labs and repeat the same process.
Ekin Doğuş Çubuk [01:09:33]: Yeah, like the kind of lab we’re trying to build, I think haven’t been built before at this scale or at this approach. So we have learned a lot from our own build-ups so far, and those lessons guide us for the next version and the next version. And then we can, as you said increase the scale, the ambition, the quality of instruments. Some instruments are very expensive. As we get a really good understand for what we get a lot of return from we can invest into those more.
Liam Fedus [01:09:58]: And it makes. Hardware engineering has become so core to it. So for example, in the early days, just for speed, we’d buy some off-the-shelf instruments. And as we kind of push the scale of the lab
Brandon [01:10:08]: You make your own.
Liam Fedus [01:10:09]: We have to make our own. So we realize, okay, for this instrument, it’s pretty fast. These components are actually pretty quick, but this weighing machine actually becomes the bottleneck. Or there can be things pertinent to data quality where, well, the resting position of this robotic arm actually is above the plate of previously mixed things, so there’s a contamination risk. So in our hardware design, we’re going to redesign it so that the resting position lies away from the-those things. These are the kind of subtle details that allow us to kind of push the noise floor down for our experimental campaign. So, basically take these learnings and scale it up.
Assembling the Team Behind Synthesis Superintelligence
Brandon [01:10:45]: You ship your org chart. You have many teams. I think the. We were surprised when we looked at your jobs page, and we were like, “Okay, we don’t think we fully map what Periodic does.”
Liam Fedus [01:10:56]: Maybe some organizing principles, so to kind of make sense of the job page. So we hire for AI research and infrastructure, computational, experimental roles, and then hardware engineering roles, and then product roles.
Ekin Doğuş Çubuk [01:11:10]: We want to achieve synthesis superintelligence, right? And we feel like for that, we need the right chemistry expertise, the right physics expertise, simulation theory, and thin films, powder. So hardware engineering these seem like very diverse roles, which they are, but they’re all actually, coherent in what they’re trying to do, which is a synthesis for intelligence. And this also applies to the LLM researchers, the infra, even the product engineers, because the product engineers kind of make sure all the research gets made into a product that the experimentals in the lab can use.
Liam Fedus [01:11:42]: Yeah. So in order to do the end-to-end loop with the physical world, it really requires your ability to have agency over the physical world to build these labs, to run the campaigns effectively, to create AI systems against that, to make them good users of the computational tools. That’s like the difficulty, but also the opportunity of Periodic. It’s like this group of people has just never been brought together before. It’s irreducibly a multidisciplinary problem.
Ekin Doğuş Çubuk [01:12:07]: Yeah. And one other guiding principle for us that is really important for us is we want the people who are very good at what they do and experienced to do it hands-on. You know the modern life has gotten us to this place where, especially in academia, when someone is really good at research in some area, we give them so much responsibility for grant writing, teaching, that they stop having the time to do research in that area hands-on. But if you look back at Bell Labs, IBM, institutions that made really good progress, it was really experienced people doing hands-on work. Like Bardeen was in the lab every day, even though he’s a theorist. Alex Müller was in the lab doing experiments, even though he was the lab lead. So we try to do that here too. So we have some of the world’s leaders in different fields, but they’re doing hands-on work.
Brandon [01:12:52]: Do you want to brag or just call them out? You know.
Ekin Doğuş Çubuk [01:12:54]: Oh, I would love to yeah. So, when I was doing my PhD, my favorite computational material scientist who’s kind of from my age group is Muratan Aykol. He did his PhD with one of our advisors, Christopher Wolverton. And he is kind of like our computational material science expert. He’s incredible. He’s sometimes, a really good experimentalist, even though he does simulations. He understands systems really well. Our lab lead is Joe Checkelsky, who is a professor at MIT, but he’s on leave to work with us full-time. And Joe is this incredible physicist, understands superconductivity really well, but he also was a professor in Japan for a bit, so learned synthesis and chemistry really well, especially for a physicist. We have Daniel Chica, who was a graduate student in Mercury Canasides’ group, which is maybe the best solid-state chemistry group in the world. And Daniel was one of his synthesis experts, so he’s like a magician with his fingers. There’s so many examples like this.
Liam Fedus [01:13:50]: No, I know. I think Dima Bahdanau, so he’s been leading a lot of our AI and LLM efforts. He was the inventor of neural network attention.
Brandon [01:13:59]: Oh, Bahdanau. The Bahdanau attention.
Liam Fedus [01:14:01]: Exactly. Yeah.
Brandon [01:14:03]: Yeah.
Liam Fedus [01:14:03]: Yeah. Yeah, so very hands-on in sort of getting into the traces, the training distribution, how to connect it to the lab.
Ekin Doğuş Çubuk [01:14:10]: The chemistry.
Brandon [01:14:11]: Chemistry.
Ekin Doğuş Çubuk [01:14:11]: He’ll go deep.
Liam Fedus [01:14:12]: He was very deep. He actually has solid-state synthesis books at his lab in Montreal. Ray Nakano, so he used to tech lead some of the operator work at OpenAI. He’s in the lab. So when we were doing some lab tours, we were initially, we’re like, “Oh, this new, scientist, this new technician really looks like Ray.” And no, it was literally this was Ray actually in the lab, with the scientists automating the machines. And it’s like this is sort of the DNA and the type of people we need to kind of pull this off.
Forward-Deployed Engineering for Semiconductor Partners
Swyx [01:14:43]: So, we also saw that you have forward-deployed engineering, roles. We happen to be spinning up a forward-deployed engineering podcast because there’s so many engineers that would do that role. They may not necessarily know that there even is a role for them at Periodic. What is it, and how do they work with customers?
Liam Fedus [01:15:01]: So we’ve been building our own products, our own tools, our own AI and computational
Swyx [01:15:06]: For yourself. Yeah.
Liam Fedus [01:15:06]: For ourselves. So we’ve been customer zero. Now what we’re doing is we’re taking these same tools to different industries, and we’we have a huge focus right now in the semiconductor industry. However, in order to do this these are very, private companies. We need to be able to operate within some of the most secure environments. And so our Periodic forward-deployed engineers and researchers will actually be on site working and actually integrating our AI systems, our computational systems to help our partners get to their goal states more quickly. So basically the tools, the know-how that we’ve built in doing our own materials discovery, materials engineering, we’re now bringing to industry. And we think this is the most effective way to get partners to their end states and to their goals, rather than just throwing some technology over the wall and telling them to figure it out. Then these forward-deployed engineers will also do the inference locally, but then also can, train on the data. So again, once we take our system that understands these different areas and we deploy, we can make it expert on the customers or the partners’ data so that they can own their own intelligence and have systems that understand this much more effectively than just hitting some API or some untrained model. And so this is a much expanded, forward-deployment engineering role than typical because this really requires some very deep, machine learning expertise. You have to be incredibly precise from a data training perspective, infrastructure perspective, but also a physics perspective. So we’re working with highly technical industries. They have to understand, different areas of chemistry, material science, some of the devices. And so that, those are the roles we’re building out right now.
Swyx [01:16:53]: Part of it is willingness to spend extended periods on site in Taiwan. It’s like, okay, well, I know that kind of customer.
Swyx [01:17:01]: It’s interesting that you’re monetizing. That is the primary way that you’re monetizing now, but it might change in the future. But like, that’s like a model that I think people don’t really get about APL Research Lab. It’s. And when you compare yourself to Bell Labs, that’s what people are thinking, right? Which is.
Commercializing Scientific Discovery
Liam Fedus [01:17:19]: I think a really great analog would be software engineering.
Swyx [01:17:22]: Right.
Liam Fedus [01:17:22]: So in the early days of software engineering we had GitHub Copilot, then we had early versions of ChatGPT. People were using these things as copilots to help them get to their solutions. And as the automation improved, now we have things like Codex, and very few of our engineers are writing code the way they used to. And we think a very similar thing could play out for Periodic as well, where you can have systems to accelerate the researchers, the material scientists, the materials engineers, the process engineers. But as the autonomy, intelligence, and capability grows, you can begin to price outcomes. So help me get to this kind of goal state, and I think that’s, a really interesting area for us.
Ekin Doğuş Çubuk [01:18:03]: Yeah. We kind of feel like our best contribution to solid-state physics and science could be if we made these tools and topics profitable, similar to how ChatGPT made CS and LLM majors way more popular in colleges before and after. We’d love to show that solid-state physics, material science research can make a big impact commercially, and then that will attract more attention. Young people will want to study physics, which will be, a dream. And Bell Labs did make a huge commercial impact, right? So they failed to commercialize some of their incredible advances. They, of course, did some research that just couldn’t be commercialized like the cosmic, background. But they did benefit a lot from the vacuum tube connecting, East Coast to West Coast by phone lines.
Liam Fedus [01:18:49]: Yeah. And I think it’s really also interesting to think that technology and capital are incredibly intertwined. So if you look at what was the progress on chatbots for many years imagine recounting, okay, what was the progress from chatbots from I’ll say 2010 to 2015? We’d be kind of at a loss to say over that five-year period, how much did they improve? Whereas if you look at the period from, 2021 to 2026, it’s night and day. And what happened was ChatGPT and these other systems were able to achieve a product market fit, and it changed the capital landscape entirely. This changes the landscape for hiring, compute, data, et cetera.
Swyx [01:19:31]: The virtuous cycle.
Liam Fedus [01:19:32]: Exactly. And so
Swyx [01:19:32]: Like success breeds success.
Liam Fedus [01:19:33]: Exactly. So technology, it’s incredibly coupled to the capital and those resources, and we want to achieve the same thing in the physical world.
Open-Source Contributions and Academic Research
Swyx [01:19:41]: Yeah.
Brandon [01:19:42]: So one thing I’ve been seeing is there’s this move from science being funded from government grants and so on to VCs and private funding. Do you all plan on making anything you’re doing, open source, like either releasing data sets or actual models? Is this something in the future even. You know, there are many models which may be not well, your state of the art, but which could still be you know, useful releases.
Ekin Doğuş Çubuk [01:20:07]: Absolutely. So, we have a couple of things. So we have been contributing to open source to kind of like PyTorch and the Materials Project, code base, Custodian, their DFT runner. TorchSim was actually created by one of the researchers here, Abhijit Ganggan, and he maintains it still. JAX-MD, Abhijit maintains it. So we do a lot of open source contributions, including LLMs.
Liam Fedus [01:20:31]: XLANG, Megatron.
Ekin Doğuş Çubuk [01:20:32]: Yeah.
Liam Fedus [01:20:32]: Yeah. So we’ve been very active contributors back into open source.
Ekin Doğuş Çubuk [01:20:36]: And we also have an academic grant program where we give academic gift grants to academic groups in universities that we feel like are really advancing this direction towards synthesis superintelligence. So that’s also been great to see. I think the first paper from this funding is about to come out, so it’ll be exciting to just get more papers come out like this. Yeah.
The Path to Better Superconductors
Brandon [01:20:57]: Okay. Let’s steelman this, say that you’ve achieved 100% lab automation, which does everything, characterizes everything correctly. You have all of the fun agents which can do everything. I still am somewhat unclear about how you actually get to superconductivity. There’s a lot of hard problems to solve. What is the path there when there is no sort of theory for model for most of the strongly correlated systems or high-temperature systems?
Ekin Doğuş Çubuk [01:21:23]: Yeah. So, our view is that it’s not hard to think of chemical spaces that would host superconductivity. What’s really hard is to synthesize them. That’s why we’re really emphasizing synthesis superintelligence. So maybe one example to give historically is Alex Müller, when he thought that transition metal oxides might be a good path to superconductivity, he was actually trying nickelates, the nickel, oxygen, and some, cations, and that didn’t work. And then he tried cuprates, which is copper, oxygen, and some cations, and that worked, and then he got a Nobel Prize. But then we know that later on turns out nickelates was a good idea. Nickel and copper are next to each other on the periodic table.
Brandon [01:21:59]: But only as two layers, right?
Ekin Doğuş Çubuk [01:22:00]: Yeah. Like there are only so many 3D transition metals.
Brandon [01:22:03]: Yeah.
Ekin Doğuş Çubuk [01:22:03]: And then, Harold Hwang, one of our, Stanford professors here who’s incredible, he realized he can make nickelates in thin film form and show the superconductivity. So, if we have a really good synthesis superintelligence in the lab and we can really scale up these things, I think we won’t run out of ideas, or the LLM won’t run out of ideas about what directions to try, like 3D transition metals, oxygen. There are very related ideas here. Many people have been trying cobaltates, which is like instead of copper or nickel, you try cobalt. So I think that’s super exciting. Maybe one historical example to also study is the Japanese group that discovered magnesium diboride. As you know, that’s the highest temperature, ambient pressure, conventional superconductor, and they discovered it by just trying a bunch of materials. I think they tried 30,000 different things. Thirty of them seemed to host interesting superconductivity, and MgB₂ was one of them. But magnesium diboride sat on people’s shelves as a precursor for decades before then, and BCS theory had been invented back in 1957. So the way people found these materials wasn’t projected from theory. It wasn’t because they couldn’t synthesize until then. It was just like they tried a bunch of things. So now imagine if we had this perfect automation that you’re describing, which sounds like a dream, and we tried the 30,000 the Taketo group tried in over his career in a month. We just really increased the surface area for luck.
Closing: Increasing the Surface Area for Discovery
Swyx [01:23:30]: No, but yeah, it’s very inspiring what you’ve done. Congrats on all your success. I feel, I feel like you’re creating, yeah, the modern playground. I imagine recruiting must be super easy for you, so I’m just a little bit jealous, but you’ve worked hard for it to get here, so yeah.
Ekin Doğuş Çubuk [01:23:44]: Yeah.
Swyx [01:23:44]: Yeah. Well, thank you so much. Yeah, it was great chatting with you both.
Ekin Doğuş Çubuk [01:23:46]: Yeah, super fun. Yeah.
Swyx [01:23:47]: Yeah. Thank you.
来源:Latent Space · latent.space