跳到正文
原文
The Decoder· Matthias Bastian·· 3 小时前AI 评分59

开源工具BootLoops支持AI模型完成精确科学计算

Open-source "BootLoops" harness supports AI models in performing precise scientific calculations

AI 导读

开源工具BootLoops支持AI模型完成精确科学计算,源代码已在GitHub公开。哈佛大学物理学家Matthew Schwartz在Anthropic客座文章中介绍,团队三个月内与19位共同作者完成覆盖18个领域的36篇手稿,Claude用该工具算出30个积分,其中15个复现已知结果、15个此前未曾算出。

正文 · 原文

Automated problem-solving and relevant research aren't the same thing.

That's the takeaway from an Anthropic guest post by Prof. Matthew Schwartz, a physicist at Harvard University. He describes how he stopped trying to use Claude like a human researcher and instead started looking for "Claude-shaped problems," tasks that play to the strengths of today's AI models. Schwartz is also currently a visiting researcher at Anthropic.

"Claude-shaped problems" sit at the overlap between what scientists want to study and what AI can actually do. | Image: Anthropic

That approach led to BootLoops, an open-source harness for exact scientific calculations. The source code is available on GitHub.

Using the harness, Claude found connections between particle physics, ecology, population genetics, economics, and linguistics, according to Schwartz. But the results often only became scientifically valuable once domain experts stepped in to set the direction.

Schwartz uses the concept of the "convex hull" to illustrate how AI can help in science. Human knowledge is fragmented, with individual fields extending in different directions while the gaps between them go unexplored. A lab might spend 20 years studying one set of genes with a single method, leaving neighboring genes and alternative approaches untouched. An AI harness like BootLoops is supposed to fill those gaps by drawing connections across the jagged frontiers of knowledge.

Human knowledge spans various disciplines while the gaps between them remain uncharted. AI harnesses like BootLoops are designed to fill those gaps. | Image: via Anthropic

36 manuscripts across 18 fields in three months

Over three months, the team produced 36 manuscripts across 18 fields with 19 co-authors. Schwartz started with particle physics calculations, specifically scattering amplitudes and elliptic integrals. Within weeks, Claude computed 30 integrals using BootLoops, fifteen reproducing known results and fifteen computed for the first time.

Then he expanded the search to other fields. In ecology, Claude solved a 20-year-old equation from neutral biodiversity theory that had previously been impossible to compute at scale. When applied to data, the results showed that tree species composition on Barro Colorado Island in the Panama Canal is changing 4.5 times faster than the theory allows. Ecologist James O'Dwyer then helped turn that finding into a better predictive model.

In population genetics, the team analyzed 5.7 billion mutation pairs from the 1000 Genomes Project and found evidence for a mechanism called gene conversion. Other projects included an AI data editor for economics journals that automatically checked 4,452 replication packages (published as an NBER Working Paper), as well as a word stress database covering 6,072 languages, built with three linguists.

"Python for engineers is now obsolete"

Schwartz describes how fast AI is changing science, to the point where planning ahead is becoming nearly impossible. Why apply for a three-year grant to fund a calculation that an AI model might solve overnight?

Training PhD students has also become an open question. Since his earlier post on Vibe Physics, that feeling has only grown stronger. In some fields like computer science, the disruption is alarming. Two years ago, he would have called a "Python for Engineers" course "essential," but today it's "unnecessary" because Claude can handle those tasks.

Building ML models to study physical phenomena is another area AI can now take over. Even deep knowledge of neural networks offers limited returns when Claude can implement the current state of ML research on command.

"Look at everything yourself"

Schwartz also warns about the models' weaknesses. Claude likes to declare victory too early, and phrases like "done, with one asterisk" often mean "not done at all." It misjudges how long tasks will take and tends to brute-force calculations instead of finding more elegant solutions. Automated checks aren't reliable, and Claude's conclusions can be wrong even when its calculations are correct.

Beyond reliability, the model gravitates toward old, heavily cited debates rather than genuinely new questions. And the projects were "compute- and token-intensive," Schwartz says.

Schwartz also sees the focus on big math problems and headlines as risky. Unrealistic expectations could distract from productive applications that already work today. Real progress will still be built on solid foundations, just faster. The scientific method itself isn't threatened, and human guidance and taste remain indispensable, Schwartz says.

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

来源:The Decoder · the-decoder.com