Open-Source BootLoops Harness Helps AI Models Perform Precise Scientific Calculations
Harvard physicist Matthew Schwartz used the open-source tool BootLoops and Claude to produce 36 manuscripts across 18 fields in three months, but human experts remained essential to make the results scientifically valuable.
A Harvard physicist has demonstrated how open-source software can help artificial intelligence models carry out precise scientific calculations, producing dozens of research manuscripts in a matter of months. Matthew Schwartz used a tool called BootLoops alongside the AI model Claude to generate 36 manuscripts spanning 18 different fields, from particle physics to linguistics, according to a report on the project.
The results, however, often only became scientifically valuable once human experts stepped in to verify and refine the output. Schwartz's experience points to both the promise and the current limits of using large language models in rigorous academic work. The BootLoops harness is designed to support AI models in performing calculations that require exactness, a task where general-purpose chatbots frequently struggle.
Schwartz's advice to researchers is blunt: «Look at everything yourself.» The warning reflects a broader reality in the scientific community, where AI-generated results can appear convincing but may contain subtle errors that only a trained eye can catch. In Schwartz's case, the manuscripts covered topics ranging from particle physics to linguistics, a breadth that would be unusual for a single researcher working without AI assistance.
The project highlights a growing trend of scientists experimenting with open-source AI tools to accelerate their work. BootLoops is available for free, allowing researchers to adapt it to their own needs. By combining such tools with human oversight, Schwartz was able to produce a volume of work that would otherwise take much longer. Yet the need for expert review remained constant, suggesting that AI is not yet ready to replace the scientific method's reliance on verification and peer scrutiny.
For the scientific community, the experiment raises questions about how AI should be integrated into research workflows. While the technology can handle calculations and draft manuscripts, the final judgment about scientific validity still rests with human experts. Schwartz's approach — using AI to generate, then personally checking every result — may serve as a model for others seeking to harness these tools without compromising rigor.
The BootLoops harness itself is open source, meaning its code can be inspected, modified, and improved by anyone. This transparency is important for scientific applications, where reproducibility and error detection are critical. As AI models become more capable, tools like BootLoops could help bridge the gap between raw computational power and the precise, verifiable results that science demands.
3
