Skip to content

Paper2Agent Turns Research Papers Into AI Agents: What the Nature Study Shows

Paper2Agent turns research papers and their code into tested AI agents. A Nature study shows how the system improves scientific method reuse while exposing major reproducibility limits.

Paper2Agent Turns Research Papers Into AI Agents: What the Nature Study Shows

On this page

A research paper can now become something you can talk to and use, rather than something you merely read. Paper2Agent, a system described in a September 2026 Nature study, converts research papers and their associated code into interactive artificial intelligence agents that can execute methods, reproduce analyses and work with new data. The interesting part is not simply that an AI can summarize a paper; it is that the system attempts to turn the paper's underlying workflow into tested software that another AI agent can call.

Paper2Agent changes what a research paper can contain

Most computational research still separates the explanation from the machinery needed to use it. A scientist may find a promising paper, then hunt for its code, install dependencies, understand the expected inputs, reproduce an example and finally work out how to adapt the method. Paper2Agent attacks that gap by treating the manuscript, supplementary material, datasets, code and analysis workflows as parts of one usable research artifact. The resulting system exposes those components through a Model Context Protocol (MCP) server, a standardized way for artificial intelligence systems to discover and call external tools and resources.

That distinction matters because Paper2Agent is not simply retrieval-augmented generation over a PDF. A normal document-based AI system can retrieve a paragraph describing an analysis, but it may still have no reliable way to run the analysis itself. Paper2Agent instead extracts executable functions from the research code, packages the relevant environment and tests the resulting tools against expected outputs. A scientist can then interact with the paper through natural language while the underlying agent invokes the original computational workflow.

Paper2Agent builds the agent before the scientist asks a question

The system uses several specialized agents rather than asking one model to understand everything at once. An environment agent prepares the software needed by the repository, while an extraction agent identifies important methods and turns them into MCP tools. A testing agent then runs those tools, compares their outputs with reference results and repairs problems before the tool is accepted. The final MCP server can be connected to a compatible AI agent, creating a paper-specific assistant that knows how to access the publication's resources and execute its methods.

That validation step is one of the most consequential parts of the design. The researchers say a tool must produce expected files, keep numerical results within defined tolerances or reproduce reference figures before it passes. Tools that repeatedly fail are left out rather than being presented as working capabilities. Each accepted tool also keeps a reference to the original research code, giving users a path back to the implementation instead of asking them to trust an unexplained AI-generated procedure.

The first large test found a surprising number of usable papers

The researchers tested Paper2Agent on 100 computational biology papers selected from bioRxiv without filtering the sample for code quality or repository maintenance. Seventy-four papers could be successfully converted into agents, producing 599 proposed tools, of which 593 passed automated validation. The remaining 26 papers exposed the less glamorous side of scientific computing: missing executable code, unavailable data or model files, broken environments and scripts that were too specific to their original example.

Those failures are as informative as the successful conversions. A paper cannot become a dependable software interface if the research artifacts needed to reproduce it are incomplete. Paper2Agent therefore exposes a practical connection between AI readiness and reproducibility: the cleaner the code, data and documentation behind a paper, the easier it becomes to turn the research into something another system can actually use. The researchers do not claim that every scientific paper can be automatically transformed into a reliable agent, and their 74-percent success rate makes that limitation visible.

The benchmark shows why tested tools beat raw repository access

The researchers also compared Paper2Agent with a coding agent that received direct access to the papers and their repositories. On 300 questions derived from the original tutorials, Paper2Agent paired with Claude Sonnet 4 reached 91.2 percent accuracy, compared with 80.3 percent for Claude Code using the same model with direct repository access. A later Sonnet 4.6 baseline reached 86.3 percent, which was still below the Paper2Agent result on this evaluation.

The difference is easier to understand when the task is viewed as software execution rather than question answering. A general coding agent has to rediscover which files matter, understand dependencies, infer how the authors intended the workflow to run and decide which functions to call. Paper2Agent performs much of that work ahead of time and validates the resulting interface. The reported benchmark therefore suggests that preparing and testing the tools can matter as much as giving an AI more access to the original material.

EvaluationPaper2Agent resultComparison
300 tutorial-derived questions91.2% accuracyClaude Code + Sonnet 4: 80.3%
42 execution tasks from 10 non-biology papers98.1% accuracyFive independent runs
100 computational biology papers74 successfully agentified599 tools proposed, 593 validated

The reported cost and speed also moved in the same direction. On the 300-question evaluation, the Paper2Agent setup averaged about $0.20 and 1.6 minutes per query, compared with about $0.38 and 4.3 minutes for direct paper-and-repository access using Sonnet 4. Those figures are measurements from the researchers' evaluation, not a universal price or speed advantage for every paper. Different repositories, models and workloads could change the balance substantially.

It also works beyond computational biology

To see whether the approach was tied to biology, the researchers tested ten computational papers from areas including artificial intelligence, statistics, econometrics, game theory and astrophysics. Across 42 execution-based tasks, Paper2Agent achieved 98.1 percent accuracy across five independent runs. That is a much smaller test than the 100-paper biology evaluation, so it should be treated as evidence that the approach can transfer across domains rather than proof that it works equally well everywhere.

There was also a fallback for papers whose code could not be converted into executable tools. For 26 data- and discovery-focused papers, Paper2Agent created a structured resource layer from the manuscripts and supporting material. On 100 synthesis questions, that layer reached 89.0 percent accuracy, compared with 82.0 percent for a browser-based baseline. The broader lesson is that a paper does not have to become fully executable to benefit from being reorganized into machine-readable research resources.

Multiple paper agents can collaborate on one research problem

The most interesting demonstration goes beyond making one paper easier to use. The researchers connected agents built from different studies so they could combine methods and datasets during a scientific investigation. One system represented AlphaGenome for genomic variant interpretation, while others represented studies using massively parallel reporter assay coupled with single-cell CRISPR interference and Perturb-seq data from human CD4-positive T cells. The agents were used together to investigate a psoriasis-associated genetic variant and prioritize a possible causal gene.

This is where the idea starts to resemble a research network rather than a collection of chatbots. One paper can contribute a prediction method, another can provide experimental evidence and a third can supply a different dataset or analysis technique. The AI system can move between those resources without requiring a researcher to manually reconstruct every software workflow. In the psoriasis case, the researchers kept humans in the loop for important decisions, so the demonstration should not be read as an autonomous scientist discovering and proving a medical mechanism by itself.

The distinction is important because scientific interpretation remains harder than executing a known method. An agent can calculate a result, compare candidates and suggest validation strategies, but deciding whether the evidence supports a biological conclusion requires domain knowledge and judgment. The researchers explicitly describe open-ended hypothesis generation and mechanistic interpretation as human-in-the-loop activities. Paper2Agent makes the machinery easier to operate; it does not remove the need to decide whether the result deserves belief.

The biggest limitation is still the research itself

Paper2Agent cannot repair every weakness in the publication process. If the original repository lacks important data, depends on unavailable software, contains undocumented assumptions or only works for one carefully prepared example, an automated agent has limited room to recover. The 26 unsuccessful conversions in the 100-paper test demonstrate this directly. Better AI tooling can expose those problems faster, but it cannot manufacture missing scientific evidence.

There is another boundary worth keeping in mind: a validated tool is not automatically a validated scientific conclusion. Paper2Agent can test whether software reproduces expected outputs, but matching an original result does not establish that the underlying experiment was well designed or that the interpretation was correct. The system's strongest promise is therefore operational rather than philosophical: it can make published computational methods easier to inspect, reproduce and apply while preserving a connection to the original implementation.

Paper2Agent points toward papers that can actually do something

The larger change may be in how researchers think about publishing computational work. Instead of treating the PDF as the final product and the repository as an optional attachment, a future paper could ship with structured resources, tested tools and workflows designed for both humans and AI systems. That would make the publication useful to an AI agent without requiring every researcher to reverse-engineer the author's software first. It would also create a stronger incentive to document dependencies, data formats, expected outputs and reproducibility steps.

Paper2Agent does not make the traditional paper obsolete, and its own results show that many research repositories are not ready for automatic agentification. What it demonstrates is more practical: once research methods are exposed as tested, callable tools, an AI system can do more than explain what scientists published. It can help people run those methods, combine them with other work and investigate new questions. The next step is likely to be less about making AI better at reading papers and more about making scientific papers better prepared for machines to use.

S

Written by

Sarah Khan

I’m fascinated by artificial intelligence and the rapid changes happening around AI tools, models, and agents. I enjoy testing new AI technologies, following important developments, and understanding how they can be useful in real life. I like explaining complex AI topics in a simple and practical way.

57 posts published

All posts by this author

0 Comments

No comments yet. Be the first to share your thoughts.

Join the conversation

Log in or create a free account to leave a comment. You can edit or delete your own comments any time.