Written in June 2026
Reproducing experiments is important to advance science, but it’s a time-consuming activity for scientists, without much reward. LLMs are a promising tool for fully automating reproduction of experiments that only need computation. This literature review examines work done in this area, by:
25 auto-reproducibility systems and 15 auto-reproducibility benchmarks have been identified. 13 out of 40 (32.5%) were published in 2026, and 24 (60%) were published in 2025. This shows that there is currently a large interest in this area. However, there is currently, as far as I’m aware, no published dedicated survey on this, which is why this literature review has been conducted.
Trisovic et al.[1] examined more than 2000 replication datasets, and found that 74% of files crashed on initial execution. Even after code cleaning, 56% crashed. This means that, currently, significant work is needed in order to reproduce computational research, even when the authors provide the code that was used in the experiments. However, being able to rerun code and get the same result is not full reproducibility. Sometimes, the paper excludes important details that are needed to reproduce the experiment. This means that these details can only be found in the associated code, preventing the reader from developing a full understanding of the experiment. It’s therefore desirable that an experiment should be reproducible based only on the paper. Edward Raff[2] tried to replicate 255 papers this way, and was able to replicate 162 of them (63.5%), where replication was defined as confirming more than 75% of the claims in the paper. An important aspect of judging whether a paper has been reproduced, is to decide whether a non-identical result leads to the same conclusion. This is very common in ML, because most algorithms have a non-deterministic element. Raff[2], for example, uses an order-of-magnitude threshold for claimed improvements (i.e. the paper claims 700x faster, but reproducer sees 300x).
Gundersen 2021[3] lays out four reproducibility types that categorize distinctions like these:
R1: Only the textual description of the experiment and raw data is provided.
This includes the reproductions in Raff 2019[2].
R2: The code for the experiment is provided in addition to the textual descriptions.
This includes the reproductions done in Trisovic et al.[1].
This is combined with three degrees of reproducibility: Outcome reproducible, Analysis reproducible, Interpretation reproducible. The type and degree can then be combined to describe both the starting point and reproducibility threshold for a reproducibility study. For example, Raff 2019[2] starts with only the paper and raw data, and compares the outcome against the results from the paper. This could be ”outcome reproducible”, but my interpretation of Gundersen 2021[3] is that the outcome must be identical for this to be the case. Nor do the evaluations assess ”analysis reproducibility”, because I interpret this to mean that something like a statistical test is done, and that is not what is done in the provided examples in Raff 2019[2]. Therefore, I will class this as ”interpretation reproducible”, and R1 since they start with the paper only. The abbreviation is therefore IR1.
The method used is ”snowballing”. Starting with a set of seed papers. Then assessing related papers recursively.
The search began with these papers
Automating Computational Reproducibility in Social Science: Comparing Prompt-Based and Agent-Based Approaches[4]
Automated Modernization of Machine Learning Engineering Notebooks for Reproducibility[5]
CompRep: A Dataset For Computational Reproducibility[6]
From AI for Science to Agentic Science: A Survey on Autonomous Scientific Discovery[7]
SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers[8]
REPRO-BENCH: Can Agentic AI Systems Assess the Reproducibility of Social Science Research?[9]
From Reproduction to Replication: Evaluating Research Agents with Progressive Code Masking[10]
Automated Reproducibility Testing in R Markdown[11]
AUTOREPRODUCE: Automatic AI Experiment Reproduction with Paper Lineage[12]
PaperBench: Evaluating AI’s Ability to Replicate AI Research[13]
Comparing Human-Only, AI-Assisted, and AI-Led Teams on Assessing Research Reproducibility in Quantitative Social Science[14]
Computational reproducibility in computational social science[15]
Reproscreener: Leveraging LLMs for Assessing Computational Reproducibility of Machine Learning Pipelines[16]
AutoAgents: A Framework for Automatic Agent Generation[17]
Automated Social Science: Language Models as Scientist and Subjects[18]
The replication crisis has led to positive structural, procedural, and community changes[19]
from earlier work by Odd Erik Gundersen, Thijs Snelleman and Peter Lawrence. Almost all papers that cited (or were cited by) these papers, were classed as ”not relevant”, ”relevant” or ”included”. Google Scholar[20] was used to find out which papers have cited the ”included” papers. Queries were made between March and June 2026. These factors were used to assess papers:
Papers that described automated reproducibility systems, or benchmarks of them, were categorized as ”included”. Other papers that were assessed to be important to automated reproducibility, based on personal judgement, were also categorized as ”included”. ”included” papers were then used to expand the search, by assessing their citations and citers.
A visualization of portions of the resulting citation graph can be seen in Figure 0. It shows how the seed papers generate the 1. generation of discovered papers, which in turn generate the 2. generation, and so on. It also shows that one generation generates papers in the next generation backwards in time through its citations, and forwards in time through papers that cite it.
Categories of papers were created and modified for relevant papers along the way. The point of this is to group similar papers together, for summarizing the literature, and for comparison. The categories used were:
Previous work - 10 papers
Literature reviews or surveys of automatic reproducibility.
Auto reproducibility - 23 papers
Systems that automatically reproduce computational research.
Research impact - 8 papers
Contributes to our understanding of the impact of automated reproducibility.
Terminology - 30 papers
Contributes by defining terminology related to reproducibility.
Validation - 60 papers
Contributes benchmarks or other validation techniques for automatic reproducibility.
Code Repair - 12 papers
Systems that automatically repair old code.
Journal policy - 22 papers
Papers that describe policies at journals related to reproducibility.
Reproducibility survey - 49 papers
Papers that survey reproducibility in a field.
Some whole categories were deemed not relevant only at the end of the process. They can be found in the appendix.
Based on the terminology from Gundersen 2021[3], I’ve created categories and abbreviations that describe the input-output behaviour of auto-reproducibility systems:
R1 → R4 (Full implementation)
The system starts from only the paper and raw data, then writes and executes the code, before outputting the code and execution data.
R1 → R3 (Full implementation + ablation)
System writes and runs code to produce results, then hides the code from its output. This somewhat pointless, and will therefore not be mentioned again in this review, but I’ve included it for completeness. Clearly any R1 → R4 system can easily be modified to do this.
R1 → R2 (Paper to code)
These systems take a paper as input and produce code that can be used by a different system to reproduce the results, enabling the transformation of an R2 system into an R1 system.
R2 → R4 (Code repair/Execution)
These systems take code that crashes or is otherwise defective and repairs it. By running repaired code, a reproducer now has access to experimental data, ending up at an R4 starting point.
R2 → R3 (Code repair/Execution + ablation)
System repairs and runs code to produce results, then hides the code from its output. Mostly pointless. Included for completeness. Won’t be mentioned again.
R3 → R4 (Full implementation with answer sheet)
The implementation agent is given the experimental results to help it produce correct code. It would ideally output the resulting data from its own code, instead of forwarding the original experimental data.
{O, A, I}R{1, 2, 3, 4} (Full reproducibility assessment)
An auto-reproducibility system can start from any reproduction starting point (R1, R2, R3, R4), and at the end of a reproduction attempt, conclude whether the experiment is reproducible according to a set reproducibility degree (Outcome, Analysis, Interpretation)
These systems are composable. One can for example take a code repair system (R2 → R4), and add on an OR4 system, by feeding the output of one into the other, and end up with an OR2 system. In the extreme, one could chain 4 systems together to make an OR1 system:
Paper to code (R1 → R2)
Code repair/Execution + ablation (R2 → R3)
Full implementation with answer sheet (R3 → R4)
OR4 judge
LLM tools were used to create figures, decide on styling, extract citations from papers (unsuccessfully), and as a conversation partner when trying to understand papers (however, this wasn’t done to any of the papers I’ve cited in this review, as far as I can recall), as well as helping with latex syntax. None of the words in this review have been generated with AI, and no proofreading or feedback has been done with AI.
1156 papers were assessed in total. Of these 28 were ”included”, 511 were ”relevant”, 617 were ”not relevant”. Papers that were assessed to be ”relevant”, are not included in this review, because there are too many of them. I was unsure how many relevant papers existed when I began this review, and many of the ”included” papers weren’t published yet at that point. It therefore made sense to cast a wide net. It turned out that the field is large, so the ”relevant” category was dismissed entirely. 6 papers that should be assessed according to the method, weren’t assessed because I couldn’t access them, and the recursive ”snowballing” ended when I ran out time to expand the review further.
Here are the statistics on how many systems end up fulfilling each classification from the method section. Systems that fulfill multiple classes are counted multiple times.
Class | Reproduction systems | Benchmarks |
R1
| 10 | 3 |
R1
| 5 | 4 |
R3
| 0 | 0 |
R2
| 5 | 4 |
OR1: outcome reproduction w/ paper only | 2 | 0 |
OR2: outcome reproduction w/ paper & code | 4 | 2 |
OR3: outcome w/ paper & experiment data | 0 | 0 |
OR4: outcome reproduction w/ paper, code & experiment data | 2 | 1 |
AR1: analysis reproduction w/ paper only | 0 | 0 |
AR2: analysis reproduction w/ paper & code | 2 | 1 |
AR3: analysis reproduction w/ paper & experiment data | 0 | 0 |
AR4: analysis reproduction w/ paper, code & experiment data | 0 | 0 |
IR1: interpretation reproduction w/ paper only | 1 | 2 |
IR2: interpretation reproduction w/ paper & code | 4 | 3 |
IR3: reproduction w/ paper & experiment data | 0 | 0 |
IR4: interpretation reproduction w/ paper, code & experiment data | 0 | 0 |
|
Significant work has been done on auto-reproduction systems. These are the systems that were identified by the literature review:
Paper2Code[21] (R1 → R2)
approaches the paper to code problem by drawing inspiration from typical software development workflows. They decompose the task into three coordinated stages: Planning, analysis and coding. Planning is then further decomposed into an overall plan, architecture design, logic design and configuration. The analysis stage deals with the behaviour of each file. Given a specific file from the planning stage, the system generates a detailed analysis of what needs to be implemented in that file. The final coding stage produces a full code repository, by generating one file at the time, using that file’s analysis. Importantly, this is done ”bottom up”, so that the agents have access to the complete code for the necessary dependencies.
Enhancing Automated Paper Reproduction via Prompt-Free Collaborative Agents[22] (R1 → R2)
builds on the Paper2Code framework, and aims to remove human-crafted system prompts. The authors validate their solution by integrating with Paper2Code[21], which enables easy testing against Paper2CodeBench[21] and PaperBench[13].
AutoP2C[23] (R1 → R2)
has four stages that imitate human experts, when implementing code repositories from papers. The first stage analyzes well-established ML code repositories, looking for common architectural patterns, in order to produce a template for the final repository. The second stage parses the pdf of the paper, extracting the necessary information for code implementation. The third stage is a hierarchical task decomposer, that takes the distilled information and the template, and creates an implementation plan. The final stage, takes task descriptions from the third stage and transforms them into an executable code repository. It does this through an implementation-verification loop. The loop terminates when the code reproduces the results in the paper.
Can GPT-4 Replicate Empirical Software Engineering Research?[24] (R1 → R2)
approaches the paper to code problem in the simplest possible way with LLMs, just prompting the LLM to write the analysis pipeline. The authors restrict the task to software engineering papers, and conduct a user study with human software engineering experts.
RepLLM[25] (R1 → R2)
restricts the paper to code problem to computer networking research. The authors claim that the field present a formidable challenge, due to long-context constraints and intricate logical dependencies. The system consists of four specialized agents: Content Parsing, Architecture Design, Code Generation, and Audit & Repair. The agents coordinate through an explicit shared memory. This decouples information storage from agent reasoning.
Deep-Reproducer[26] (R1 → R2)
approaches the paper to code problem by searching the web for existing repositories that can be used as a starting point for the coding task. It does this after first having done a comprehensive paper analysis. This summarizes six critical dimensions: Data requirements, Methodological approaches, Task objectives, Research methodology, Related work and Foundational methods.
RePro[27] (R1 → R2)
approaches the paper to code problem by extracting what authors call a paper fingerprint. This is a set of atomic pass-or-fail criteria that act as supervisory signals to the coding agents, guiding them to a correct implementation. A verifier agent uses the pass-or-fail criteria to give succinct feedback to the coding agents.
NERFIFY[28] (R1 → R2)
The authors restrict the paper to code problem to neural radiance field computer vision papers. This restriction lets them constrain the agents’ actions to producing architecturally correct code, and use a specialized critique agent.
PaperRepro (a)[29] (R1 → R2)
has as its contribution PaperRepro, which must not be confused with PaperRepro by Zhang et al.[30]. The system attempts to solve the paper to code problem through distinct stages. The first stage builds and prunes the citation graph, in order to surface necessary tacit knowledge. The next stage uses the paper and surfaced tacit knowledge to create the code implementation, through execution-feedback refinement.
Sci-Reproducer[8, pages 6–7] (R1 → R2)
is the default agent for SciReplicate-Bench[8].
AutoReproduce[12] (R1 → R4)
approaches the full implementation problem by first establishing the necessary implicit domain knowledge. The idea is that a key limiting factor on other auto-reproducibility systems is that LLMs lack this knowledge by default, and that the author’s paper lineage algorithm can establish it. A multi-agent system then writes the code and executes the experiments.
From paper to benchmark[31] (R1 → R4)
takes on full implementation within the Prognostics and Health Management (PHM) field, or as they refer to the problem: ”We formulate framework-grounded agentic PHM reproduction as a benchmarkable paper-to-implementation task in a scientific ML, focusing on a domain with severe reproducibility constraints and limited code availability.” [31, page 2] They use a framework-coupled agent workflow. The framework in question provides fixed task contracts, dataset adapter interfaces, transformation operations, windowing operations, and metric definitions. These are benchmark conditions that are held fixed across all reproductions. The aim of this is to make code generation verifiable.
AI Copilots for Reproducibility in Science[32] (R1/R2 →* R4)
approaches full implementation and code repair/execution as an AI-human collaboration, where the AI system generates a Jupyter notebook, based on the paper (and existing code if available/desired). The idea is that the agent shouldn’t try to resolve ambiguities itself, but instead prompt the human researcher on how these should be resolved. I’ve not classed this system as an assessment system, because the paper has nothing to say on how the researcher should decide if an experiment has been successfully reproduced.
ResearchCodeAgent[33] (R1/R2 → R4)
approaches full implementation and code repair/execution problem with a multi-agent system. It consists of a planner and a worker. These each have a bespoke set of actions that they can take to mutate their environments. Several actions that the planner can take invoke the worker to do a specified task. The result of that task is returned to the planner when done.
Jailbreak Foundry[34] (R1/R2 → R4)
approaches the full implementation and code repair/execution problem, restricted to papers on LLM jailbreak attacks, by using a multi-agent system, which does iterative planning, implementation and fidelity checks. The target for code is a module, that is compatible with JBF-LIB, which is a library by the authors, that sets a standard interface to produce attack benchmarks. The authors’ intent is not to assess the reproducibility of proposed jailbreaks. They are instead focusing on shortening the time between the publication of a jailbreak, and when mitigations can be tested, by automatically creating a benchmark to develop against. Jailbreak Foundry does not qualify to be a reproducibility system, because it doesn’t explicitly test the jailbreaks against the models used in the original paper. However, I suspect that this would be a trivial extension to add. The system can act both as R1 → R4 and R2 → R4, depending on whether existing code is available.
Automating Computational Reproducibility in Social Science[4] (R2 → R4)
approaches the code repair/execution problem in two ways: A prompt based workflow and an agent based workflow. The prompt based workflow works by feeding code, error logs, and a prompt to an LLM. The LLM then attempts a single shot repair which is then executed. This goes back to the beginning, with the new code and errors, if the code fails to execute correctly.
The agent workflow uses commercial off the shelf coding agents, specifically OpenCode[35] and Claude Code[36].
Computational Reproducibility of R Code Supplements on OSF[37] (R2 → R4)
approaches the code repair/execution task with a light touch. It does not use AI, or modify code. It only attempts to construct an environment where the code will run unmodified. It then publishes the environment as a Docker image, along with the execution logs.
LockForge[38] (OR1)
restricts the OR1 problem to logic locking research in integrated circuits. Logic locking refers to techniques for designing integrated circuits, that restrict functionality, unless the user is authorized. This is meant to prevent IP piracy, counterfeiting, malicious modification, and reverse engineering, among other concerns. The multi-agent system consists of 5 stages: forethoughts, implementation, refinement loop, smoke tests, and validation. Reproducibility is a part of the final validation. It works by measuring the fidelity of code execution on paper-specified benchmarks.
Comparing Human-Only, AI-Assisted, and AI-Led Teams on Assessing Research Reproducibility in Quantitative Social Science[14] (OR1*/OR2*)
tests the effects on outcome reproducibility assessment, from human teams using AI, and human teams led by AI.
ReplicatorAgent[39, pages 5–6] (IR1/IR2)
is the reference test subject for ReplicatorBench[39]. It’s a multi-agent system, similar to others. However, it has a novel option of forcing the agent to translate all non-Python scripts into Python in order to run them. The rationale for this is that LLMs prefer working with, and are more proficient in, Python[40], so the authors hypothesize that forcing the agent to work in Python, will improve performance.
REPRO-Agent[9, page 9] (OR2/AR2/IR2)
is the default agent for REPRO-Bench[9]. It follows an organized template, which is based on successful reproducibility assessment cases to improve planning. Input-Output behaviour is discussed in the benchmark section on REPRO-Bench.
PaperRepro (b)[30] (OR2/AR2/IR2)
not to be confused with PaperRepro by Li et al.[29], approaches the problem of assessing reproducibility of papers and available code, with a two stage, multi-agent system. The system consists of an execution stage, and an evaluation stage. A setup agent configures the environment and creates a plan for the reproduction attempt. An execution agent follows the plan and runs the script it has written, which produces research artifacts. In the evaluation stage, a scoring agent evaluates reproducibility, using the reproduced artifacts, the paper’s reported results, and an execution summary. A report agent aggregates everything into a reproduction report. The reason why all reproducibility levels are applicable, is because PaperRepro[30] conforms to REPRO-Bench[9], which requires the system to specify which level has been attained. See the discussion of REPRO-Bench in the benchmark section.
Reproscreener[16] (IR2) is a reproducibility assessment software tool. It approximately has the same input-output behaviour as an IR2 agent, and therefore complements well with LLM methods, as dissimilar redundancy.
Artisan[41] (OR2/OR4)
approaches the code repair/execution task, which they refer to as artifact evaluation, by using a single coding agent, that is guided by two judges. One judge is an LLM that gives feedback on how well the agent’s code represents the paper’s methodology. The other judge is entirely hand-coded, and compares the computed output from the coder’s script, to that of the original artifact. It gives a simple pass/fail for each value in the result. The system’s preprocessing has obfuscated the results from the paper and other artifacts, in order to prevent the agent from simply copying the results. However, the authors had to implement even further measures to mitigate this to an acceptable level. It is unclear whether this should be classified as OR2 or OR4, since the experimental data is required for it to operate. However, the idea is to produce a reproduction script, which doesn’t depend on it.
Supporting Artifact Evaluation with LLMs[42] (OR4*)
restricts the OR1 problem to computer security papers.
Partial automation of computer security reproduction.
Some papers need significant modification to fit into the classification:
Natural Language to Code Generation in Interactive Data Science Notebooks[43]
The main difference between existing auto-reproducibility systems, is not the models they’re using, but system prompts and scaffolding. Most use a multi-agent approach, where the task is decomposed sequentially, and specialized agents execute each task.
The field has not agreed on a common problem statement for automated reproducibility, and some authors who have contributed to the field had slightly different problems in mind. This makes it difficult to compare performance across systems. Due to the differing problem statements, most papers primarily (and sometimes exclusively) develop their own dataset and scoring system.
The same classification system, based on Gundersen 2021[3], can be used to classify benchmarks for automated reproducibility systems. The classification is decided by which classes of systems the benchmark can test.
Paper2CodeBench[21, pages 6–8] (R1 → R2)
is the benchmark for Paper2Code[21]. Reproducibility is not their main focus, but they run a small evaluation on 10 papers with human judges, and find that their system reproduces the papers 29.46% of the time.
It consists of 90 papers from ICLR, ICML, and NeurIPS 2024, filtered to have code repositories smaller than 70 000 tokens, and by an LLM ranking.
Evaluation is done with an LLM judge, both with and without the original repository as reference, in addition to a human evaluation by the original paper’s authors.
Paper2Repo[23, pages 7–8] (R1 → R2)
is only used by AutoP2C[23]. It consists of 8 papers and covers training strategy optimisation, node classification, model compression, parameter-efficient fine-tuning, network pruning, and image classification. See AutoP2C[23] in the auto-reproducibility system, for input-output behaviour.
SciReplicate-Bench[8] (R1 → R2)
benchmarks paper-to-code agents on 100 tasks from 36 NLP papers.
PaperBench[13] (R1 → R4)
is a full implementation benchmark.
It consists of 20 papers from ICML[44], across 12 different ICML topics.
ReplicationBench[45] (R1 → R4)
aims to benchmark agents on the full implementation task, on 20 astrophysics papers. The agents are graded on answers to questions that require various levels of implementation. The task’s difficulty is measured on a logarithmic scale from 1 to 9, based on an estimate of how long a human non-author domain expert would take to complete the task. 1 corresponds to a few minutes. 9 corresponds to several weeks.
ReproduceBench[12, pages 5–8] (R1 → R4)
Only used by AutoReproduce[12].
13 curated papers across diverse domains, including knowledge distillation and Partial differential equations.
Evaluates whether the generated code accurately reflects the paper and code by the authors, using an LLM judge.
Also evaluates whether the generated code reproduces the original result, using the original metric from the paper.
AutoExperiment[10] (R1/R2 → R4) is a benchmark which blurs the line between full implementation and code repair/execution. This is an intentional feature, where the benchmark operator can choose a continuously varying starting point, between R1 and R2. The authors call their approach to accomplish this ”progressive code masking”. The benchmark software is configured with a integer n. It masks out n functions from the dataset, which means removing a function’s implementation. Setting n to 0 corresponds to an (R2 → R4), since the code is passed along to the agent unaltered. Setting n to its max value corresponds to an (R1 → R4), since all of the code is masked out. One could argue that the remaining function signatures (function name, argument names, and return types) don’t match up with the requirement that no code is provided, but I interpret R2 to mean that the original authors made a good faith attempt to share the code used to produce the original results. A code repository consisting only of function signatures would obviously not make one believe this. However, a single missing function body is an oversight that can easily happen with bad version control.
Super[46] (R2 → R4)
is a code repair/execution benchmark, that focuses on environment setup as its core challenge. It contains 45 end to end problems, 152 sub-problems, and 604 automatically generated problems.
CORE-Bench[47] (R2 → R4)
is testing agent’s abilities to run code from research papers, but sees its role primarily to be testing abilities the authors suspect to be important for reproducibility testing, but benchmarking these abilities on real research code. The benchmark has three difficulty levels: CORE-Bench-Easy, CORE-Bench-Medium, and CORE-Bench-Hard. The easy task does not conform to the classification scheme, because the agent only has to answer questions about the pre-existing code output. The hard task conforms well to the R2 → R4 class, if the question-answering is merely viewed as the judgement mechanism for how well the agent executed the code. The medium task also conforms to R2 → R4, but is easier, because the agent is handed a Dockerfile. I have not classified CORE-Bench[47] as a reproducibility assessment benchmark1 , because none of the task questions I could see in the dataset[48] on github asked the agent to assess whether any experiment was reproduced. The first task is close: ”Report the average AUC score of Had using the CTGCN-C method on the UCI dataset.” This question leaves the assessment for reproducibility to someone/something else.
EnvBench[49] (R2 → R4)
is a benchmark for code repair/execution agents. It only checks the agent’s capability to set up a correct execution environment. It encompasses 994 code repositories, and uses only generic static analysis tools to judge performance, which is what enables the size of the benchmark.
MMReview[50] (IR1)
is a benchmark for automated peer review. Peer review is not a reproducibility assessment, but it fits the input-output behaviour of IR1, since it is a judgement of paper quality. I’ve identified one (of potentially many) techniques from the paper that are transferable to benchmarks for automated reproducibility assessment. The paper uses adverserially constructed inputs, to test whether LLM agents can overcome their sycophantic tendencies[51]. This involves inserting fake strengths and weaknesses into papers before giving it to the agent.
ReplicatorBench[39] (IR1/IR2)
is one of a handful of reproducibility assessment benchmarks, that includes negative examples, and by that I mean irreproducible claims. It is crucially important that an automatic assessment system is able to confidently declare a paper irreproducible, when it is possible (and correct to do so). Otherwise, irreproducible papers will be stuck in an in-between state, where it’s unclear whether it is irreproducible, or if the agent is not capable enough to reproduce it. This problem persists no matter how well the agent performs on benchmarks with only positive examples, such as Artisan-Bench[41, pages 9–17], because out-of-distribution performance degradation is always possible.
ReplicatorBench doesn’t test whether a system can reproduce the same outcome as the original study, with the same data. It aims to instead test whether the system can figure out if the paper’s findings generalize to new equivalent experimental settings. The agent is handed the new data by the benchmark, instead of having to produce it itself, because the benchmark is focused on social science, which is not a purely computational field.
SocSciRepro-Bench[52] (IR2)
has as its goal, to widen the coverage of benchmarks on social science. It does so by including 221 tasks from 54 papers. The dataset covers four disciplines, and has both confirmed reproducible and irreproducible papers. The paper is under Review at Nature Machine Intelligence.
REPRO-Bench[9] (OR2/AR2/IR2)
is a benchmark for automated reproducibility in social science. The agent is given:
The agent must output a single number between 1 and 4 (inclusive), which is the paper’s reproducibility score. These are the definitions of what each score represents, as well as the corresponding reproducibility levels from Gundersen 2021[3]:
”major findings are irreproducible.” (not reproducible)
”there are minor inconsistencies and/or errors in the provided code. [...] A score of 2 indicates identifiable inconsistencies in the code that do not alter the paper’s major findings.” (Interpretation reproducible)
”there are minor issues at the display and reporting level, e.g., rounding errors.” (Analysis reproducible)
”major findings are fully reproducible.” (Outcome reproducible)
Artisan-Bench[41, pages 9–17] (OR2/OR4)
is only used by Artisan[41], and consists of 23 curated software engineering papers with available and evaluated artifacts. The benchmark has examples in python, java, rust, scala, ocaml and bash.
The literature for automatic reproducibility has, only recently, become extensive. However there are still gaps which should be filled. The plurality of the attention has gone to the paper to code problem, and much less attention has been given to fully automating reproducibility assessments (see Figure 0). This is reasonable, given that writing code and, to a similar extent, successfully running code is the biggest obstacle with manual reproduction. It is good then, that code repair/execution and full implementation has received a lot of attention, and some full assessment systems exist for specialized research topics. However, I believe that fully automated and general computational reproducibility systems, that are as capable as current specialized systems, are within reach.
Ana Trisovic et al. “A large-scale study on research code quality and execution”. In: CoRR abs/2103.12793 (Mar. 2021). arXiv: 2103.12793. url: https://arxiv.org/abs/2103.12793.
Edward Raff. A Step Toward Quantifying Independently Reproducible Machine Learning Research. Sept. 2019. arXiv: 1909.06674 [cs.LG]. url: https://arxiv.org/abs/1909.06674.
Odd Erik Gundersen. “The fundamental principles of reproducibility”. In: Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences 379.2197 (Mar. 2021). issn: 1471-2962. doi: 10.1098/rsta.2020.0210. url: http://dx.doi.org/10.1098/rsta.2020.0210.
Syed Mehtab Hussain Shah, Frank Hopfgartner, and Arnim Bleier. “Automating Computational Reproducibility in Social Science: Comparing Prompt-Based and Agent-Based Approaches”. In: Companion Proceedings of the ACM Web Conference 2026. ACM, May 2026, pp. 989–998. doi: 10.1145/3774905.3795485. url: http://dx.doi.org/10.1145/3774905.3795485.
Bihui Jin, Kaiyuan Wang, and Pengyu Nie. Automated Modernization of Machine Learning Engineering Notebooks for Reproducibility. 2026. arXiv: 2602.07195 [cs.SE]. url: https://arxiv.org/abs/2602.07195.
Lázaro Costa, Susana Barbosa, and Jácome Cunha. “CompRep: A Dataset For Computational Reproducibility”. In: Proceedings of the 3rd ACM Conference on Reproducibility and Replicability. ACM REP ’25. Association for Computing Machinery, Oct. 2025, pp. 168–178. isbn: 9798400719585. doi: 10.1145/3736731.3746160. url: https://doi.org/10.1145/3736731.3746160.
Jiaqi Wei et al. From AI for Science to Agentic Science: A Survey on Autonomous Scientific Discovery. Aug. 2025. arXiv: 2508.14111 [cs.LG]. url: https://arxiv.org/abs/2508.14111.
Yanzheng Xiang et al. SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers. Aug. 2025. arXiv: 2504.00255 [cs.CL]. url: https://arxiv.org/abs/2504.00255.
Chuxuan Hu et al. REPRO-Bench: Can Agentic AI Systems Assess the Reproducibility of Social Science Research? July 2025. arXiv: 2507.18901 [cs.CL]. url: https://arxiv.org/abs/2507.18901.
Gyeongwon James Kim et al. From Reproduction to Replication: Evaluating Research Agents with Progressive Code Masking. June 2025. arXiv: 2506.19724 [cs.AI]. url: https://arxiv.org/abs/2506.19724.
Andreas Brandmaier and Aaron Peikert. “Automated Reproducibility Testing in R Markdown”. In: Collabra: Psychology 11 (June 2025). doi: 10.1525/collabra.138638.
Xuanle Zhao et al. AutoReproduce: Automatic AI Experiment Reproduction with Paper Lineage. May 2025. arXiv: 2505.20662 [cs.AI]. url: https://arxiv.org/abs/2505.20662.
Giulio Starace et al. PaperBench: Evaluating AI’s Ability to Replicate AI Research. Apr. 2025. arXiv: 2504.01848 [cs.AI]. url: https://arxiv.org/abs/2504.01848.
Gunther Bensch and Belay Moges. “Comparing Human-Only, AI-Assisted, and AI-Led Teams on Assessing Research Reproducibility in Quantitative Social Science”. In: (Jan. 2025).
David Schoch et al. Computational Reproducibility in Computational Social Science. July 2023. arXiv: 2307.01918 [cs.CY]. url: https://arxiv.org/abs/2307.01918.
Adhithya Bhaskar and Victoria Stodden. “Reproscreener: Leveraging LLMs for Assessing Computational Reproducibility of Machine Learning Pipelines”. In: Proceedings of the 2nd ACM Conference on Reproducibility and Replicability. ACM REP ’24. Rennes, France: Association for Computing Machinery, July 2024, pp. 101–109. isbn: 9798400705304. doi: 10.1145/3641525.3663629. url: https://doi.org/10.1145/3641525.3663629.
Guangyao Chen et al. AutoAgents: A Framework for Automatic Agent Generation. Sept. 2024. arXiv: 2309.17288 [cs.AI]. url: https://arxiv.org/abs/2309.17288.
Benjamin S. Manning, Kehang Zhu, and John J. Horton. Automated Social Science: Language Models as Scientist and Subjects. Apr. 2024. arXiv: 2404.11794 [econ.GN]. url: https://arxiv.org/abs/2404.11794.
Max Korbmacher et al. “The replication crisis has led to positive structural, procedural, and community changes”. In: Communications Psychology 1.1 (July 2023), p. 3. issn: 2731-9121. doi: 10.1038/s44271-023-00003-2. url: https://doi.org/10.1038/s44271-023-00003-2.
Google Scholar. url: https://scholar.google.com (visited on 06/08/2026).
Minju Seo et al. Paper2Code: Automating Code Generation from Scientific Papers in Machine Learning. Apr. 2026. arXiv: 2504.17192 [cs.CL]. url: https://arxiv.org/abs/2504.17192.
Zijie Lin et al. Enhancing Automated Paper Reproduction via Prompt-Free Collaborative Agents. Dec. 2025. arXiv: 2512.02812 [cs.AI]. url: https://arxiv.org/abs/2512.02812.
Zijie Lin et al. AutoP2C: An LLM-Based Agent Framework for Code Repository Generation from Multimodal Content in Academic Papers. Apr. 2025. arXiv: 2504.20115 [cs.SE]. url: https://arxiv.org/abs/2504.20115.
Jenny T. Liang et al. Can GPT-4 Replicate Empirical Software Engineering Research? Oct. 2024. arXiv: 2310.01727 [cs.SE]. url: https://arxiv.org/abs/2310.01727.
Yining Jiang et al. Leveraging Large Language Models for Automated Reproduction of Networking Research Results. Sept. 2025. arXiv: 2509.21074 [cs.NI]. url: https://arxiv.org/abs/2509.21074.
Pengcheng Chen et al. “Deep-Reproducer: From Paper Understanding to Code Generation”. In: NeurIPS 2025 Fourth Workshop on Deep Learning for Code. Sept. 2025. url: https://openreview.net/forum?id=zw2DpSxnXn.
Mingyang Zhou et al. RepRo: Reflective Paper-to-Code Reproduction Enabled by Fine-Grained Verification. Aug. 2025. arXiv: 2508.16671 [cs.SE]. url: https://arxiv.org/abs/2508.16671.
Seemandhar Jain et al. NERFIFY: A Multi-Agent Framework for Turning NeRF Papers into Code. Feb. 2026. arXiv: 2603.00805 [cs.CV]. url: https://arxiv.org/abs/2603.00805.
Lehui Li et al. What Papers Don’t Tell You: Recovering Tacit Knowledge for Automated Paper Reproduction. Mar. 2026. arXiv: 2603.01801 [cs.AI]. url: https://arxiv.org/abs/2603.01801.
Linhao Zhang et al. PaperRepro: Automated Computational Reproducibility Assessment for Social Science Papers. Feb. 2026. arXiv: 2603.00058 [cs.CY]. url: https://arxiv.org/abs/2603.00058.
Raffael Theiler et al. From paper to benchmark: agentic, framework-based reproduction of under-specified methods in machine health intelligence. May 2026. arXiv: 2605.28371 [cs.AI]. url: https://arxiv.org/abs/2605.28371.
Adrien Bibal et al. AI Copilots for Reproducibility in Science: A Case Study. June 2025. arXiv: 2506.20130 [cs.AI]. url: https://arxiv.org/abs/2506.20130.
Shubham Gandhi et al. ResearchCodeAgent: An LLM Multi-Agent System for Automated Codification of Research Methodologies. Apr. 2025. arXiv: 2504.20117 [cs.SE]. url: https://arxiv.org/abs/2504.20117.
Zhicheng Fang et al. Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking. Feb. 2026. arXiv: 2602.24009 [cs.CR]. url: https://arxiv.org/abs/2602.24009.
opencode: The open source AI coding agent. url: https://opencode.ai (visited on 06/08/2026).
Claude Code. url: https://code.claude.com (visited on 06/08/2026).
Lorraine Saju et al. Computational Reproducibility of R Code Supplements on OSF. US: ICWSM, June 2025. doi: 10.36190/2025.49. url: https://doi.org/10.36190/2025.49.
Akashdeep Saha et al. LockForge: Automating Paper-to-Code for Logic Locking with Multi-Agent Reasoning LLMs. Nov. 2025. arXiv: 2511.18531 [cs.CR]. url: https://arxiv.org/abs/2511.18531.
Bang Nguyen et al. ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences. Feb. 2026. arXiv: 2602.11354 [cs.AI]. url: https://arxiv.org/abs/2602.11354.
Lukas Twist et al. A Study of LLMs’ Preferences for Libraries and Programming Languages. Mar. 2026. arXiv: 2503.17181 [cs.SE]. url: https://arxiv.org/abs/2503.17181.
Doehyun Baek and Michael Pradel. Artisan: Agentic Artifact Evaluation. Feb. 2026. arXiv: 2602.10046 [cs.SE]. url: https://arxiv.org/abs/2602.10046.
David Heye et al. “Supporting Artifact Evaluation with LLMs: A Study with Published Security Research Papers”. In: 2025 IEEE International Conference on Big Data (BigData). IEEE, Dec. 2025, pp. 5077–5085. doi: 10.1109/bigdata66926.2025.11401815. url: http://dx.doi.org/10.1109/BigData66926.2025.11401815.
Pengcheng Yin et al. Natural Language to Code Generation in Interactive Data Science Notebooks. Dec. 2022. arXiv: 2212.09248 [cs.CL]. url: https://arxiv.org/abs/2212.09248.
ICML 2024. url: https://icml.cc/Conferences/2024 (visited on 06/09/2026).
Christine Ye et al. ReplicationBench: Can AI Agents Replicate Astrophysics Research Papers? Oct. 2025. arXiv: 2510.24591 [cs.CL]. url: https://arxiv.org/abs/2510.24591.
Ben Bogin et al. SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories. Sept. 2024. arXiv: 2409.07440 [cs.AI]. url: https://arxiv.org/abs/2409.07440.
Zachary S. Siegel et al. CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark. Sept. 2024. arXiv: 2409.11363 [cs.CL]. url: https://arxiv.org/abs/2409.11363.
CORE-Bench github. url: https://github.com/siegelz/core-bench/blob/main/benchmark/dataset/core_train.json (visited on 06/09/2026).
Aleksandra Eliseeva et al. EnvBench: A Benchmark for Automated Environment Setup. Mar. 2025. arXiv: 2503.14443 [cs.LG]. url: https://arxiv.org/abs/2503.14443.
Xian Gao et al. MMReview: A Multidisciplinary and Multimodal Benchmark for LLM-Based Peer Review Automation. Aug. 2025. arXiv: 2508.14146 [cs.CL]. url: https://arxiv.org/abs/2508.14146.
Ryan Liu et al. “Large Language Models Assume People are More Rational than We Really are”. In: The Thirteenth International Conference on Learning Representations. Jan. 2025. url: https://openreview.net/forum?id=dAeET8gxqg.
Meysam Alizadeh et al. “Evaluating AI Coding Agents in Social Science Reproducibility”. In: (June 2026). url: https://malizad.github.io/Alizadeh_et_al_Agents_Reproducibility.pdf.
Agentic Science
AI systems that can conduct scientific tasks autonomously. - 80 papers
Pre-publication reproducibility
Techniques that authors can use to make their research more reproducible. - 94 papers
Hypothesis Generation
Systems that generate scientific hypotheses. - 18 papers
Automatic Software Engineering
Systems that can do software engineering tasks autonomously. - 34 papers
Science-specialized LLM
LLMs that are fine-tuned to be better at science. - 10 papers
Multi-Agent
AI systems which consist of multiple collaborating agents. - 56 papers