Peer Review Simulation

The Peer-Review Black Box: Why Science Cannot Audit the System That Audits Science

Science demands evidence for almost everything—except, strangely, for the process that decides what counts as science.

A researcher submitting a manuscript is expected to disclose methods, report limitations, preserve data, explain analytical choices and provide enough information for others to examine the work. The paper may be rejected because an experiment cannot be reproduced, a statistical decision is insufficiently justified or an important claim lacks an evidentiary trail.

These expectations are reasonable. Science cannot function without scrutiny.

But the scrutiny itself is usually delivered through a system that would fail many of the standards it imposes on authors.

Peer-review reports are often confidential. Editorial reasoning is rarely disclosed. Rejected manuscripts may leave no accessible record. The original version evaluated by reviewers may disappear from view once the revised paper is published. Author rebuttals, reviewer disagreements and the grounds for the final decision are frequently unavailable.

We call peer review the quality-control system of science. Yet in most cases, science cannot reconstruct how that quality-control decision was produced.

Peer review demands reproducibility from science while remaining largely irreproducible itself.

The Contradiction at the Heart of Scientific Publishing

The contradiction should trouble anyone who believes scientific authority must be earned rather than assumed.

A manuscript may be accepted after two favourable reviews. Another may be rejected after one hostile report. A third may undergo extensive revision because an editor considered one criticism decisive and another irrelevant.

But who can later inspect whether those judgments were reasonable?

Usually, only the people already inside the process.

The wider research community cannot see whether reviewers examined the correct manuscript version, whether their criticisms were supported by the text, whether an author corrected the identified problems or whether the editor’s decision followed the evidence contained in the reports.

This is not merely a transparency inconvenience. It is an epistemic weakness.

Scientific publishing converts private judgments into public authority. Once a journal accepts a paper, that decision affects hiring, promotion, funding, clinical practice, policy and the future direction of research. Once a paper is rejected, a potentially valuable contribution may be delayed, redirected or abandoned.

Decisions with such consequences should leave a defensible record.

Publishing Review Reports Is Not Enough

Open peer review is often presented as the solution. Publishing reports is certainly better than preserving nothing, but a review report alone is not an audit trail.

To evaluate a review, we need to know exactly what the reviewer saw.

Consider a published review stating that the manuscript failed to explain its sampling method. If readers can access only the final revised paper, they may find a complete sampling description and conclude that the reviewer was careless. In reality, the missing explanation may have been added precisely because the reviewer identified the problem.

The reverse is also possible. A reviewer may criticise something that was already explained clearly in the original submission. Without that original version, the criticism cannot be checked.

A review separated from the reviewed manuscript is evidence without context.

The authors of the Multidisciplinary Open Peer Review Dataset, known as MOPRD, identified this problem directly. They argued that a complete peer-review dataset should contain the initial submission, review comments, meta-reviews, revisions, author rebuttals and editorial decisions.

Most existing resources, they found, do not cover the entire process. Some include review comments but not the original manuscripts on which those comments were based. Others concentrate almost entirely on computer science, limiting the ability to study how peer review operates across disciplines.

MOPRD was constructed to preserve the full review history of 6,578 papers, including manuscript versions, reviews, rebuttals and decisions. Its significance lies not merely in its size, but in the recognition that peer review can be understood only when the entire chain is available.

Science Cannot Correct What It Cannot Observe

We routinely discuss bias, inconsistency and poor reviewing. But how can those failures be measured reliably when the underlying records are missing?

Without complete review histories, researchers cannot answer basic questions with confidence:

  • Did reviewers identify genuine methodological weaknesses?
  • Did authors correct those weaknesses?
  • Were similar manuscripts judged by similar standards?
  • Did editorial decisions reflect the content of the reviews?
  • Were unconventional methods rejected because they were flawed or merely unfamiliar?
  • Were researchers from particular institutions, countries or disciplines treated differently?
  • Did peer review materially improve the published paper?

We often behave as though peer review has already been validated simply because it has existed for generations. Longevity, however, is not evidence of reliability.

A system cannot learn systematically from its mistakes if it does not preserve inspectable records of how those mistakes occurred.

Who Reviews the Reviewers?

Authors are expected to answer reviewers. Reviewers may answer editors. Editors may answer publishers.

But the process usually ends there.

There is rarely an independent mechanism for examining whether a reviewer’s claims were accurate, whether an editor applied consistent standards or whether the same criticism would have been considered decisive if the manuscript had come from a more prestigious institution.

This does not mean that every reviewer should be publicly named. Anonymity can protect reviewers from retaliation, particularly when they criticise senior or powerful researchers. Confidentiality may also be necessary for unpublished commercial, clinical or security-sensitive work.

But anonymity is not the same as unaccountability.

A system can protect identities while preserving an authorised audit trail. It can store immutable manuscript versions, timestamped reports, declared reviewer roles, author responses and documented decision grounds without exposing confidential material to the public.

The false choice between total secrecy and total openness has allowed the status quo to continue. The real requirement is controlled auditability.

The Minimum Record Science Should Demand

A defensible peer-review process should preserve, at minimum:

  1. The exact manuscript version reviewed. A report must remain linked to the document that produced it.
  2. Timestamped reviewer reports. Later changes should not overwrite the original evaluation.
  3. Traceable manuscript references. Criticisms should identify the claim, section, figure, table or method being challenged.
  4. Declared roles and conflicts. The record should distinguish reviewers, editors, statistical assessors and automated tools.
  5. Author rebuttals and revision decisions. It should be possible to see how each major concern was answered.
  6. Editorial decision grounds. A verdict should not exist without a recorded rationale.
  7. Explicit missing data and exclusions. Absent reports and failed processing steps should be disclosed rather than silently ignored.
  8. An appeal or correction path. Authors should be able to challenge demonstrably incorrect or procedurally unfair assessments.

None of these requirements prevents human judgment. They simply require judgment to leave evidence behind.

Artificial Intelligence Makes the Black Box More Dangerous

The arrival of generative AI has made auditability urgent.

Researchers are already using AI to improve language, organise arguments and analyse manuscripts. Reviewers may use it to summarise papers or draft comments. Editors and publishers are exploring automated screening, reviewer selection and manuscript evaluation.

The important question is no longer whether AI will enter peer review. It already has.

The question is whether its use will be visible and reconstructable.

When an automated system contributes to a review, the record should show which system was used, what manuscript version it received, what task governed the analysis, what output it produced and which human accepted, rejected or modified its recommendations.

Without these records, AI will not solve the peer-review black box. It will place another black box inside it.

An undisclosed automated review is not transparency. It is opacity at machine speed.

What Auditable AI-Assisted Review Could Look Like

The open xPeer benchmark offers one example of the evidence infrastructure that automated peer-review research should provide.

The benchmark released 1,108 manuscript-level records and defined a strict cohort of 271 manuscripts containing two human reports and two xPeer reports. It documented exclusions, preserved version-linked inputs, separated human reviews from the generation workflow and exposed concern-level and manuscript-level analytical outputs.

All 837 exclusions were reported. All 22 prespecified computational quality checks passed. The study was inspectable at the level of records, reports, concern units, categories, matching thresholds, statistical tests and figures.

Just as importantly, the benchmark did not pretend that computational consistency proved scientific correctness. It explicitly separated reproducibility from construct validity and reserved the scientific utility of individual concerns for expert adjudication.

That distinction is essential.

An auditable system does not claim to be infallible. It makes its behaviour inspectable enough for others to find out where it succeeds and where it fails.

Peer Review Must Become Reproducible Without Becoming Performative

There is a risk that calls for transparency become exercises in appearances. Journals may publish a few reports, add an “open review” label and declare the problem solved.

But openness without structure can produce more information without producing more accountability.

Thousands of disconnected documents scattered across publisher websites do not create a reproducible system. Neither does a published review that cannot be connected to the manuscript version it assessed.

What science needs is not transparency theatre. It needs evidence architecture.

That means standardised records, persistent identifiers, version control, traceable concerns, documented human responsibility and governance rules that distinguish analytical support from editorial authority.

The Future of Peer Review Is an Audit Trail

Peer review will always involve judgment. No dataset, checklist or artificial intelligence system can eliminate disagreement about novelty, significance or methodological adequacy.

But disagreement is not an excuse for opacity.

Science has spent decades strengthening expectations for research transparency. Data should be available. Methods should be documented. Analytical choices should be disclosed. Results should be reproducible where possible.

It is time to apply the same intellectual discipline to the process that decides which research becomes visible.

We do not need to publish every confidential manuscript or expose every reviewer’s identity. We do need a review system that can demonstrate what was assessed, what concerns were raised, how authors responded and why the final decision was made.

Until then, peer review will continue to demand evidence from everyone except itself.

The system that audits science must finally become auditable.


Sources and Further Reading