xPeerd Benchmark Study 2026: AI Peer Review Simulation vs Human Review

Open benchmark study · August 2026

An operational and human-reference benchmark of the xPeer engine for simulating scholarly peer review.

The study combines operational behavior, a public same-manuscript human-reference resource, explicit cohort accounting, concern-level diagnostics, recommendation correspondence and reproducibility controls. The evaluated engine is xPeer; its web front end is xPeerd.com.

352
valid operational simulation reports

1,108
public manuscript-level records

271
strict paired manuscripts

15,563
extracted concern units

Executive benchmark summary

What the benchmark establishes

Observable xPeer profile

Longer reports, more extracted concerns, more explicit manuscript targets, broader category representation and more explicit revision actions.

Observable human profile

More explicit rationale cues, stronger lexical manuscript attestation, slightly higher taxonomy-based scientific relevance and lower mean redundancy.

Inference boundary.
The study quantifies textual structure, analytical availability, detector-recognized concern patterns, source correspondence and reproducibility. Scientific correctness, severity, novelty and editorial utility remain expert-adjudication questions.

Study architecture

Two evidence layers, one benchmark

Operational component
352 / 500

Stable-task simulation reports used to characterize disciplinary breadth, task-conditioned workload, simulated decisions and page anchoring.

Human-reference component
271 paired manuscripts

Two human reports and two usable xPeer reviewer reports on the same version-1 manuscript.

Workflow-level review withholding. Human-review text, recommendations, decisions and reference reviewer metadata remained outside the generation input. xPeer outputs were persisted before source joining.

Operational benchmark

352 valid simulations across five disciplinary groups

The operational component retained 70.4% of the original 500 records. Physical Sciences and Health Sciences were the largest subject groups. Every retained report exceeded the declared 0.20 classification-confidence threshold.

Disciplinary coverage
Exact counts among the 352 stable-task reports.
Health Sciences
113
Physical Sciences
109
Humanities
70
Social Sciences
46
Life Sciences
14
Multidisciplinary
0

Simulated decisions

Revision formed more than half of simulated outcomes in every represented disciplinary group. Rejection was approximately 42% in Life Sciences and 45% in Health Sciences. Acceptance was rare.

Task-conditioned issue load

Conventional critique concentrated around 4–10 issues; data-analysis around 6–12; double-blind simulation had the widest spread with a median near 10; repeated review concentrated around 1–3 terminal issues.

Approximate mean page-anchor fraction by field
Overall mean issue-level page-anchor fraction: 0.29.
Physical Sciences
0.34
Social Sciences
0.30
Life Sciences
0.29
Health Sciences
0.29
Humanities
0.20
Report length had a weak positive association with anchor fraction: Spearman ρ = 0.13, p = 0.014.

Human-reference resource

Explicit cohort accounting from source corpus to strict paired analysis

The resource was built from Re3-Sci2.0 F1000Research records restricted to manuscript version 1 with at least two linked human reports.

Cohort funnel

1,146
eligible version-1 manuscripts
2,661 linked human reports
1,108
persisted HTTP-success records
802
exactly two human reports
271
strict paired manuscripts

Incomplete xPeer reviewer-field states

Observed packaged state Records
All four parsed fields empty 125
Text only in Recommendation 294
Editorial summary + Recommendation only 2
Reviewer 1 only 106
Reviewer 2 only 4
Total incomplete exact-two-human records 531

Complete-pair availability: 33.8%. The strict cohort is 271 of the 802 exact-two-human records. The paired analysis used released reviewer fields exactly as stored, with zero imputation, reconstruction or reassignment.

Paired human-reference analysis

xPeer reports were longer and contained more detector-recognized concerns

The strict cohort contains 271 manuscripts, 542 human reports and 542 xPeer reports. Concern extraction produced 15,563 units across 1,023 of 1,084 reports, or 94.4% report-unit coverage.

Report scale
Human shown in grey; xPeer shown in black.
Median report words
Human

763

xPeer

1,889

Median concern count
Human

13

xPeer

41

Mean concerns per 1,000 words
Human

18.981

xPeer

21.597

TRACE-R source profile
Mean manuscript-level observables; scale 0–1.
Targeting
Human

0.279

xPeer

0.414

Explicit reasoning
Human

0.081

xPeer

0.018

Attested alignment
Human

0.119

xPeer

0.096

Category coverage
Human

0.324

xPeer

0.546

Executability
Human

0.606

xPeer

0.669

Scientific relevance
Human

0.932

xPeer

0.886

Observable Human xPeer Difference 95% CI Rank-biserial
Median report words 763 1,889 940.5* 832.9 to 1,040.3 0.836
Median concern count 13 41 25.4* 22.9 to 27.9 0.904
Mean concerns / 1,000 words 18.981 21.597 2.616 1.395 to 3.921 0.262
Mean targeting 0.279 0.414 0.135 0.105 to 0.164 0.557
Mean explicit reasoning 0.081 0.018 −0.063 −0.076 to −0.051 −0.720
Mean attested alignment 0.119 0.096 −0.023 −0.030 to −0.017 −0.525
Mean category coverage 0.324 0.546 0.222 0.195 to 0.250 0.829
Mean executability 0.606 0.669 0.063 0.031 to 0.098 0.270
Mean relevance 0.932 0.886 −0.046 −0.074 to −0.015 −0.319
Mean redundancy 0.037 0.063 0.026 0.005 to 0.046 0.340

*For report words and concern count, source values are medians while the reported difference is the mean paired difference.

Scientific-category prevalence

Broader detector-recognized coverage across the prespecified taxonomy

xPeer showed higher prevalence in statistics, study design, methods and reproducibility, data and results, interpretation and claims, literature context, ethics and reporting, and presentation and clarity. All reported differences remained significant after false-discovery-rate correction.

Methods & reproducibility

Human · 58.3%
xPeer · 95.6%

Data & results

Human · 72.0%
xPeer · 95.2%

Higher prevalence represents broader detector-recognized coverage, not scientific correctness. Report length, templated structure, explicit headings and action-oriented wording may contribute to the observed differences.

Source correspondence

Low lexical concern matching and low recommendation agreement

Human and xPeer concern units were assigned one-to-one and accepted above a prespecified lexical-similarity threshold. Median matched fraction was zero in both source-normalized views.

0.026
mean human recovery
0.009
mean xPeer alignment
56
manuscripts with ≥1 accepted pair

Recommendation correspondence

Human recommendation metadata were available for all 542 human reports. Normalized recommendation language was extracted from 380 of 542 xPeer reports, giving 70.1% report-level coverage. At manuscript level, 240 cases had usable source consensus values.

Rounded exact agreement 43.75%
Spearman association 0.170
Lin concordance 0.164
Quadratic weighted kappa 0.137
Mean ordinal error 0.465
Human \ xPeer Reject Revise Approve
Reject 2 15 1
Revise 6 56 20
Approve 9 84 47

The supported inference is source non-equivalence. Editorial decision use remains under accountable human authority; scientific value requires expert adjudication of individual concerns.

System process

Design-level process mapped to measured outputs

Phase 1

Input & deconstruction

Claims, evidence, pages
Phase 2

Argument framing

Structured analytical target
Phase 3

Manuscript evaluation

Quality and integrity constructs
Phase 4

Decision synthesis

Aggregate findings
Phase 5

Final outcome

Reports, summary, recommendation

The benchmark separates design-level abstractions from directly observed implementation variables.

Quality control & reproducibility

22 of 22 prespecified computational checks passed

The checks covered cohort counts, report balance, nonempty report text, concern-unit coverage, metric bounds, paired-row counts, sensitivity thresholds, category reconciliation, recommendation auditing and output existence.

22 / 22
computational checks passed
2,000
bootstrap replicates
1,999
permutation replicates
Reproducibility parameter Value
Version-pinned archive xpeerd_benchmark_study_2026_v1.0.0.zip
Archive SHA-256 0ab7cb88d2b2db687b586ad303e017a9db0a8104e10f1b8e18c30b8f6a75129c
Random seed 20260723
Manuscript chunks 160 words with 40-word overlap
Minimum concern-unit length 5 words
Primary lexical matching threshold 0.35
Frozen repository commit 99e602873ddb1f7dca8a08d8aa05979e1fce643e

Strategic implications

Use xPeer as an additional scrutiny layer, not an autonomous editorial authority

Researchers & authors

Use xPeer as a pre-submission stress test for methodological detail, reporting omissions, unsupported interpretation, statistical issues and presentation barriers.

Editors & publishers

Use xPeer as a standardized methodological and reporting sweep before or alongside human review while retaining accountable human publication authority.

Benchmark designers

Use the released resource as a common test bed for aligned future evaluation with expert adjudication, cost, latency and governance measures.

Product-development priority identified by the study: high-impact concerns should explicitly connect the observed issue, its methodological or evidential consequence and the requested revision.

Limitations

Twelve principal inference boundaries

  1. The strict paired cohort contains 271 of 1,108 released records and 271 of 802 exact-two-human records, creating complete-case selection risk.
  2. The 531 incomplete exact-two-human records include blank, recommendation-only and partial-reviewer states; raw transport responses are absent.
  3. The manuscripts and human reports derive from F1000Research-linked data; transfer to anonymous pre-publication review and other venue types requires external validation.
  4. Human reports are manuscript-linked references, not ground truth for scientific correctness.
  5. Concern extraction uses deterministic lexical and structural rules that may interact with source style.
  6. No blinded source-stratified expert-annotation subset was available for precision, recall, category accuracy and inter-annotator agreement.
  7. Sixty-one reports yielded zero extracted concern units: 22 human and 39 xPeer.
  8. Attested alignment quantifies lexical attestation, not factual correctness, citation validity or domain-grounded reasoning.
  9. Cross-source concern correspondence is threshold-sensitive and requires expert adjudication for utility claims.
  10. System recommendation labels were observable in 70.1% of xPeer reports.
  11. Workflow-level review withholding does not measure prior model exposure to public manuscripts or review text.
  12. KNOWDYN produces the system, and the author declares a controlling interest; public data, code, exclusions, hashes, independent replication and external adjudication form the conflict-management framework.

Methods

Key analytical definitions

Formal review object and strict cohort

A manuscript is represented as M = ⟨C, E, P⟩, with claims, evidential units and available location indices. The intended output is δ(M,τ) = ⟨R₁, R₂, Sed, L⟩.

Iᵢ = 1(nᵢH = 2) · 1(nᵢX = 2) · 1(mᵢ = 1), with Σ Iᵢ = 271.

Concern-unit extraction

A concern unit was a sentence of at least five words matching explicit concern language, an explicit revision action, a question form, or concern/recommendation section context without praise-only language.

Dᵢₛᵣ = 1000 · Kᵢₛᵣ / max(1, W(Rᵢₛᵣ)).

TRACE-R observables

TRACE-R comprises Targeting, explicit Reasoning, Attested alignment, category Coverage, Executability and scientific Relevance. Each dimension is reported separately as an observable text property.

Redundancy and cross-source lexical matching

Concern units were compared by TF–IDF cosine similarity using unigram and bigram features and one-to-one assignment. The default threshold was 0.35, with sensitivity analysis from 0.25 to 0.50.

MᵢH(η) = mᵢ(η)/max(1,KᵢH);   MᵢX(η) = mᵢ(η)/max(1,KᵢX).

Recommendation normalization and agreement

Human recommendations were mapped from source metadata. xPeer recommendations were mapped from explicit recommendation, decision or verdict language. Ordinal encoding used reject = 0, revise/reservations = 1 and approve = 2. Association and absolute agreement were reported separately.

Data & code availability

Public benchmark resources

Version-pinned archive

Zenodo DOI 10.5281/zenodo.21479700

Evaluation repository

Frozen GitHub commit

Copyright. © 2026 Khalid M. Saqr and KNOWDYN LTD, as applicable. All rights reserved except where a separate license is expressly stated.

Disclaimer. This publication is provided for research and informational purposes only. It does not constitute legal, editorial, scientific or professional advice; users remain responsible for independent evaluation and decisions.

Intellectual property. xPeer and xPeerd (.com, .online), including their software, methods, prompts, workflows, interfaces and confidential technical information, are proprietary to KNOWDYN LTD.