Open benchmark study · August 2026
An operational and human-reference benchmark of the xPeer engine for simulating scholarly peer review.
The study combines operational behavior, a public same-manuscript human-reference resource, explicit cohort accounting, concern-level diagnostics, recommendation correspondence and reproducibility controls. The evaluated engine is xPeer; its web front end is xPeerd.com.
Executive benchmark summary
What the benchmark establishes
Observable xPeer profile
Longer reports, more extracted concerns, more explicit manuscript targets, broader category representation and more explicit revision actions.
Observable human profile
More explicit rationale cues, stronger lexical manuscript attestation, slightly higher taxonomy-based scientific relevance and lower mean redundancy.
The study quantifies textual structure, analytical availability, detector-recognized concern patterns, source correspondence and reproducibility. Scientific correctness, severity, novelty and editorial utility remain expert-adjudication questions.
Study architecture
Two evidence layers, one benchmark
Stable-task simulation reports used to characterize disciplinary breadth, task-conditioned workload, simulated decisions and page anchoring.
Two human reports and two usable xPeer reviewer reports on the same version-1 manuscript.
Operational benchmark
352 valid simulations across five disciplinary groups
The operational component retained 70.4% of the original 500 records. Physical Sciences and Health Sciences were the largest subject groups. Every retained report exceeded the declared 0.20 classification-confidence threshold.
Exact counts among the 352 stable-task reports.
Simulated decisions
Revision formed more than half of simulated outcomes in every represented disciplinary group. Rejection was approximately 42% in Life Sciences and 45% in Health Sciences. Acceptance was rare.
Task-conditioned issue load
Conventional critique concentrated around 4–10 issues; data-analysis around 6–12; double-blind simulation had the widest spread with a median near 10; repeated review concentrated around 1–3 terminal issues.
Overall mean issue-level page-anchor fraction: 0.29.
Human-reference resource
Explicit cohort accounting from source corpus to strict paired analysis
The resource was built from Re3-Sci2.0 F1000Research records restricted to manuscript version 1 with at least two linked human reports.
Cohort funnel
eligible version-1 manuscripts
2,661 linked human reports
persisted HTTP-success records
exactly two human reports
strict paired manuscripts
Incomplete xPeer reviewer-field states
| Observed packaged state | Records |
|---|---|
| All four parsed fields empty | 125 |
| Text only in Recommendation | 294 |
| Editorial summary + Recommendation only | 2 |
| Reviewer 1 only | 106 |
| Reviewer 2 only | 4 |
| Total incomplete exact-two-human records | 531 |
Paired human-reference analysis
xPeer reports were longer and contained more detector-recognized concerns
The strict cohort contains 271 manuscripts, 542 human reports and 542 xPeer reports. Concern extraction produced 15,563 units across 1,023 of 1,084 reports, or 94.4% report-unit coverage.
Human shown in grey; xPeer shown in black.
763
1,889
13
41
18.981
21.597
Mean manuscript-level observables; scale 0–1.
0.279
0.414
0.081
0.018
0.119
0.096
0.324
0.546
0.606
0.669
0.932
0.886
| Observable | Human | xPeer | Difference | 95% CI | Rank-biserial |
|---|---|---|---|---|---|
| Median report words | 763 | 1,889 | 940.5* | 832.9 to 1,040.3 | 0.836 |
| Median concern count | 13 | 41 | 25.4* | 22.9 to 27.9 | 0.904 |
| Mean concerns / 1,000 words | 18.981 | 21.597 | 2.616 | 1.395 to 3.921 | 0.262 |
| Mean targeting | 0.279 | 0.414 | 0.135 | 0.105 to 0.164 | 0.557 |
| Mean explicit reasoning | 0.081 | 0.018 | −0.063 | −0.076 to −0.051 | −0.720 |
| Mean attested alignment | 0.119 | 0.096 | −0.023 | −0.030 to −0.017 | −0.525 |
| Mean category coverage | 0.324 | 0.546 | 0.222 | 0.195 to 0.250 | 0.829 |
| Mean executability | 0.606 | 0.669 | 0.063 | 0.031 to 0.098 | 0.270 |
| Mean relevance | 0.932 | 0.886 | −0.046 | −0.074 to −0.015 | −0.319 |
| Mean redundancy | 0.037 | 0.063 | 0.026 | 0.005 to 0.046 | 0.340 |
*For report words and concern count, source values are medians while the reported difference is the mean paired difference.
Scientific-category prevalence
Broader detector-recognized coverage across the prespecified taxonomy
xPeer showed higher prevalence in statistics, study design, methods and reproducibility, data and results, interpretation and claims, literature context, ethics and reporting, and presentation and clarity. All reported differences remained significant after false-discovery-rate correction.
Methods & reproducibility
Data & results
Source correspondence
Low lexical concern matching and low recommendation agreement
Human and xPeer concern units were assigned one-to-one and accepted above a prespecified lexical-similarity threshold. Median matched fraction was zero in both source-normalized views.
Recommendation correspondence
Human recommendation metadata were available for all 542 human reports. Normalized recommendation language was extracted from 380 of 542 xPeer reports, giving 70.1% report-level coverage. At manuscript level, 240 cases had usable source consensus values.
| Rounded exact agreement | 43.75% |
| Spearman association | 0.170 |
| Lin concordance | 0.164 |
| Quadratic weighted kappa | 0.137 |
| Mean ordinal error | 0.465 |
| Human \ xPeer | Reject | Revise | Approve |
|---|---|---|---|
| Reject | 2 | 15 | 1 |
| Revise | 6 | 56 | 20 |
| Approve | 9 | 84 | 47 |
System process
Design-level process mapped to measured outputs
Input & deconstruction
Argument framing
Manuscript evaluation
Decision synthesis
Final outcome
The benchmark separates design-level abstractions from directly observed implementation variables.
Quality control & reproducibility
22 of 22 prespecified computational checks passed
The checks covered cohort counts, report balance, nonempty report text, concern-unit coverage, metric bounds, paired-row counts, sensitivity thresholds, category reconciliation, recommendation auditing and output existence.
| Reproducibility parameter | Value |
|---|---|
| Version-pinned archive | xpeerd_benchmark_study_2026_v1.0.0.zip |
| Archive SHA-256 | 0ab7cb88d2b2db687b586ad303e017a9db0a8104e10f1b8e18c30b8f6a75129c |
| Random seed | 20260723 |
| Manuscript chunks | 160 words with 40-word overlap |
| Minimum concern-unit length | 5 words |
| Primary lexical matching threshold | 0.35 |
| Frozen repository commit | 99e602873ddb1f7dca8a08d8aa05979e1fce643e |
Strategic implications
Use xPeer as an additional scrutiny layer, not an autonomous editorial authority
Researchers & authors
Use xPeer as a pre-submission stress test for methodological detail, reporting omissions, unsupported interpretation, statistical issues and presentation barriers.
Editors & publishers
Use xPeer as a standardized methodological and reporting sweep before or alongside human review while retaining accountable human publication authority.
Benchmark designers
Use the released resource as a common test bed for aligned future evaluation with expert adjudication, cost, latency and governance measures.
Limitations
Twelve principal inference boundaries
- The strict paired cohort contains 271 of 1,108 released records and 271 of 802 exact-two-human records, creating complete-case selection risk.
- The 531 incomplete exact-two-human records include blank, recommendation-only and partial-reviewer states; raw transport responses are absent.
- The manuscripts and human reports derive from F1000Research-linked data; transfer to anonymous pre-publication review and other venue types requires external validation.
- Human reports are manuscript-linked references, not ground truth for scientific correctness.
- Concern extraction uses deterministic lexical and structural rules that may interact with source style.
- No blinded source-stratified expert-annotation subset was available for precision, recall, category accuracy and inter-annotator agreement.
- Sixty-one reports yielded zero extracted concern units: 22 human and 39 xPeer.
- Attested alignment quantifies lexical attestation, not factual correctness, citation validity or domain-grounded reasoning.
- Cross-source concern correspondence is threshold-sensitive and requires expert adjudication for utility claims.
- System recommendation labels were observable in 70.1% of xPeer reports.
- Workflow-level review withholding does not measure prior model exposure to public manuscripts or review text.
- KNOWDYN produces the system, and the author declares a controlling interest; public data, code, exclusions, hashes, independent replication and external adjudication form the conflict-management framework.
Methods
Key analytical definitions
Formal review object and strict cohort
A manuscript is represented as M = ⟨C, E, P⟩, with claims, evidential units and available location indices. The intended output is δ(M,τ) = ⟨R₁, R₂, Sed, L⟩.
Concern-unit extraction
A concern unit was a sentence of at least five words matching explicit concern language, an explicit revision action, a question form, or concern/recommendation section context without praise-only language.
TRACE-R observables
TRACE-R comprises Targeting, explicit Reasoning, Attested alignment, category Coverage, Executability and scientific Relevance. Each dimension is reported separately as an observable text property.
Redundancy and cross-source lexical matching
Concern units were compared by TF–IDF cosine similarity using unigram and bigram features and one-to-one assignment. The default threshold was 0.35, with sensitivity analysis from 0.25 to 0.50.
Recommendation normalization and agreement
Human recommendations were mapped from source metadata. xPeer recommendations were mapped from explicit recommendation, decision or verdict language. Ordinal encoding used reject = 0, revise/reservations = 1 and approve = 2. Association and absolute agreement were reported separately.
Data & code availability
Public benchmark resources
Study-level dataset
Version-pinned archive
Evaluation repository
Copyright. © 2026 Khalid M. Saqr and KNOWDYN LTD, as applicable. All rights reserved except where a separate license is expressly stated.
Disclaimer. This publication is provided for research and informational purposes only. It does not constitute legal, editorial, scientific or professional advice; users remain responsible for independent evaluation and decisions.
Intellectual property. xPeer and xPeerd (.com, .online), including their software, methods, prompts, workflows, interfaces and confidential technical information, are proprietary to KNOWDYN LTD.