2026-08-19 · PRDT Deepresearch Series №5 · generated from the P-Research corpus pipeline Method: corpus retrieval + 90-day targeted arXiv sweep + GraphRAG communities + top-venue odds. All numbers are re-derivable from data/papers.jsonl; abstract-level evidence is a research signal, not a verified fact.
| Metric | Value |
|---|---|
| Eval/Benchmark-related papers | 363 (63% of the 574-paper corpus) |
| Papers containing “benchmark” | 154 |
| Papers containing “evaluation” | 209 |
| By year | 2023: 73 → 2024: 68 → 2025: 84 → 2026 (first half): 138 |
Observation 0 — Eval is the largest theme in the whole corpus. More than six in ten papers touch evaluation in some form. Eval is not one of the five elements; it is the common substrate under the other four — and it is still accelerating. (abstract-level signal)
This report reads the corpus through five observations, cross-checked against the open problems of the four earlier reports in this series. Where benchmarks are saturated, where they are missing, and where the frontier is rebuilding evaluation as infrastructure are the three axes of what follows.
Corpus benchmark leaders: MATH (29) · GSM8K (8) · DROP (6) · MMLU (3) — math and commonsense benchmarks dominate, while agent benchmarks are nearly absent (AgentBench appears only once). Combined with the survey’s §6.2 finding (hard benchmarks hide in full texts), the benchmark ecosystem carries a double bias — a category bias plus a visibility bias. (abstract-level signal)
The open problems of Reports #1–#4 line up: RAG lacks a retrieve-or-abstain benchmark, Memory lacks a neural-memory-module benchmark, Tool Use lacks a financial-tool-audit benchmark, Planning lacks a financial-planning benchmark. Each report closed with “benchmark missing” — not a coincidence but a structural pattern: capability runs ahead of evaluation, and evaluation runs ahead of infrastructure.
→ Benchmarks are evolving from “static test sets” into “executable, dynamic environments” — precisely the form self-evolution infrastructure needs, because evolution requires an unlimited supply of tasks.
→ Eval moves from “final exam” to “health check”: diagnose, localize, repair.
The recursion loop is: act → eval → diagnose → update. Every self-evolving system on the frontier — SPADE, EvoTS-Agent (2608.17933), AutoSR (2608.16876) — embeds some eval mechanism. As the program’s Eval discussion zone puts it: self-evolution without eval is Brownian motion. The implication is direct: eval is not a gate at the end of a run but the feedback channel that makes adaptation directed rather than random — which is why the eval loop, and the benchmarks that feed it, sit at the center of self-evolution infrastructure rather than at its edge.
Eval is not a fifth element of the five-element curriculum; it is the loop that closes the other four. How does RAG know it retrieved correctly? How does Memory know it remembered? How does Tool Use know it used the right tool? How does Planning know it planned well? — all of it routes back to eval. The program’s self-evolution loop (governed adaptation and publication) is eval-driven.
The frontier’s current moves — executable environments, diagnostic eval, eval-loop-driven self-evolution — line up with the P-Research program’s existing practice:
| Frontier claim | The P-Research program’s practice |
|---|---|
| Environment-as-benchmark | Weekly real arXiv sweeps = an unlimited supply of real tasks |
| Eval-as-diagnosis | presearch audit (0 FAIL; has caught 24 mechanical errors) |
| Precision measurement | 99% label match rate (the structured pipeline’s self-evaluation) |
| Hand-computed baselines | dsh-quant’s tradition of hand-computed baselines (174 tests) |
| Benchmark card | benchmark card culture across the program’s pipelines (PRDT / PAT) |
| Admission prediction | top-venue odds (using evaluation features to predict evaluation outcomes) |
The program’s position is that evaluation engineering is not a scoreboard appended at the end of the pipeline; it is the closed loop that makes the pipeline self-evolving — act → eval → diagnose → update. What the frontier is only now formalizing as a research program, the P-Research pipeline already practices as routine: eval as the substrate that lets a corpus pipeline know how well it is doing and where to evolve next.
Eval is the largest theme in the corpus, the common gap across the four earlier reports, and the infrastructure of self-evolution. Evaluation engineering is the fuel of recursion — solid agent-eval infrastructure is what lets self-evolution close its loop — and the program’s audit, top-venue odds, and hand-computed baselines are already the first segment of that road.