p-research

Frontier Deepresearch: Planning

2026-08-19 · PRDT Deepresearch Series №4 · generated from the P-Research corpus pipeline Method: corpus retrieval + 90-day targeted arXiv sweep + GraphRAG communities + top-venue odds. All numbers are re-derivable from data/papers.jsonl; abstract-level evidence is a research signal, not a verified fact. Pipeline note: this cycle’s 90-day targeted arXiv sweep was deferred due to an upstream service interruption; the scheduled weekly sweep will backfill it. Corpus counts below are unaffected.

Key numbers

Metric Value
Planning-related papers (574-paper corpus) 42
Full-text verified 16 (the highest in the series)
Year distribution 2023: 2 → 2024: 5 → 2025: 12 → 2026 (half year): 23

Observation 0 — planning research is exploding. The 2026 half-year count is nearly double the 2025 full-year count, and this subset carries the series’ highest full-text verification rate. By corpus volume, planning is becoming the largest active front of agent research.

Findings

1. Planning × Memory (co-occurrence 10)

Claim: the quality ceiling of planning equals the quality ceiling of memory — a thread that echoes the Memory deepresearch (Series №2).

Evidence (abstract-level signal):

2. Planning × RL (RLVR co-occurrence 10)

Claim: planning is no longer a script — it is a policy that gets trained. The confluence of RLVR and planning is the strongest signal in this sweep.

Evidence (abstract-level signal): ten planning papers carry RLVR / GRPO tags; this is the densest cross-product in the planning subset.

3. Planning × self-evolution

Claim: planning is the execution layer of recursion — self-evolution edits planning, and planning determines action.

Evidence (abstract-level signal):

4. Planning × monitorability and steerability

Claim: “visualize + intervene” is becoming a theme — setting boundaries and goals requires a human in the loop.

Evidence (abstract-level signal):

5. Planning × Quant

Claim: financial-scenario agent planning is the second cross-pillar convergence observed in this series (after the Tool Use deepresearch), landing on the program’s Quant × AI pillar.

Evidence (abstract-level signal):

6. Planning × verification (Verifier/PRM co-occurrence 3)

Claim: planning needs verification — the plan → verify → replan loop is becoming the standard structure.

Evidence (abstract-level signal): BATON’s verifier-based governance; AutoSR’s research-state verification.

Top-venue odds (planning papers)

BATON (0.609) · HarnessEval-W (0.650 band) · AutoSR (0.8 novelty) — “memory-enabled planning” and “eval-enabled planning” carry the top-venue profile.

Implications for the P-Research program

Read against the program’s five-element capability curriculum:

Element Frontier signal
Planning planning explosion + RL-trained planning + human-in-the-loop monitoring
Memory the quality ceiling of planning (co-occurrence 10)
Eval HarnessEval / StartupBench: planning needs benchmark companions
Tool Use the execution surface of planning (connects to the composite tool-call thread of the Tool Use deepresearch, Series №3)
RAG the memory supply for planning

As program positions: planning is the element whose research mass is growing fastest, so the corpus pipeline treats it as a first-class sweep target and planning quality as a measurable quantity rather than a prompt-crafting art; memory remains the ceiling, so raising planning quality means raising memory quality; eval support should cover planning workflows (the plan → verify → replan loop) rather than single turns; tool use is planning’s execution surface; and the corpus pipeline’s retrieval layer is the memory supply that feeds the planner.

Open problems

  1. Evaluation benchmarks for financial agent planning. EvoTS-Agent and StartupBench bracket the space, but the middle is empty: how should “planning for quant tasks” be evaluated?
  2. Can the plan → verify → replan loop become a standard component of harnesses?
  3. The relation between formal boundary setting in planning and safety — what is a verifiable boundary for a goal?

One-line takeaway

In 2023, planning was a prompting trick; in 2026, planning is a core capability trained by RL, supplied by memory, checked by verifiers, and monitored by humans — and financial-scenario planning (the EvoTS-Agent line) is the intersection the program’s Quant × AI pillar occupies.