Protected collaborator review · hard-200 upgrade v1

200 measured-hard seeds, expanded to long-horizon research tasks

Each of the 200 high-value seed requests (0–12.5% measured pass rate) was processed through the same pipeline that produced the 179 corpus: construct preserved, request rewritten as a natural decision-led long-horizon task, clause partition, prompt-bound rubric atoms, judge rubric, and live public-web feasibility probes. Seed pass rates describe the seeds; every final task requires re-measurement.

378 tasks (both corpora)3515 clauses3924 rubric atoms135 feasible243 explicit shortfalls100 pass the quality gate0 human acceptedneeds re-measurement
dimensionseed meanfinal mean
naturalness2.943.63
quality2.713.72
depth3.164.72
breadth2.964.52
verification_level3.834.65
1 · Blind prompt review Final tasks only — no seeds, rubrics, or scores. Judge naturalness, scope, and completion boundaries unbiased.Start here → 2 · Full seed-vs-final review Side-by-side seed and final, measured seed stats, scorecards, clauses, bound rubric atoms, source packets, and feasibility dispositions. Open full corpus → 3 · Annotate tasks Review every prompt and rubric, edit proposed language, record CMU or Browser-Use decisions side by side, and export a mergeable annotation file. Open annotation workspace →
Status boundary: internal QA artifact for human review — not a scored benchmark release, and no task here has human acceptance yet. Source seed file remains unmodified.