Referenced Papers — Download Index
All 35 references from Lilian Weng, Harness Engineering for Self-Improvement (2026-07-04).
Numbering matches the post's References section.
- ✅ downloaded · ⚠️ metadata stub only (source blocked / no free full text)
| # | Status | File | Title | Source |
|---|---|---|---|---|
| 1 | ⚠️ | ref01_good1965-ultraintelligent-machine.md |
Good, I. J. "Speculations Concerning the First Ultraintelligent Machine." Advances in Comp | link |
| 2 | ✅ | ref02_yudkowsky2008-recursive-self-improvement.html (1.0 MB) |
Yudkowsky, Eliezer. "Recursive Self-Improvement." LessWrong, 2008. | link |
| 3 | ⚠️ | ref03_anchored-self-play-code-repair.md |
Choi, et al. "Anchored Self-Play for Code Repair." ICML 2026. | link |
| 4 | ✅ | ref04_absolute-zero.pdf (5.1 MB) |
Zhao, et al. "Absolute Zero: Reinforced Self-play Reasoning with Zero Data." 2025. | link |
| 5 | ✅ | ref05_self-rewarding-language-models.pdf (1.1 MB) |
Yuan, et al. "Self-Rewarding Language Models." 2024. | link |
| 6 | ✅ | ref06_spin-self-play-finetuning.pdf (1.4 MB) |
Chen, et al. "Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Model | link |
| 7 | ✅ | ref07_ace-agentic-context-engineering.pdf (2.2 MB) |
Zhang, et al. "Agentic Context Engineering: Evolving Contexts for Self-Improving Language | link |
| 8 | ✅ | ref08_mce-meta-context-engineering.pdf (7.0 MB) |
Ye, et al. "Meta Context Engineering via Agentic Skill Evolution." 2026. | link |
| 9 | ✅ | ref09_meta-harness.pdf (1.0 MB) |
Lee, et al. "Meta-Harness: End-to-End Optimization of Model Harnesses." 2026. | link |
| 10 | ✅ | ref10_lu2026-e2e-automation-ai-research.html (384 KB) |
Lu, et al. "Towards end-to-end automation of AI research." Nature, 651:914-919, 2026. | link |
| 11 | ✅ | ref11_scientistone.pdf (4.9 MB) |
Meng, et al. "ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence. | link |
| 12 | ✅ | ref12_autodata.pdf (7.2 MB) |
Kulikov, et al. "Autodata: An agentic data scientist to create high quality synthetic data | link |
| 13 | ✅ | ref13_adas-automated-design-agentic-systems.pdf (783 KB) |
Hu, Lu, and Clune. "Automated Design of Agentic Systems." ICLR 2025. | link |
| 14 | ✅ | ref14_self-refine.pdf (1.9 MB) |
Madaan, et al. "Self-Refine: Iterative Refinement with Self-Feedback." NeurIPS 2023. | link |
| 15 | ✅ | ref15_aflow.pdf (1.3 MB) |
Zhang, et al. "AFlow: Automating Agentic Workflow Generation." ICLR 2025. | link |
| 16 | ✅ | ref16_stop-self-taught-optimizer.pdf (944 KB) |
Zelikman, et al. "Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation | link |
| 17 | ✅ | ref17_self-harness.pdf (4.1 MB) |
Zhang, et al. "Self-Harness: Harnesses That Improve Themselves." 2026. | link |
| 18 | ✅ | ref18_promptbreeder.pdf (800 KB) |
Fernando, et al. "Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution." | link |
| 19 | ✅ | ref19_gepa.pdf (2.8 MB) |
Agrawal, et al. "GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning." | link |
| 20 | ✅ | ref20_alphaevolve.pdf (3.4 MB) |
Novikov, et al. "AlphaEvolve: A coding agent for scientific and algorithmic discovery." 20 | link |
| 21 | ✅ | ref21_shinkaevolve.pdf (3.8 MB) |
Lange, Imajuku, and Cetin. "ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program | link |
| 22 | ✅ | ref22_thetaevolve.pdf (3.8 MB) |
Wang, et al. "ThetaEvolve: Test-time Learning on Open Problems." 2025. | link |
| 23 | ✅ | ref23_darwin-godel-machine.pdf (3.6 MB) |
Zhang, et al. "Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents." 2025. | link |
| 24 | ✅ | ref24_hyperagents.pdf (8.6 MB) |
Zhang, et al. "Hyperagents." 2026. | link |
| 25 | ✅ | ref25_learning-to-discover-at-test-time.pdf (1.2 MB) |
Yuksekgonul, et al. "Learning to Discover at Test Time." 2026. | link |
| 26 | ✅ | ref26_epistemic-uncertainty-test-time-discovery.pdf (813 KB) |
Riaz, et al. "Epistemic Uncertainty for Test-Time Discovery." 2026. | link |
| 27 | ✅ | ref27_sia-self-improving-ai.pdf (970 KB) |
Hebbar, et al. "SIA: Self Improving AI with Harness & Weight Updates." 2026. | link |
| 28 | ✅ | ref28_why-llms-arent-scientists-yet.pdf (969 KB) |
Trehan and Chopra. "Why LLMs Aren't Scientists Yet: Lessons from Four Autonomous Research | link |
| 29 | ✅ | ref29_early-science-acceleration-gpt5.pdf (2.7 MB) |
Bubeck, et al. "Early science acceleration experiments with GPT-5." 2025. | link |
| 30 | ✅ | ref30_paperbench.pdf (1.6 MB) |
Starace, et al. "PaperBench: Evaluating AI's Ability to Replicate AI Research." ICML 2025. | link |
| 31 | ✅ | ref31_re-bench.pdf (14.5 MB) |
Wijk, et al. "RE-Bench: Evaluating frontier AI R&D capabilities of language model agents a | link |
| 32 | ✅ | ref32_mle-bench.pdf (799 KB) |
Chan, et al. "MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineerin | link |
| 33 | ✅ | ref33_scienceagentbench.pdf (1.9 MB) |
Chen, et al. "ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Dr | link |
| 34 | ✅ | ref34_core-bench.pdf (2.8 MB) |
Siegel, et al. "CORE-Bench: Fostering the Credibility of Published Research Through a Comp | link |
| 35 | ✅ | ref35_kernelbench.pdf (2.9 MB) |
Ouyang, et al. "KernelBench: Can LLMs Write Efficient GPU Kernels?" 2025. | link |
Totals: 33 PDFs/pages downloaded, 2 metadata stubs, out of 35 references.
Notes
- ref01 (Good 1965): a 1965 book chapter (Advances in Computers); no free full-text PDF exists. See
ref01_*.md. - ref03 (Choi et al., OpenReview): PDF blocked by OpenReview's anti-bot challenge (HTTP 403). Abstract + direct URL captured in
ref03_*.mdfor manual download. - ref02 (Yudkowsky): saved as HTML (LessWrong post).
- ref10 (Lu et al., Nature): saved as HTML landing page (full text is paywalled by Nature).
- All other 31 references are arXiv PDFs, verified to start with the
%PDFmagic header.