๐ก The Core Idea
A self-evolving agentic loop where an LLM iteratively rewrites its output guided solely by external quality scores โ no access to model internals needed.
The oracle only needs to return a score. It does not need to be differentiable, open-source, or even documented internally. This makes the loop applicable to grammar checkers, readability formulas, factuality verifiers, and AI detectors โ all with the same code, zero modification.
โก Why Prompt-Only? The Generalization Argument
Existing gradient-based and token-level methods are fundamentally tethered to a single oracle type. Here is why.
| Oracle Type | Differentiable? | Gradient Methods Work? | Our Loop Works? |
|---|---|---|---|
| ๐ค AI Detection | Only with proxy | โ with retraining | โ |
| ๐ Readability Flesch-Kincaid, etc. |
No โ rule-based formula | โ cannot apply | โ |
| โ๏ธ Grammar LanguageTool, etc. |
No โ rule-based | โ cannot apply | โ |
| ๐ Factuality | No โ LLM-as-judge | โ prohibitively expensive | โ |
| โ๏ธ Bias / Toxicity | Marginally | Requires model access | โ |
Prompt-only is not a simplified substitute โ it is the only general-purpose architecture that works across the full space of real-world quality oracles.
๐ฌ Connection to Automated Scientific Discovery
This is an instance of a paradigm shift happening across AI: using agentic loops with external feedback signals when gradient flow is unavailable.
When the evaluation function is non-differentiable โ which is the norm, not the exception, in real-world settings โ prompt-based feedback loops are the architecturally natural choice. This work applies that paradigm to text quality optimization and is the first to study it systematically across multiple oracle types.
โ Research Questions
- RQ1 Does a prompt-only agentic loop successfully reduce AI detection scores without gradient access or training?
- RQ2 What are the convergence dynamics โ how many iterations are needed, does quality degrade as scores improve?
- RQ3 Does the same unmodified loop successfully optimize text against non-detection oracles (readability, grammar, toxicity)?
- RQ4 Can the loop simultaneously satisfy multiple quality constraints using a combined feedback signal?
- RQ5 Does loop performance generalize across text domains (news, essays, Wikipedia, medical)?
- RQ6 What is the performance cost of oracle-agnosticism vs. gradient methods on the detection task specifically?
๐งช Experiment Plan
Eight progressive experiments, from minimal to comprehensive. Each builds on the previous. Run E1 first โ everything depends on it.
๐ฏ Result-Agnostic Design
The study is designed to be publishable regardless of outcome direction. The contribution is the systematic study itself.
๐ป What the Loop Looks Like in Practice
All LLM calls use the Claude CLI โ no API key setup, no model weights, no GPU required for the core loop.
The oracle is a black box that returns a score. Swap DetectionOracle for ReadabilityOracle or GrammarOracle โ the loop code does not change.
๐ Paper Positioning
Framed as a text quality framework paper โ not an AI detection evasion paper.
"Bypassing AI detection systems" โ dual-use concerns, likely desk-rejected at top NLP venues.
"Oracle-agnostic text quality optimization" โ in the tradition of Self-Refine, FunSearch, AI Scientist. AI detection is one application among several.
Target venues: EMNLP 2026, ACL 2026, or NAACL 2026 (main track).