Accepted at CAISc 2026 ยท See publication โ†—

Oracle-Agnostic Agentic Loops
for Text Quality Optimization

A prompt-only, training-free framework that iteratively improves LLM-generated text against any quality criterion โ€” no gradients, no fine-tuning, no model access required.

No Training
No Gradients
Any Quality Oracle
Claude CLI
Venue target: EMNLP / ACL 2026

๐Ÿ’ก The Core Idea

A self-evolving agentic loop where an LLM iteratively rewrites its output guided solely by external quality scores โ€” no access to model internals needed.

Step 1
LLM Generates
textt
โ†’
Step 2
Oracle Scores
score = f(textt)
โ†’
Step 3
Inject Feedback
score โ†’ prompt
โ†’
Step 4
LLM Rewrites
textt+1
โ†ฉ
repeat until all oracles pass
โœ“   Loop terminates when score crosses the target threshold โ€” or max iterations reached
Key Insight

The oracle only needs to return a score. It does not need to be differentiable, open-source, or even documented internally. This makes the loop applicable to grammar checkers, readability formulas, factuality verifiers, and AI detectors โ€” all with the same code, zero modification.

โšก Why Prompt-Only? The Generalization Argument

Existing gradient-based and token-level methods are fundamentally tethered to a single oracle type. Here is why.

Oracle Type Differentiable? Gradient Methods Work? Our Loop Works?
๐Ÿค– AI Detection Only with proxy โœ“ with retraining โœ“
๐Ÿ“– Readability
Flesch-Kincaid, etc.
No โ€” rule-based formula โœ— cannot apply โœ“
โœ๏ธ Grammar
LanguageTool, etc.
No โ€” rule-based โœ— cannot apply โœ“
๐Ÿ” Factuality No โ€” LLM-as-judge โœ— prohibitively expensive โœ“
โš–๏ธ Bias / Toxicity Marginally Requires model access โœ“

Prompt-only is not a simplified substitute โ€” it is the only general-purpose architecture that works across the full space of real-world quality oracles.

๐Ÿ”ฌ Connection to Automated Scientific Discovery

This is an instance of a paradigm shift happening across AI: using agentic loops with external feedback signals when gradient flow is unavailable.

AI Scientist ยท Sakana AI (2024)
Generate paper โ†’ evaluate (reviewer score, experiments) โ†’ refine. Evaluator is non-differentiable.
FunSearch ยท DeepMind (2023)
Generate program โ†’ run fitness function โ†’ score feeds back to LLM. No gradient through evaluator.
Our Work ยท Text Quality
Generate text โ†’ score quality oracle(s) โ†’ inject score as prompt feedback โ†’ rewrite.
The Pattern

When the evaluation function is non-differentiable โ€” which is the norm, not the exception, in real-world settings โ€” prompt-based feedback loops are the architecturally natural choice. This work applies that paradigm to text quality optimization and is the first to study it systematically across multiple oracle types.

โ“ Research Questions

๐Ÿงช Experiment Plan

Eight progressive experiments, from minimal to comprehensive. Each builds on the previous. Run E1 first โ€” everything depends on it.

E1
Feasibility
Does the loop work at all?
100 ChatGPT texts from HC3, single AI detection oracle, 10 iterations. Establishes minimum viable result. Start here.
Start Here
E2
Convergence Analysis
How many iterations? Does quality degrade?
500 samples across 5 domains, 15 iterations. Identifies the "prompt ceiling" and optimal stopping point via Pareto analysis.
Core
E3
Domain Generalization
Does it work across text domains?
HC3 (5 domains), OUTFOX essays, M4 Wikipedia. Tests whether performance is consistent or domain-specific.
Core
E4
Multi-Oracle Optimization
Can the loop satisfy multiple criteria at once?
AI detection + readability + grammar simultaneously. Combined feedback signal. Measures metric conflicts and joint convergence.
Core Claim
E5
Oracle Agnosticism Test
Same code, all oracle types, zero modification?
Run identical loop_runner.py on 4 oracle types independently. Validates the central novelty claim empirically.
Key Novelty
E6
Baseline Comparison
What do we trade away vs. gradient methods?
Compare to DIPPER paraphraser and single-shot rewriting on detection task. Quantifies the cost of oracle-agnosticism.
Core
E7
Prompt Ablations
Which feedback format works best?
4 feedback styles: numeric-only, prescriptive, chain-of-thought, combined. Identifies optimal prompt template.
Engineering
E8
Multi-Model Study
Does model capability drive loop performance?
LLaMA, Mistral, Claude, Gemini via OpenRouter. Tests whether instruction-following quality predicts convergence speed.
Future Work

๐ŸŽฏ Result-Agnostic Design

The study is designed to be publishable regardless of outcome direction. The contribution is the systematic study itself.

If Results Are Positive
Prompt-only loops are surprisingly competitive
We don't need gradients for this class of problem. Loops match or approach gradient methods and uniquely generalize to non-differentiable oracles.
If Results Are Negative
We characterize the "prompt ceiling"
We identify the fundamental limit of prompt-only optimization and specify which oracle types require gradient methods โ€” a principled guide for practitioners.
If Results Are Mixed
A taxonomy of oracle amenability
Some oracles respond well to natural-language feedback; others don't. We build a framework mapping oracle types to their amenability โ€” itself a novel contribution.
Unconditional contributions regardless of any outcome: (1) formal framework definition, (2) open-source implementation, (3) convergence characterization, (4) comparison with gradient baselines.

๐Ÿ’ป What the Loop Looks Like in Practice

All LLM calls use the Claude CLI โ€” no API key setup, no model weights, no GPU required for the core loop.

# The loop calls the oracle, builds a prompt, and calls Claude CLI # Same 3 lines regardless of oracle type for iteration in range(max_iterations): result = oracle.score(current_text) # any oracle prompt = build_prompt(text, result, iteration) # inject score as text current_text = claude_cli(prompt) # claude -p "..." if result.passed: break # threshold crossed

The oracle is a black box that returns a score. Swap DetectionOracle for ReadabilityOracle or GrammarOracle โ€” the loop code does not change.

๐Ÿ“„ Paper Positioning

Framed as a text quality framework paper โ€” not an AI detection evasion paper.

Avoid This Framing

"Bypassing AI detection systems" โ€” dual-use concerns, likely desk-rejected at top NLP venues.

Use This Framing

"Oracle-agnostic text quality optimization" โ€” in the tradition of Self-Refine, FunSearch, AI Scientist. AI detection is one application among several.

Target venues: EMNLP 2026, ACL 2026, or NAACL 2026 (main track).