Agentic Evolution Systems

What Makes Agentic Evolution Work?

A controlled study of the optimization loop behind evolving vision-language model inference programs.

Seongryong Jung

Chung-Ang University

Read the full paper

Abstract

LLM-driven code optimization repeatedly mutates, evaluates, and selects programs, but the design choices that make improvements accumulate are not well understood. We introduce PRISMEVOLVE, a controlled framework for studying candidate selection, evaluation feedback, initialization, and reuse of prior search experience in vision-language model inference optimization. Across experiments on Qwen3-VL-2B-Instruct and Qwen3-VL-2B-Thinking, we find that strong evolution depends on selectively propagating promising descendants and giving the search policy feedback it can act on. Initial accuracy alone does not predict evolvability, and lineage-specific memory helps in our setting while random archive inspiration does not consistently help.

What makes evolutionary improvements accumulate?

Agentic optimization can generate many candidate programs, but a mutation only matters if the search can identify it, preserve it, and build on it. This study isolates four design choices that govern that process: candidate selection, evaluation feedback, the starting program, and reuse of prior search experience.

A controlled program-evolution loop

At each iteration, the sampler selects programs from the archive, an LLM proposes modifications, and evaluators score the candidates. Their scores and feedback determine which programs are retained and used as inspiration in later rounds.

PRISMEVOLVE framework diagram showing the human-defined task and evaluation criteria, prompt sampler, mutation LLM, program archive, and evaluator pool
PRISMEVOLVE framework. The task and evaluation criteria guide a loop of prompt sampling, LLM-based mutation, program evaluation, and archive-based reuse.

Establishing the starting point

Before comparing evolutionary strategies, we manually optimize an interpretable program for each target model. The ablations below select its inference backend, batching, response format, generation budget, and image resolution.

Plots comparing inference backend, response-generation settings, and image resolution for Qwen3-VL-2B models
Manual optimization. The three panels compare backend and batching, response configuration and output budget, and image resolution. Runtime includes model initialization and evaluation over the 128-example optimization set.

Four factors shape the search

We vary one part of the loop at a time under a fixed mutation budget, then compare how each choice changes the search trajectory and final program quality.

Selection shapes the search tree

Fixed, FIFO, Greedy, and Beam use the same mutation budget but choose candidates differently. Greedy repeatedly expands high-scoring programs, while Beam preserves several promising candidates. The genealogy plot shows how these policies allocate attempts and which program lineages are inherited.

Program genealogies for Fixed, Greedy, Beam, and FIFO search policies
Search genealogies. Nodes denote mutation attempts, colors indicate accuracy, and bold edges trace the selected lineage. The policies differ in how broadly they explore and how often they extend promising descendants.

Feedback helps when a policy can use it

Aggregate feedback reports a program-level score; diagnostic feedback also identifies per-question predictions and errors. The trajectories connect this added information to search behavior: diagnostic feedback yields much larger gains for Greedy than for FIFO, so feedback quality and candidate selection work together.

Accuracy trajectories comparing aggregate and diagnostic feedback under four search policies
Feedback granularity. Best-so-far accuracy under aggregate feedback (F0) and diagnostic feedback (F1). The reported F1–F0 gain is +10.9 points for Greedy and +1.6 points for FIFO.

Initialization: starting quality is not evolvability

The initialization experiment compares baseline, manually optimized, and previously evolved programs. A higher initial score does not guarantee a larger improvement: under Greedy search, the manual start makes larger gains than the already evolved program.

Best-so-far accuracy under Greedy and Beam search from baseline, manually optimized, and previously evolved initial programs
Initialization and evolvability. Search trajectories show how the starting program changes both the initial score and the improvement reached within 36 mutation attempts.

Qwen3-VL-2B-Instruct

Correct answers out of 128

Optimization: manual start → evolved88 → 110
Disjoint validation: manual → evolved79 → 95

Qwen3-VL-2B-Thinking

Correct answers out of 128

Optimization: manual start → evolved80 → 103
Disjoint validation: manual → evolved69 → 92

Search experience: reuse the lineage, not just any memory

The reuse experiment compares no additional context, ancestry memory from the selected parent’s lineage, and an unrelated archive example. Ancestry memory reaches 118/128 with the manually optimized start; random archive inspiration does not consistently improve the endpoint.

Final accuracy comparing no reuse, ancestry memory, and archive inspiration for baseline and manually optimized initializations
Search-experience reuse. Final best-so-far accuracy for the reuse conditions in the baseline and manually optimized settings. Ancestry memory gives the strongest endpoint in the manually optimized condition.

Final scores do not show when those differences emerge. The complete trajectories below follow best-so-far accuracy across the mutation budget and show the effect of each reuse condition over time.

Best-so-far accuracy trajectories comparing no reuse, ancestry memory, and archive inspiration over 36 mutation attempts
Search trajectories. Best-so-far accuracy over 36 attempts for each reuse condition. These results describe the search-experience experiment in Section 6.4.