Optimizers reference¶
The library ships four optimizers. They all implement the same interface, so once you know how to drive one, you know how to drive all of them:
public interface AgentOptimizer<Input, Output, InputLabel> {
public suspend fun train(session: TrainingSession<Input, Output, InputLabel>): TrainingResult
public fun loadOptimizedAgent(baseAgent: GraphAIAgent<Input, Output>): GraphAIAgent<Input, Output>
}
Each optimizer learns from your training runs and writes a JSON artifact to its
storagePath (a ResilientPath). Usage is always two-phase: train learns and
persists the artifact, then loadOptimizedAgent rebuilds your agent with that
artifact applied.
optimizer.train(session) // learns + saves an artifact to storagePath
val optimized = optimizer.loadOptimizedAgent(agent) // rebuilds the agent with the artifact applied
val answer = optimized.run(input)
This page assumes you already have a TrainingSession. Set one up as shown in
Getting Started (build the agent, dataset, metric,
serializers, and call trainingSession(...)), then pick an optimizer below.
Attribution & provenance
None of these optimizers are original work. They are non-official re-implementations of algorithms published by others, ported to Koog. Our implementations may diverge from the original references — due to technical constraints of the Koog runtime, differences in the agent model, or pragmatic simplifications — but we have tried to preserve the core semantics of each. For the authoritative description of any algorithm, consult its original source below.
ACE optimizer deviates the most. Its current implementation differs from the original algorithm in ways that can affect results. Treat it, and other optimizers, as experimental, and validate on your own task before relying on them. Closing the gap is future work.
| Optimizer | Original authors | Reference |
|---|---|---|
| BootstrapFewShot | Omar Khattab et al. (DSPy) | DSPy BootstrapFewShot |
| MIPROv2 | Krista Opsahl-Ong et al. | DSPy MIPROv2 · arXiv:2406.11695 |
| GEPA | Lakshya A. Agrawal et al. | DSPy GEPA tutorial · arXiv:2507.19457 |
| ACE | Qizheng Zhang et al. | arXiv:2510.04618 |
Choosing an optimizer¶
| Optimizer | What it learns | Best for | Needs ground-truth label? |
|---|---|---|---|
| BootstrapFewShot | few-shot demonstrations mined from successful runs | quick wins when you already have good runs to mine | optional — judged by the metric |
| MIPROv2 | instructions + demos, via proposal and search | when you want both instruction tuning and demos | optional — judged by the metric |
| GEPA | module instructions, via reflective evolution over a Pareto frontier | ambiguous or underspecified instructions; instruction-only | yes — via gepaFeedback |
| ACE | an evolving "playbook" of insights appended to the system prompt | accumulating reusable strategy and knowledge | yes — via a labelExtractor (string label) |
For models, the examples below use OpenAIModels.Chat.GPT4oMini as the example
LLModel; substitute any Koog model. storagePath is always a ResilientPath,
e.g. ResilientPath("build/artifacts/foo.json").
BootstrapFewShot¶
Runs each training item up to maxRounds times, mines the successful runs into
few-shot demonstrations, and applies them to your agent. The simplest optimizer:
no extra LLM calls beyond running your own agent.
BootstrapFewShotOptimizer<Input, Output, InputLabel>(
maxBootstrappedDemos: Int,
maxRounds: Int,
maxTotalDemos: Int,
includeLabeledExamples: Boolean,
storagePath: ResilientPath,
randomSeed: Long,
)
Key params:
maxBootstrappedDemos— how many successful traces to collect and use as demos.maxRounds— retry budget per training item while hunting for a success.maxTotalDemos— total demo slots. If bootstrapped demos don't fill them andincludeLabeledExamplesistrue, the remaining slots are filled from labeled dataset items.includeLabeledExamples— whether to backfill empty demo slots with labeled examples.
import ai.koog.agents.optimization.optimizers.fewShot.BootstrapFewShotOptimizer
import ai.koog.agents.optimization.utils.common.ResilientPath
val optimizer = BootstrapFewShotOptimizer<String, String, Double>(
maxBootstrappedDemos = 4,
maxRounds = 3,
maxTotalDemos = 8,
includeLabeledExamples = true,
storagePath = ResilientPath("build/artifacts/fewshot.json"),
randomSeed = 42L,
)
optimizer.train(session)
val optimized = optimizer.loadOptimizedAgent(agent)
val answer = optimized.run("What is the weather at 9:00? Answer with a single number.")
MIPROv2¶
Jointly proposes and searches over instructions and demos. A meta-LLM proposes candidate instructions; demos are bootstrapped; candidates are scored and the best combination is kept.
MIPROv2Optimizer<Input, Output, InputLabel>(
metaModel: LLModel,
autoMode: AutoRunMode = AutoRunMode.LIGHT, // LIGHT | MEDIUM | HEAVY
numCandidatesOverride: Int? = null,
numTrialsOverride: Int? = null,
maxBootstrappedDemos: Int = 4,
maxTotalDemos: Int = 8,
includeLabeledExamples: Boolean = true,
proposerConfig: InstructionProposerConfig = InstructionProposerConfig(),
storagePath: ResilientPath, // named arg, has no default
randomSeed: Long = 42,
parallelism: Int = 1,
)
Key params:
metaModel— the LLM used for instruction proposal and dataset summarization.autoMode— a preset budget:LIGHT,MEDIUM, orHEAVY. Higher modes try more candidates and trials (and cost more). Start withLIGHT.numCandidatesOverride/numTrialsOverride— manual overrides for the search budget when the preset doesn't fit.storagePath— note this has no default, so pass it as a named argument even though it sits after defaulted params.parallelism— concurrent runs during bootstrapping.
A run records its search into the training records: the modules it found and the size of each dataset
split, the size of every generated demo set, every proposed instruction, and, for each trial, the
candidate indices it drew alongside its score. Step 3: Grid Search names the trial the run kept, so
the saved artifact can be traced back to the combination that produced it. Long values are clipped to
TrainingResources.actionLogTruncation, which a run can raise to keep more of them; the run log carries each proposed instruction in full.
import ai.koog.agents.optimization.optimizers.mipro.MIPROv2Optimizer
import ai.koog.agents.optimization.optimizers.mipro.AutoRunMode
import ai.koog.agents.optimization.utils.common.ResilientPath
import ai.koog.prompt.executor.clients.openai.OpenAIModels
val optimizer = MIPROv2Optimizer<String, String, Double>(
metaModel = OpenAIModels.Chat.GPT4oMini,
autoMode = AutoRunMode.LIGHT,
storagePath = ResilientPath("build/artifacts/mipro.json"),
)
optimizer.train(session)
val optimized = optimizer.loadOptimizedAgent(agent)
val answer = optimized.run("What is the weather at 14:00? Answer with a single number.")
GEPA¶
Instruction-only. GEPA evolves module instructions with reflective LLM feedback over agent traces, maintaining a Pareto frontier of candidates and optionally merging complementary ones (crossover).
GEPAOptimizer<Input, Output, InputLabel>(
gepaFeedback: GEPAFeedback<Input, Output, InputLabel>,
feedbackValSplitFn: (TrainSet<Input, InputLabel>) -> GEPATrainSetSplit<Input, InputLabel>,
reflectionModel: LLModel,
storagePath: ResilientPath,
numRollouts: Int,
feedbackBatchSize: Int = 3,
skipPerfectFeedbackBatches: Boolean = true,
perfectScoreThreshold: Double = 1.0,
randomSeed: Long = 42,
moduleSelectionStrategy: GEPAModuleSelectionStrategy = GEPAModuleSelectionStrategy.ROUND_ROBIN,
useStructuredOutput: Boolean = false,
requireThinkingFieldInOutput: Boolean = false,
mergeConfig: GEPAMergeConfig = GEPAMergeConfig(),
failureScore: Double = 0.0,
feedbackFailureRateThreshold: Double = 1.0,
validationFailureRateThreshold: Double = 0.9,
seedValidationFailureRateThreshold: Double = 0.0,
abortOnFailureRateExceeded: Boolean = true,
)
Key params:
reflectionModel— the LLM that reflects on traces and proposes new instructions.numRollouts— the required total rollout budget. It includes the initial validation-set evaluation and must be greater than the validation-set size; otherwise GEPA would have no budget left for optimization. A practical starting point is 6–20× the input training dataset size.feedbackBatchSize— the number of feedback items sampled per reflection attempt.moduleSelectionStrategy— which optimizable modules to update each iteration.ROUND_ROBINupdates one module at a time in rotation;ALLupdates every module each iteration.mergeConfig— whether and how to merge complementary candidates. Experimental: tests reach its proposal step and no further, merging has never run end to end, and it may diverge from GEPA's original implementation more significantly. Merging also needs an agent with at least two optimizable modules — it combines the module one lineage changed with the module the other changed, so enabling it for a single-module agent does nothing.failureScore— score assigned when an agent rollout still fails after retries.feedbackFailureRateThreshold— maximum failed-item ratio in parent and candidate feedback batches. The default1.0disables the cap: a failed parent batch is skipped, while failed candidate rollouts retainfailureScorein the acceptance average.validationFailureRateThreshold— maximum failed-item ratio in candidate validation runs. WithabortOnFailureRateExceededenabled, exceeding it terminates the entire optimization.seedValidationFailureRateThreshold— maximum failed-item ratio in the initial seed validation. With aborting enabled, the default0.0stops after the first failure, preventing an artificially weak baseline.abortOnFailureRateExceeded— whether exceeding any rollout failure-rate threshold aborts GEPA. Whenfalse, GEPA evaluates the full rollout set and records threshold breaches without terminating the optimization; terminal agent failures still receivefailureScore.
Thresholds use a strict > comparison: 0.0 permits no failures, while 1.0 disables the cap.
Failed rollout items remain visible as failed stages in the training record even when their rollout
set stays within its configured threshold.
A run records its search into the training records: the seed and every accepted candidate with their
validation scores and parents, each batch's parent, module, scores and outcome, and every proposed
instruction. Long values are clipped to TrainingResources.actionLogTruncation, which a run can raise
to keep more of them; the run log carries each proposed instruction in full.
import ai.koog.agents.optimization.optimizers.gepa.GEPAOptimizer
import ai.koog.agents.optimization.optimizers.gepa.GEPAModuleSelectionStrategy
import ai.koog.agents.optimization.optimizers.gepa.GEPATrainSetSplit
import ai.koog.agents.optimization.optimizers.gepa.strategyOnlyGepaFeedback
import ai.koog.agents.optimization.utils.common.ResilientPath
import ai.koog.prompt.executor.clients.openai.OpenAIModels
val optimizer = GEPAOptimizer<String, String, String>(
gepaFeedback = strategyOnlyGepaFeedback { input, output, gold, _, _ ->
val score = if (output == gold) 1.0 else 0.0
score to if (score == 1.0) "Solved." else "Expected '$gold', got '$output'."
},
feedbackValSplitFn = { items ->
val splitAt = items.size / 2
GEPATrainSetSplit(feedbackSet = items.take(splitAt), validationSet = items.drop(splitAt))
},
reflectionModel = OpenAIModels.Chat.GPT4oMini,
storagePath = ResilientPath("build/artifacts/gepa.json"),
numRollouts = 100,
feedbackBatchSize = 3,
moduleSelectionStrategy = GEPAModuleSelectionStrategy.ROUND_ROBIN,
failureScore = 0.0,
feedbackFailureRateThreshold = 1.0,
validationFailureRateThreshold = 0.9,
seedValidationFailureRateThreshold = 0.0,
abortOnFailureRateExceeded = true,
)
optimizer.train(session)
val optimized = optimizer.loadOptimizedAgent(agent)
val answer = optimized.run("Route this support ticket: my invoice is wrong")
ACE¶
Curates an evolving "playbook" of bullet-point insights that is appended to the system prompt. Useful when you want the agent to accumulate reusable strategy and knowledge across runs rather than just demos or a reworded instruction.
ACEOptimizer<Input, Output, InputLabel>(
playbookStoragePath: ResilientPath,
onExistingPlaybook: OnExistingPlaybookAction, // OVERWRITE | OPTIMIZE_FURTHER | THROW
reflectorModel: LLModel,
curatorModel: LLModel,
labelExtractor: (datasetItem: TrainSetItem<Input, InputLabel>) -> String,
)
Key params:
playbookStoragePath— where the playbook artifact is read from and written to.onExistingPlaybook— what to do if a playbook already exists at that path:OVERWRITEdeletes it and starts fresh,OPTIMIZE_FURTHERloads it and keeps refining,THROWrefuses to start to avoid clobbering prior results.reflectorModel— diagnoses trajectories into insights.curatorModel— turns those insights into playbook deltas (additions/edits).labelExtractor— extracts the ground-truth label string for an item, fed to the reflector. Required.
import ai.koog.agents.optimization.optimizers.ace.ACEOptimizer
import ai.koog.agents.optimization.optimizers.ace.OnExistingPlaybookAction
import ai.koog.agents.optimization.utils.common.ResilientPath
import ai.koog.prompt.executor.clients.openai.OpenAIModels
val optimizer = ACEOptimizer<String, String, Double>(
playbookStoragePath = ResilientPath("build/artifacts/ace-playbook.json"),
onExistingPlaybook = OnExistingPlaybookAction.OVERWRITE,
reflectorModel = OpenAIModels.Chat.GPT4oMini,
curatorModel = OpenAIModels.Chat.GPT4oMini,
labelExtractor = { item -> item.itemLabel.toString() },
)
optimizer.train(session)
val optimized = optimizer.loadOptimizedAgent(agent)
val answer = optimized.run("Route this support ticket: my invoice is wrong")
Artifacts & reuse¶
train writes the learned artifact to the optimizer's storagePath
(playbookStoragePath for ACE). loadOptimizedAgent reads that artifact back and
applies it to your base agent.
The two steps are decoupled by the file on disk. If the artifact already exists —
for example from a previous training run, or one you committed — you can skip
train entirely and just call loadOptimizedAgent(agent) to get the optimized
agent. This is how you deploy an optimized agent without re-running optimization.
Next¶
- Making agents optimizable — declare optimizable
modules with
optimizableSubgraphWithTask, and strategy- vs subgraph-level optimization. - Writing a custom optimizer — implement the
AgentOptimizerinterface with theStageScopetraining DSL. - API reference — full Dokka docs.