If you are about to fund a fine-tune to close a quality gap, run GEPA against your own labeled examples first. You pay for LLM calls instead of a training job, it works from 20 to 100 examples, and in the published head-to-heads it beat RL fine-tuning (the GEPA paper) and matched supervised fine-tuning (Databricks' test on GPT-4.1). Use DSPy's dspy.GEPA if your pipeline is already in DSPy and the standalone gepa library if it is not. Fine-tune only when you have a reason a prompt cannot satisfy: token cost at high volume, latency, a behavior no prompt produces, or a requirement to own the weights.
Two things change the usual advice in October 2026. OpenAI's dataset-backed prompt optimizer shuts down with its Evals platform on November 30, and OpenAI's fine-tuning platform is no longer open to new users. So if you were planning either OpenAI path, you need a different plan. The second thing is the one buyers miss: an optimized prompt is compiled against one model version, and the next upgrade can undo the gain.
| Option | Optimizes against | Data you need | Cost (checked Oct 11, 2026) | Pick it if | Skip it if |
|---|---|---|---|---|---|
| DSPy GEPA | Your metric's score plus its written feedback | 20-100 labeled examples, train/val split | Free (MIT); you pay LLM calls | Your pipeline is in DSPy or you can wrap it | Your metric can only return a number |
| DSPy MIPROv2 | A scalar metric; tunes instructions and few-shot demos | A validation set of at least 35 with default minibatching | Free (MIT); you pay LLM calls | You only have a scalar metric | You can write feedback; GEPA beat it in the paper |
| gepa (standalone) | Any text artifact with a feedback metric | Same as GEPA | Free (MIT), v0.1.4 | A single system prompt, an MCP tool description, a LangChain chain | You need a UI for non-engineers |
| MLflow / Opik / Arize AX | GEPA or meta-prompting over data you already log | A dataset built from logged traces | Opik Pro $19/mo; Arize AX Pro $50/mo; MLflow open source | Your traces and scorers already live there | The tool isn't in your stack yet |
| OpenAI Prompt Optimizer | Graders and annotations on a dataset | At least 3 graded rows | Shuts down Nov 30, 2026 | Nobody, from today | You want it to exist next quarter |
| Anthropic Prompt Improver | No metric; it rewrites from best practices | Optional examples and feedback | No separate price | You want a first structured draft | You need a measured improvement |
| Fine-tuning (OpenAI for existing users, Together AI for open weights) | Weights, from training examples | A labeled training set | gpt-4.1-mini: $5.00/1M training tokens, inference at 2x base; Together gpt-oss-120B: $2.50/1M | Volume, latency, owned weights | You have not run an optimizer yet |
The loser is the OpenAI Prompt Optimizer, for a reason that has nothing to do with quality: OpenAI's own guide says it is deprecating the dataset-backed optimizer as part of the Evals platform, which goes read-only on October 31, 2026 and shuts down on November 30. Anything you build on it now is a migration you have already scheduled.
What Is Prompt Optimization, and How Is It Different From Fine-Tuning?
Prompt optimization is an automated search over the text you send a model (instructions, few-shot examples, tool descriptions) that scores each candidate against a metric on labeled examples and keeps the best one. Fine-tuning changes the model's weights instead. The optimizer's output is a string you can read, diff and version. The fine-tune's output is a model artifact you have to host, or pay a provider to host.
The two optimizers that matter here are both from the DSPy project. MIPROv2 bootstraps candidate few-shot demos from your training set, has a model propose instruction candidates, then uses Bayesian optimization to pick the best combination. GEPA (Genetic-Pareto) reads full execution traces, has a "reflection" model diagnose in plain language why an attempt failed, proposes an edited instruction, and keeps a Pareto set of candidates that each win on some examples instead of one global best.
The practical difference is the metric. DSPy's optimizer guide says GEPA requires a metric that returns Prediction(score, feedback), while COPRO and MIPROv2 treat the metric as a single scalar. If your evaluator can say "the answer cited the 2024 contract instead of the 2025 amendment," GEPA can use that sentence. MIPROv2 only sees a 0.
What the GEPA Paper Actually Measured
GEPA beat reinforcement learning and MIPROv2 on six tasks across two models, and the headline margin over RL shrank between versions of the paper. The current revision (v2, February 14, 2026) of "GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning," accepted as an ICLR 2026 oral, reports:
- GEPA beats GRPO, the RL fine-tuning baseline, by 6% on average and up to 20%, using up to 35x fewer rollouts. The first version from July 25, 2025 said 10% on average. If someone quotes you "10% over RL," they are quoting the older draft.
- GEPA beats MIPROv2 by over 10%, including +12% accuracy on AIME-2025, with prompts up to 9.2x shorter than MIPROv2's.
- The tasks were HotpotQA, IFBench, HoVer, PUPA, AIME-2025 and LiveBench-Math. The models were Qwen3 8B (open weights) and GPT-4.1 Mini.
The protocol detail that matters to a budget: GRPO ran 24,000 rollouts per benchmark. GEPA on Qwen3 8B used between 1,839 and 7,051, and the authors note that most of that went to validation. The train-set rollouts needed to reach peak performance were 79 to 737. In the paper's own cost appendix, every GEPA run in the main results table cost $86 in total on GPT-4.1 mini, against $76 for MIPROv2.
The caveats are in the paper too. GRPO was run with LoRA, and full fine-tuning appears only in a figure. The Merge variant hurt Qwen3 8B on some tasks. These are academic benchmarks with clean metrics, which your task probably does not have.
The Eval Set Is the Actual Project
The optimizer is the cheap part. The labeled examples and the metric decide its ceiling, and they take most of the work.
Three production writeups say the same thing from different angles.
Decagon ran more than 19 ablations of GEPA on a binary classification supervisor (March 25, 2026) and found that 20 to 100 training examples beat 500. At 500, performance dropped 2%, compute rose 10x and prompts grew 75% longer than at 50, because GEPA encoded every edge case. A small reflection model (GPT-4o-mini) left the prompt essentially unchanged; GPT-4.1, GPT-5.2, Claude Sonnet and Claude Opus each gave about 5 to 6%. Unconstrained, GEPA wrote prompts over 5,000 characters, and a 1,500-character cap cost only 0.8%.
Simon Willison ran GEPA on Datasette Agent's SQL system prompt (July 2, 2026) with 20 training and 10 held-out questions. His results table shows train accuracy rising from 90.0 to 95.0 and held-out accuracy falling from 95.0 to 85.0. GEPA had added advice to "query the distinct statuses first," which collided with a display rule already in the prompt. Two of the three apparent baseline failures turned out to be bugs in his scorer; fixing them moved baseline test accuracy from 81.7 to 95.0, a bigger move than the optimizer made.
Databricks' IE Bench work (September 24, 2025) found GEPA needed roughly three times the LLM calls of MIPROv2, about 2 to 3 hours of run time against 1.
So budget the eval set as the project. You need 20 to 100 examples split into train and validation, a third untouched test set the optimizer never sees, and a metric that explains failures in words. If your metric is an LLM judge, test the judge before you trust what it optimizes. Our piece on why 200 hand-written questions beat synthetic ones covers how to build that set.
Who Should Not Pick Each Option
Every option here has a buyer it fails.
DSPy GEPA and MIPROv2
Don't pick GEPA if your only metric is exact-match and you will not write feedback. It still runs, but you throw away the thing that separates it from MIPROv2. Don't pick MIPROv2 in a new project unless you are stuck with a scalar metric; the paper measured GEPA ahead of it by over 10%. Don't pick DSPy at all if your team will not restructure its pipeline into DSPy modules. DSPy is MIT-licensed and the code is free, but adopting it is a framework decision.
DSPy's own guide is blunt on the fine-tuning question: the lift from fine-tuning over GEPA or MIPROv2 is "usually small and sometimes negative," and it treats fine-tuning as the last lever, after prompt-only optimization plateaus.
The standalone gepa library
The gepa library is MIT-licensed, and version 0.1.4 shipped July 15, 2026. It optimizes any text artifact, with adapters for a single-turn system prompt, a generic RAG pipeline, MCP tool descriptions and system prompts, LangChain pipelines and whole DSPy programs. This is the right choice when your agent is not written in DSPy and you do not want it to be. Don't pick it if the people who own the prompt are not engineers. It is a library with a 0.1 version number, with no UI and no review workflow.
MLflow, Opik and Arize AX
These run optimizers inside the platform where your traces already live. MLflow's optimize_prompts (MLflow 3.5.0 or later) offers a GEPA optimizer and a meta-prompting optimizer, takes prompts from its registry by URI and writes optimized versions back to it. Opik's Agent Optimizer (our Opik page) lists six optimizers including GEPA, MetaPrompt and an Evolutionary one, and works from "the datasets, metrics, and traces you already log to Opik." Opik is Apache-2.0; its hosted Pro tier is $19 per month for 100k spans. Arize AX's Prompt Learning (our Arize AX page) is meta-prompting over evaluations, natural-language feedback and production traces; AX Pro is $50 per month for 50k spans, and the pricing page does not state which tier includes Prompt Learning.
The advantage is versioning: the optimized prompt lands in a registry next to the traces that produced it. Don't pick one of these just to get the optimizer if you are not already on the platform. Moving observability vendors to get a feature that the free gepa library has is the wrong trade. Note also that Arize's Prompt Learning is meta-prompting, a different algorithm from GEPA, and the GEPA paper's numbers do not transfer to it.
OpenAI Prompt Optimizer
Its design was sound. It accepted Good/Bad annotations, text critiques and grader results, and its guide warned that an optimized prompt can do worse than the original on some inputs. The problem is the shutdown date. Don't pick it. If you have datasets in it, export them before October 31.
Anthropic Prompt Improver
The Prompt Improver rewrites a prompt from best practices and runs no search against a metric. Anthropic's October 14, 2024 launch post describes it adding a reasoning section, standardizing examples into XML and adding a prefilled Assistant message, and cites a 30% accuracy gain on a multilabel classification test with Claude 3 Haiku (a vendor-reported figure). Don't pick it when you need a measured improvement, and watch the prefill. Anthropic's current guidance says prefilled responses on the last assistant turn are no longer supported starting with Claude 4.6 models. A prompt the improver produced in 2024 can break on a current Claude model. As of October 11, 2026, the improver's old documentation URL redirects to that general best-practices page.
What Fine-Tuning Still Wins
Fine-tuning wins outright in four cases, and you should be able to name which one you are in before you sign off on the spend.
- Token cost at volume. A fine-tune can replace a long instruction block and a dozen few-shot examples. If that removes most of your input tokens and you serve millions of requests, the training bill pays back.
- Latency. A shorter prompt means faster time to first token. A small fine-tuned model can stand in for a large prompted one.
- Behaviors you cannot prompt. A strict output dialect, a domain vocabulary or a style the base model resists no matter what you write.
- Owning the weights. Data residency, air-gapped deployment, or a policy that a model you depend on cannot be deprecated under you.
You don't have to choose one or the other. In Databricks' test on GPT-4.1, supervised fine-tuning added 1.9 points, GEPA added 2.1, and GEPA on top of the fine-tune added 4.8. The GEPA-optimized GPT-4.1 cost about 20% less to serve than the fine-tuned one. Their larger result: GEPA-optimized gpt-oss-120b beat baseline Claude Opus 4.1 on IE Bench at roughly 90x lower serving cost. Databricks sells a platform that runs both, so read that as a vendor result, but the GPT-4.1 comparison has the structure you want to copy: same task, same eval, both methods, and the combination.
Cost on the Same Workload
Here is one workload, priced both ways at list rates checked October 11, 2026: gpt-4.1-mini, 1 million requests a month, 1,500 input tokens and 200 output tokens per request. The arithmetic is ours; the rates are OpenAI's pricing page.
- Base model, prompted: $0.40 per 1M input and $1.60 per 1M output gives $600 + $320 = $920 a month.
- Optimized prompt, 500 tokens longer: $800 + $320 = $1,120 a month, plus a one-off optimization run that cost the GEPA paper $86 across its whole main table. Cached input at $0.10 per 1M lowers this if the static prefix hits the cache.
- Fine-tuned, same 1,500-token prompt: fine-tuned gpt-4.1-mini bills $0.80 input and $3.20 output, exactly double, so $1,840 a month. Training is $5.00 per 1M training tokens: 1,000 examples of 1,000 tokens for 3 epochs is $15.
- Fine-tuned, prompt cut to 300 tokens: $240 + $640 = $880 a month.
The training run costs $15 once; the serving premium on the same prompt costs $920 a month. On these rates a fine-tune only beats the plain prompted base model if it lets you delete around 80% of the input, because output tokens do not get shorter and their price doubles. Most teams budget the training job and forget the premium.
On open weights the math moves. Together AI lists supervised fine-tuning at $2.50 per 1M training tokens for gpt-oss-120B and $0.40 for gpt-oss-20B, with per-job minimums of $6 and $4. Its pricing page shows no separate rate for serving your adapter, so get that number in writing before you model the year. Our earlier fine-tuning vs RAG cost breakdown covers the serving contract in detail.
Multi-Agent Systems Break the Single-Agent Numbers
Every number above comes from single-agent or single-pipeline tasks. In multi-agent systems, prompt optimization helps or hurts depending on the topology, and the swing is large.
MAS-PromptBench (Juyang Bai and Laixi Shi, June 22, 2026) adapted GEPA and MIPRO to multi-agent systems across four frameworks, nine tasks and five topologies, optimizing each agent's system prompt while keeping the communication structure fixed. The best gain was +24.0 points (BFCL tool calling, sequential topology, CrewAI). The worst was -16.0 (MATH, independent topology, LangGraph, 76.0 to 60.0). Averaged by topology, the single-agent baseline gained +4.2 while independent agents lost 0.5. Average gains for multi-agent GEPA fell from +2.4 with two agents to -2.1 with ten.
The authors' explanation for the independent topology is that parallel agents revise their prompts without coordinating, and the revisions cancel. Their protocol also only kept an optimized configuration if it beat the original on validation. Copy that guard. If you run an orchestrator with sub-agents, optimize and measure per configuration, and keep the seed prompts as the fallback.
The Optimized Prompt Expires With the Model
An optimized prompt is a compiled artifact. It was searched against one model's behavior and it is tuned to that model's quirks. When the model changes, the quirks change. None of the studies above measured how much of an optimized gain survives a version bump.
The evidence that this happens is on the record. Anthropic's prefill change broke a technique its own 2024 tool inserted. We measured few-shot prompting that stopped paying on GPT-4o while a different model still wanted it. And model retirements arrive on a schedule you do not set; our deprecation notice comparison lists the windows.
That is the cost nobody puts in the business case. Treat it like code:
- Version the prompt with the model ID it was optimized against, in the same commit or registry entry.
- Re-run the optimizer as a step in every model migration, against the same held-out set, before cutover.
- Budget it as recurring work. One run per model upgrade per optimized prompt, plus the time to review the diff. A fine-tune has the same problem in a heavier form, because a new base model means a new training run.
How to Decide
The criteria that predict regret are not the ones on the feature list.
- Do you have 20 to 100 labeled examples and a metric that explains failures? If not, build that first. No optimizer and no fine-tune will rescue a task you cannot measure.
- Is the gap a quality gap or a cost gap? Quality gap: optimize first. Cost gap at high volume: price the fine-tune including the inference premium, then compare.
- How often will you change models? If you move every quarter, an optimizer you can re-run in an afternoon beats a training pipeline you must rebuild.
- Single agent or many? Many agents: measure per topology, and assume a regression is possible.
- Must you own the weights? If yes, the answer is open-weight fine-tuning, and you should run GEPA on top of it anyway.
What changes the answer: a provider removing fine-tuning (OpenAI already has for new users), a model upgrade that erases the optimized gain, or volume that grows past the point where the serving premium dominates.
What to Do Next
This Week:
- Pull 100 real examples from production traces, label them, and split them 50 train, 25 validation, 25 untouched test.
- Rewrite your metric to return a reason with every score. If it is an LLM judge, hand-check 20 of its verdicts.
- If you have data in OpenAI's Evals or Prompt Optimizer, export it before October 31.
This Month:
- Run
dspy.GEPAor thegepalibrary on thelightbudget with a frontier reflection model, and cap the prompt length. - Score the result on the untouched test set. If train improves and test does not, you have Willison's result; shrink the training set or fix the metric.
- Only if a gap remains, price a fine-tune with the inference premium included, and test GEPA on top of it.
Before Your Next Model Migration:
- Store each optimized prompt with the model ID it was compiled for.
- Add "re-run optimizer, compare on held-out set" to the migration runbook, with an owner and a line in the budget.
The Bottom Line
This is the compiler pattern applied to prompts. Compiled code is fast and specific to one target, and you rebuild it when the target changes. GEPA gives you most of what a fine-tune would, for the price of a labeled set and a few hundred LLM calls, and it ties the result to one model version. Fine-tuning ties the result to one version too, and charges double for every output token while it does.
Build the eval set, run the optimizer, and put the re-run in the migration plan before you open a training job.
Continue Reading
- Fine-Tuning vs RAG Cost: The Training Bill Isn't the Bill
- Few-Shot Stopped Paying on GPT-4o. Qwen Still Wants It.
- RAG Eval Dataset: 200 Hand-Written Questions Beat Synthetic Ones
- Braintrust vs Langfuse vs Promptfoo: Don't Pay for the Gate
- Model Deprecation: Bedrock Keeps Claude 4 Months Past Anthropic
- Only 3 of 36 Model Gaps Were Real. Size Your Eval Set.
