Automatic Prompt Optimization, Proven on Your Data
Give Auto-Refine your labeled data and a starting prompt. It explores dozens of prompt variants, measures the accuracy of every one, and hands you the top 3 candidates — ranked by accuracy and cost.
How Auto-Refine Works
From labeled data to a production-ready prompt in four steps
Choose Your Labeled Data
Pick a folder with 20+ labeled items — images, PDFs, or text. The output schema is derived automatically from your data.
Set the Task and Models
Write a starting prompt and choose the model to optimize. Sensible defaults cover the rest — adjust budget and targets only if you want to.
Run Automatic Optimization
Auto-Refine generates and evaluates prompt variants against your data in the background, learning from failures to write better instructions each round.
Pick a Winner and Save
Review the Top-3 candidates — most accurate, cheapest, and recommended — and save the winner as a reusable prompt in one click.
Accuracy You Can See — and Trust
Every prompt variant is evaluated on your own labeled data with statistical confidence bounds — so improvements are real, not lucky.
Tested on Your Data
Accuracy is measured against your ground truth labels, not synthetic benchmarks.
Improvement Over Time
Watch accuracy climb round by round with an improvement chart and live progress.
Every Variant, Compared
Browse every explored prompt with its accuracy, cost, and lineage — nothing is a black box.
Images, PDFs, and Text
Optimize extraction prompts for multimodal inputs, not just plain text.
Better Prompts Without Runaway Costs
Stay in control of spend from start to finish: see the estimate before you launch, cap the budget, and stop early the moment your target is reached.
Upfront Estimates
See the estimated cost and duration before you start — no surprises.
Hard Budget Cap
Set a spending limit in USD. Optimization stops automatically before exceeding it.
Stop at Good Enough
Set a target accuracy and Auto-Refine stops as soon as it is reached — don't pay for perfection you don't need.
Cancel Anytime
Stop a run whenever you like and keep every variant found so far.
Pick the Trade-Off That Fits: Accuracy, Cost, or Both
Auto-Refine surfaces three winners so you can choose the trade-off that fits your use case.
Most Accurate
The highest measured accuracy — for when every extraction has to be right.
Cheapest
Nearly the same accuracy at a fraction of the token cost — ideal for high-volume workloads.
Recommended
The best balance of accuracy and cost — the sensible default for production.
Save as Prompt
Promote the winner to a reusable prompt in one click, then keep validating it with statistical experiments.
Auto-Refine is available on the Team plan — up to 20 optimization runs per month.
View pricingReady to Improve Your LLM Prompts?
Start testing with statistical rigor today
Enterprise SSO • Data Isolation • Cloud-Native