PromptProof
Auto-Refine

Automatic Prompt Optimization, Proven on Your Data

Give Auto-Refine your labeled data and a starting prompt. It explores dozens of prompt variants, measures the accuracy of every one, and hands you the top 3 candidates — ranked by accuracy and cost.

How Auto-Refine Works

From labeled data to a production-ready prompt in four steps

01

Choose Your Labeled Data

Pick a folder with 20+ labeled items — images, PDFs, or text. The output schema is derived automatically from your data.

02

Set the Task and Models

Write a starting prompt and choose the model to optimize. Sensible defaults cover the rest — adjust budget and targets only if you want to.

03

Run Automatic Optimization

Auto-Refine generates and evaluates prompt variants against your data in the background, learning from failures to write better instructions each round.

04

Pick a Winner and Save

Review the Top-3 candidates — most accurate, cheapest, and recommended — and save the winner as a reusable prompt in one click.

Measured Improvement

Accuracy You Can See — and Trust

Every prompt variant is evaluated on your own labeled data with statistical confidence bounds — so improvements are real, not lucky.

Tested on Your Data

Accuracy is measured against your ground truth labels, not synthetic benchmarks.

Improvement Over Time

Watch accuracy climb round by round with an improvement chart and live progress.

Every Variant, Compared

Browse every explored prompt with its accuracy, cost, and lineage — nothing is a black box.

Images, PDFs, and Text

Optimize extraction prompts for multimodal inputs, not just plain text.

Cost Control

Better Prompts Without Runaway Costs

Stay in control of spend from start to finish: see the estimate before you launch, cap the budget, and stop early the moment your target is reached.

Upfront Estimates

See the estimated cost and duration before you start — no surprises.

Hard Budget Cap

Set a spending limit in USD. Optimization stops automatically before exceeding it.

Stop at Good Enough

Set a target accuracy and Auto-Refine stops as soon as it is reached — don't pay for perfection you don't need.

Cancel Anytime

Stop a run whenever you like and keep every variant found so far.

Top-3 Candidates

Pick the Trade-Off That Fits: Accuracy, Cost, or Both

Auto-Refine surfaces three winners so you can choose the trade-off that fits your use case.

Most Accurate

The highest measured accuracy — for when every extraction has to be right.

Cheapest

Nearly the same accuracy at a fraction of the token cost — ideal for high-volume workloads.

Recommended

The best balance of accuracy and cost — the sensible default for production.

Save as Prompt

Promote the winner to a reusable prompt in one click, then keep validating it with statistical experiments.

Auto-Refine is available on the Team plan — up to 20 optimization runs per month.

View pricing

Ready to Improve Your LLM Prompts?

Start testing with statistical rigor today

Enterprise SSO • Data Isolation • Cloud-Native