PRICING

Your whole app, booted and attacked, once a week.

Every plan is priced in runs. Start with one a month and pay nothing. Unlimited people on every paid plan, because charging per seat would mean charging you for the person who reads the report.

WHAT A RUN IS

We build your repository into a sandbox and start it the way your own machine would. Then we write customers who want the things your customers want, hold real conversations with your running app on your own model key, and grade every reply against the rules you wrote. You get back what broke, where in your code it broke, and a change that closes it. Nothing is written to your repository unless you ask for it.

Free

$0forever

Enough to see your own app break once, on your own code, before you decide anything.

Connect your repo
  • 1 run a month
  • 1 app
  • The full report and every fix it proposes
  • 1 fix written and proved against your own tests
  • Cases written from your own code and content
  • Results kept 7 days

  • 25 questions to Sandy a month
  • No check on your commits

NO CARD ASKED FOR

Ship

$99a month

Find the behavior that costs you customers. Test a repair and keep the proof.

Start Ship
  • 10 runs a month
  • 2 apps, unlimited people
  • 50 questions to Sandy a month
  • 30 repair stages, tested against your own suite
  • Behavior Memory: production logs become test cases; your ratings shape your judge
  • A check on every push to your main branch, written back as a commit comment
  • Results kept 30 days

  • Cannot stop a merge
  • Sandy does not open pull requests

ONE RUN A WEEK, EVERY WEEK

Prove

MOST TEAMS

$990a month

Make behavioral checks part of release review, with verified fix pull requests from Sandy.

Start Prove
  • Your dedicated senior engineer. Someone who learns your business and takes ownership of improving your AI.
  • Weekly working calls. Review what broke, agree on priorities, and track the fixes together.
  • 30 runs a month, 3 at once
  • 5 apps, unlimited people
  • 200 questions to Sandy a month
  • 120 repair stages, tested against your own suite
  • Behavior Memory: learn from production, each run, and your team’s ratings
  • A merge gate on every pull request. A behavior your app already had cannot be lost quietly.
  • Sandy opens the fix as a pull request on a branch of her own
  • Production drift alerts
  • Your own test suites and journeys run beside ours
  • What changed between two releases, side by side
  • Results kept 90 days

NOTHING SHIPS UNTESTED

Past the plan

Nothing stops on its own. When a plan runs out you are asked once, and you decide whether to keep going at these rates.

WhatOn ShipOn Prove
One more run$14$11
Ten more runs, bought together$119$95
A pull-request check, scoped to the diff$4.67$3.67
Production learningBehavior MemoryBehavior Memory
Your evaluation standardsJudge shaped by your team’s ratingsJudge shaped by your team’s ratings
100 more questions to Sandy$25$20
One more app$29 a month$19 a month
One more run at the same time$49 a month$49 a month
Keep results for 180 days$39 a month$39 a month
Keep results for a yearProve only$99 a month

Bigger than this

From $1,500 a month: single sign-on, SCIM, roles, an audit log, the runner inside your own cloud, unlimited apps, your own retention and residency, a signed agreement and our SOC 2 report, a named engineer, and committed run volume at $8 a run.

Talk to us

Short answers

Whose model key does a run spend?

Yours. Your app answers on the key you sealed, so the replies we grade are the replies your customers would get. We pay for the grading, the customers we write, and the plan.

Can you see my code or my keys?

Your environment is opened inside the sandbox and nowhere else. The service that unseals it holds no copy and we hold no key. Your repository is read, never written, unless you press a button that says it will write.

What happens when I run out of runs?

The next run asks first. Nothing runs on a card you did not expect, and unused runs do not roll into next month.

Do you charge for a run that failed to start?

No. If our sandbox could not boot your app, that run costs you nothing. If it booted and your app broke, that is a result and it counts.

How long does a run take?

Between four and twelve minutes on the apps we have measured, most of it the first build. After that your app is restored from a snapshot in seconds.

What is Behavior Memory?

A growing record of how your users and your AI behave. Categorize real requests and replies, turn recurring problems into test cases, and track what improves or regresses across runs. Use that evidence to test the next release against the situations your customers actually face.

A judge with your standards.

Rate answers and explain what makes them great or unacceptable. Those expert reviews build a rubric for your product and shape how Sandy judges the next run. Each review adds examples of what your team values, so the judge can distinguish a polished answer from one that actually gets the job done.

Do I need to write tests first?

No. The cases are written from your own prompts, tools and code. If you already have an eval suite, Prove runs it beside ours.

Can I cancel?

Any time, from your billing page. You keep the current month and nothing renews.