PRICING
Every plan is priced in runs. Start with one a month and pay nothing. Unlimited people on every paid plan, because charging per seat would mean charging you for the person who reads the report.
WHAT A RUN IS
We build your repository into a sandbox and start it the way your own machine would. Then we write customers who want the things your customers want, hold real conversations with your running app on your own model key, and grade every reply against the rules you wrote. You get back what broke, where in your code it broke, and a change that closes it. Nothing is written to your repository unless you ask for it.
$0forever
Enough to see your own app break once, on your own code, before you decide anything.
Connect your repoNO CARD ASKED FOR
$99a month
Find the behavior that costs you customers. Test a repair and keep the proof.
Start ShipONE RUN A WEEK, EVERY WEEK
$990a month
Make behavioral checks part of release review, with verified fix pull requests from Sandy.
Start ProveNOTHING SHIPS UNTESTED
Nothing stops on its own. When a plan runs out you are asked once, and you decide whether to keep going at these rates.
| What | On Ship | On Prove |
|---|---|---|
| One more run | $14 | $11 |
| Ten more runs, bought together | $119 | $95 |
| A pull-request check, scoped to the diff | $4.67 | $3.67 |
| Production learning | Behavior Memory | Behavior Memory |
| Your evaluation standards | Judge shaped by your team’s ratings | Judge shaped by your team’s ratings |
| 100 more questions to Sandy | $25 | $20 |
| One more app | $29 a month | $19 a month |
| One more run at the same time | $49 a month | $49 a month |
| Keep results for 180 days | $39 a month | $39 a month |
| Keep results for a year | Prove only | $99 a month |
From $1,500 a month: single sign-on, SCIM, roles, an audit log, the runner inside your own cloud, unlimited apps, your own retention and residency, a signed agreement and our SOC 2 report, a named engineer, and committed run volume at $8 a run.
Talk to usYours. Your app answers on the key you sealed, so the replies we grade are the replies your customers would get. We pay for the grading, the customers we write, and the plan.
Your environment is opened inside the sandbox and nowhere else. The service that unseals it holds no copy and we hold no key. Your repository is read, never written, unless you press a button that says it will write.
The next run asks first. Nothing runs on a card you did not expect, and unused runs do not roll into next month.
No. If our sandbox could not boot your app, that run costs you nothing. If it booted and your app broke, that is a result and it counts.
Between four and twelve minutes on the apps we have measured, most of it the first build. After that your app is restored from a snapshot in seconds.
A growing record of how your users and your AI behave. Categorize real requests and replies, turn recurring problems into test cases, and track what improves or regresses across runs. Use that evidence to test the next release against the situations your customers actually face.
Rate answers and explain what makes them great or unacceptable. Those expert reviews build a rubric for your product and shape how Sandy judges the next run. Each review adds examples of what your team values, so the judge can distinguish a polished answer from one that actually gets the job done.
No. The cases are written from your own prompts, tools and code. If you already have an eval suite, Prove runs it beside ours.
Any time, from your billing page. You keep the current month and nothing renews.