Bonsai — a bonsai forming a checkmark inside a gold enso circle
Bonsai self-improving eval harness
See it live →

We grow trust, not just models.

A self-improving eval harness: it watches an agent fail, mints a check to catch that failure forever — and shows it stayed honest.

1Catcha failure — a claim with no real support in its source
2Clustersimilar failures by meaning — Atlas vector search
3Mintone general check that catches the whole class
4Growthe rubric improves itself

The rubric is the harness — each branch of the bonsai is one check.

The honesty rail

A build-failing test bars the loop that writes the checks from referencing the answer key it’s graded against — a guard against the obvious leak, not a sandbox. Improvement is reported as direction + a confidence interval, never a bare %.

Where it goes

Today it’s a SaaS vendor’s trust & security copilot, answering prospects’ security questionnaires with citations. The same loop owns the contract for any agent — a video or image generator, a contract analyzer, a support bot. A domain expert defines what “good” means; Bonsai grows the harness. The mechanism doesn’t change — only the contract does.

Watch a check get born →

Even the brand is Google AI: the Bonsai mark was drawn with Nano Banana and its animation with Veo — a wink, not the build; the product itself runs on Gemini 3.5.