We grow trust, not just models.
A self-improving eval harness: it watches an agent fail, mints a check to catch that failure forever — and shows it stayed honest.
The rubric is the harness — each branch of the bonsai is one check.
The honesty rail
A build-failing test bars the loop that writes the checks from referencing the answer key it’s graded against — a guard against the obvious leak, not a sandbox. Improvement is reported as direction + a confidence interval, never a bare %.
Where it goes
Today it’s a SaaS vendor’s trust & security copilot, answering prospects’ security questionnaires with citations. The same loop owns the contract for any agent — a video or image generator, a contract analyzer, a support bot. A domain expert defines what “good” means; Bonsai grows the harness. The mechanism doesn’t change — only the contract does.
Even the brand is Google AI: the Bonsai mark was drawn with Nano Banana and its animation with Veo — a wink, not the build; the product itself runs on Gemini 3.5.
