Bailey is a whole firm of AI agents — intake, delegation, drafting, review, and the business around it — that runs behind gates, budgets, ethical walls, and hash-chained receipts. Not a chatbot. A firm, with controls — where every step is on the record and nothing leaves without a human.
the walled courtyard of a castle, where the work happens: inside the walls, behind the gate.
Building a good agent is hard — and it's only the beginning. Before an agent can touch legal work, it needs walls, budgets, gates, and receipts around it. Bailey wraps every agent — whatever model is behind it — in the controls a law practice actually needs.
Work is handed to a firm, not a chat window. Every request becomes a matter with an owner; every handoff between agents is a sub-matter you can open, read, and audit.
Every agent runs under a hard budget cap it cannot exceed, with every token metered — per agent, per matter, per company. The spending stops before you get surprised.
Outbound actions — a court filing, an e-signature, a document upload — stop at an egress gate and wait, hash-only, for a human to approve or reject. No exceptions, no override.
Every gated decision writes a hash-chained receipt you can verify offline with openssl — plus optional RFC 3161 timestamp anchoring. The spec is public; a standalone verifier ships in the repo.
Ethical walls run as separate companies with separate gates and receipt chains. A screened matter says it is screened — denied is never dressed up as nonexistent.
Create new agents with custom skills, connectors, and tools — shaped to your business and your practice of law. Swap the models underneath: cheaper where it's routine, stronger where it counts.
A layer, not a fork: Bailey packages a complete legal operation on top of the open-source paperclip agent runtime — pinned, never modified. The governance pieces are standalone and portable.
Bailey began as a bet that atomic work, decomposed and reassembled by an orchestrated firm, would beat one strong model. On the affordable, non-frontier tiers we tested, it didn't clear that bar: on Harvey's public LAB benchmark — long-horizon graded tasks, scored by Harvey's own grader — orchestration showed a higher ceiling (every top-scoring run in the campaign was orchestrated) but an equal median. What did hold is the governance layer. The gates, budgets, walls, and receipts work identically no matter which model is underneath.
So that's how you use Bailey: keep building and keep testing. Improve the skills. Improve the tools. Swap in stronger models as they arrive, and re-run the harness — it ships in the repo, with all of our data, so your next result replaces ours.