At mixus, three engineers and one lawyer spent two years working toward a result an attorney could send to a client without fixing the document first.
Getting a plausible edit was easy. Getting it into Word without breaking the document took much longer. We built checks around the model's edits and a way to return work to an attorney when those checks fail.
We focused on four tasks in venture and corporate work: cap table pro formas, NVCA redlines, term sheets, and NDAs. Claude runs behind an abstraction layer right now, so the model can change while the checks and the firm's accumulated corrections stay put.
A redline updated a defined term in one clause and left the six others that reference it unchanged; a tracked change read correctly and landed twelve words away from where Word thinks it belongs. Those are the failures the loop has to catch, and each has its own check: the rule check looks for the inconsistent term, the validator checks the offset, and the run fixes the problem or returns it for review.
The task and the review
There are two separate questions: can the agent finish the task, and should its result go to a client? The inner loop works on the task: think, act, observe. The outer loop checks the result and decides whether to retry or hand it to a person. Build the harness, put a reviewer where a mistake would cost something, close the loop with tests, watch everything the agent does, and feed what happened back in. Each product runs that same loop around its own task.
The interface is email, where the lawyers we work with already discuss their documents. Send a document to the right address and a task starts. The result comes back in the same thread. Copy another lawyer on the email and they join the same conversation, and their feedback reaches the agent in that thread.
The harness and the playbook
Each product uses a shared agent engine, with its own instructions, guardrails, and output format. Every proposed edit is checked against the firm's playbook, so a required clause goes on the checklist even when the input lacked it, and when a firm's rule conflicts with a system default the firm's rule takes priority and the disagreement is logged.
Known text replacements run in code before the model starts, and the resulting document is still checked. Each decision to accept, skip, or counter a rule goes into the decisions store with a rationale.
Attorney corrections flow back through a Word plugin and a web reviewer, with a version history the firm can inspect. A firm's corrections tell us which edits its lawyers will accept, and that history makes the next review more useful.
The redline engine
Native Word tracked changes fail on details of the document structure. Word stores text across XML runs that rarely line up with sentence boundaries; position offsets shift as soon as an edit lands; formatting and footnotes can corrupt when two edits touch the same paragraph. We have rebuilt the redline engine through more than ten generations to handle those failures.
In the newest generation, one writer drafts each edit, parallel reviewers check it independently, and a single patcher fixes whatever the reviewers reject.
A second validator sends work back when there is too much redline, because an attorney reviewing a hundred documents should not have to inspect changes that add no substantive value.
What we can measure
The cap table pro forma gives us a result we can compare across runs: a financing model built from a cap table export and the deal terms.
We ran it on a real cap table export, the kind of input a firm already has on hand, and checked the result against a lawyer's own independent hand calculation. Five consecutive runs produced identical price per share, identical pool increase, and identical fully diluted total, landing within 45 shares of that hand calculation, a rounding-level gap from share-rounding conventions rather than model drift. Every model ships with eight financial identity checks the agent reports on itself, and a live-formula verification sheet a lawyer can open and watch recalculate if any input changes. The model's only job is translating a messy input into a fixed schema. It never touches the math.
The system can propose a fix from recorded failures, apply it in an isolated worktree, score it on separate training and holdout sets, and run adversarial critique and review before the promotion gates. Our tests block known regressions and flag failures overnight. They cannot cover mistakes we have not learned to test for.
A lawyer should be able to ask why a clause changed and see the rule and reasoning behind that edit. The action trace records the steps of the run, and each redline carries a receipt linking it to a rule and a rationale.
Two things here are still open, not solved. Is there real ground truth to check the cap table pipeline against, beyond a lawyer's hand calculation? Not yet; the check above runs against that calculation, not a machine oracle. Will a rule hierarchy that resolves cleanly today keep resolving cleanly as a firm's playbook grows past what one reviewer can audit by reading it top to bottom? We don't know yet.
Technical appendix: the engine
# the outer loop, once per request while not sendable: # inner loop: think, act, observe proposal = agent.run(task, playbook, context) checks = validators.run(proposal) # deterministic gates if checks.passed: return review(proposal) # human checkpoint if checks.patchable: task = patch(task, checks.fails) # fix, retry continue return review(checks.report) # stop, show receipts
The engine runs Python, Bash, and Node code and renders PDFs to inspect its output. Long jobs can exhaust the context window before the document is finished. We clear stale tool results once a run passes about 50,000 tokens while keeping the reasoning still needed for the task. Turn limits and stall detection stop runs that keep repeating work. Retries use backoff, and we record token usage per run.
Some jobs run for hours, occasionally into a second day for a full NVCA suite across five interlocking documents. Dedicated workers take them from a durable queue. The workers run LibreOffice at the OS level; in our batch document editing, that was roughly 144 times faster than round-tripping the same files through a hosted conversion API. Retrieval uses a separate vector store and object storage for each firm. We have SOC 2 Type II and ISO 42001 and have passed security review at firms with five thousand people.
Rules and document edits
A law firm's institutional knowledge lives as a set of machine-readable rules the engine executes directly, rather than as free text buried in a prompt. Put the same rules in a prompt instead and the model has to re-derive the hierarchy from natural language on every run, with no guarantee it resolves the same conflict the same way twice. As code, each rule carries an intensity. Required means the clause must exist; the agent adds it if missing. Preferred means the agent pushes for that position but adapts to context. Preferred if addressed is a softer preference: it will not override reasonable language already on the page.
Rules resolve through a strict hierarchy. Explicit instructions from the user come first, then the firm's own playbook, then system defaults for the document type, then the agent's general legal knowledge. The NVCA suite, the stock purchase agreement, the charter, and three rights agreements bundled with it (investor rights, right of first refusal, voting), is the clearest example of why that hierarchy has to be mechanical. Five interlocking documents, thousands of pages of context once a firm's playbook layers on top, and a single defined term can reach across every one of them. When a firm's rule contradicts a system default on the same clause, the firm's rule wins. The disagreement gets logged with a rationale, not silently overwritten.
A document text index extracts dates, amounts, and quoted text from the XML before the model runs. Batch pre-validation checks the targets before editing starts and fails the run if too many are invalid. Cross-run replacement maps a target back to the underlying XML nodes. Edits apply from the end of the document toward the start so an earlier edit does not shift the later targets. Matching falls back from exact to normalized text, then a Bitap diff, then n-gram similarity to handle OCR noise and typography differences. Validation checks the content and legal risk, section compliance, formatting, and rendered pages, with a separate table-integrity check.
Calculation and test cases
The pipeline separates judgment from arithmetic across five layers. A chat model generates the instructions. Code-level detection fixes the one parameter that moves the result more than any other: whether the option pool target applies to available or total shares. That decision never reaches the model. A constrained inner agent reads the workbook and fills a strict JSON schema: shareholders, investors, terms. A code validation layer enforces integer shares, normalizes percentages, applies the pool-type override, and checks the investor split. Only then does an algebraic closed-form solver compute the result: no goal-seek, no circular references, the same inputs producing the same outputs every single time.
We record test cases with an input document, an optional expected answer or playbook, and the behavior to check. Different graders check whether an edit stays small, matches known changes, follows the playbook, or cites the rule behind it. NVCA tests also score the document against its checklist. A score below its recorded baseline counts as a regression.