A legal tech company used OpenAI's new flagship model to compress a financial-statement review into minutes, the same week OpenAI classified the model as its riskiest yet.

A financial-statement tie-out is the unglamorous grind sitting underneath a lot of legal and audit work. Every figure in a draft set of accounts gets checked against a trial balance, a consolidation schedule, and the prior year's numbers until each line agrees. Legora, a legal technology company used by more than 100,000 professionals across over 1,800 law firms and in-house legal departments, says that work "can take an entire evening, sometimes days" for one professional. Using OpenAI's newly released GPT-6 Astra model, an agent inside Legora completed a tie-out across 41 documents in a single run, in minutes, and caught all four errors the company had planted to test it, including a 500,000 pound gap hidden in a revenue note (OpenAI).

OpenAI released GPT-6 Astra on September 3, 2026, describing it as the most capable model the company has broadly deployed (OpenAI). For Legora, the practical difference showed up in one specific workflow. The company evaluated Astra against its own Benchmark for Agentic Reasoning and reported close to a 40% improvement on the financial-statement tie-out task, well above the roughly 3% average gain the benchmark showed across all legal tasks. Legora's legal engineer, Percevale Perks, credits the jump to raw processing capacity: the ability to ingest a large number of documents, hold complex figures in context, and track every line item without losing the thread.

The same release day, OpenAI published a second business case study. Playco, which builds an AI-assisted game development environment, reported 50% fewer manual fixes when prototyping games with Astra compared with the previous model, and said the model produced three distinct themed prototypes from one shared foundation in a single pass (OpenAI). Different workflow, similar shape: a task that used to require a person stepping in to fix mistakes by hand now needs less of that correction.

For a business reading past the model name, the useful signal is what moved. A first-pass review that used to consume a chunk of an evening, or stretch into days, can now surface every discrepancy in minutes, with a record of each check attached. That does not mean the review is finished. Legora is explicit that the agent handles the exhaustive comparison while the legal professional keeps the judgment call on what each flagged item means. The company frames this pattern, faster ingestion paired with a human decision at the end, as the model it is counting on as it extends beyond legal work into audit, tax, and compliance.

The same release carries a caveat that is easy to miss if you only read the customer stories. OpenAI says Astra is the first model it has classified as reaching the "Critical" level of cybersecurity capability under its own Preparedness Framework, meaning that with the right access, it can find previously unknown security flaws and build working exploits against hardened systems without a person guiding each step (OpenAI). OpenAI is restricting Astra's most advanced capabilities to a small group of testers for now, and its own safety writeup acknowledges that the model has become somewhat harder for the company's monitoring systems to fully track under adversarial conditions, even as it behaves more safely in ordinary use. None of that changes what Legora or Playco reported. It is a reminder that the same jump in raw capability that speeds up a document review does not arrive free of new risk to manage, and that one agent run on one company's workflow is not yet a general claim about how every review will go.

It is worth sitting with that contrast for a moment. The model that gave one legal team back an evening is also the model OpenAI is being the most careful about it has ever released. Both things are true at once, and the second one deserves as much attention as the case study does.