Harvey opens 10 legal-agent environments with up to 80M tokens
Harvey is open-sourcing 10 environments where AI agents must perform real due-diligence work across huge virtual data rooms. The largest contains 80 million tokens, and Harvey says reviewing one of these rooms manually could consume $300,000 to $1 million in lawyer time. The bigger signal is that legal AI is moving beyond chat answers and into long, testable workflows where one missed clause can change an entire deal.
We're open sourcing 10 diligence RL environments to expand @harvey's legal agent bench:
— Gabe Pereyra (@gabepereyra) July 17, 2026
- The largest one is 80M tokens
- Human review of a dataroom costs ~$300K-$1M lawyer time
- Each one has 100-1000 unit tests for automated evaluation
This dataset will allow us to train… https://t.co/0yJ0rKrrfI pic.twitter.com/NgNojvSlZn
Q1What actually happened?
Harvey released 10 open-source reinforcement-learning environments for legal due diligence. Each environment gives an AI agent a data room, a set of legal tasks and roughly 100 to 1,000 tests that check whether its work is correct. Harvey says the largest room contains 80 million tokens.
Q2What is a diligence environment?
Think of it as a practice deal for an AI lawyer. The agent must search large collections of contracts and company files, find the right evidence, connect facts across documents and produce a useful answer. It is much harder than asking a chatbot one legal question because the agent has to manage an entire workflow without losing important details.
Q3Why does the 80 million-token figure matter?
Because this is far beyond a neat document that fits comfortably inside a normal prompt. An 80 million-token data room may contain thousands of files, repeated versions, hidden risks and conflicting clauses. The challenge is not simply reading the text. It is finding the few facts that could change the price, structure or risk of a deal.
Q4Why mention $300,000 to $1 million in lawyer time?
That number shows the economic value of the workflow Harvey is trying to reproduce. Large transactions can require teams of lawyers to review documents, flag unusual terms and build detailed reports. Harvey is not claiming that one agent has already replaced all of that work. It is turning an expensive human assignment into something AI systems can repeatedly practise and be tested on.
Q5How is this different from older legal benchmarks?
Many older benchmarks ask models legal questions and compare their final answers. Newer projects are becoming more interactive, but they remain difficult. A 2025 benchmark called J1-ENVS used six legal scenarios and found that even its strongest tested agent scored below 60%. Harvey is pushing the idea toward much larger transaction rooms with hundreds of concrete checks per task.
Q6Why would Harvey give this away?
Because better public environments can attract researchers, reveal where agents fail and speed up training techniques that Harvey can later use. It also helps define what good legal-agent performance should look like. The strategic bet is that Harvey can benefit from making the training ground open while keeping its customer data, product workflow and production system private.
Q7So what is the real signal?
Legal AI is leaving the simple chatbot phase. The next competition is about agents that can work through massive files, use tools, preserve evidence and survive hundreds of tests. Models will still matter, but the training environments may become just as important. Harvey is effectively opening a gym for legal agents, built around work that previously required a very expensive team of humans.
