Research
Every method is public.
The numbers on this site come from runs anyone can inspect. Here is the benchmark, the code behind it and the writing that explains it.
Legal Agent Benchmark v1.0
Better answers from a cheaper model, because of the system around it.
On the legal AI leader's own public benchmark, Irys completed nearly one in three complex legal tasks in full. The leader's own agent completed about one in five.
Strict all-pass, Irys on a low-cost off-the-shelf model
The legal AI leader's own agent, on its own benchmark
Per task, across the full run
Tasks, across 27 practice areas
How the run was done
No fine-tuning, no custom training data and no domain-specific scaffolding.
Every task starts from empty state, so the run does not learn across tasks.
Strict all-pass: a task counts only when every rubric criterion passes. Across individual criteria the run reached 91.44%, and 62.7% of tasks cleared 95%.
The harness, the full run and every output are open-sourced under the MIT licence, so the numbers can be rescored independently.
Strict all-pass: a task counts only when every rubric criterion passes, with no partial credit. Legal Agent Benchmark v1.0. Irys ran the full public set of 2,010 tasks across 27 practice areas on a low-cost off-the-shelf model, with no fine-tuning and every task starting from empty state. The benchmark owner reports its own agent on a private held-out split of roughly 1,200 tasks, so the comparison is across the same benchmark family rather than identical task instances. Methodology, caveats and every output are public on GitHub.
Open source
The code is the methodology.
The benchmark harness, the complete run and every raw output, in one repository.
irys-stateful-swarms
The stateful multi-agent harness behind the benchmark result: the comparison table, the caveats and every output from the 2,010-task run. MIT licence.
Published work
Devansh, CTO and co-founder.
Built and open-sourced the reasoning infrastructure under Irys, and writes Artificial Intelligence Made Simple, a technical newsletter on how AI systems actually work.
Run it on your own work.
The benchmark is a public test. The real one is a workflow inside your own environment.
