Building and evaluating legal AI requires a holistic approach that integrates domain-specific expertise with rigorous, multi-layered testing. Harvey, a platform for legal professionals, addresses the inherent complexity of legal documents—which often feature dense, non-standard formats—by employing a "lawyer-in-the-loop" development model. This strategy ensures that human judgment remains the primary quality signal, supplemented by automated benchmarks like BigLawBench and model-based evaluations. Effective evaluation involves breaking down multi-step agentic workflows into discrete, verifiable components. Beyond technical metrics, success relies on qualitative feedback and professional taste, as legal accuracy often hinges on nuanced interpretations rather than simple binary outcomes. Future advancements in agentic systems will likely depend on capturing tacit "process data"—the unwritten playbooks and institutional knowledge that define how complex legal tasks are executed in practice.
Sign in to continue reading, translating and more.
Open full episode in Podwise
