In this article, we’ll cover the main models Harvey uses and how they compare.
Last updated: Aug 22, 2026
Overview
To help you maximize your impact, Harvey leverages different advanced AI Large Language Models (LLMs), each designed with unique strengths. When you ask Harvey for assistance, our multi-model system will break down the request into sub-tasks, select a model to use, then synthesize the outputs.
By default, Harvey’s Auto mode will select a model for you. The current models we use are a mix of the following:
Anthropic Sonnet/Opus 4 model suite
Anthropic Sonnet/Opus 5 model suite
OpenAI GPT-5 model suite
OpenAI o3 model suite
Google Gemini 2.5 Pro model suite
If Model Selection has been enabled in your workspace, you will have the option to choose from specific models before running a query, including opt-in options such as:
Refer to our Model Comparison table to view your model selection options.
Our Model Evaluation Process
Harvey’s model evaluation methodology is comprehensive so that we understand not only raw model performance, but also the safety and reliability of each model before including it in our system. The pillars of our model evaluation are BigLaw Bench (BLB), Legal Agent Benchmark (LAB), product performance, and unstructured evaluation.
Our benchmark that measures the ability of LLMs to complete real-world legal tasks. It evaluates both general-purpose LLMs and Harvey’s specialized agentic systems using detailed, lawyer-designed rubrics that score answer quality (accuracy, completeness, legal reasoning) and source reliability (verifiable citations to legal documents).
This approach ensures the models are judged not just on linguistic output, but on their ability to perform trusted, billable legal work with traceable, high-fidelity results. To learn more, read our blog article on Introducing BigLaw Bench.
Our open-source benchmark that measures how well AI agents complete longer, multi-step legal work from start to finish. Each task mirrors real firm work: the agent receives a short instruction and a client matter of relevant and peripheral documents, then must produce reviewable legal work product.
Outputs are graded against expert-written rubrics using "all-pass" grading, where a task is complete only if every criterion is met, reflecting how high-stakes legal work is reviewed in practice. LAB assesses model performance against 1,700+ tasks across 20+ practice areas. To learn more, read our blog article on Introducing Harvey's Legal Agent Benchmark.
To view LAB scores for each model, view the LAB Leaderboards maintained by Vals and Artificial Analysis.
In a test environment, we integrate the model into key product surfaces–Assistant, Vault, and Workflows–and measure system performance through both human preference and traditional machine learning metrics (accuracy, precision, recall, hallucination rate, latency, and more) over product datasets.
Researchers test the model in open-ended, less predictable ways to uncover negative features such as toxicity or bias, as well as identify major, sudden improvements in reasoning ability.
After we evaluate a model, we revisit the evaluation throughout its lifecycle to ensure we’re offering our users the best functionality.
Model Comparison
To help you navigate the models we offer, we’ve put together a high-level comparison table. Use the horizontal scrollbar at the bottom of the table to view all columns, including availability by region.
Note: If model selector is enabled in your workspace but you’re not seeing a particular model, it may not be available for your region. Ask your Customer Success Manager to confirm what’s available to you.
Model
Developer
Model Release Date
Strengths
Weaknesses
Knowledge Cut-off
Regional Availability
GPT-5 (reasoning)
OpenAI
August 7, 2025
Analysis detail and quality
Hard problem solving, particularly long-form writing and agentic behavior
Can overthink, providing overly complicated answers to straightforward problems
Formatting, particularly structured use of headers and lists
September 2024
US
EU
AU
GPT-5.4
OpenAI
March 5, 2026
Gets straight to task at hand
Intuitive structure and appropriate detail level
Pending further evaluation
August 2025
US
EU
GPT-5.5
OpenAI
April 24, 2026
Substantive accuracy
Organizational structure
Audience calibration
Pending further evaluation
December 2025
US
EU
GPT-5.6 Sol
OpenAI
July 9, 2026
Substantive accuracy
Document review and analysis
Pending further evaluation
February 2026
US
EU
Gemini 2.5 Pro
Google
March 25, 2025
Drafts longer, detailed outputs
Multi-step analysis and outputs
Can overthink straightforward problems
Occasional misplaced or stiff tone
February 2025
US
EU
Claude Sonnet 4.5
Anthropic
September 29, 2025
Long-context reasoning
Numerical reasoning
Pending further evaluation
January 2025
US
EU
AU
Claude Sonnet 4.6
Anthropic
February 17, 2026
Structured workflows
Instruction-driven tasks
Efficient problem-solving
Pending further evaluation
May 2025
US
EU
AU
Claude Sonnet 5
Anthropic
June 30, 2026
Legal accuracy
Drafting
Cost efficiency
Pending further evaluation
January 2026
US
EU
Claude Opus 4
Anthropic
May 22, 2025
Strong formatting and clarity
Grounding responses in underlying documents
Occasionally misses key factual or legal elements or over-simplifies
Sometimes too rigid in formatting
Tasks requiring specific formats
January 2025
US
EU
Claude Opus 4.5
Anthropic
Nov 24, 2025
Excels at agentic tasks that require planning and iteration
Pending further evaluation
May 2025
US
EU
Claude Opus 4.6
Anthropic
February 5, 2026
Excels at deep research, nuanced reasoning, and analytical tasks
Pending further evaluation
May 2025
US
EU
AU
Claude Opus 4.7
Anthropic
April 16, 2026
Improved reasoning calibration
Strong performance on ambiguous tasks
Substantive accuracy
Pending further evaluation
January 2026
US
EU
Claude Opus 4.8
Anthropic
May 28, 2026
Legal accuracy
Review-and-revise drafting
Tone and length calibration
Pending further evaluation
January 2026
US
EU
AU
Claude Opus 5
Anthropic
July 24, 2026
Transactional and disclosure-heavy work
Litigation work
Pending further evaluation
May 2026
US
EU
AU
Fable 5
Anthropic
July 1, 2026
Drafting and markup analysis
Multi-document consistency tracking
Long-horizon agentic task performance
Pending further evaluation
January 2026
US
EU
AU
Mistral Medium 3.5
Mistral
June 3, 2026
Analysis and review
Research and extraction
Pending further evaluation
October 2024
EU
Note: Learn more about regional data processing in the FAQs section.
Looking forward, the landscape of AI technology is continuously evolving, and so is Harvey. We are constantly evaluating the latest models and their performance to ensure we are always providing the most advanced and effective solutions. Stay up-to-date on model availability and feature enhancements by following our Release Notes. Customers can also discover additional technical details on topics like security, data handling, and prompt caching in the Harvey Trust Center.
FAQs
Terminology
Generative AI refers to a class of artificial intelligence systems designed to generate new content, ideas, or solutions based on patterns and data it has learned from. It uses machine learning models to create text, images, music, and even code or designs, mimicking human creativity. Unlike traditional AI, which focuses on recognizing patterns or making decisions based on existing data, generative AI can produce novel outputs, often by learning from vast datasets.
When you input information into Harvey, or when Harvey generates a response, it's processed in tokens. A token is the smallest unit that a model can understand. Tokens have specific statistical properties that make them better to use than just using words.
Think of a "context window" as the amount of information an AI model can 'see' and process at one time. When you're working with Harvey on a document, the context window determines how much of that document the model can consider to generate its response.
A larger context window means the model can review more text at once, leading to more comprehensive and accurate analysis, especially for lengthy and complex documents like credit agreements.
Yes, LLMs are trained on data up to a certain point in time, known as a "knowledge cutoff date." This means the model's knowledge of events or information beyond that date may be limited. We recommend always querying with attached files and legal research knowledge sources so that your responses are high-quality and grounded in context supplemental to the model's knowledge.
If you select a model via the Model Selector, the knowledge cutoff would depend on the model used in response to a query. Most of the models we support have similar knowledge cutoffs.
When you use Harvey, we combine AI's broad knowledge with specific, real-time information from your documents or trusted data sources, to optimize performance.
If you have model selector enabled, you’ll notice we note whether each model is best for complex or general-purpose tasks.
General-purpose tasks: Summarizing and analyzing documents, writing, ideation, and more everyday work.
Complex tasks: Multi-step reasoning, long-form drafting, complex workflows, and more.
Models tagged for “complex tasks” are reasoning models that take more time to think before answering — this usually lets them perform better on tasks that are at the edge of model capabilities. They may be slower to process, and sometimes overthink the easy things.
If you’re not sure where your task falls, choose Auto — Harvey will line up your task with the best model for the job.
Security, Privacy, & Data
Harvey is built by and for the world's best legal and financial companies, and data security is paramount. We engineer our platform with robust privacy and security measures. Sensitive data is handled with the utmost confidentiality and is not used to train the underlying public AI models, nor is it shared externally without your explicit permission. We prioritize practical solutions that honor established practices while using technology to enhance professional work.
Harvey accesses models through a combination of the model developer's own API and third-party cloud hosting platforms, which allows us to offer regional processing options that may not be available directly from the model developer.
For example, for workspaces with European processing elected, Customer Data and Content is processed within Europe, including for models where European processing is only available through a hosting platform rather than the model developer's API directly.
Performance
While AI models are incredibly powerful, they’re designed to be sophisticated reasoning tools. Sometimes, models might provide a very comprehensive response that seems to "overthink" a simple query. This occurs because these models are built for deep reasoning on hard problems.
If you encounter an unexpected response, consider rephrasing your question or, if available for the task, trying a different model that is optimized for more succinct answers.
Updates to the models powering Harvey occur regularly, aimed at continuous improvement and enhancing your experience without requiring you to manage complex technical configurations.
Although most users will benefit from Harvey's default selection, power users might like the option to choose themselves due to personal model preferences. You may find that you prefer the results or tone of certain models depending on the use case.
Not at this time, our multi-model use makes it challenging to break this down clearly per query, however this request is on our development list for future consideration.
Harvey specializes in complex professional tasks, particularly in legal and tax workflows. This means that model results and built-in features are optimized and benchmarked against daily legal workflows, with testing done by our internal bench of lawyers to ensure high-quality results.
No, Harvey is built so you can work with simple, clear instructions and minimizes the need for detailed prompt engineering. We also provide example libraries and features to automatically refine prompts for better outputs.
Think of Harvey as a digital associate—a fast and effective thought partner and generator of first drafts. While the output requires verification, similar to the work of a junior colleague, it saves significant time and enhances overall quality. Harvey also allows you to save favorite prompts for future use, expediting workflows and saving you time.
For more on getting the most out of your prompts, visit our Prompt Writing article.
Excessive use may result in temporary throttling of Customer's access to the Service.
Hallucinations are a type of error that occurs when LLMs (Large Language Models) generate inaccurate or fabricated information. This can happen because LLMs are trained on large and diverse corpora of text, but may lack sufficient domain knowledge or logical reasoning.
Harvey minimizes hallucinations by using domain-specific models and knowledge bases:
Domain-Specific Models: These models are trained on large datasets of specialized legal documents, capturing the nuances and complexities of the law. This narrows the gap between general-purpose models, which lack domain expertise, and human experts with specialized knowledge.
Knowledge Bases: Many of Harvey’s tools use domain-specific resources, such as statutes, case law, and legal ontologies, to ground responses in authoritative sources. For legal tasks, Harvey references selected case law and statutes, and when documents are uploaded, the output will include citations from the user-provided materials.
Technical
What are the models trained on? What makes Harvey's models "fine-tuned" or "tailored" for legal work? We’ve built our own Reinforcement Fine-Tuning (RFT) processes and invest in several areas that go beyond training a single model:
Composable systems: Harvey combines many small model “calls” together. This means a single query might use dozens—or even hundreds—of steps behind the scenes to give you a complete answer with supporting references.
Multi-model strategy: We partner with more than one AI provider. Each model has different strengths, so leveraging them all makes our results more accurate for different types of legal tasks, all while optimizing efficiency.
Orchestration: As our systems get more complex, we route your request to the right workflow. This helps ensure you get the best possible result without extra effort.
Harvey uses domain expertise to convert professional processes into high-quality AI agents that produce expert-quality work product:
Each request is routed to a cascading series of LLMs tuned for legal synthesis, RAG systems that incorporate public or user-provided data, and powerful reasoning models that help orchestrate the work end to end.
This allows Harvey to both reduce hallucinations and improve response output quality.
Our system is designed to break up a larger task into smaller tasks that AI models are more likely to execute correctly. With each query, Harvey infers the goal, relevant sources needed, and the desired outcome.
The type of work being completed determines the model that is used. Based on these factors, Harvey routes the query under the hood to various models to complete sub-tasks, optimizing for quality and efficiency.