How Luminance selects AI models for Contract Intelligence

Luc Frachon and the Luminance Evals Team

LLM providers release new models at an increasing pace, and it becomes harder and harder to know which model to use.  This makes model selection a demanding engineering problem, but also an exciting one. We are fortunate to have so many capable models competing to solve increasingly complex problems. The challenge is understanding where each model performs best, and how those strengths translate into real-world usefulness. A model optimized for an academic benchmark may not be the best choice for a specific contract task. Model evaluation must therefore be grounded in realistic documents, realistic workflows, and the use cases that matter to users. 

Over the last two weeks, OpenAI and Anthropic have released four major models: GPT 6 Sol, GPT 6.1 Sol, Claude Sonnet 5.5, and Claude Opus 5.5. This is the perfect opportunity to illustrate our testing ethos while providing some real insight about their respective strengths and weaknesses. 

In this blog, we explain our general philosophy around AI model evaluation and selection, and provide a worked example based on a recent batch of new AI model releases. 


Why Luminance takes a model-opinionated approach 

The Luminance Platform covers many legal workflows. For instance, reviewing the Repository might involve finding agreements, organizing information, surfacing and explaining contract details, etc. In Negotiation, Luminance’s AI might be asked to compare a clause with a preferred position, redline positions, find precedents, etc. There is no one-size-fits-all approach to model selection across such a wide range of tasks. The right model for a document retrieval task may not be the right model for contract interpretation, drafting, negotiation, or portfolio analysis. 

Some software providers make a point of letting the user choose between all new models. At Luminance, we don’t believe in this approach.  

Choosing the right model for a given task is among the hardest problems in the industry and requires detailed knowledge of every feature and model. We find it unfair to the customer to leave the decision to them.  

As we are not pricing by usage, our customers are guaranteed optimal performance and consistent pricing. This is what we call being model-opinionated by design: Luminance selects the model best suited to each Contract Intelligence task, rather than asking customers to become experts in model selection themselves. 

The models are available to everyone. The judgment is not. 


How we evaluate AI models for contract work  

Our evaluation framework is based on real documents and real tasks a user would ask the product to perform. We constantly evolve it to cover the new features developed by our engineers to ensure that evaluation metrics closely correlate with actual user experience. It makes this framework a lot more informative and grounded than generic legal benchmarks.  

Every model is assessed over contract workflows for legal accuracy, usefulness, tone, speed, and consistency, on the same set of tasks and against the same evaluation criteria. There are two main categories of tasks that we assess: Negotiation and Repository, with subcategories in each.  

We compare each LLM to two established reference models for which we have reliable metrics and report accuracy and speed over multiple repeats. This allows us to assess whether a new model performs well in isolation and whether it represents a meaningful improvement for a particular task. 

How this cohort performed: AI model evaluation results across Repository and Negotiation  

Figure 1. Accuracy and time differences versus reference models (GPT 5.6 Terra and GPT 5.4), with higher accuracy scores indicating stronger performance.

Figure 1. Accuracy differences from reference models (GPT 5.6 Terra and GPT 5.4, respectively), with higher values indicating better performance against our assessment criteria. Time is also relative to reference models. 

Our latest test campaign pitches four recently released models against each other: OpenAI’s GPT 6 Sol and GPT 6.1 Sol, and Anthropic’s Claude Sonnet 5.5 and Claude Opus 5.5.  

  • On Repository tasks, Opus performed best overall, followed by 6.1 Sol, 6 Sol, and Sonnet 5.5.  

  • On Negotiation, Opus came first again by a short margin over Sonnet. However, the differences with the reference model are much narrower than on Repository. 

Completion times showed large differences, ranging from 0.7x to 1.6x in the reference models’ times. These differences are very noticeable to the user and must be considered when selecting the right model for the job. 

Going into more detail, we see interesting patterns emerge. On Negotiation, every tested model brings a clear uplift in document retrieval, a mechanical, agent-calling task, with 6.1 Sol improving by 20pp. Results are less consistent on interpretation and understanding questions such as market standards or drafting and amendments. The other lesson here is that bigger is not always better, with Sonnet 5.5 punching above its weight in most categories.  

Figure 2. Negotiation accuracy differences versus reference models, showing performance across different types of negotiation work.

Figure 2. Negotiation accuracy differences from reference models, in percentage points. This breakdown helps us make a more specific choice than simply selecting the highest overall score. We can see which kinds of work improve, and whether those improvements fit the way customers use each part of Luminance. 


How Model performance compounds throughout a task 

The Luminance Repository view enables users to ask a question across their entire repository and get detailed per-document answers. For example, they might ask, “Analyze my termination rights over all my master service agreements.”  The agent decides what columns to generate and writes column descriptions. Those are then used by a subagent to populate the columns against each document. 

Figure 3. Repository preparation accuracy differences versus the reference model. Opus performed particularly strongly in column selection and description quality.

Figure 3. Repository preparation, shown as accuracy differences from reference model in percentage points. Our metrics show each stage independently. Opus was particularly strong at choosing the correct columns, as well as writing their description. Sonnet and 6 Sol were less impressive on column creation, but when they chose the right columns, their descriptions were still an improvement. 

An interesting effect is that the next stage, cell population, can be significantly improved without changing its own model, just from the better column descriptions generated by the top-level agent. This shows that choosing the right model can bring benefits to a whole sequence of agentic tasks, not just the specific steps where it intervenes. In a well-designed AI product, improvements compound over a workflow, and better performance in the early steps provides the user with better accuracy over the whole system. 


What AI model selection means for customers 

A lot more goes into proper model testing than just running them on a public benchmark.  

For Contract Intelligence, the relevant question is not “Which model is best?” in the abstract. It’s “Which model is best for this task, in this workflow, using this contractual context?” 

From this round of testing, we learned that Claude Opus 5.5 looks particularly strong as the Repository’s main agent, while Claude Sonnet 5.5 seems to be a good choice for some Negotiation applications. The GPT Sol models deliver consistently good performance but come with a speed impact. Final validation by our legal experts will confirm which of these models are worth deploying into production. Our users benefit from these choices automatically without having to worry about cost. 

This funnel supports Luminance’s model-opinionated approach: the platform selects the right model for each Contract Intelligence task, combining frontier models with purpose-built legal intelligence, contextual retrieval, agentic workflows, and differentiated validation.   

The model can change. The standard does not. 

For customers, this means that advances in the wider AI ecosystem can improve the product without creating another model-selection problem for the enterprise. Luminance continuously evaluates the field, routes each task to the intelligence best suited to it, and validates the result against the standards required for contract work. 

GET STARTED

Get more from your contracts

See how Luminance can help your teams negotiate smarter, surface what matters, and keep contracts moving across the enterprise. 

GET STARTED

Get more from your contracts

See how Luminance can help your teams negotiate smarter, surface what matters, and keep contracts moving across the enterprise. 

GET STARTED

Get more from your contracts

See how Luminance can help your teams negotiate smarter, surface what matters, and keep contracts moving across the enterprise.