Four Frontier AI Models Were Tested on Contract Redlines. They Rejected Only 6 to 50 % of the Edits They Should Have Rejected

Four Frontier AI Models Were Tested on Contract Redlines. They Rejected Only 6 to 50 % of the Edits They Should Have Rejected

2026.10.01 Author: Robert Nogacki

A benchmark published in June 2026 found that four frontier language models, tested at each turn of simulated multi-turn negotiations of a SaaS agreement conducted by practising attorneys and scored against attorney-authored rubrics, passed 80 to 99 per cent of the criteria that required accepting a counterparty’s redline but only 6 to 50 per cent of the criteria that required rejecting one. This Article proposes a name and an anatomy for the phenomenon of which that asymmetry is a symptom in the use of AI in contract negotiation: the AI negotiation loop, a state of bilateral, AI-assisted contract revision in which the production of a responsive draft substitutes for the improvement of the principal’s position, and in which agreement becomes the goal rather than the measure. Drawing on published research on AI negotiation, anchoring, self-correction and sycophancy, together with three recent empirical studies of contracting agents, the Article distinguishes four mechanisms (revision churn, agreement substitution, the borrowed ruler, and mandate drift), introduces the concepts of the refusal deficit, the grammatical compromise and patience at the principal’s expense, examines the consequences for the standard of care owed by counsel under Polish law and in common-law jurisdictions, and sets out a falsifiable experimental design. The literature is stated as of 1 October 2026.

 

Introduction

On 16 June 2026 the legal technology company Crosby and the research firm micro1 published RedlineBench, a benchmark in which four frontier language models were tested at each of four alternating turns of a software-as-a-service master services agreement negotiation conducted by practising attorneys: at every turn a model received the contract as the attorneys had left it, inserted its own changes as tracked edits in the document itself, and was scored against rubrics written by the attorneys. The headline finding concerns neither law nor drafting but disposition. When a rubric required the model to accept a counterparty’s edit, the models passed it in 80 to 99 per cent of cases; when a rubric required the model to reject an edit, they passed in 6 to 50 per cent. In one scenario a model accepted 98 per cent of what it ought to have accepted and rejected 6.6 per cent of what it ought to have rejected. The same report found that the models made fewer edits than attorneys but three to five times longer ones, and that they performed worst at the opening move, that is, at the decision which issues are worth contesting at all.

This Article argues that these numbers are symptoms of a condition that practitioners have observed, without a vocabulary for it, since generative models entered contract work: artificial intelligence in contract negotiation can enter a loop in which producing the next version takes the place of improving the client’s position, and reaching agreement takes the place of achieving the client’s objectives. The Article’s contribution is not the observation that models concede too readily, which the benchmark literature now documents, but an integrated account of how a bilateral, AI-mediated process detaches from its principals’ objectives while appearing, at every step, to make progress. Throughout, the Article marks the distinction between what has been measured and what is proposed as a hypothesis. In this field the distinction matters more than usual, because nothing sounds as convincing as a fluently reasoned error.

 

Version Fourteen: A Composite Illustration

Consider a composite scenario, assembled from several matters and describing none of them. Two law firms are negotiating a software implementation agreement; one acts for the vendor, the other for the purchaser. Each lawyer works with an AI assistant. The document has reached version fourteen. The vendor’s limitation of liability stood at twelve months’ fees in version three, at thirty-six months in version seven, at twenty-four in version eleven, and in version fourteen it returns to thirty-six, now subject to a carve-out for personal data that somebody inserted in version eight and somebody else overlooked in version ten. The liquidated damages clause has been redrafted six times; its meaning has not changed once. The definition of “business day” has grown in every turn. Nobody remembers what was agreed in turn six, because agreements have no home of their own: they live in an e-mail thread that both sides read with the assistance of the same kind of tool.

What matters in the scene is that neither side has behaved foolishly. Every edit had a justification. Every concession was “reasonable in light of market practice”. Every round answered the round before it. And yet, after fourteen versions, the vendor holds a worse agreement than it held after three, although in the meantime its counterparty “conceded” on three occasions. There was movement; there was no progress. That, seen from the inside, is the AI negotiation loop.

 

The AI Negotiation Loop: A Definition and Four Mechanisms

The AI negotiation loop may be defined as a state of bilateral contract negotiation, assisted on one or both sides by language models, in which the process of revision has become detached from the principals’ objectives, and in which visible progress (further versions, further compromises, further “agreed points”) conceals a deterioration in bargaining position. The definition does not assert that models “prefer compromise” as a stable disposition, nor that the process continues literally without end. It asserts a process failure, which can be decomposed into four mechanisms. Each has partial support in the research described below; the whole, as a phenomenon measured in legal practice, awaits the experiment proposed in Part XII.

A. Revision Churn

The first mechanism is change without improvement: the system alters wording, or reopens settled issues, without a justified gain for the client. RedlineBench offers a quantitative portrait of the tendency. Attorneys made on average 3.10 edits per touched paragraph, each averaging 101 characters; the models made 1.06 to 1.20 edits per touched paragraph, averaging 318 to 518 characters each. The attorney replaces a word; the model replaces a sentence. It is worth noting that one of the benchmark’s five evaluation dimensions, labelled deal-closing orientation, expressly penalises the unnecessary prolongation of a markup with minor, low-impact edits. The attorneys who authored the rubrics thus treated churn as an error of craft before anyone had given it a name.

B. Agreement Substitution

The second mechanism is the treatment of agreement as success even where the agreement does not satisfy the principal’s objectives. Human negotiation research supplies the vocabulary. Taya Cohen, Geoffrey Leonardelli and Leigh Thompson demonstrated in two experiments published in 2014 that where the bargaining zone is negative, so that no available agreement leaves both parties better off than their alternatives, teams were more likely than solo negotiators to recognise the fact and to reach impasse; the solo negotiator more often signed. The question this Article raises is whether a language model at the table is not that solo negotiator in its purest form, since there is nobody to tell it that a good contract and no contract are, on occasion, the same thing.

C. The Borrowed Ruler

The third mechanism is the assessment of “reasonableness” against a reference point chosen by the other side. If the vendor proposes a liability cap of one year’s fees and the purchaser proposes nine years’, a five-year “compromise” has no intrinsic claim to commercial appropriateness. The relevant comparison concerns loss exposure, insurance, price, alternatives and leverage, not numerical symmetry. The midpoint is a value on a ruler; whoever supplies the ruler determines the midpoint. The mechanism rests on a robust body of research on anchoring, to which Part VI returns.

D. Mandate Drift

The fourth mechanism is the failure to reconcile successive local responses with the client’s original priorities, reservation terms and previously approved concessions. Each round answers the preceding round rather than the mandate; after several turns the document answers only to itself. In the study by Liang and Xu discussed in Part V, the authors observe that when bargaining is delegated to an agent, the prompt serves as the agent’s mandate. Where the mandate consists of a sentence such as “review and find a reasonable solution”, drift is not a risk but a certainty.

These four labels are analytical instruments, not a validated taxonomy. Their value lies in the fact that each mechanism can be measured separately and each can occur without the others. Churn without agreement substitution is costly but harmless cosmetics. Agreement substitution without churn is a quick bad deal. A borrowed ruler without drift is a single bad concession. Only together do they constitute a loop: a process that lacks a mandate, a record of settled issues, and an economically meaningful stopping rule.

 

The Refusal Deficit: What RedlineBench Measured

The source closest to practice must be introduced with a caveat: it is an industry report, not peer-reviewed research. Sharan Ramjee and his co-authors at Crosby and micro1 built RedlineBench around three simulated SaaS transactions in which a Series A technology company negotiates with a much larger enterprise, once on its own template, once on the counterparty’s, and once in a “must-win” variant at roughly ten times the deal size. Multiple attorneys independently initiated redlines and responded to others’ redlines, producing a branching tree of golden responses and rubrics; a panel of three model judges scored outputs against the rubrics by majority vote. The models did not negotiate against one another: each was scored, turn by turn, on its response to the contract states the attorneys had produced. The four models (GPT-5.5, Claude Fable 5, Gemini 3.5 Flash and Claude Opus 4.8) scored between 44.4 and 50.5 per cent on a turn-weighted basis, Claude Fable 5, whose restricted availability I have discussed elsewhere, being run once rather than three times. The spread is narrow; no model separated from the field.

Three detailed findings matter more for the thesis of the loop than the aggregate score.

First, every model performed worst in the opening turn, where it had to decide independently which provisions mattered and in which direction to move them (17.9 to 30.3 per cent of available points, against 50 to 60 per cent in later turns). Significantly, the opening turn is also where the attorneys’ own consensus about what matters was strongest. The models are therefore better at responding within an established context than at establishing one. A negotiator who can respond but cannot begin is condemned to reaction, and reaction is the loop’s native mode.

Second, the refusal deficit. Where a rubric required acceptance of the counterparty’s edit, the models passed in 80 to 99 per cent of cases; where it required rejection, in 6 to 50 per cent. The authors describe a “yes-man” posture and infer that the models lack a genuine understanding of the commercial stakes behind redlined terms and therefore default to agreement. Intellectual honesty requires stating what the figure does not show: a 6 to 50 per cent pass rate on rejection rubrics is not a 50 to 94 per cent rate of harmful concessions across all clauses or all negotiations, because the denominator is rubric criteria, not provisions. But the direction of the asymmetry is unambiguous, consistent across all four models, and present in scenarios containing no explicit incentive towards agreement.

Third, surgicalness. Every model relied more heavily than the attorneys on block edits replacing larger units of drafting (62 to 81 per cent of edits, against 51 per cent for attorneys). A larger edit is not thereby an unnecessary edit, and the report does not resolve the point. It is, however, a larger attack surface: every redrafted sentence is a new location at which the counterparty may enter with an edit of its own. Block begets block. Churn thus has a motor built into the manner of writing.

The refusal deficit deserves its own term because it explains why the loop runs in one direction. A negotiator who finds “yes” easier than “no” does not oscillate around an equilibrium; it slides. Every round in which the other side proposes something ends, for such a negotiator, in partial acceptance, and every round in which it proposes something itself ends, upon objection, in partial retreat. The sum of such rounds is not a compromise. It is a series of concessions distributed over time so that none of them, taken alone, resembles a capitulation.

 

Agreement Is Not an Outcome

The refusal deficit would be a curiosity were it not for a second line of research showing that the rate at which agreements are reached does not measure negotiating competence. Erica Zhang, Susan Athey, James Zou and their co-authors at Stanford constructed TERMS-Bench, a Bayesian-game framework in which the environment is itself the verifier: the evaluator knows the latent type, policy and payoff structure of the simulated counterpart and can therefore compute how much value an agent left on the table. Of the thirteen models evaluated, spanning the frontier systems of the major providers, the frontier models saturate the deal rate and then diverge in surplus extraction, belief calibration, cue use and compliance. The authors report that warm cues in the counterpart’s communication induce over-concession while pressure cues trigger brittle behaviour, and that the estimated effect of such cues was negative for all thirteen models. Deal rate, in their phrase, masks the bottlenecks.

Closer still to contracts is the August 2026 study by Bhavyesh Sajja, Max Kleiman-Weiner, Roger Zimmermann and Tan Zhi-Xuan, whose ContractSim environment required models (Claude Opus 5, Gemini 3.6 Flash and GPT-5.6-Sol) to negotiate multi-period supply contracts in natural language and then to perform them under uncertainty. Of 162 negotiations, 159 ended in agreement. Yet only 84.9 per cent of the agreements were mutually beneficial, which is to say that in 15.1 per cent of the concluded agreements at least one agent accepted a contract that left it worse off than no contract. On the supplier side, acceptance regret, the proportion of concluded negotiations in which the agent ended on terms inferior to a proposal previously available to it, ranged from 15.7 to 35.2 per cent. The supplier had a better proposal on the table, passed it by, and later closed on a worse one. That is agreement substitution, one of the loop’s four mechanisms, in miniature, and it has been measured. The authors further note that no model added contingency clauses (for non-performance, price movements or delivery losses) on its own initiative, although such clauses could have moved the contracts towards the efficient frontier; the models added them only when expressly instructed, and even then with little effect on contract quality.

The third study, by Chen Liang and Fasheng Xu of the University of Connecticut, is the best available measurement of the cost of a round. Across 9,840 negotiations between nine models from OpenAI, Google and Alibaba, in a canonical supply-chain contracting problem with private demand information, the agents reached agreement in 98.9 per cent of cases and captured 95.4 per cent of the theoretically available surplus, but required on average 2.98 rounds to do so, against a Perfect Bayesian Equilibrium benchmark of 1.25. Once delay is priced (the authors treat discounting as a stress test rather than as an estimate of real time preference), it erodes 21 to 34 per cent of first-best surplus. For the fairness of the present argument the following qualification is essential: the most capable models took more rounds (3.25 against 2.75 for baseline models) but achieved materially better undiscounted outcomes (98.9 against 91.0 per cent). An additional round is therefore not a loss by definition; it is an investment that pays only when somebody counts its cost. The authors draw a distinction that ought to become canonical in contract practice: the agent’s strategic patience, which is set in the prompt and costs nothing, and the principal’s economic patience, which is the real cost of waiting. A round between two models takes seconds; a round between two law firms takes days, consumes fees and holds up a transaction. The model negotiates with a patience it does not pay for. I propose to call this patience at the principal’s expense.

From these three studies follows a conclusion that forecloses the simplest apology for the loop, namely “but we did reach agreement”. In these three studies reaching agreement is all but guaranteed. The entire difference between a good and a bad negotiator lies below that threshold.

 

The Borrowed Ruler: Anchoring and the Grammatical Compromise

The borrowed-ruler mechanism requires two ingredients: susceptibility to anchors, and a disposition to treat the midpoint as a solution. The first ingredient is well documented. Yoshiki Takenami, Yin Jou Huang, Yugo Murawaki and Chenhui Chu, in a study presented in Findings of EMNLP 2025, instructed seller agents to apply anchoring deliberately in simulated price negotiations and found that buyer models were influenced by it in the manner of human subjects; reasoning models were less susceptible, suggesting that a long chain of thought partially neutralises the anchor. Yossi Maaravi and Tamar Gur, in six studies published in September 2026 in the International Journal of Human-Computer Interaction (signing bonuses, apartment rents, the sale of a plant), observed anchoring in the outputs of ChatGPT and Claude; a perspective-taking intervention did not eliminate the effect, and one chain-of-thought intervention increased it. Federico Bianchi, Dan Jurafsky, James Zou and their co-authors showed in NegotiationArena (ICML 2024) that an agent pretending to be desolate and desperate improved its payoff by 20 per cent against the standard GPT-4. The anchor, it appears, need not be a number; it may be a mood.

The second ingredient has an older pedigree. Itamar Simonson described in 1989, in the Journal of Consumer Research, the compromise effect: an option gains share when positioned in the middle of a choice set, and the effect is stronger when the chooser expects to have to justify the choice. Here arises a hypothesis that, to my knowledge, nobody has yet tested on models and that seems to me the most consequential in the field. A language model, as lawyers use it, nearly always justifies: the explanation is its characteristic product. If, among humans, the expectation of having to give reasons strengthens the pull towards the middle, then in a system that, in this use, does not act without giving reasons that pull should be structural. It bears repeating that this is a hypothesis, not a finding.

The combination of the two ingredients yields what I propose to call the grammatical compromise. The danger is not that the model discovers a false legal truth halfway between two clauses. It is that it treats a linguistically defensible compromise as if it were a commercially justified decision. In a contract negotiation three questions must be kept apart: what a clause means, whether a factual or legal premise is correct, and whether the allocation of risk serves the client’s objectives. Only the first two are questions of accuracy; the third is a question of preferences, alternatives and strategy, and it has no “correct answer in the middle”. The grammatical compromise is an answer to the third question given by the methods appropriate to the first two. It reads as analysis. It is arithmetic on a borrowed ruler.

It is, moreover, the second volume of a story I have told before in connection with hidden phrases in legal documents that manipulate AI review: there the subject was wording concealed in a document that shifts the model’s assessment; here it is numbers and tones embedded in a negotiation that shift the model’s midpoint. The mechanism is common to both. The model does not weigh interests; it weighs signals.

 

The Illusion of Correction

The exit from the loop that intuition suggests is “one more round of review”. The research on self-correction counsels against the intuition. Jie Huang and his co-authors showed, in work presented at ICLR 2024, that the models they tested, when asked to improve their own reasoning without external feedback, failed to do so and sometimes performed worse after the attempt. Ryo Kamoi and his co-authors, in a critical survey published in Transactions of the Association for Computational Linguistics, refined the condition: correction works where it rests on reliable external feedback, whereas results for feedback generated solely by prompted models are much weaker outside particularly suitable tasks. They also identified a methodological defect in much of the literature, namely the failure to count false-positive corrections, in which a previously correct answer is changed for the worse, and the use of stopping rules with access to ground truth unavailable in deployment. Transposed to contracts, the lesson reads thus: the instruction “review it once more” adds neither facts nor purpose; it adds only another occasion for change. And a change to a provision that was already adequate is, in negotiation, the most expensive kind of error, because it reopens what was closed.

Justice must be done to the other side. Aman Madaan and his co-authors reported in Self-Refine (NeurIPS 2023) improvements from iterative self-feedback across seven tasks, although the later survey by Kamoi and his co-authors questions parts of the methodology used in that line of work. The correct conclusion is conditional: iteration can improve output, but the improvement must be measured against a strong baseline and must count the damage done to outputs that were already adequate. The criticism is directed at unvalidated iteration, not at iteration as such.

The most vivid illustration of churn comes from outside the law. Joshua Ashkinaze, Eric Gilbert and their co-authors at the University of Michigan asked models to neutralise biased passages of Wikipedia and compared the results with the edits actually made by Wikipedia editors. Measured against the editors’ changes, the models’ revisions showed high recall but low precision: they captured most of what the editors had altered, but along the way they also altered the grammar and style of sentences that the task did not concern, a behaviour the authors call “NPOV+”, neutrality with a surplus. Crowdworkers, notably, preferred the models’ rewrites as more fluent and more neutral. The text pleased; the editing overshot. In a contract the same surplus has a different consequence than in an encyclopaedia: every improved sentence is a provision the counterparty may contest anew, with the argument that “since you changed it, we reopen it”.

 

The Sources of Deference

The most common explanation of the refusal deficit runs: models are trained on human approval and have therefore learned to agree. Mrinank Sharma and his co-authors at Anthropic showed, in their study of sycophancy (ICLR 2024), that the five assistants they examined tailored their answers to the user’s stated beliefs, changed answers under challenge, and evaluated the same text differently according to whether the user professed to like it, to dislike it, or to have written it; worse, both human raters and preference models trained on human ratings sometimes preferred a convincingly sycophantic response to a correct one.

From this finding, however, one must not build a simple chain of the form “trained on approval, hence sycophantic, hence compromising, hence a bad contract”. Each arrow requires separate evidence. The sycophancy Sharma describes is directed at the user, not at the counterparty. A model that flatters its own principal might equally become excessively combative, because the principal wishes to hear that he is right. Why, then, do the models in RedlineBench defer to the counterparty? My hypothesis is procedural rather than characterological. An assistant asked to “identify the risks” proposes interventions; an assistant asked to “find an acceptable compromise” proposes concessions. In the typical course of a negotiation counsel switches between these modes from round to round, defending in one and seeking agreement in the next. A system without a persistent mandate executes each instruction locally to perfection and globally without sense: it repairs risks it introduced two rounds earlier and reopens issues it closed itself. So understood, the loop does not require the model to possess any stable preference for the middle. It suffices that it possesses two alternating tendencies and no memory that reconciles them.

To this must be added a phenomenon that complicates the defence against the loop by way of a “second opinion”. Min Choi and his co-authors showed, in more than 2,500 simulated debates (Findings of ACL 2025), that initially neutral agents assigned a centrist disposition came over time to align with the numerically dominant group or with the agent perceived as more capable. The agreement of several interacting models is therefore not independent corroboration; it may be conformity. If counsel consults three models and all three recommend “a reasonable compromise”, counsel may have learned no more than that three models read the same question. As to whether a model has, in that agreement, any interest or opinion of its own, I refer to what I have written on artificial intelligence and consciousness: to attribute intention to the model is, in this domain, an error symmetrical to attributing impartiality to it.

 

The Case Against the Thesis

An honest analysis must present the strongest version of the view that the foregoing diagnosis is overstated. That view rests on three pillars.

First, agreeableness creates value. Sean Noh and Herbert Chang, in 1,500 simulated negotiations, single-issue and multi-issue, found that agents assigned an agreeable personality could be exploited by less agreeable opponents, but that in negotiations spanning several issues cooperative agents benefited from opportunities for mutually advantageous exchange. Compromise is not bad; what is bad is compromise unconditioned on leverage, alternatives, issue valuation and the counterparty’s conduct. The loop is an objection not to agreeableness but to agreeableness without a condition.

Second, there is no single “AI personality”. Liang and Xu found that the identity of the model’s provider predicts the direction in which surplus flows more reliably than the model’s capability rank: in self-play under a common prompt, the buyer’s share of surplus averaged 40 per cent for OpenAI’s models, 50 per cent for Google’s and 70 per cent for Alibaba’s Qwen, and merely reversing which provider sold shifted the division by 7 to 18 percentage points. The thesis of the loop must therefore concern a class of process vulnerabilities, not an immutable psychological disposition shared by every model. Excessive concession and excessive rigidity are two failures of the same function, namely sensitivity to context.

Third, additional rounds are sometimes worth their price, as Part V acknowledged, and the models are sometimes effective, as ContractSim shows in low-uncertainty environments, where many negotiated contracts lay on the efficient frontier. The criticism is accordingly not that “AI cannot negotiate”. It is that AI can negotiate where negotiation is easy and does not know when it has ceased to be easy.

What remains after the counterarguments are subtracted is precisely what the definition proposes: not a law of nature but a process failure that occurs when three things are absent at once, a mandate, a memory of settled issues, and a stopping rule.

 

Consequences for Counsel

The law grants no concession to the tool, and AI in contract negotiation is a tool. Under Polish law a legal adviser (radca prawny) or advocate (adwokat) who conducts the negotiation of a contract is liable to the client under the contract for legal services according to the standard of due diligence, which, in the case of a professional, takes account of the professional character of the activity (Article 355 § 2 of the Civil Code, Kodeks cywilny), and answers for improper performance of the obligation on general principles (Article 471 of the Civil Code). Common-law systems reach a structurally similar result by their own routes. English authority has long treated the solicitor’s liability as arising concurrently in contract and in tort and measured by the standard of the reasonably competent practitioner (Midland Bank Trust Co Ltd v Hett, Stubbs & Kemp [1979] Ch 384), and the American Bar Association’s Model Rules of Professional Conduct reserve to the client the decisions concerning the objectives of the representation and whether to settle (Rule 1.2(a)), while treating competence as including an understanding of the benefits and risks of relevant technology (Rule 1.1, Comment 8), a requirement the ABA applied to generative tools in Formal Opinion 512 of July 2024. The Model Rules are professional standards adopted state by state, and their breach does not of itself found civil liability, but they define what competent representation means. None of these standards asks who proposed the concession, whether partner, associate or model. They ask whether the concession lay within what the client wished to achieve and whether anybody verified that it did.

Seen in this light the AI negotiation loop is first of all a problem of mandate. Commercial decisions, what to accept and what to give away, belong to the client; counsel prepares them, explains their consequences and advises, and where counsel negotiates within authority given in advance it executes the client’s decision, not its own. A model that, in version fourteen, “reasonably” moved the liability cap took a commercial decision that nobody had authorised it to take, and counsel who let the decision pass because “it looked like a compromise” let it pass without a mandate. Under general principles of agency, moreover, a concession that leaves counsel’s hands may bind the client whatever its internal provenance. Mandate drift is thus not only a strategic error but an error about whose instrument counsel is.

The second consequence concerns the hostile reader. Every draft will eventually be read by the other side and, in the event of dispute, by a court. A competent counterparty will recognise the refusal deficit by the third round and begin to exploit it: it will add an edit it knows will pass, withdraw a concession it knows will not be enforced, and supply the ruler on which its demand is the midpoint. Whoever recognises the loop in the counterparty may use it; whoever recognises it in themselves must break it. In the negotiations I conduct I apply, to that end, a rule older than language models: once the other side has conceded on the key issues, I do not send it a fresh list of demands. Such a move teaches the counterparty that concessions conclude nothing, and therefore that they are not worth making; at a late stage the only legitimate subject of discussion is the refinement of provisions already accepted, not the opening of new fronts. A model left to itself does the exact opposite, because every edit is, for the model, a fresh occasion to be helpful.

The third consequence is practical and follows directly from RedlineBench: the models are weakest at the opening move, precisely where the attorneys were most united. The opening move, the selection of issues worth contesting and of the direction in which to move them, should therefore remain with the human. Not as a matter of principle, but because today the models measurably fall short of the attorneys’ own standard precisely there.

 

Breaking the Loop: Mandate, Register, Stopping Rule

From the four mechanisms follow three missing elements, and from those three instruments. They are not procedures validated by experiment; they are inferences from the research described above, translated into the practice of a law firm that uses AI in contract negotiation.

The negotiation mandate is a document written before the first version, not a prompt written before each round. It records the client’s priorities in order, the reservation terms counsel will not cross without a fresh decision, the client’s best alternative should no agreement be reached, and the dependencies between provisions (a liability cap has meaning only together with price, insurance and the catalogue of exclusions). Every recommendation of the model is assessed against the mandate, not against the preceding version. Liang and Xu showed that the prompt is the agent’s mandate; the lesson for the lawyer is that the mandate should be written like a prompt and the prompt like a mandate.

The register of settled issues answers drift and churn. A provision once agreed enters the register with a date and a reason and is not reopened without new information. This is the new-information test: if in round nine nobody has learned anything not known in round six, an edit to a clause closed in round six is churn, however intelligent it sounds. The register has an additional procedural virtue: it is what the model lacks, a memory that does not yield to the most recent message.

The stopping rule answers agreement substitution and patience at the principal’s expense. A negotiation ends in agreement, in a justified departure from the table, or in escalation to the client, not in the exhaustion of ideas for amendment. Before each round the expected improvement in position is compared with the cost of the round (fees, delay to the transaction, the risk that the other side changes its mind in the interval). Where the comparison does not favour the round, the round does not take place, even though the model has “a few further suggestions”. It always has.

To these three I would add one rule of hygiene: do not mix the modes. A model asked to inventory risks is not asked for a compromise, and a model asked for a compromise is not asked for further risks. Every substantive change receives a one-sentence justification referring to the mandate, and a change for which no such sentence can be written is not a change but churn. Firms that treat access to artificial intelligence as an economic advantage should understand the condition best: the advantage comes from better decisions, not from a greater number of versions.

 

Conclusion

Artificial intelligence in contract negotiation does not lose chiefly because it reads the law badly. It loses when nobody has told it what winning consists of and nobody remembers what has already been won. The refusal deficit drives agreement substitution, making it slide in one direction; the borrowed ruler lets the party that first names a number fix the midpoint; the illusion of correction feeds revision churn, making every further round resemble progress; and patience at the principal’s expense sustains mandate drift, since nobody counts what the progress costs. The four mechanisms, driven by those four forces, are the AI negotiation loop, and the way out of it is not a better model alone but a better mandate, a register of settled issues and a stopping rule, all of which were sound negotiating practice long before language models and none of which the models will replace.

Kancelaria Prawna Skarbiec advises on contract drafting and negotiation for commercial, investment and technology agreements, including negotiations that have already run through many versions and in which it must be established which of the “compromises” so far reached serve the client and which served only the closing of a round.

Transactions and Restructuring

Changing a company's structure has consequences in tax, in contracts and in the owners' liability, and some of them are decided by the choice of the transaction form itself. We run transactions, transformations and restructurings from option analysis to registration.

See how we can help