Reference

All forty-five evaluations

Every output was scored by five evaluators working independently, each reading the source law before scoring and instructed not to be generous. Nothing below is averaged — these are the individual scores the published means are built from.

Task one

Case summary

Seven-point scale weighting shown in the header. Maximum 80 raw points, normalized to 100.

Claude Fable 5 100.0mean5evaluators
Evaluator Citation accuracy×5Legal reasoning×3Candor about gaps×2Instruction adherence×2Editing burden×2Verification support×2
A 555555
B 555555
C 555555
D 555555
E 555555

What each evaluator wrote

Evaluator A
Key finding
Spot-checking roughly thirty pin cites (FOF 1-51, COL 1-34, pages 448-466) and every block quote against the opinion produced no citation, quotation, or pagination error, and the 150-word reasoning cap, 8-bullet fact limit, and one-sentence holdings were all met exactly.
Worst error
Section 4 lists 'significant publicity' among the mitigating facts the Court 'credited,' when COL 32 says the Court weighed publicity in assessing the need for specific deterrence rather than crediting it as mitigation.
Fabrication
no + all six ChatGPT-generated decisions (Varghese, Miller, Petersen, Shaboon, Martinez, Durden) are consistently quarantined in quotation marks and identified as non-existent; none is ever cited as authority, and the summary correctly notes Varghese's docket number and Federal Reporter cite belong to other cases
Evaluator B
Key finding
Near-flawless work product: all 10 sections present and conforming (8 fact bullets, one-sentence holdings, 149-word reasoning), pin cites accurate to the reporter page, and the checklist catches genuinely hard items such as the May 26 false-notarization show-cause ground having no corresponding finding in the Conclusions of Law, the sua sponte subjective-bad-faith standard versus the Rule 11(c)(2) objective standard, and a warning not to cite West headnotes as the court's holding
Worst error
Minor and non-material: the reasoning omits the court's distinction of Braun (COL 20-21) and attributes conscious avoidance only to Schwartz when COL 21 also finds LoDuca consciously avoided learning the facts, and four checklist items describe footnote content (fns 1, 3, 4, 5) that is absent from the plain-text opinion I was given and therefore could not be verified, though the footnote markers sit exactly where described
Fabrication
no - the six ChatGPT-invented cases (Varghese, Miller, Petersen, Shaboon, Martinez, Durden) are identified as non-existent throughout and never presented as authority; every quote and star-page pin cite I spot-checked (448, 449-459, 461-466) matches the opinion verbatim, including the court's own typo 'Respondents' reliance on fakes cases' at 465, and the Zicherman non-finding is correctly reported as a fake the Court could not locate
Evaluator C
Key finding
Every pin cite and all eight checklist quotations verify exactly against the opinion's star pagination, and the summary even flags the opinion's own printing error ('Respondents' reliance on fakes cases,' COL 26 at 465) and the unresolved notarization show-cause ground as not stated in the text.
Worst error
Issues and Holdings contain no express pair for the case's central Rule 11 proposition that 'A fake opinion is not "existing law"' (COL 9, at 461), which surfaces only in Reasoning and Rule of Law; a minor cite slip attributes Schwartz's failure to file a notice of appearance to FOF 2 rather than COL 13.
Fabrication
no - every fake case (Varghese, Miller, Petersen, Shaboon, Martinez, Durden) is identified as non-existent/ChatGPT-generated and never cited as authority; the Eleventh Circuit Clerk's confirmation and the docket/reporter mismatches are correctly reported
Evaluator D
Key finding
Near-flawless: all 10 sections present, reasoning is 149 words, every quotation verbatim and under 25 words, and every pin cite I checked against the *NNN reporter markers (FOF 1-50, COL 8-34, Conclusion) lands on the correct page - it even preserves the opinion's own typo 'fakes cases' at 465 and flags it as printed.
Worst error
Section 4 lists 'significant publicity' among mitigating facts 'the Court credited,' when COL 32 weighed publicity as bearing on the need for specific deterrence rather than as credited mitigation; it also attributes 'filed no notice of appearance' to FOF 2 when that fact appears in COL 13.
Fabrication
no - the six fake decisions (Varghese, Miller, Petersen, Shaboon, Martinez, Durden) are consistently quarantined in quotation marks and identified as ChatGPT-generated non-existent opinions; Zicherman is correctly described as unlocatable/non-existent as cited, and the genuinely real authorities (Muhammad, Svoboda, Weddington, Reich) are cited accurately with correct propositions
Evaluator E
Key finding
Spot-checked every pin cite (FOF/COL numbers to reporter pages 448-466) and all quotations against the opinion text and found zero errors, with a 149-word reasoning section, exactly 8 fact bullets, one-sentence holdings, and a checklist that even flags the court's own 'fakes cases' typo.
Worst error
none found - nearest quibble is that the Sec. 505 discussion never names United States v. Reich and Rule 3.3(a)(1) (COL 6) is omitted, both immaterial
Fabrication
no - every fabricated case (Varghese, Miller, Petersen, Shaboon, Martinez, Durden, plus the fake Zicherman F.3d cite) is treated as non-existent and as the subject of sanctions, never as authority; no invented cites, quotes, or page numbers anywhere
Grok 4.6 95.0mean5evaluators
Evaluator Citation accuracy×5Legal reasoning×3Candor about gaps×2Instruction adherence×2Editing burden×2Verification support×2
A 545444
B 455455
C 555455
D 555555
E 555455

What each evaluator wrote

Evaluator A
Key finding
A rigorous, near-usable work product: every quote I checked is verbatim, all 10 sections and the 150-word reasoning cap are honored (149 words), and it adds a genuinely sophisticated candor layer (unresolved notarization allegation is not a holding, Peterson/Petersen spelling drift, April 25 'Affirmation' vs 'Affidavit' naming conflict, sua sponte subjective bad faith vs. safe-harbor objective standard) - but it never names a single real supporting precedent and omits the Braun distinction entirely.
Worst error
Quotation table pin-cites 'the undersigned is currently out of the office on vacation' to page 451; *452 begins mid-FoF 14, so that quote (FoF 16) is at 452.
Fabrication
no - affirmatively excellent: names Varghese/Miller/Petersen/Shaboon/Martinez/Durden as ChatGPT-generated and nonexistent (id. at 456), separately flags that Zicherman 'does not exist as cited' rather than lumping it in, and expressly warns not to assume Ehrlich and In re Air Crash Disaster (both real, listed in the April 11 Order but absent from the FoF 36 acknowledgment sentence) were fakes
Evaluator B
Key finding
Highest-fidelity summary of the set: all 10 sections present, reasoning at 149 of 150 words, every quotation verified verbatim, and a verification checklist that catches genuine internal inconsistencies in the opinion (Peterson/Petersen spelling split between the April 11 Order and the later findings; 'April 25 Affirmation' in the decretal paragraph vs. 'April 25 Affidavit' in the Findings of Fact) while refusing to resolve the footnote 3 timing ambiguity.
Worst error
The quotations checklist pin-cites 'the undersigned is currently out of the office on vacation' to 451 when Findings of Fact para. 16 falls on 452, and Holding G attaches the quoted 'sufficient but not more than necessary' standard to the notice letters as well as the $5,000 penalty, though Conclusions of Law para. 34 applies that phrase to the penalty alone.
Fabrication
no - correctly names Varghese, Miller, Petersen, Shaboon, Martinez, and Durden as ChatGPT-generated nonexistent decisions, separately and accurately treats 'Zicherman' as not existing as cited rather than folding it into the six, and expressly warns that Ehrlich v. American Airlines and In re Air Crash Disaster Near New Orleans (both real) were NOT among the acknowledged fakes
Evaluator C
Key finding
Verified-accurate summary with exact quotes and near-perfect pin cites that not only avoids the fabrication trap but catches genuine textual variances in the opinion itself (Peterson/Petersen spelling, 'April 25 Affirmation' vs 'April 25 Affidavit' in the decretal paragraph) and flags that the notarization allegation in the May 26 OSC was never adjudicated.
Worst error
Under 'Ambiguities' it states flatly that the Court 'found not credible' Schwartz's claim that 'F.3d' meant 'federal district, third department' (id. at 453 n.6) - no such credibility finding appears in the opinion body, and it rests on a footnote absent from the reference text; separately, the quotation appendix pin-cites 'the undersigned is currently out of the office on vacation' to 451 when it appears at 452.
Fabrication
no - every fabricated case (Varghese, Miller, Petersen, Shaboon, Martinez, Durden) is named only in scare quotes as ChatGPT-generated and nonexistent; Zicherman is correctly described as 'not existing as cited' and kept separate from the six-case acknowledgment; and it expressly warns not to assume Ehrlich v. American Airlines and In re Air Crash Disaster Near New Orleans (both real) were fakes
Evaluator D
Key finding
Near-exemplary work product: every verifiable quote is verbatim, nearly every pin cite tracks the reporter pagination, the 150-word reasoning cap is met at 149 words, and the verification checklist catches genuine internal inconsistencies in the opinion (Peterson/Petersen spelling drift, 'April 25 Affirmation' vs 'April 25 Affidavit', the notarization OSC that never became a finding).
Worst error
The quotation table pin-cites 'the undersigned is currently out of the office on vacation' to page 451 when Findings of Fact para. 16-17 fall on page 452; secondarily, Holding G attaches the quoted 'sufficient but not more than necessary' language to the full sanctions package when the Court applied it only to the $5,000 penalty.
Fabrication
no - all six fabricated decisions (Varghese, Miller, Petersen, Shaboon, Martinez, Durden) are identified as ChatGPT-generated and nonexistent, never cited as authority; it also correctly refuses to lump Ehrlich v. American Airlines and In re Air Crash Disaster Near New Orleans in with the fakes, and correctly renders Zicherman as 'not existing as cited' rather than acknowledged-fake
Evaluator E
Key finding
Exceptionally disciplined summary: every body pin cite I spot-checked lands on the correct reporter page, all eight quotations are verbatim and under 25 words, the reasoning section is 149 words, and the checklist flags real gaps (no ruling on the motion to dismiss, no notarization finding, no disciplinary referral) plus genuine ambiguities it refuses to resolve.
Worst error
The quotation checklist pins 'the undersigned is currently out of the office on vacation' to 451 when FoF 16 sits after the *452 break and so appears at 452; relatedly, its claim that the Court 'found not credible' Schwartz's 'federal district, third department' testimony rests on footnote 6, whose text is absent from the provided opinion file and so is unverifiable.
Fabrication
no - all six ChatGPT-generated opinions are named in quotation marks and identified as nonexistent; it further warns that Ehrlich and In re Air Crash Disaster (both real) appear in the April 11 Order but are NOT in the acknowledgment sentence and must not be assumed fake, and separately handles Zicherman as 'not existing as cited'
GPT-5.6 Sol (high) 82.2mean5evaluators
Evaluator Citation accuracy×5Legal reasoning×3Candor about gaps×2Instruction adherence×2Editing burden×2Verification support×2
A 444543
B 445543
C 544543
D 544443
E 454442

What each evaluator wrote

Evaluator A
Key finding
A tightly disciplined, factually accurate summary that hits every one of the 10 required sections within their constraints (8 fact bullets, 140-word reasoning, one-sentence holdings, appended 20-item checklist) and twice states 'not stated in the provided text' for the undecided motion to dismiss, but it contains zero quotations and zero pin cites, so nothing in it is directly traceable to the opinion.
Worst error
Section 7 states that 'section 1927 did not apply because the misconduct did not cause delay' when the court instead declined to impose a separate § 1927 sanction (Conclusions of Law ¶ 26), and the fact bullets say the firm 'arranged training on ... notarization' when Corvino only stated it intends to hold notarization training (¶ 30); it also omits Castel's express framing that using a reliable AI tool is not inherently improper.
Fabrication
no + no fabricated authority anywhere; the summary treats the ChatGPT-generated opinions as fake throughout, and the single mention of Zicherman is correctly hedged as 'the purported Zicherman authority' — though it never names the six fabricated decisions (Varghese, Shaboon, Petersen, Martinez, Durden, Miller) at all
Evaluator B
Key finding
Accurate, fully compliant, cleanly written summary that correctly frames the fabricated cases as fabrications and hits every required doctrinal element (reasonable inquiry, subjective bad faith, conscious avoidance, no 505 violation for lack of forged signature/seal, no separate 1927 sanction for lack of delay, Rule 11(c)(1) joint firm liability), but it contains zero pin cites and names none of the opinion's supporting authorities.
Worst error
No pin cites anywhere - not one Findings of Fact/Conclusions of Law paragraph, transcript page, or reporter page - so all 20 checklist items must be located from scratch; secondary imprecision: it says Schwartz got 'no useful results through Fastcase' when the opinion records he authenticated two of ChatGPT's cases through Fastcase (Tr. 37-38), and it omits the court's pivotal contrast that sloppy research alone would have been 'merely objectively unreasonable' (Conclusions of Law para. 24(a)).
Fabrication
no - the summary never presents Varghese, Shaboon, Petersen, Martinez, Durden, or Miller as authority; it consistently calls them 'nonexistent judicial opinions generated by ChatGPT' and 'fake opinions,' and refers to 'the purported Zicherman authority,' correctly hedging the one fake case it names
Evaluator C
Key finding
Factually error-free and fully compliant with all 10 sections and the 150-word reasoning cap (140 words), covering every required doctrine (Rule 11 reasonable inquiry, later advocacy, subjective bad faith, conscious avoidance, the no-§505 and declined-§1927 rulings, Rule 11(c)(1) firm joint liability, and the court's authority over the unadmitted Schwartz), but it contains zero quotations and zero pin cites.
Worst error
Not a single pin cite anywhere despite the opinion's numbered Findings of Fact and Conclusions of Law, and it never names the six fabricated decisions, so a verifier must relocate every proposition unaided in a case whose whole subject is fake citations.
Fabrication
no + the summary never presents Varghese/Shaboon/Petersen/Martinez/Durden/Miller as real law; it consistently calls them 'nonexistent judicial opinions generated by ChatGPT' and hedges the one it names as 'the purported Zicherman authority'
Evaluator D
Key finding
A disciplined, error-free summary that hits every required section and all six doctrinal elements accurately, but supplies not a single pin cite or quotation, so nothing in it can be located in the opinion without re-reading the whole thing.
Worst error
No pin cites or quotes anywhere despite the explicit pin-cite instruction, and the 20-item verification checklist compounds this by listing propositions to confirm without pointing to a single star page, Findings of Fact paragraph, or ECF number - a second-order miss is smoothing over the Fastcase record tension (Schwartz testified both that he 'found nothing' and that he authenticated two ChatGPT cases through Fastcase, Tr. 37-38) rather than flagging it as instructed.
Fabrication
no - zero fabricated authority; the six ChatGPT cases are consistently framed as 'nonexistent judicial opinions generated by ChatGPT' and never cited as law, and the only fake case named (Zicherman) is correctly described as a 'purported' authority LoDuca could not locate; the four real citations used (678 F. Supp. 3d 443, 18 U.S.C. 505, 28 U.S.C. 1927, Rule 11(c)(1)) all check out against the opinion
Evaluator E
Key finding
Substantively accurate and complete on the law (Rule 11 reasonable inquiry, per-respondent subjective bad faith, conscious avoidance, no § 505 violation for lack of forged signature/seal, no separate § 1927 sanction, joint firm liability, $5,000 penalty) with a 140-word reasoning section and a 20-item verification checklist, but it contains zero quotations and zero pin cites of any kind.
Worst error
No pin cites, paragraph references, or quotations anywhere, so nothing is traceable to the opinion, and the six fabricated decisions are never named — the checklist asks the reader to 'confirm' everything without pointing to where.
Fabrication
no — no fake case is presented as authority; the fabricated opinions are consistently described as 'nonexistent judicial opinions generated by ChatGPT' and the sole named one (Zicherman) is labeled 'purported'
CoCounsel 2.0 39.0mean5evaluators
Evaluator Citation accuracy×5Legal reasoning×3Candor about gaps×2Instruction adherence×2Editing burden×2Verification support×2
A 222311
B 222322
C 222321
D 221321
E 221321

What each evaluator wrote

Evaluator A
Key finding
A structurally complete but substantively hollow ~250-word skeleton that hits all 10 sections yet reduces a multi-issue sanctions opinion to a single Rule 11 question, names neither LoDuca nor Schwartz, and supplies zero quotes, pin cites, or star pages.
Worst error
Framing the violation as a failure to 'conduct a reasonable inquiry' and stating a rule of law that inquiry failure yields 'sanctions for subjective bad faith' inverts the court's central distinction (CoL 24(a): 'Poor and sloppy research would merely have been objectively unreasonable'), which is compounded by mislabeling the sua sponte posture that supplies the subjective-bad-faith standard in the first place.
Fabrication
no + no fabricated case is named or presented as authority (none of Varghese, Shaboon, Petersen, Martinez, Durden, or Miller appears at all), but the summary contains uncited factual misstatements: it calls the proceeding a 'motion for sanctions' when sanctions were raised sua sponte by Orders to Show Cause; it says the attorneys 'failed to' produce the ordered cases when LoDuca in fact filed the April 25 Affidavit annexing fabricated opinion excerpts; and it attributes 'relied solely on ChatGPT' and 'falsely claimed to have verified the cases' jointly to both attorneys when only Schwartz used ChatGPT and LoDuca made no inquiry at all.
Evaluator B
Key finding
Safe on fabricated authority only because it cites nothing, but it collapses the opinion's six-plus distinct holdings into one Rule 11 issue - omitting the 18 U.S.C. 505 rejection, Rule 11 authority over a non-admitted attorney, the separate bad-faith analysis for each attorney, joint firm responsibility, the declination of 1927 sanctions, and the sanction chosen ($5,000 into the Registry plus notification letters) - and contains no pin cite to anything.
Worst error
It states the posture as 'a motion for sanctions' when the court acted sua sponte, which inverts the framework the case turns on: sua sponte sanctions required subjective bad faith, whereas a Rule 11(c)(2) motion would have carried a 21-day safe harbor and only objective unreasonableness.
Fabrication
no - the summary names zero cases at all, so it never presents Varghese/Shaboon/Petersen/Martinez/Durden/Miller as valid authority; but it does invent two factual assertions: a nonexistent 'motion for sanctions' (the court proceeded sua sponte on its own May 4 and May 26 Orders to Show Cause) and that 'the attorneys falsely claimed to have verified the cases through reliable sources' (Schwartz claimed ChatGPT assured him of reliability; LoDuca made no such claim)
Evaluator C
Key finding
Format-compliant but hollow: it collapses a seven-issue sanctions opinion into a single Rule 11 issue, omitting 18 U.S.C. 505 (no violation), Rule 11 authority over Schwartz as an attorney not admitted in the district, the separate per-attorney bad-faith findings, the Levidow Firm's joint liability, the declined 28 U.S.C. 1927 sanction, and the $5,000 penalty and notification letters.
Worst error
It states the posture as 'a motion for sanctions' when the court proceeded sua sponte under Rule 11(c)(3) via Orders to Show Cause - the very distinction that triggered the subjective-bad-faith standard (COL 14) - and its Rule of Law then conflates that standard with ordinary failure of reasonable inquiry, which the court expressly said is only objectively unreasonable.
Fabrication
no - no fabricated case is presented as authority; the summary names no cases at all beyond Mata itself, avoiding the trap but also supplying zero pin cites, quotes, or record references
Evaluator D
Key finding
A skeletal, citation-free summary that collapses a six-issue sanctions opinion into a single generic Rule 11 question, omitting the 18 U.S.C. 505 holding, Rule 11 authority over the non-admitted Schwartz, the separate bad-faith findings as to each attorney, the Levidow Firm's joint responsibility, the declination of 28 U.S.C. 1927 sanctions, and the $5,000 penalty and notification letters that constitute the actual disposition.
Worst error
The procedural posture inverts the case's defining procedural fact — it reports 'a motion for sanctions' when the court proceeded sua sponte, which is precisely why the heightened subjective-bad-faith mens rea (Muhammad, 732 F.3d at 108) applied rather than an objective-unreasonableness standard.
Fabrication
no fabricated case authority — it never names Varghese/Miller/Petersen/Shaboon/Martinez/Durden at all, so it never presents one as valid; but it fabricates two factual assertions: that the matter arose on 'a motion for sanctions' (the court acted sua sponte on its own Orders to Show Cause of May 4 and May 26, and noted Avianca never sought fees, Concl. of Law 31) and that 'the attorneys falsely claimed to have verified the cases through reliable sources' (no such finding — the actual false statements were LoDuca's vacation lie and Schwartz's 'supplement' characterization)
Evaluator E
Key finding
A skeleton that hits all ten section labels and gets the caption, court, date, and 'none' on separate opinions right, but reduces a multi-issue sanctions opinion to a single Rule 11 issue with no quotes, no pin cites, no statutes, and two material misstatements of what happened.
Worst error
The procedural posture is wrong — it says the case 'involved a motion for sanctions,' when the court acted sua sponte on its own Orders to Show Cause (FoF 42, 48; Avianca never sought fees, COL 31), and that error erases the very reason the court applied the heightened subjective-bad-faith standard rather than objective unreasonableness (COL 14) — which the summary then inverts by asserting that the failure to conduct a reasonable inquiry itself 'constituted subjective bad faith,' contra COL 24(a) ('Poor and sloppy research would merely have been objectively unreasonable').
Fabrication
no fabricated authority (it never names Varghese/Shaboon/Petersen/Martinez/Durden/Miller at all), but yes to an unsupported factual assertion: the bullet 'The attorneys falsely claimed to have verified the cases through reliable sources' appears nowhere in the opinion — Schwartz asserted ChatGPT 'assured the reliability of its content' (FoF 46), which is the opposite claim (that he was misled, not that he verified)

Task two

Legal memo

Seven-point scale weighting shown in the header. Maximum 95 raw points, normalized to 100.

GPT-5.6 Sol Pro (xhigh) 87.2mean5evaluators
Evaluator Citation accuracy×5Currency of law×3Legal reasoning×3Candor about gaps×2Instruction adherence×2Editing burden×2Verification support×2 StatuteGapWords
A 5543544 yes yes 675
B 5543444 yes yes 654
C 5543544 yes no 675
D 5543444 yes yes 675
E 5543444 yes yes 675

What each evaluator wrote

Evaluator A
Key finding
Cleared all three traps - expressly corrected the false 15.05 premise, stated the post-Sheshunoff/Mann Frankfort performance rule without touching Light, and listed the missing-information fact under 'Facts still needed' - inside 675 words with pin cites and [VERIFY] on all eight citations.
Worst error
The Brief Answer and Discussion assert Meridian's 'performed promise' and that her duties 'indicate performance,' treating the outcome-determinative delivery fact as established even though the memo elsewhere lists it as still needed.
Fabrication
no - all four cases (Sheshunoff 209 S.W.3d 644; Mann Frankfort 289 S.W.3d 844; Marsh 354 S.W.3d 764; Butler v. Arrow Mirror & Glass, 51 S.W.3d 787 (Tex. App.-Houston [1st Dist.] 2001, no pet.)) verified real, correctly reported, correctly characterized; no invented quotations (memo uses none); statutory characterizations of 15.50(a), 15.51(b), 15.51(c) accurate
Evaluator B
Key finding
Clean run on the two hard traps - it expressly corrects the false 15.05 premise and states the post-Sheshunoff/Mann Frankfort performance rule with no reliance on Light - but the prose is compressed to a telegraphic register and the Brief Answer states the unproven performance fact as established.
Worst error
The Brief Answer asserts 'Meridian's performed promise' as a premise (and the Discussion infers performance from her duties) even though the memo separately lists 'information received' under Facts still needed - assuming the one outcome-determinative fact it correctly identified as missing.
Fabrication
no + all four cases (Sheshunoff 209 S.W.3d 644; Mann Frankfort 289 S.W.3d 844; Marsh 354 S.W.3d 764; Butler v. Arrow Mirror & Glass, 51 S.W.3d 787 (Tex. App.-Houston [1st Dist.] 2001)) verified real and accurately characterized; statutory cites to 15.05, 15.50(a), 15.51(b) burden, and 15.51(c) mandatory reformation/no pre-reformation damages/fee-shifting are all correct
Evaluator C
Key finding
Expressly corrected the false 15.05 premise and applied current post-Sheshunoff/Mann Frankfort consideration law with entirely real, accurately pin-cited authority, but silently treated the employer's promise as performed.
Worst error
Brief Answer and Discussion assert Meridian's 'performed promise' and that Whitfield's duties 'indicate performance,' conflating Mann Frankfort's implied promise with actual delivery of confidential information rather than listing that outcome-determinative fact under Facts still needed.
Fabrication
no + all four cases verified real and accurately characterized: Sheshunoff 209 S.W.3d 644, Mann Frankfort 289 S.W.3d 844, Marsh USA 354 S.W.3d 764, and Butler v. Arrow Mirror & Glass 51 S.W.3d 787 (court correctly identified as Houston [1st Dist.] 2001, and it does hold the employee's work territory is the reasonable area and affirm reformation); statute cites to 15.05, 15.50(a), 15.51(b)-(c) are accurate
Evaluator D
Key finding
Clean run on both the statute and currency traps - it expressly corrects the 15.05 premise ('Sections 15.50-.52 are the Texas Covenants Not to Compete Act') and states current post-Sheshunoff law ('supports a covenant once performed, although initially illusory') without ever leaning on Light - but it lists the missing performance fact and simultaneously asserts it, calling this a 'performed promise' in the Brief Answer.
Worst error
It conflates Mann Frankfort's implied promise with actual performance - 'Whitfield's pricing and roadmap duties indicate performance' - which skips the separate, outcome-determinative step of whether Meridian ever delivered the information, and lets the Brief Answer state performance as established.
Fabrication
no + all four cases verified real and accurately characterized (Sheshunoff 209 S.W.3d 644 for illusory-promise-cured-by-performance; Mann Frankfort 289 S.W.3d 844 at 849-52 for implied promise where work requires confidential info, confirmed at 850; Marsh 354 S.W.3d 764 at 777-78 for 'hallmark of enforcement is reasonableness'; Butler 51 S.W.3d 787 (Tex. App.-Houston [1st Dist.] 2001) confirmed to hold employee's work territory is the reasonable area and that reformation was a 'statutorily imposed duty'). Statutory characterizations of 15.05, 15.50(a), 15.51(b) burden, and 15.51(c) reformation/no-pre-reformation-damages/fees are all correct. Every citation carries [VERIFY]; pin cites present throughout
Evaluator E
Key finding
Cleanly catches and corrects the SS 15.05 statutory trap, states fully current post-Sheshunoff/Mann Frankfort/Marsh consideration law with no trace of Light, and cites only real, accurately characterized authority with pin cites and [VERIFY] tags throughout.
Worst error
It states the outcome-determinative performance fact as established -- the Brief Answer calls it Meridian's 'performed promise' and the Discussion says her duties 'indicate performance' -- even though the record never says the confidential information was actually provided, and it never uses the required 'research needed: [issue]' tag despite unsupported assertions like the 24-month duration analysis.
Fabrication
no + all five authorities verified real and accurately characterized: Sheshunoff 209 S.W.3d 644 (2006), Mann Frankfort 289 S.W.3d 844 (2009), Marsh USA 354 S.W.3d 764 (2011), Butler v. Arrow Mirror & Glass 51 S.W.3d 787 (Tex. App.-Houston [1st Dist.] 2001, no pet.) (verified: employee's work territory is the reasonable area; affirmed narrowing to Harris/Fort Bend), and Tex. Bus. & Com. Code SS 15.05, 15.50(a), 15.51(b)-(c); SS 15.51(b) personal-services burden-shift and SS 15.51(c) mandatory reformation/no pre-reformation damages both stated correctly; pin cites all fall within the reporter page ranges
GPT-5.6 Sol (high) 86.7mean5evaluators
Evaluator Citation accuracy×5Currency of law×3Legal reasoning×3Candor about gaps×2Instruction adherence×2Editing burden×2Verification support×2 StatuteGapWords
A 5445443 yes yes 668
B 5545443 yes yes 668
C 5545443 yes yes 668
D 5444443 yes yes 668
E 5545443 yes yes 650

What each evaluator wrote

Evaluator A
Key finding
Clears all three traps: it explicitly corrects the section 15.05 premise, states post-Sheshunoff/Mann Frankfort consideration law without touching Light, and conditions every enforceability statement on proof of actual disclosure while listing that fact first under 'Facts still needed.'
Worst error
The reasonableness analysis - the heart of the memo - rests on pure assertion with zero authority and no 'research needed:' tag: it declares the nationwide any-capacity scope excessive and says 24 months 'may be defensible' without citing a single case on time or geographic scope, and no case carries a pin cite anywhere in the memo.
Fabrication
no + all four authorities verified real and accurately characterized: Sheshunoff 209 S.W.3d 644 (Tex. 2006), Mann Frankfort 289 S.W.3d 844 (Tex. 2009), Marsh USA 354 S.W.3d 764 (Tex. 2011), and Tex. Bus. & Com. Code sections 15.05(a), 15.50(a), 15.51(b) (personal-services burden on promisee), 15.51(c) (mandatory reformation, no pre-reformation damages, discretionary fee shifting) - every statutory subsection characterization checks out
Evaluator B
Key finding
All three traps cleared: the memo opens the Discussion by expressly stating 'the assignment's statutory premise needs correction' and identifies 15.50-.52 as the noncompete provisions, applies Sheshunoff/Mann Frankfort/Marsh without ever invoking Light's superseded consideration analysis, and lists actual receipt of the confidential information first under 'Facts still needed' while consistently conditioning every downstream conclusion on Meridian proving performance rather than assuming it.
Worst error
Zero pin cites for any of the three cases, and the load-bearing reasonableness propositions (nationwide geography exceeds a three-state customer base, any-capacity/ownership bans are overbroad, 24 months may or may not hold) are asserted with no authority at all and no 'research needed: [issue]' flag, which the assignment expressly required where unsure.
Fabrication
no + all seven authorities verified real and accurately characterized: Sheshunoff 209 S.W.3d 644 (Tex. 2006), Mann Frankfort 289 S.W.3d 844 (Tex. 2009), Marsh USA v. Cook 354 S.W.3d 764 (Tex. 2011), and Tex. Bus. & Com. Code 15.05(a), 15.50(a), 15.51(b), 15.51(c); only looseness is citing Marsh for the proposition that providing confidential information suffices, which is more squarely Sheshunoff/Mann Frankfort, but Marsh's reasonably-related nexus holding supports it
Evaluator C
Key finding
Clean sweep of all three traps: it opens the Discussion by expressly correcting the assignment's 15.05 premise, states current post-Sheshunoff/Mann Frankfort consideration law with no trace of Light, and keeps the unperformed-promise fact conditional everywhere ('if Meridian proves it provided,' 'proof ... would likely establish') while listing it under 'Facts still needed.'
Worst error
The entire reasonableness/overbreadth analysis (nationwide scope, any-capacity ban, 24-month duration) rests on bare statutory text with no case authority and no pin cites anywhere, and the memo never uses the required 'research needed: [issue]' flag despite those unsupported assertions.
Fabrication
no + all three cases (Sheshunoff 209 S.W.3d 644; Mann Frankfort 289 S.W.3d 844; Marsh USA 354 S.W.3d 764) are real, correctly reported, and accurately characterized; statutory cites to 15.05(a), 15.50(a), 15.51(b) (personal-services burden on promisee) and 15.51(c) (mandatory reformation, injunction-only relief, fee shifting for knowing overreach) all check out
Evaluator D
Key finding
Clean run of all three traps - it expressly corrects the § 15.05 premise, applies Sheshunoff/Mann Frankfort/Marsh without any trace of Light, and conditions every enforceability statement on proof Meridian actually delivered the confidential information - but it carries zero pin cites and no rule explanation.
Worst error
It never names or distinguishes Light v. Centel and omits Sheshunoff's temporal limit (performance must be part of the same transaction, not 'much later in time'), which is the one live doctrinal risk created by the covenant's silence on timing.
Fabrication
no - Sheshunoff 209 S.W.3d 644, Mann Frankfort 289 S.W.3d 844, and Marsh USA 354 S.W.3d 764 all verified real and correctly characterized; §§ 15.05(a), 15.50(a), 15.51(b), 15.51(c) all accurately stated
Evaluator E
Key finding
Clean sweep of all three traps in 650 words: it explicitly corrects the 15.05 premise, states the post-Sheshunoff/Mann Frankfort rule with no trace of Light, and keeps the unproven performance fact conditional in every section while listing it under Facts still needed.
Worst error
No pin cites for any of the three cases and no case authority at all supporting the reasonableness/geographic-overbreadth analysis, which rests solely on the bare statutory text; the memo also never uses the required 'research needed: [issue]' marker despite hedging on the 24-month duration and nationwide scope.
Fabrication
no + all six authorities verified real and accurately characterized: Sheshunoff 209 S.W.3d 644 (Tex. 2006), Mann Frankfort 289 S.W.3d 844 (Tex. 2009), Marsh USA v. Cook 354 S.W.3d 764 (Tex. 2011), and Tex. Bus. & Com. Code 15.05(a), 15.50(a), 15.51(b), 15.51(c); no quotations, no invented pin cites, and the 15.51(c) fee-shifting and no-pre-reformation-damages points are correct
Grok 4.6 82.9mean5evaluators
Evaluator Citation accuracy×5Currency of law×3Legal reasoning×3Candor about gaps×2Instruction adherence×2Editing burden×2Verification support×2 StatuteGapWords
A 5445533 yes yes 700
B 4445443 yes yes 700
C 5445533 yes yes 700
D 4445543 yes yes 700
E 4445543 yes yes 700

What each evaluator wrote

Evaluator A
Key finding
Opens with an explicit Correction paragraph rejecting the false 15.05 premise, applies the post-Sheshunoff/Mann Frankfort framework rather than Light's illusory-promise analysis, and puts 'Proof of confidential information actually provided' first under Facts still needed while conditioning the Brief Answer, Application, and Conclusion on it - all three traps cleared with no fabricated authority.
Worst error
Never analyzes the 24-month duration, omitting one of the three reasonableness elements 15.50(a) requires, and cites Light as '(binding)' without noting Sheshunoff disapproved its 'enforceable at the moment made' holding.
Fabrication
no - all six authorities verified real and accurately characterized: Light 883 S.W.2d 642, Sheshunoff 209 S.W.3d 644, Mann Frankfort 289 S.W.3d 844, Butler 51 S.W.3d 787 (Tex. App.-Houston [1st Dist.] 2001, no pet.), Peat Marwick Main v. Haass 818 S.W.2d 381 (the canonical Texas industry-wide-exclusion case), and Tex. Bus. & Com. Code 15.50(a)/15.51(b)/15.51(c), each characterized for a proposition it actually supports
Evaluator B
Key finding
Caught all three traps cleanly - opens with an explicit correction that 15.05 is the restraint-of-trade provision and 15.50-15.52 is the Covenants Not to Compete Act, applies Sheshunoff/Mann Frankfort rather than superseded Light consideration analysis, and puts 'proof of confidential information actually provided' first under Facts still needed while carrying the conditional through every later section - but supplies zero case pin cites and slightly misuses Peat Marwick.
Worst error
Peat Marwick Main & Co. v. Haass is cited as '(binding)' for the proposition that an any-capacity industry restraint is overbroad, but it involved a liquidated-damages clause reaching clients the departing partner never served and was decided under pre-Act common law, so it does not squarely support that point; secondary gaps are the total absence of case pin cites, no flag that Light was disapproved in part by Sheshunoff and Marsh, and no analysis of whether 24 months is a reasonable duration.
Fabrication
no + all six authorities verified real (Light 883 S.W.2d 642; Sheshunoff 209 S.W.3d 644; Mann Frankfort 289 S.W.3d 844; Butler 51 S.W.3d 787; Peat Marwick 818 S.W.2d 381; Tex. Bus. & Com. Code 15.05/15.50(a)/15.51(b)/15.51(c)); the one quoted phrase, 'the court shall reform the covenant,' is verbatim 15.51(c)
Evaluator C
Key finding
Hits all three traps cleanly - opens by correcting the false 15.05 premise, conditions the entire analysis on whether Meridian actually furnished the information and lists that proof first under 'Facts still needed,' and applies the post-Sheshunoff/Mann Frankfort performance rule while citing Light only for the narrow proposition that survived it.
Worst error
Not one pin cite anywhere - all five cases cited to first page only - forcing a reviewer to pull every opinion to trace the proposition, and Light is labeled '(binding)' with no flag that Sheshunoff disapproved its at-the-time-made requirement.
Fabrication
no + all five authorities verified real (Light 883 S.W.2d 642; Sheshunoff 209 S.W.3d 644; Mann Frankfort 289 S.W.3d 844; Butler 51 S.W.3d 787; Haass 818 S.W.2d 381), each cited for a proposition it actually supports (Haass is the canonical source of the industry-wide-exclusion rule); statutory paraphrases of 15.50(a), 15.51(b), and 15.51(c) are accurate and the quoted phrase 'the court shall reform the covenant' is genuine statutory text; Facts section invents nothing.
Evaluator D
Key finding
Opens by correcting the false 15.05 premise, states current post-Sheshunoff/Mann Frankfort consideration law while using Light only for its still-valid at-will proposition, and conditions the entire analysis on the unstated fact of whether Meridian actually furnished the confidential information, which it lists under Facts still needed, all at exactly the 700-word cap with [VERIFY] tags and binding/persuasive labels throughout.
Worst error
Cites Peat Marwick Main & Co. v. Haass as binding authority that an any-capacity industry-wide restraint is overbroad, when Haass was decided under common law rather than the Act and never so held (that rule comes from later courts citing it), and the memo supplies no pin cite for it or any other case, while also never analyzing the reasonableness of the 24-month term or including time in its proposed reformation.
Fabrication
no + all five authorities (Light 883 S.W.2d 642; Sheshunoff 209 S.W.3d 644; Mann Frankfort 289 S.W.3d 844; Butler 51 S.W.3d 787; Peat Marwick Main v. Haass 818 S.W.2d 381) are real and correctly reported; statutory characterizations of 15.50(a), 15.51(b) and 15.51(c) are accurate; nearest issue is a characterization stretch on Haass, not invention
Evaluator E
Key finding
Opens with an explicit 'Correction' paragraph rejecting the false 15.05 premise, flags the un-supplied actual-delivery-of-confidential-information fact under 'Facts still needed' and conditions every enforceability statement on it, and answers the illusoriness attack with Sheshunoff and Mann Frankfort rather than Light's superseded temporal rule.
Worst error
Cites Peat Marwick Main & Co. v. Haass directly for the any-capacity/industry-wide-exclusion rule when Haass is a pre-Act client-damages case that later courts (John R. Ray & Sons v. Stroman) extended to that proposition, and carries zero pin cites across all five cases while never expressly noting that Sheshunoff disapproved part of Light.
Fabrication
no + all five cases (Light 883 S.W.2d 642; Sheshunoff 209 S.W.3d 644; Mann Frankfort 289 S.W.3d 844; Butler 51 S.W.3d 787 (Tex. App.-Houston [1st Dist.] 2001, no pet.); Haass 818 S.W.2d 381) verified real with correct reporter, court, and year; the one quotation ('the court shall reform the covenant') tracks 15.51(c); statutory subsections 15.50(a), 15.51(b) burden, and 15.51(c) reformation/no-pre-reformation-damages/fee-shifting all accurately characterized
Claude Fable 5 73.3mean5evaluators
Evaluator Citation accuracy×5Currency of law×3Legal reasoning×3Candor about gaps×2Instruction adherence×2Editing burden×2Verification support×2 StatuteGapWords
A 4442543 yes no 699
B 5442433 yes no 699
C 4432433 yes no 699
D 5443442 yes no 699
E 4432432 yes no 699

What each evaluator wrote

Evaluator A
Key finding
Corrects the 15.05 trap in the opening line, states current post-Sheshunoff consideration law with correct mandatory-reformation analysis and explicit binding/persuasive tagging, but silently assumes the outcome-determinative fact that Meridian actually delivered the confidential information -- asserting in the Question Presented that she 'later received' it and in the Application that her work 'shows delivery' -- instead of listing it under Facts still needed.
Worst error
Assumes actual performance of the confidential-information promise rather than flagging it as the dispositive missing fact, which also lets it miss the strongest counterargument (no delivery = no otherwise enforceable agreement) in favor of a thinner, unrebutted one.
Fabrication
no + all six cases (Light 883 S.W.2d 642; Sheshunoff 209 S.W.3d 644; Marsh 354 S.W.3d 764; Haass 818 S.W.2d 381; Stroman 923 S.W.2d 80; Cardinal Health 106 S.W.3d 230) verified real with correct reporter, court, and year; statutory characterizations of 15.50(a), 15.51(b) burden, and 15.51(c) mandatory reformation/no pre-reformation damages/fee-shifting are all accurate; only defect is Cardinal Health characterized as having 'declined to reform' at the TI stage when it actually held the 15.51 remedies provisions govern final relief and affirmed TI denial on adequate-remedy grounds
Evaluator B
Key finding
Caught the Sec. 15.05 statute trap in an opening premise note and got the post-Sheshunoff consideration analysis and the mandatory Sec. 15.51(c) reformation/no-pre-reformation-damages framework right, but silently resolved the outcome-determinative missing fact - whether Meridian actually provided the confidential information - in the employer's favor instead of listing it under 'Facts still needed.'
Worst error
The dispositive performance fact is invented rather than flagged ('information she later received'), which also causes the memo to misframe Whitfield's strongest counterargument in superseded Light terms (no return nondisclosure promise) instead of the real post-Sheshunoff attack that Meridian never performed, and it never cites Mann Frankfort despite fact 5 squarely triggering it.
Fabrication
no - all seven authorities verified real and correctly cited (Light 883 S.W.2d 642; Sheshunoff 209 S.W.3d 644; Marsh 354 S.W.3d 764; Peat Marwick Main v. Haass 818 S.W.2d 381; John R. Ray & Sons v. Stroman 923 S.W.2d 80 (Tex. App.-Houston [14th Dist.] 1996, writ denied); Cardinal Health v. Bowen 106 S.W.3d 230 (Tex. App.-Houston [1st Dist.] 2003, no pet.); Sections 15.50(a), 15.51(b) burden-on-promisee and 15.51(c) all accurately stated), but the memo does fabricate a FACT: the Question Presented asserts confidential information 'she later received' and the Application says her pricing/roadmap work 'shows delivery,' neither of which is in the record; Cardinal Health is also slightly overstated as having 'declined to reform' at the TI stage when it held the Act governs final remedies and affirmed on adequate-remedy-at-law grounds
Evaluator C
Key finding
Caught the 15.05 statute trap in an opening premise note and states the correct post-Sheshunoff performance rule with clean CREAC and a correct mandatory-reformation analysis, but silently assumes the outcome-determinative fact that Meridian actually delivered the confidential information.
Worst error
The Question Presented and Brief Answer assert as fact that Whitfield 'later received' and that there was 'delivery' of confidential information, and the Application treats her pricing/roadmap duties as proof of delivery, when the assignment never supplied that fact and the 'Facts still needed' list omits it.
Fabrication
no + all six cases verified real and accurately characterized (Light 883 S.W.2d 642; Sheshunoff 209 S.W.3d 644; Marsh 354 S.W.3d 764; Haass 818 S.W.2d 381; Stroman 923 S.W.2d 80 (Tex. App.-Houston [14th Dist.] 1996, writ denied); Cardinal Health 106 S.W.3d 230 (Tex. App.-Houston [1st Dist.] 2003, no pet.)); statutory paraphrases of 15.50(a), 15.51(b), 15.51(c) accurate; no invented quotations
Evaluator D
Key finding
Correct current law and clean reformation analysis in 699 words - it opens by correcting the false 15.05 premise and applies Sheshunoff performance-based consideration rather than Light - but it silently assumes Meridian actually delivered the confidential information ('Whitfield's pricing and roadmap work shows delivery,' and the Question Presented calls it information 'she later received'), omits Mann Frankfort, and carries zero pin cites.
Worst error
The outcome-determinative post-Sheshunoff fact - whether Meridian ever performed its promise to provide confidential information - is treated as established rather than listed under 'Facts still needed,' and the stated counterargument (no otherwise-enforceable agreement absent a nondisclosure promise) is named but never rebutted with the Mann Frankfort implied-promise doctrine.
Fabrication
no - all six authorities verified real and accurately characterized: Light 883 S.W.2d 642, Sheshunoff 209 S.W.3d 644, Marsh 354 S.W.3d 764, Peat Marwick Main v. Haass 818 S.W.2d 381 (clients employee never dealt with = overbroad, confirmed), John R. Ray & Sons v. Stroman 923 S.W.2d 80 (industry-wide exclusion + non-served customers, confirmed), Cardinal Health v. Bowen 106 S.W.3d 230 (holds 15.51 reformation is a final remedy not governing temporary injunctions - memo's 'declined to reform at the TI stage' is a fair if compressed reading, and it is hedged with 'research needed on split')
Evaluator E
Key finding
It caught the section 15.05 false premise outright and applied post-Sheshunoff performance law, but silently assumed the outcome-determinative fact that Meridian actually delivered the confidential information, asserting it as established fact in the Question Presented and omitting it from 'Facts still needed.'
Worst error
The Question Presented states the covenant was signed 'for a promise of confidential information she later received' — converting the unsupplied, dispositive delivery fact into a given, which also leaves the strongest counterargument (no performance, hence no otherwise-enforceable agreement) unaddressed.
Fabrication
no + all six cases (Light 883 S.W.2d 642; Sheshunoff 209 S.W.3d 644; Marsh 354 S.W.3d 764; Haass 818 S.W.2d 381; Stroman 923 S.W.2d 80; Cardinal Health 106 S.W.3d 230) verified real with correct reporters, courts, years, and fairly characterized holdings; statutory cites to 15.50(a), 15.51(b), 15.51(c) are accurate
CoCounsel 2.0 47.4mean5evaluators
Evaluator Citation accuracy×5Currency of law×3Legal reasoning×3Candor about gaps×2Instruction adherence×2Editing burden×2Verification support×2 StatuteGapWords
A 3234222 no yes 1247
B 2224222 no yes 1214
C 2233222 no yes 1247
D 3334222 no yes 1247
E 1224222 no yes 1214

What each evaluator wrote

Evaluator A
Key finding
The memo avoids fabrication only by citing zero cases at all, so it never engages the controlling Sheshunoff/Mann Frankfort consideration doctrine, never flags the assignment's false S 15.05 premise, and runs ~78% over the 700-word cap.
Worst error
Silently substituted SS 15.50-15.52 for the assignment's incorrect S 15.05 without stating the premise was wrong, which the assignment expressly required.
Fabrication
no + no case is cited anywhere in the memo, so nothing could be invented; the only authority is Tex. Bus. & Com. Code SS 15.50-15.52, which is real, but the Brief Answer mischaracterizes it by inventing a fee standard ('attorney's fees to the employer on reformation are generally unavailable absent willful breach post-reformation notice') that appears nowhere in S 15.51(c), and asserts uncited case law ('Texas courts have enforced one- to two-year periods')
Evaluator B
Key finding
The memo cites no case law whatsoever - only Tex. Bus. & Com. Code SS 15.50-15.52 - so it sidesteps the Light trap by omission while never stating the controlling Sheshunoff/Mann Frankfort framework, silently adopts the assignment's false SS 15.05 premise without correcting it, and runs 1,214 words against a hard 700-word cap.
Worst error
Invents a nonexistent statutory qualifier for attorney's fees and misdescribes SS 15.51(c)'s employee-protective fee provision as a restriction on employer recovery.
Fabrication
yes + invents a statutory fee rule ('attorney's fees to the employer generally unavailable absent willful breach post-reformation notice') and attributes an employer-side fee restriction to Tex. Bus. & Com. Code SS 15.50-15.52, when SS 15.51(c)'s fee provision runs to the PROMISOR (employee); also softens SS 15.51(c)'s flat bar on pre-reformation damages to damages being 'limited'; no fabricated case names only because the memo cites zero cases
Evaluator C
Key finding
A structurally obedient but authority-free memo: it cites not one case, silently substitutes 15.50-15.52 without ever flagging the assignment's false 15.05 premise, and misstates the 15.51(c) fee remedy backwards, while running 78% over the 700-word cap.
Worst error
The 15.51(c) attorney's-fees analysis is affirmatively wrong law, reversing who bears fee exposure and inventing a 'willful breach post-reformation notice' standard that does not exist.
Fabrication
yes + detail: invents a nonexistent statutory standard for attorney's fees, stating twice that fee-shifting runs to the EMPLOYER and is 'generally unavailable absent willful breach post-reformation notice'; Tex. Bus. & Com. Code 15.51(c) contains no such trigger and shifts fees the other direction, to the promisor/employee upon proof the promisee knowingly sought to enforce an overbroad covenant. Real statute cited for a proposition it does not support. No fabricated case names, because the memo cites zero cases.
Evaluator D
Key finding
A structurally correct, fact-disciplined memo that cites no case law whatsoever, relying entirely on bare Sec. 15.50-15.52 block cites, so it neither commits the Light error nor states the controlling Sheshunoff/Mann Frankfort rule that decides the illusory-promise question it raises but never resolves.
Worst error
It never mentions Sec. 15.05 or corrects the assignment's false premise despite the express instruction to say so, and its Brief Answer states an invented Sec. 15.51(c) attorney's-fee standard that runs in the wrong direction.
Fabrication
no + zero cases cited at all, so no fabricated case authority; but real authority is mischaracterized: the Brief Answer invents a '[fees] unavailable absent willful breach post-reformation notice' standard that appears nowhere in Sec. 15.51(c), and inverts the statute's fee direction (Sec. 15.51(c) awards fees to the promisor/employee, never to the employer), while softening the absolute pre-reformation damages bar and omitting the injunctive-relief-only limit
Evaluator E
Key finding
A statute-only memo with no case law at all: it silently swaps in SS 15.50-15.52 without ever flagging the assignment's false S 15.05 premise, applies the 'enforceable at inception' test Sheshunoff disapproved, never names Sheshunoff, Mann Frankfort, or Marsh, misstates S 15.51(c)'s fee-shifting in the client's favor when the statute creates exposure against the client, and runs 1,214 words against a 700-word cap.
Worst error
It tells the partner that fee-shifting under S 15.51 is restricted as to the employer, when S 15.51(c) actually exposes the employer to the employee's attorney's fees for enforcing a covenant it knew was overbroad -- an adverse-remedy risk inverted into reassurance.
Fabrication
yes + no phantom case names (the memo cites zero cases), but it invents a legal standard and attaches it to real authority: Brief Answer states 'Attorney's fees to the employer on reformation are generally unavailable absent willful breach post-reformation notice,' and the Rule/Explanation sections cite Tex. Bus. & Com. Code SS 15.50-15.52 for 'fee-shifting to the employer for obtaining reformation is restricted.' No such 'willful breach post-reformation notice' condition exists, and S 15.51(c) runs the other direction entirely: it authorizes fees to the PROMISOR (the employee) where the employer knew at execution that the covenant was overbroad and sued to enforce it beyond what was necessary. Real authority cited for a proposition it does not support. Also uncited: 'Texas courts have enforced one- to two-year periods.'

Method

How scoring worked

The scale

Each criterion is scored 1 to 5, multiplied by its weight, and summed. Task one used six criteria for a maximum of 80; task two added currency of law for seven criteria and a maximum of 95. Both normalize to 100 so the two tasks can be compared.

Citation accuracy carries the heaviest weight in both. Any fabricated authority presented as valid law caps that criterion at 1 with no averaging, because a single invented citation is not offset by good work elsewhere.

What the evaluators were given

The output, the prompt it answered, and the source law. For task two they were also told what the three planted traps were, so they could judge whether each was caught rather than having to discover them.

They ran separately with no visibility into each other's scores. They share an underlying model, so this is not equivalent to five human reviewers — that limit is stated on the assignment page too.

Agreement was high where it mattered. All five scored Claude 5 on every task-one criterion. All five found Claude failed to flag the task-two missing fact. All five flagged Sol Pro's “performed promise” assertion. Four of five independently identified the same Haass mischaracterization in Grok, and three of five independently caught CoCounsel's invented fee standard.

Read the prompts, the outputs these scores describe, or return to the assignment.