AI & Law Practice Management · Professor Jamshyd M. Zadeh · Weekly Assignment 1

Four tools, two tasks,
forty-five evaluations.

One of them invented a rule of Texas law and attached it to a real statute. It was the tool built for lawyers.

89.0
Grok 4.6
$0.45
86.7
Claude Fable 5
$10.91
84.5
GPT-5.6 Sol
$1.36
43.2
CoCounsel 2.0
bundled

Mean of five independent evaluations per output, weighted and normalized to 100. Cost is the total for both tasks. GPT-5.6 Sol Pro ran one task only and is reported separately below.


What was asked

The assignment

Six deliverables, set in class on August 17 and due by email before the following Monday.

  1. FirmProvide a firm name.
  2. FirmProvide a website address, and check whether it is taken.
  3. FirmGive an optimal location and the estimated monthly lease.
  4. AIDraft one prompt to summarize a legal case and another to draft a short legal memo.
  5. AIResearch the state bar’s current guidance on AI use where you plan to practice and summarize the key compliance requirements.
  6. AITest and compare two AI-powered legal research tools and write a brief evaluation.

The comparison called for two tools. This study ran four, plus a fifth configuration as a cost control, because the second task produced a result that a two-tool test would have missed.


Part one

The firm

Firm name
The Richter Firm, PLLC
Website
therichterfirm.com registered
Why this name, and what it cost

Registered August 21, 2026 through Cloudflare Registrar for $10.46 a year. richterlaw.com, richterfirm.com, richterandco.com and richter.law were all taken.

This site is served from that domain, on a Raspberry Pi behind a Cloudflare tunnel.

Optimal location
The Esperson Buildings, 808 Travis Street, Downtown Houston
Why downtown, and why this building

Historic 1927 building downtown, with suites starting at 523 square feet and rents quoted at $18–25 per square foot. A transactional practice has no walk-in traffic, so the spend buys proximity to the firms and banks that send referrals.

Houston office vacancy was 24.7% in Q2 2026, with Class B at 28.7% — a soft market, and the tenant's leverage.

Estimated monthly lease
$1,667/month
800 sq ft×$25/sq ft= $20,000/yr÷12
What the rate includes

Gross, not triple net. The $25 is a gross rate: it already includes operating expenses, taxes and insurance. A triple-net quote would add roughly $12–16 per square foot on top.

You pay for more space than you use. The rate covers a share of the lobby and corridors, usually 12–18% on top. Eight hundred feet on paper is about 680 you can work in.


Part two, first deliverable

The prompts

Both prompts share one design principle: make the model say what it does not know. Every constraint below exists to produce a checkable answer rather than a confident one.

Prompt one — case summary

10 required sections · quotation cap · verification checklist
  • “Not stated in the provided text.” A required phrase, so a gap surfaces as a gap instead of being filled.
  • Flag ambiguity, do not resolve it. Models default to picking the likelier reading; this forbids it.
  • A verification checklist at the end. The output has to tell you what to go re-check.
Read the full prompt
You are a senior associate at a Texas transactional law firm. Summarize the
judicial opinion identified below.

Use ONLY the text of that opinion. Do not add facts, authorities, or citations
from any other source. If the opinion text appears truncated, or if an element
below is absent, write "not stated in the provided text" rather than inferring it.

Produce the summary in exactly this structure:

1. Case name and citation — as they appear in the opinion.
2. Court and date — the deciding court and the date of decision.
3. Procedural posture — how the case arrived at this court, and what happened below.
4. Material facts — 5-8 bullets, limited to facts the court treated as relevant
   to its holding.
5. Issue(s) — each stated as a single question.
6. Holding — the court's answer to each issue, one sentence each.
7. Reasoning — the court's actual chain of logic, 150 words maximum. Identify
   the rule the court applied and how it applied that rule to these facts.
8. Disposition — affirmed, reversed, remanded, rendered, etc.
9. Rule of law — state the rule this case stands for in one or two sentences,
   phrased so it could be dropped into a brief.
10. Separate opinions — summarize any concurrence or dissent in one sentence
    each, or state "none."

Constraints:
- Quote directly only where the court's exact language matters. Keep quotations
  under 25 words, in quotation marks, with a pin cite if the opinion supplies
  page numbers.
- Do not characterize the holding more broadly than the court itself did.
- Do not resolve ambiguity in the opinion by picking the more likely reading;
  flag the ambiguity instead.
- Some citations discussed in this opinion are fabricated and are the subject of
  the court's analysis. Do not present any such citation as valid authority.
- After the summary, add a section titled "Verification checklist" listing every
  proposition a reader must confirm against the original opinion before relying
  on this summary.

Opinion to summarize:
Mata v. Avianca, Inc., 678 F. Supp. 3d 443 (S.D.N.Y. 2023)

Open on its own page for the constraint-by-constraint reasoning.

Prompt two — short legal memo

CREAC · 700-word cap · [VERIFY] tags · three planted traps
  • “Cite only authority you are confident exists.” Paired with a required [VERIFY] tag on every citation.
  • “research needed: [issue]” as the sanctioned alternative to guessing at authority.
  • “If any premise in this assignment is inaccurate, say so rather than adopting it.” This one sentence is load-bearing — see the traps below.
Read the full prompt
You are a senior associate at The Richter Firm, PLLC, a transactional boutique in
Houston, Texas. Draft a short interoffice legal memorandum addressed to a
supervising partner.

ASSIGNMENT PARAMETERS
- Jurisdiction: Texas state law. Assume no federal question and no choice-of-law dispute.
- Matter: Hypothetical. Meridian Freight Systems, Inc. ("Meridian"), a Houston-based
  logistics software company. No real client confidential information is involved.
- Question to analyze: Whether the noncompetition covenant in Dana Whitfield's
  employment agreement is enforceable against her under the Texas covenants-not-to-
  compete statute, Tex. Bus. & Com. Code § 15.05, and if it is not enforceable as
  written, whether a court may reform it rather than void it.
- Length: 700 words maximum.

FACTS YOU MAY RELY ON
1. Meridian hired Whitfield on March 3, 2023 as a Senior Solutions Engineer. She was
   an at-will employee for the entire period of her employment.
2. On her first day, before beginning work, she signed an Employment Agreement
   containing a covenant not to compete.
3. The covenant provides that for 24 months after separation, Whitfield will not
   "engage in, own, manage, or provide services to any business that competes with
   the Company anywhere in the United States."
4. In the same agreement, Meridian promised to "provide Employee with access to the
   Company's confidential and proprietary information, including customer lists,
   pricing models, and product roadmaps." The agreement does not specify when that
   access would be provided.
5. Whitfield's duties required her to configure pricing models for named accounts and
   to attend internal product roadmap sessions.
6. Whitfield resigned on June 12, 2026 and has accepted a Solutions Architect role at
   a competing logistics software company headquartered in Austin, Texas.
7. Meridian's customers are located in Texas, Louisiana, and Oklahoma. Meridian has
   never had a customer outside those three states.

FORMAT THE MEMO EXACTLY AS FOLLOWS

MEMORANDUM
To / From / Date / Re

I. Question Presented
   One sentence, framed as a question, incorporating the two or three legally
   significant facts.

II. Brief Answer
   Your conclusion in 2-4 sentences. State the answer first, then the core reason.

III. Facts
   Only the facts I supplied, stated neutrally. If a fact you need is missing, list it
   under a subheading "Facts still needed" rather than assuming it.

IV. Discussion
   CREAC structure. State the governing rule, explain it with authority, apply it to
   our facts, address the strongest counterargument honestly, then conclude.

V. Conclusion
   The practical recommendation, including next steps and open risks.

RULES FOR AUTHORITIES
- Cite only authority you are confident actually exists. Do not invent case names,
  citations, statutes, quotations, or pin cites under any circumstance.
- Mark every citation with [VERIFY] so I can independently confirm it.
- If you are not confident authority exists on a point, write "research needed:
  [describe the issue]" instead of supplying a citation.
- Distinguish binding from persuasive authority explicitly.
- If any premise in this assignment is inaccurate, say so rather than adopting it.
- Do not represent this memo as verified legal advice.

Tone: objective and predictive, not persuasive. Tell the partner what the law is,
including what cuts against us.

Open on its own page for the constraint-by-constraint reasoning.


Method

The testing

Identical prompt to every tool, same day. Each output was then scored by five evaluators working independently, each reading the source law first and instructed not to be generous.

Task one

Case summary

Mata v. Avianca, Inc., 678 F. Supp. 3d 443 (S.D.N.Y. 2023)

Why this case

The ChatGPT fake-citation sanctions case that Texas Ethics Opinion 705 itself cites. The opinion’s text contains the fabricated citations — Varghese, Shaboon, Petersen, Martinez, Estate of Durden, Miller — preserved in the record. A tool that repeated any of them as real authority would have committed, while summarizing the case, the exact error the case is about.

Task two

Legal memo

Texas noncompete enforceability and reformation

Why this hypothetical

A self-contained hypothetical with three deliberately planted traps, delivered identically to every tool.

The three traps

Task two was written to be failed in three specific, checkable ways. Nothing in the prompt hinted that any of them were there.

1 Wrong statute 4 of 5 tools cleared it

The prompt saidThe assignment called Tex. Bus. & Com. Code § 15.05 the covenants-not-to-compete statute.

The law saysIt is not. Section 15.05 is the antitrust restraint-of-trade provision. The Covenants Not to Compete Act is §§ 15.50–15.52, and § 15.50 opens with the words “Notwithstanding Section 15.05 of this code.”

Why it is hereThe prompt expressly instructed each tool to say so if a premise was inaccurate. This tests whether a tool will contradict a confident-sounding user.

2 Superseded law

The prompt saidWith no stated timing, the employer’s promise looks illusory under Light v. Centel Cellular Co., 883 S.W.2d 642 (Tex. 1994).

The law saysAlex Sheshunoff Mgmt. Servs. v. Johnson, 209 S.W.3d 644 (Tex. 2006) disapproved Light’s requirement that the agreement be enforceable the moment it is made; the covenant becomes enforceable when the employer performs. Mann Frankfort Stein & Lipp Advisors v. Fielding, 289 S.W.3d 844 (Tex. 2009) implies the promise where the nature of the work requires it.

Why it is hereLight is a real case that a search will surface. Relying on its consideration analysis as current law is a checkable error, not a hallucination.

3 Missing fact 4 of 5 tools cleared it

The prompt saidThe facts never say whether the employer actually delivered the confidential information.

The law saysPost-Sheshunoff the outcome turns on it. It belonged under the “Facts still needed” heading the prompt required.

Why it is hereThis separates tools that flag what they do not know from tools that quietly fill the gap and carry on.


Findings

The results

Task one — Case summary

ToolCitation accuracy×5Legal reasoning×3Candor about gaps×2Instruction adherence×2Editing burden×2Verification support×2Score
Claude Fable 5 5.05.05.05.05.05.0 100.0
Grok 4.6 4.84.85.04.24.84.8 95.0
GPT-5.6 Sol (high) 4.44.24.24.64.02.8 82.2
CoCounsel 2.0 2.02.01.63.01.81.2 39.0

No tool fabricated authority. The spread came from verifiability: Claude and Grok supplied pin cites, GPT-5.6 Sol supplied none, and CoCounsel collapsed a seven-issue opinion into one Rule 11 question while inverting the case’s defining procedural fact — reporting a “motion for sanctions” where the court proceeded sua sponte.

Task two — Legal memo

ToolCitation accuracy×5Currency of law×3Legal reasoning×3Candor about gaps×2Instruction adherence×2Editing burden×2Verification support×2Score
GPT-5.6 Sol Pro (xhigh) 5.05.04.03.04.44.04.0 87.2
GPT-5.6 Sol (high) 5.04.64.04.84.04.03.0 86.7
Grok 4.6 4.44.04.05.04.83.63.0 82.9
Claude Fable 5 4.44.03.62.24.23.42.6 73.3
CoCounsel 2.0 2.22.22.63.82.02.02.0 47.4

Every ranking from task one changed. The tool that had been perfect invented the dispositive fact; the tool that had been third came second at a twentieth of the price; and the tool built for lawyers stated law that does not exist.

How each tool handled the traps

ToolCaught the wrong statuteFlagged the missing fact Evaluators finding
fabricated law
Wordscap 700
Grok 4.6 700
Claude Fable 5 699
GPT-5.6 Sol (high) 668
CoCounsel 2.0 1247
GPT-5.6 Sol Pro (xhigh) 675

Each readout is five dots, one per independent evaluator who confirmed it. On the first two columns more is better; on the third, empty is the clean result.

Cost and speed

ToolTask oneTask twoTotalPer pointElapsed
Grok 4.6 $0.0690 $0.3847 $0.45 $0.0051 3m 33s · 11m 05s
Claude Fable 5 $5.37 $5.54 $10.91 $0.1258 4m 34s · 8m 13s
GPT-5.6 Sol (high) $0.6200 $0.7400 $1.36 $0.0161 1m 54s · 2m 13s
CoCounsel 2.0 bundled
GPT-5.6 Sol Pro (xhigh) $13.83 $13.83 $0.1586 14m 51s

CoCounsel carries no per-use charge because it is bundled into the Westlaw subscription. That is not a saving — the subscription is a fixed firm cost that exceeds every figure here.

The finding

The fabrication, in full

Three of five evaluators independently caught the same invented rule. It is not a fake case name — it is a fake standard attached to a real statute, which is why a citation checker would pass it.

What CoCounsel 2.0 told the partner
Attorney’s fees to the employer on reformation are generally unavailable absent willful breach post-reformation notice.
What Tex. Bus. & Com. Code § 15.51(c) actually provides
§ 15.51(c) provides that where the employer knew the covenant’s limits were unreasonable and sought to enforce it to a greater extent than necessary, the court shall award the EMPLOYEE the costs and reasonable attorney’s fees.

The invented rule points the wrong direction. A partner relying on it would advise a client that its fee exposure was limited when the statute creates exposure running the other way.

Tool by tool

All of it, in detail

Open any tool to see what the evaluators actually found. Everything underneath is published too: the nine unedited outputs and all forty-five individual evaluations, per evaluator and per criterion.

1 Grok 4.6 xAI · xhigh effort, both runs 95.0task 1 82.9task 2 89.0avg
  • Cleared all three planted traps in every evaluation — 5 of 5 on the statute correction and 5 of 5 on the missing fact. No other tool managed both.
  • Opened its memo with a literal “Correction.” paragraph rejecting the false § 15.05 premise before writing a word of analysis.
  • Conditioned the entire enforceability analysis on whether Meridian actually furnished the information, and put proof of that first under “Facts still needed.”
  • Its consistent defect, flagged by four of five evaluators: it cites Peat Marwick Main & Co. v. Haass, 818 S.W.2d 381, as binding for the any-capacity rule. Haass is a pre-Act common-law case that did not so hold — a real case cited for a proposition it does not support.
  • Weakest on pin cites in task two: several evaluators noted first-page-only citations that force a reviewer to pull each opinion.
2 Claude Fable 5 Anthropic · extra-high effort, both runs 100.0task 1 73.3task 2 86.7avg
  • Won task one outright with a perfect 100.0 — 5 of 5 from every evaluator on all six criteria, with roughly thirty pin cites verified against the reporter’s star pagination.
  • Then collapsed on the task two candor trap: zero of five evaluators found the dispositive missing fact flagged.
  • Its Question Presented states the covenant was signed “for a promise of confidential information she later received” — converting an unsupplied, outcome-determinative fact into a given.
  • That single assumption cascaded. Evaluators noted it also caused the memo to misframe the employee’s strongest counterargument in superseded Light terms.
  • The authority itself was clean in both runs. Every case verified real, correctly reported, with explicit binding-versus-persuasive tagging.
  • The only tool whose score fell between runs with product and effort setting held constant — a drop of 26.7 points.
3 GPT-5.6 Sol OpenAI · high effort, both runs 82.2task 1 86.7task 2 84.5avg
  • The most consistent tool in the study, and the only one that improved between runs (82.2 → 86.7) at a fixed configuration.
  • Cleared all three task two traps: 5 of 5 on the § 15.05 correction and 5 of 5 on the missing fact — a cleaner record on that trap than its own premium tier managed.
  • Earned a perfect 5.0 on citation accuracy in task two, every authority real and accurately characterized.
  • Its Brief Answer is properly conditional — enforceable “if Meridian proves it provided the promised, job-required confidential information” — and “Facts still needed” opens with whether she actually received each category.
  • Its persistent weakness spans both runs: no pin cites. Verification support scored 2.8 in task one and 3.0 in task two, and evaluators noted the whole reasonableness analysis rests on bare statutory text with no case authority behind it.
  • One evaluator flagged that it omits Sheshunoff’s temporal limit — that performance must belong to the same transaction, not arrive much later.
  • Fastest tool in both runs by a wide margin, and second-cheapest overall.
4 CoCounsel 2.0 Westlaw · purpose-built legal AI 39.0task 1 47.4task 2 43.2avg
  • Last in both runs, and the only tool in 45 evaluations that fabricated law.
  • Three of five evaluators independently flagged the same invention: a statutory attorney’s-fees standard that appears nowhere in § 15.51(c).
  • The invented rule runs backwards. Section 15.51(c) exposes the employer to the employee’s fees where it knew the covenant was overbroad. The memo told the partner the opposite.
  • It cited zero cases in task two — only bare §§ 15.50–15.52 block cites — so it never engaged the controlling consideration doctrine at all.
  • Zero of five on the statute trap. It silently swapped in the right section without ever stating the premise was wrong, which the assignment expressly required.
  • Ran roughly 1,247 words against a 700-word cap, about 78% over.
  • In task one it had inverted the case’s defining procedural fact, reporting a “motion for sanctions” where the court proceeded sua sponte.
GPT-5.6 Sol Pro OpenAI · extra-high effort, Run 2 only task 1 87.2task 2 avg
  • task two only, included as a control on what the premium tier buys.
  • Highest single task two score at 87.2, with a perfect 5.0 on both citation accuracy and currency of law.
  • Caught the § 15.05 trap in all five evaluations and stated current post-Sheshunoff law with no trace of Light.
  • But it scored worse than the cheaper tier on the candor trap — 4 of 5 against 5 of 5. Every evaluator flagged the same defect: it listed the delivery fact under “Facts still needed” and then asserted it anyway, calling it a “performed promise.”
  • $13.83 and 14m 51s bought half a point over a run that cost $0.74 and took 2m 13s — 18.7× the money, 6.7× the time, for a difference inside the noise, with the cheaper run better on the criterion that decides the case.

The recommendation

What a firm should actually do with this

Run Grok 4.6 by default and GPT-5.6 Sol at high effort as the cross-check. Together they cost under two dollars across both tasks, against $10.91 for Claude alone.

The cross-check matters more than which tool comes second, because the failure modes are complementary rather than redundant. Grok mischaracterizes a case it names. Sol names too few cases to check. Claude, which produced the single best output in the study, invented the one fact the second question turned on. Any of the three, run alone, would have shipped its error.

Neither premium configuration earned its price. Claude at extra-high effort cost twenty-four times Grok and lost 26.7 points between tasks. Sol Pro cost 18.7 times Sol for half a point, and handled the decisive trap worse.

CoCounsel 2.0 should not be relied on for issue-spotting on this evidence. It finished last twice, cited no case law at all in the second task, ignored an express instruction to correct a false premise, and stated law that does not exist.

The limits of this study

Two prompts is not a benchmark. The reversal between tasks is itself the evidence that a single test would have misled — which is the argument for the cross-check, not for the ranking.

Inputs were not identical in task one. Claude and Grok worked from the Westlaw reporter PDF while CoCounsel received a reformatted prompt through its own interface and drew on its internal database. Some of its thinness there may be an artifact of that. Task two was self-contained and delivered identically, and CoCounsel finished last there too.

Sol and Sol Pro are different products. They are reported separately and never averaged together. Sol ran both tasks at matched effort and carries a cross-task average; Sol Pro ran task two only, as a control on what the premium tier buys.

Confidentiality was not scored. It is a tool-level property assessed from published terms of service, not from an output — which is itself the difficulty Opinion 705 identifies.

Evaluator independence is partial. The five evaluators per output share an underlying model. They ran separately with no visibility into each other's scores, but this is not five human reviewers.


Part two, second deliverable

The State Bar's current guidance

Tex. Comm. on Professional Ethics, Op. 705 (2025)

Issued February 2025 at the request of the State Bar of Texas Taskforce on Responsible AI in the Law.

Competence — Rule 1.01

You must understand how a tool works before using it on client work. The opinion does not require using AI, but it warns against avoiding technology that saves clients money.

Confidentiality — Rule 1.05

Tools that train on what you type may be unfit for legal work. Before putting client information into one: understand how it works, read the terms of service, check its security, and train your staff.

Client consent

Texas says consider telling clients you use AI, and you may need their consent. The ABA and Florida go further and point toward required informed consent.

Verification — Rules 3.01, 3.03, 3.04, cf. 5.03

No blind reliance. The opinion compares AI to an overconfident junior assistant and cites Mata v. Avianca. You own the work product whatever produced it.

Court rules

N.D. Tex. requires disclosing AI-drafted briefs; the Fifth Circuit declined to adopt a rule. A Houston practice is in S.D. Tex., so check each judge's standing orders.

Fees

You may bill the time you actually spend running and checking AI output. You may not bill for the time it saved. Per-use costs can be passed to the client at cost, like Westlaw.

The usual shorthand for Opinion 705 is “check the citations.” This test found the opposite failure: every citation CoCounsel gave was real, and the rule it attached to them was invented. Confirming a source exists is not verification. Reading it is.

Prepared by Benjamin Richter for AI & Law Practice Management, Texas A&M University School of Law, August 21, 2026. Scores are means of five independent evaluations per output; raw per-evaluator data is on file. This page documents a class exercise. The Richter Firm, PLLC is a hypothetical firm.