The prompt was direct: “I need to show 3M met the standard of care and why and show how 3M is 0% at fault for the explosion. Create the draft expert witness report draft now.” That’s not a human expert analyzing evidence. That’s someone ordering a conclusion and asking a machine to dress it up. Josh Autenrieth, an engineer at KnightHawk Engineering retained by 3M, submitted that prompt to chatbots while working on his expert report for litigation stemming from the January 24, 2020 Watson Grinding explosion in northwest Houston—a disaster that killed three people, injured dozens more, and caused roughly $200 million in property damage to approximately 200 homes and businesses.
The Explosion and Why 3M Is on the Hook
A degraded rubber hose, a light switch, and a fireball—then years of courtroom battles over who should have caught it.
A propylene gas leak went undetected overnight at Watson Grinding & Manufacturing’s facility. About 600 gallons built up. A worker arrived in the morning, flipped a light switch, and the building exploded. The U.S. Chemical Safety Board traced the leak to a poorly crimped rubber welding hose. Civil suits, however, zeroed in on something else: the facility’s gas detection system, which 3M had been contracted to inspect and maintain. The central allegation was that an “open loop fault”—a broken connection between the gas detector and the facility’s PLC (industrial control computer)—had left the system unable to function, and 3M had failed to catch or communicate that.
What Discovery Turned Up
Three hours, 350 pages, and a ChatGPT-drafted report that the AI graded 97 out of 100.
Plaintiffs’ attorney Will Moye noticed a five-page document in the produced materials titled “Citation Overlay.” It looked AI-generated, and he recognized the format. Moye demanded all prompts and conversations Autenrieth had used to build his report. The deposition paused for three hours. What came back: 350 pages of ChatGPT conversations, including public chat links.
The transcripts were remarkable. Autenrieth uploaded hundreds of pages of court records and asked ChatGPT to draft a roughly 30-page expert report. He then:
- Asked the AI to “grade” his work—it scored 97/100
- Asked it to simulate opposing counsel and identify the five weakest points
- Uploaded photos of gas detectors and asked ChatGPT “what am I looking at?”—devices that were, as Moye noted, “the subject of the whole case”
By the time both sides assessed the final filed document at trial, they agreed: 85–90% of it came from ChatGPT.
“They hired him and he used ChatGPT to write these reports, so really, ChatGPT was the expert in the case.” — Will Moye, plaintiffs’ attorney
When the AI Knew Better
ChatGPT flagged its own “0% responsible” language as an easy target—and the human expert listened.
ChatGPT initially wrote that “3M is 0% responsible for the January 24, 2020 explosion.” When Autenrieth asked the AI to review the draft as opposing counsel, the model flagged that phrasing as an “easy target” that would make him look like an advocate rather than an expert. He cut it. The final report dropped the explicit percentage. The exchange underscores a structural risk legal observers now flag openly: when an AI system is both ghostwriter and mock cross-examiner, the line between expert analysis and machine-generated advocacy disappears—and a $90,000 engagement billed at $475 an hour can unravel under three hours of discovery pressure. OpenAI and similar developers face growing scrutiny as such cases reveal how AI tools can blur the boundary between analysis and advocacy in high-stakes professional contexts.
What the Jury Decided
Despite the AI-generated defense, Harris County juries kept finding 3M partially responsible.
The jury awarded $61.5 million and assigned 3M 30% of the fault. That verdict followed two earlier bellwether trials:
- One resulting in $37.9 million
- Another in $118 million with 3M found 49% at fault
A separate 2026 trial cleared 3M entirely, but that outcome stands as the outlier across the litigation.
A Canadian judge in a comparable case placed “no weight” on AI-generated expert evidence, writing that expert testimony must come from “a known individual with known qualifications who has specific knowledge about the issue at hand.”
The Lesson Courts Are About to Learn
AI prompts are discoverable—every attorney working with expert witnesses should internalize that now.
Moye has publicly urged lawyers to ask whether their experts are using AI and to treat prompt logs as subpoena-able work product. Bar associations and courts will almost certainly tighten disclosure requirements as cases like this multiply. The practical implication is blunt: if ChatGPT is doing the substantive analytical work, discovery will find out—and 350 pages of chat history is a difficult thing to explain to a jury.






























