OpenAI Is Firing Contractors for Using AI to Train Its AI

Contractors earning over $50 an hour score real ChatGPT conversations by hand, and getting caught using AI means immediate removal

C. da Costa Avatar
C. da Costa Avatar

By

Image: Deposit Photos

Key Takeaways

Key Takeaways

  • OpenAI fires contractors instantly for using AI tools to evaluate ChatGPT responses.
  • Recursive AI training on synthetic data narrows model accuracy, driving OpenAI’s strict human-only rule.
  • ChatGPT users are enrolled in Project Lily data training by default; opt out via data controls.

OpenAI promotes AI as a tool for every workplace task, from drafting emails to writing code. Behind that pitch, hundreds of contractors are doing the opposite, by strict order. These workers read your real ChatGPT conversations and score the model’s responses. Use AI to help with that job, and you’re out.

What Project Lily Actually Is

The program is more hands-on with your data than most users realize.

Codenamed Project Lily, it recruits contractors through a firm called Crossing Hurdles and pays them via Mercor, an AI-training company. Internal documents obtained by 404 Media show some contractors earn more than $50 per hour.

Reviewers read an anonymized ChatGPT conversation, summarize what the user wanted, then score multiple responses on a 1–7 scale. They are specifically instructed to flag sycophantic behavior, meaning instances where ChatGPT flatters users excessively or acts overly human, and to push evaluations toward a more restrained, professional standard.

Your conversations feed this pipeline by default. For Free, Plus, and Pro plan users, the training opt-in is on by default; turn it off in ChatGPT’s data controls.

OpenAI says prompts pass through a Privacy Filter model before reaching reviewers. Its own documentation reportedly acknowledges that sensitive personal details can still get through, including information from user memory summaries.

Why the No-AI Rule Exists

The ban on AI tools traces directly to a well-documented failure mode in machine learning.

The contractor guidelines are explicit, according to 404 Media’s reporting: no LLMs, no Grammarly, no AI translation tools, and no AI-detection software including GPTZero. Internal guidance reportedly instructs reviewers to distrust detection tools as unreliable and to identify AI-assisted work through pattern recognition instead.

A paper titled “The Curse of Recursion” and a separately published Nature study both document the same failure mode: when AI models train repeatedly on AI-generated content rather than fresh human output, they lose accuracy and diversity over time. The distribution narrows, and the model starts mis-perceiving reality.

If the humans scoring ChatGPT’s responses are quietly using ChatGPT to write those scores, the problem compounds. The human signal OpenAI depends on becomes synthetic noise.

The Labor Reality

Enforcement is active, and the consequences for violations are swift.

One contractor described people using AI in labeling work as being “let go all the time,” calling it “pretty much the one thing that will get you kicked off ASAP.” Internal Slack channels reportedly host ongoing threads where workers share samples and flag AI-assisted writing by pattern: repetitive phrasing, heavy punctuation, and unusually fast completion times.

One anonymized contractor received a termination letter citing concerns about the “authenticity” of their work. They described having “needed a little boost and turned to AI.”

The sabotage problem is harder to quantify. At least one contractor working across multiple AI companies admitted to sometimes deliberately selecting the worst responses when rating outputs, describing the experience as “getting paid to make AI worse.” How much noise that introduces into a dataset with hundreds of raters is unclear, but the incentive problem it represents is real.

“Our experts are hired for their expertise and judgement, which is essential to the ongoing advancement of AI. Our contracts strictly prohibit the use of LLMs to complete projects and we enforce that. We invest heavily in our tools and systems to detect misuse and ensure our experts comply with project rules and contract terms. When we confirm an expert has used AI to complete a task, we immediately remove them from the project.” , Mercor spokesperson

The contradiction at the center of Project Lily is not going away. OpenAI’s training pipeline depends on keeping AI out of one specific workflow: the human judgment that shapes what ChatGPT tells you next. If that tension concerns you, the opt-out is in ChatGPT’s settings under data controls.

Share this

At Gadget Review, our guides, reviews, and news are driven by thorough human expertise and use our Trust Rating system and the True Score. AI assists in refining our editorial process, ensuring that every article is engaging, clear and succinct. See how we write our content here →