❯ OpenAI’s reported Project Lily uses human chat review, anonymization may miss personal information
Review processOpenAI uses Project Lily to have reviewers read real ChatGPT conversations and assess relevance, tone, sycophancy and anthropomorphic language, according to ITHome’s account of a 404 Media investigation. One reviewer reported earning more than $50 an hour. The investigation drew on operating documents, workplace messages and conversation samples.
Privacy boundariesIn the current process described by the report, reviewers cannot see usernames, but materials may include user memory summaries. OpenAI acknowledged to 404 Media that anonymization may miss some personal data. This does not establish that every conversation is read by a person, and a model’s promises of confidentiality cannot substitute for the product’s actual data-handling rules.
Quality reviewSuch reviews primarily assess response style and quality, flagging only obvious factual errors, and are separate from human safety review. Users and enterprise buyers need to examine training use, human access conditions and retention periods separately. Turning off model improvement should not be interpreted as withdrawing data already incorporated into training.
▪ SIGNALConversation improvement still relies on human labor; anonymization, stopping future training use and deleting historical data are distinct steps that cannot replace one another.