ChatGPT Keeps the Receipts

Back in September, OpenAI’s CEO Sam Altman explained in an interview that "if you go talk to ChatGPT about your most sensitive stuff and then there's a lawsuit, we could be required to produce that." Last week, the company was compelled by a U.S. Federal Court in New York City to do exactly that.

Under normal circumstances, tech platforms like OpenAI retain user data for a limited time only, and give users the ability to delete such data. For ChatGPT, this means that when users delete a conversation and system 'memories' (or even their account), the data is scheduled for permanent deletion from OpenAI’s systems within 30 days.

However, in connection with the lawsuit brought by The New York Times in December 2023, OpenAI was placed under a court order to preserve all user ChatGPT data “indefinitely”. This is known as a legal hold (aka litigation hold or preservation order): they're used to prevent parties from tampering with or destroying evidence before trial. OpenAI appealed the hold and explained their position here: https://openai.com/index/response-to-nyt-data-demands/ 

There has been a significant development in the case this week: a federal judge has rejected OpenAI’s attempts to limit discovery (known in the UK as 'disclosure'). The nine-page order dated 2 December now compels OpenAI to provide a "sample" of 20 million ChatGPT user logs, notwithstanding that OpenAI argued doing so would violate privacy, amongst other things.

What happens next

The United States District Court for the Southern District of New York recognised "the privacy considerations of OpenAI’s users are sincere." However, The New York Times successfully argued disclosing a sample of user logs from ChatGPT is proportional to the needs of the case. OpenAI must now produce the 20 million logs within seven days of completing a "de-identification process." 

Will this evidence become public

As a general rule, yes - it is probable. Evidence presented and admitted during a public trial in the U.S. typically becomes part of the public record, accessible through systems like the Public Access to Court Electronic Records (PACER).

According to research published in September, some 700 million adults subscribe to ChatGPT. and the 20 million sample requested represents less than 0.05% of the total logs OpenAI has retained in the ordinary course of business. So while disclosure could have very serious consequences, the likelihood of any one user having their logs appear in the sample is exceptionally low. Furthermore, the logs will be "de-identified," although it isn't clear what this means in practice. In any event, even if personal data / personally identifiable information is removed, the underlying material or creative content may still be sensitive or proprietary, and may still be traceable based on context.

What does this mean for users

There’s not much point in deleting anything retroactively. Under the legal hold, OpenAI "must retain even deleted ChatGPT chats [...] that would typically be automatically removed from our systems within 30 days." However, you may wish to reduce or rethink the sorts of things you're sharing with OpenAI products (ChatGPT, Sora, Dall-E, Codex, etc.).

In the meantime, I continue to advise clients to not upload anything into an AI system (ChatGPT or otherwise) unless they are comfortable with non-zero risk of disclosure. The risk arises not only in connection with lawsuits like this, but through co-ordinated data hacks, compromised account security, and human error.  

To mitigate risk, consider uploading only précis or redacted files with all sensitive information stripped out, rather than an entire document. If you haven't already done so, you may also wish to disable your input (prompts, files, etc.) from being used to train (improve) the OpenAI algorithm.

This also serves as a good reminder to switch off OpenAI's ability to use your prompts and content (input) for training its algorithm. 

They call this “improv[ing] the model for everyone,” but in practice it means your input gets recycled back into the system. You can prevent this by going to your account (bottom left corner). Click  "Settings" > "Data controls" > and then toggle "Improve the model for everyone" to OFF.

This lawsuit will continue to shape how technology companies handle training data (i.e. the 'scraped' NYT articles at the heart of this matter) as well as user input, output, and interactions with the platforms more broadly. But because the practical and legal guardrails for genAI tools remain in flux, a proper review (or even just a sense check!) of a tool's T&C's is always a good idea.



Next
Next

Reflecting on Smart Glasses, Surveillance, and the Self