Conversation

RE: https://mastodon.social/@tristanbuckmaster/117233413705701198

When you share thoughts with AI, and don't have Zero Data Retention (ZDR), you are feeding into their training data.

OpenAI admits that they cannot rule out that de-identified data derived from Levent Alpöge and Tristan Buckmaster's work was used to train their models and therefore used in their Navier-Stokes result.

https://xcancel.com/OpenAI/status/2097375276384567642

1
1
0

Original thoughts are what frontier AI needs to be fed with to continue to improve.

Just like we have laws in the EU that forbid publicly funded research to end up behind paywalls, should we require publicly funded research to be conducted on inferencing infrastructure that is publicly owned and operated to avoid feeding the beast?

1
1
0

@fj The only issue seems to be the research exclusivity. If the research is truly Open, the beast will be fed.
So I really think it should be the institution or researchers who decide: do you want to keep exclusivity before publishing or are you OK with your inference provider rushing to publish your ongoing work first? (works with any service (like storage) provider, really)

1
0
0

@fj For long-horizon researchers, the leaks could even be accidental by leaking your unfinished insights to other coopeting researchers, if the model training to release cycle completed before you are done.

1
0
0
@Aissen @fj For those of us in the "Find software vulnerabilities in this code" business, we've known and seen this happen for a very long time (and we keep telling everyone this every chance we get.)

Glad to see other people/groups/companies also realizing that anything you send to an external site, should be considered public at that point in time.
0
0
2