RE: https://mastodon.social/@tristanbuckmaster/117233413705701198
When you share thoughts with AI, and don't have Zero Data Retention (ZDR), you are feeding into their training data.
OpenAI admits that they cannot rule out that de-identified data derived from Levent Alpöge and Tristan Buckmaster's work was used to train their models and therefore used in their Navier-Stokes result.
Original thoughts are what frontier AI needs to be fed with to continue to improve.
Just like we have laws in the EU that forbid publicly funded research to end up behind paywalls, should we require publicly funded research to be conducted on inferencing infrastructure that is publicly owned and operated to avoid feeding the beast?
@fj The only issue seems to be the research exclusivity. If the research is truly Open, the beast will be fed.
So I really think it should be the institution or researchers who decide: do you want to keep exclusivity before publishing or are you OK with your inference provider rushing to publish your ongoing work first? (works with any service (like storage) provider, really)
@fj For long-horizon researchers, the leaks could even be accidental by leaking your unfinished insights to other coopeting researchers, if the model training to release cycle completed before you are done.