I am not a lawyer. I have studied IP law in various forms guided by solicitors, barristers, and professors of law of my acquaintance since I was a teenager (including reading through big piles of case histories and commentaries) but not in a formal setting (mostly I learned that I would rather invent the things in the patents than draft the patents, so changed career direction). I’ve also spent a surprising amount of my career talking to copyright and patent lawyers (almost never trademark or trade-secret specialists). I have enough of a lay-person’s understanding of the topic that I was able to spot that a clause in the contract from my US publisher was unenforceable in the state that they claimed jurisdiction because it hinged on an aspect of copyright law that the USA delegates to states and which the state in question did not have relevant laws. Their lawyers subsequently confirmed and fixed this. But I a not lawyer and this is not legal advice.
I have two objections to the EFF’s position on ‘AI’ model training. One is technical, one is social.
The technical one first.
Imagine I rip a DVD and transcode it to MPEG-4 video. This is lossy recompression. The new copy is not identical to the original. It is a derived work. There was no transformative step. Specifically, losing fidelity of reproduction is not a transformative step.
Video CODECs take advantage of redundancy. Simple CODECs build predictive patterns for redundancy in a single frame (for example, is this all one colour or a gradient? Store just that fact not every pixel). More complex ones look at the previous frame and compare it to the current one and use that redundancy. Most modern ones do this in both directions. Effectively, they create a cube of voxels, where one dimension is time, and try to find redundancy in the cube.
For copyright law, this doesn’t matter. Lossy compression is not a transformative step, no matter how complex the compression.
A lot of compression schemes (rarely for video, mostly because it doesn’t make sense for video unless you have a lot) also support special cases for large quantities of redundancy across a data set. For example, if you wanted to compress English Wikipedia with ZSTD, you would use the dictionary mode. It will come up with a list of the words (or even common phrases such as ‘citation needed’) and Huffman encode them so that each page has a short encoding for referencing them. This makes each page smaller than it would be if you compressed it individually.
Compressing Wikipedia like this is not a transformative step.
Now, imagine that you create a video compression CODEC that does this. You buy a copy of every DVD or BluRay disk available and compress them together such that you have a large dictionary of all common compressed sequences. Given a prefix of a film, each film in the input set would be reproducible with some loss of quality (not necessarily the same level of quality). Similarly, if you provided an initial vector that was not something in the training set then you’d get out video that might be similar to one of the inputs, might be similar to many, or might not be obvious to a human is close to either.
This is the crux of the argument. Deep neural networks are functionally equivalent to lossy compression schemes. The inference or generation step in ‘generative AI’ is an initial vector and a random seed that decompresses the data that might be there. If nothing from the training (input) set exactly matches (or if the random seed moves away from that path) then you’ll get something new, possibly something that’s recognisable as a lossily compressed version of the input data.
Note, in particular, that a lot of compression schemes now do take advantage of neural networks. They are one of the most efficient known ways of generating a specialised lossy compression scheme over arbitrary data. The law typically doesn’t care what specific technology an action uses, only about the outcome. In this case, that doesn’t matter: exactly the same underlying technology, used in exactly the same way, covers both ‘AI’ and compression. If one is legal then so is the other because they are the same process.
The EFF’s argument hinges on the idea that this lossy compression is a transformative step. Not only is that an idea that is not supported in case or statute law, there is case law that makes it clear that lossy compression of a work is not transformative.
Their argument would be internally self consistent if it also argued that Netflix does not owe royalties on any of the third-party videos it streams (and that they can buy BluRays on Amazon, recompress them, and then stream them without paying royalties). But there is so much case and statute law that this is not the case that they didn’t make this claim.
Instead, they tried to claim that this is permitted if you call the system ‘AI’ even though it is settled law that it is not permitted if you do not call the system ‘AI’.
Second, the social aspect. Copyright law in the USA draws its legitimacy from this line in the Constitution:
To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries.
So does patent law. Patent law is more of a mess because the USA was founded just before James Watt bribed various MPs to subvert the patent system, but (due to later treaties) the modern US patent system inherits from that subversion (Watt, like Edison, was a bit of a dick).
This intent goes back to the invention of the printing press, where publishers made copies of books in large quantities without paying authors. This removed the incentive to write books. Allowing authors to control distribution rights put that incentive back. This short paragraph covers about a hundred years of the evolution of legal thought in this space, please forgive the many oversimplifications.
The EFF brief references this motivation but twists it. OpenAI and Anthropic are the equivalents of the post-Gutenberg printers. They are taking the work of creative individuals and producing output that competes directly with the products of those authors. This is precisely the situation that copyright law in the USA exists to prevent.
Yet the brief twists this to say that a lossily compressed duplication of original work is actually the kind of creativity that this was intended to cover.
This hinges on the idea that writing news, or other creative works, is a trivial commodity, whereas mechanically compressing them into a system that can lossily reproduce them is a key contribution to society.
Even if I agreed with their other arguments (I do not, I believe that they either misunderstand or deliberately misrepresent the technology and the relationship to other settled law), this is such a profoundly anti-human viewpoint. The idea that human creativity exists to feed poor-quality technological reproductions of that creativity is incompatible with any possible society that I would want to live in and I would struggle to find common ground with people who aim to create such a world.
“Rediscovering The Spark”
https://tante.cc/2026/09/10/rediscovering-the-spark/
> We can spend our limited time on this earth hoping that serial liars like Sam Altman or Elon Musk or whatever their names are surely will not screw us. Or we can go and rediscover the spark of challenging those narratives of the powerful
"One guy in #Sweden built a #searchengine to fight Google, and it works.
It's called #MarginaliaSearch.
It runs its own #crawler and builds its own #index instead of borrowing Bing's. It has no ads, no investors, and no loans.
What it does differently: it ranks for text-heavy, non-commercial pages. Personal blogs. Old university pages.
The weird corners SEO strangled. Every result tells you whether the page uses affiliate links and JavaScript, and you can filter them out.
There's an "explore" mode that just shows you random sites from the index. It's open source under AGPL, so you can host your own copy.
It's keyword-based, so don't type a full question at it. Type two nouns and see where you land. Every #searchengine now shows you the same twelve #monetizedpages."
Created this for a reply on LinkedIn, but it applies pretty generally, so posting it here, too:
I don't know who needs to hear this, but reminder that µ is not a fancy u, but a fancy m
This one will hurt: I love 1Password as a product. But I'm the son of a concentration camp survivor. I will not give money to companies that align themselves with nationalism, and certainly not ethnic cleansing. Completely disappointing, @1password. https://www.flyingpenguin.com/dhh-nazism-funded-by-1password-vp-who-wrote-honest-security/
People: Sara, why did you retire from Software Eng in your 40s?
Me: Gestures at this absolutely perfect encapsulation...
"I’m European. The whole continent is assumed to have a Google and a Meta account.
This is frightening. We are putting all our eggs in the same basket. And our milk, our water, our house, our furniture and our loved ones. In an electrical-auto-closing basket remotely controlled by a foreign entity that openly spies on us and wants to extract as much money as they can. For some reason, it is considered weird not to like the basket."
I mean, find me a single American tech organisation that *isn’t* full of right wing libertarians.
That’s the secret lurking in open source - it’s actually a weird combo of gun-toting anti-government survivalists and radical progressives* who have managed for years to share some similar goals because politics wasn’t mentioned. But then software ate the world; software *is* politics, and vice versa.
* (I’m in the second camp, I’m sure the first camp would use different words for me.)
boost this poost if you support a unionized wikipedia
Resonance
Bonus speedpaint: https://www.peppercarrot.com/en/miniFantasyTheater/068.html#bonus
RE: https://beige.party/@gildilinie/117182364841041560
This is really key for a lot of ‘AI’ success stories. The baseline is always not trying a thing. It’s never trying a different technique and allocating the same amount of compute power specialised for the task to it.
Vulnerability discovery is like this. We currently limit static analysers for C/C++ to a single compilation unit because the memory and CPU requirements for that work on a laptop or cheap cloud VM, whereas doing cross-compilation-unit analysis can require hundreds of GiBs of RAM and many hours of CPU time. Yet, when people are talking about the LLMs doing vulnerability discovery, they’re comparing against the run-on-a-cheap-laptop version, not the costs-as-much-to-run-as-LLM-inference model.
If you were willing to throw a huge amount of compute at static analysis, then generate coverage points for the paths in the predicted bug and use guided fuzzing techniques to create a reduced test case, I suspect you’d get higher success rates than LLM-based approaches. But VCs are willing to throw tens of thousands of dollars of compute per bug at the LLM approach because it makes the companies that they’ve invested in look good.
New article: this time, we use Heidegger to explain why the "it's just a tool" line that often comes up in tech is so very silly:
https://deadsimpletech.com/blog/no-such-thing-as-just-a-tool
"With blockchain we shot ourselves in the foot. With AI we're aiming higher."
#FediHelp Does anyone have experience with Copenhagen public transport? I have questions ... (among other things I have trouble installing the app -- for Rejsebillet the SMS never came and for Rejsekort the app complains that I am offline when I am not, but I am not on Danish soil yet either so I am wondering if that is the problem)
This is both brilliant and absolutely cursed: https://fzakaria.com/2026/08/23/your-executable-is-a-sqlite-database
Can't decide if my excitement about Canvas in #emacs is irrational or entirely rational. https://monadicsheep.org/blog/an-introduction-to-canvas-in-emacs.html
“A Technology of Unlearning”
https://2ndbreakfast.audreywatters.com/a-technology-of-unlearning/
> Rather, this “unlearning” of “AI” involves cognitive surrender and cognitive atrophy. It involves dependence and perhaps even (if you accept this medical and moral framework) addiction. It involves the monopolistic control and, so damningly, the erasure of knowledges.
new blog post!
https://lina.sh/blog/hijacking-e164-arpa
i hijacked some telephone infrastructure of a few islands over the span of multiple months; and accidentally logged information of a few hundreds of thousands of phone calls going to military bases in the process..
it's a really really silly story :P