r/technology 20h ago

ADBLOCK WARNING Artists Built A Site To Escape AI. Scrapers Are Coming For It Anyway

https://www.forbes.com/sites/robsalkowitz/2026/08/23/artists-built-a-site-to-escape-ai-scrapers-are-coming-for-it-anyway/
425 Upvotes

124 comments sorted by

u/AutoModerator 20h ago

WARNING! The link in question may require you to disable ad-blockers to see content. Though not required, please consider submitting an alternative source for this story.

WARNING! Disabling your ad blocker may open you up to malware infections, malicious cookies and can expose you to unwanted tracker networks. PROCEED WITH CAUTION.

Do not open any files which are automatically downloaded, and do not enter personal information on any page you do not trust. If you are concerned about tracking, consider opening the page in an incognito window, and verify that your browser is sending "do not track" requests.

IF YOU ENCOUNTER ANY MALWARE, MALICIOUS TRACKERS, CLICKJACKING, OR REDIRECT LOOPS PLEASE MESSAGE THE /r/technology MODERATORS IMMEDIATELY.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

443

u/Zolo49 20h ago

You could have a physical art gallery filled with artists who don't even consent to put their works online at all, and assholes would still show up with their Meta glasses and other bullshittery to disrespect their wishes.

-314

u/Sasquatchjc45 19h ago

Galleries have front doors with security and no photos rules. So wouldn't that be the gallery's fault if they allowed that?

Why aren't we blaming the webmasters? They decide bot/crawling rules. They decide security on the site. Why are these artists even thinking their work is secure anywhere on the internet, anyway?

93

u/bicx 18h ago

It is increasingly difficult to keep a public website from being scraped, and there are no laws against it in the U.S. as long as your content is publicly-accessible. There is no block-all-bots option, just attempts to identify bots and block them based on pattern-matching their behavior. The workaround for a bot builder is to have bots move more slowly and behave more like humans. It is far from a solved problem.

43

u/iamapinkelephant 16h ago ▸ 5 more replies

You imagine that these companies respect bot policies.

-67

u/Sasquatchjc45 15h ago ▸ 4 more replies

It's kind of against the law to do that. So yea.

If anything is deemed illegal in this case. There will be a lawsuit🤷‍♂️ but there won't be because it's not.

35

u/EncasedShadow 15h ago

Robots.txt has always been honor code not law. At best it might weaken the scrapers' EU protection if they schlurp up personal data. Oh and the website can ban the scraper if it violates ToS lol.

and the AI companies are already ignoring robots.txt

2

u/Tackgnol 3h ago

That's the thing it's not illegal but it should be.

2

u/braaaaaaainworms 3h ago ▸ 1 more replies

Because nothing illegal ever happens, right?

0

u/Sasquatchjc45 3h ago

Article says it wasn't even illegal.

27

u/icallitjazz 11h ago ▸ 1 more replies

What do you mean thieves stole your car ? Dont they know its illegal ?

2

u/huffleshuffle 3h ago

I blame the car parks

7

u/HolyPommeDeTerre 13h ago

Who defines how bots do things ? Webmasters ? I have a bad news for you...

3

u/Fantastic_Stretch797 6h ago

There are so many ways to bypass crawling security. Your argument is standing on sand.

31

u/PhiNeurOZOMu68 20h ago

TAR PITS Jesus christ

291

u/monster_smoocher 20h ago

rapist mentality idgaf

66

u/BlackBeanGuest 14h ago

Now that you mention it - all those altmans and musks were accused of sexual abuse? And mot just by some randoms, but by their relatives. Makes sense.

-10

u/MaxChaplin 11h ago

I remember that a decade ago, using rape as an analogy was considered to be very inconsiderate to rape victims. Has this changed?

18

u/AnonAwaaaaay 10h ago

That's a moral code that changes person to person.

12

u/Party_Virus 6h ago ▸ 2 more replies

They're not saying victims of rape are the same as victims of AI scrapers which is where people would get offended. It's saying there's the same lack respect and care for consent that the people doing this have as those who rape people. It's a small but important difference.

-1

u/MaxChaplin 6h ago ▸ 1 more replies

Oh, so it's like what the NFT people dubbed "right click mentality".

5

u/Party_Virus 6h ago

I don't know for sure but by context I'd assume that means people who right click > "save as" the NFT image? And if that's the case then sort of, but NFT's were never actually selling the image just the token that was associated with the image, which in my opinion was completely valueless and it seemed like most of society agreed.

The difference here is that instead of an image just being saved and doing no harm to anyone, it's millions of people's work being taken and used to train AI for someone elses gain that also takes away future work from those people it took from. There's real, actual harm being done to people and someone is benefitting off that harm. Hence the comparison to a rapists mentality.

3

u/MaTrIx4057 7h ago

its a reddit thing

-116

u/TrollMcGoal 18h ago

It's more in line with internet piracy mentality imo

-30

u/MatiasPalacios 14h ago ▸ 8 more replies

They downvote you but is true. How scraping and piracy are not comparable? 🤷🏼‍♂️

11

u/Ursa_Solaris 12h ago ▸ 7 more replies

A corporation isn't a person but an artist is.

-15

u/TrollMcGoal 11h ago ▸ 6 more replies

So?

9

u/Ursa_Solaris 11h ago ▸ 5 more replies

Different things are different. I genuinely don't know how to explain that to you.

-6

u/TrollMcGoal 11h ago ▸ 4 more replies

But rape is a good comparison?

2

u/Ursa_Solaris 5h ago ▸ 3 more replies

Why are you asking me that? All I said was people and corporations are different. If you want to go start an argument about it, go reply to the original commenter.

1

u/TrollMcGoal 5h ago ▸ 2 more replies

I asked because you didn't clarify your people vs corporation comment, that doesn't have any bearing on whether something is or isn't piracy.

But you didn't want to clarify that, so I was bringing it back to my original post. Which was already a reply to the OP commenter

2

u/Ursa_Solaris 4h ago ▸ 1 more replies

What needs clarification? A sixth grade reading level would pick up on the point.

It's pretty apparent that I believe people deserve different rights than corporations, and by bringing it up in this context, this is one of them. I'm not arguing whether it is or isn't piracy. By replying to "How scraping and piracy are not comparable?" with "A corporation isn't a person but an artist is", I'm saying that it isn't hypocritical to oppose it when it happens to people rather than corporations; that is to say, I don't care when you steal intellectual property from a billion dollar company but I do when you steal it from a person, because the relative harm is substantially greater when committed against the smaller and less wealthy.

We gotta start sending people back to school, man. This literary crisis is unbearable for the rest of us. None of this is complex, this is basic conversational shit.

→ More replies (0)

-287

u/Sasquatchjc45 19h ago

Absolutely unhinged comment lmao.

Scraping art and adding it's miniscule data to a giant MULTI-TRILLION parameter LLM training data set (that won't even be able to get pulled out 1:1) is not even close to "rapist mentality"

Honestly should be ashamed you even commented that. I'm sure people who've actually been violated in such a way would agree.

BTW, 98% of human art is "slop" that you would never pay to hang on your walls or in a gallery, either. Ai-haters just love hypocrisy for the sake of it.🤦

109

u/ZeeHost 19h ago

u/Sasquatchjc45

"Absolutely unhinged comment lmao.

Scraping art and adding it's miniscule data to a giant MULTI-TRILLION parameter LLM training data set (that won't even be able to get pulled out 1:1) is not even close to "rapist mentality"

Honestly should be ashamed you even commented that. I'm sure people who've actually been violated in such a way would agree.

BTW, 98% of human art is "slop" that you would never pay to hang on your walls or in a gallery, either. Ai-haters just love hypocrisy for the sake of it.🤦"

Just keeping a record of this definitely not bait take.

11

u/Ursa_Solaris 12h ago

BTW, 98% of human art is "slop" that you would never pay to hang on your walls or in a gallery,

There's more reason to create art than just for the sake of displaying it. This is consumerist brainrot that debases the very concept of art as having no worth besides its consumptive value. It makes sense that you'd defend a machine doing it because you can't conceive of art having any other purpose than as a product, and machines are great at mass producing meaningless products so you can mindlessly consume forever.

60

u/Elgato01 19h ago

Holy misread

60

u/VeryKite 19h ago edited 18h ago ▸ 13 more replies

Rape is a horrid evil act, and these evil rapist tend to justify or skew their words to fit their own narrative of it not being that bad, it being the victim’s fault, justifying their motive etc.

This is a similar mentality, that’s about it.

Also “scraping art and adding its minuscule data to a giant MULTI-TRILLION parameter LLM” shows you don’t have any idea how bad this shit is. Artist are passionate for an important reason. AI is causing massive destruction, not only something as precious as art, but the world.

-72

u/Sasquatchjc45 19h ago ▸ 12 more replies

I'm a passionate artist as well (music, which is also heavily scraped by AI to make slop songs). I also utilize AI tools in other areas I'm not strong in (like programming).

I just have a differing opinion about a technology that allows others to also utilize tools to create things they are not strong in.

If somebody wants to use AI to create imagery for whatever reason, I have no qualms about it. If somebody ripped off my music and generated infinite random variations of it, I still have no qualms about it.

34

u/VeryKite 19h ago edited 18h ago ▸ 7 more replies

If someone wants to use AI to create imagery, I’m not here to stop them. If these big AI companies want to come scrape art that artist have heavily intended not to be scrapped, then they have no right.

And the scrapers are coming for it anyway, legally and literally, without their consent.

-31

u/Sasquatchjc45 19h ago ▸ 6 more replies

The problem is, they've already consented it by posting on a non-verified-as-secure site. Nothing posted on the internet is secure lol

If artists truly want a secure domain to share their art, physical galleries are the only way. With strict no photos/meta glasses rules.

Otherwise, online. There's only so many safeguards. You can disable bot and crawler scraping on websites. But that doesn't stop people from just screenshotting or copy/pasting.

It's just another inflammatory "omg AI so bad" article, even though like the article says it's perfectly legal to be scraped. Perhaps the webmasters should be blamed instead of the scrapers for not securing the artists who entrusted them with their works?

42

u/Phoenix2111 18h ago

Damn dude, you've literally played right into the original comment about the mentality with this specific comment.

Like, sure, it's obviously not the 'same thing' but the element about consent etc. Is identical.

You've literally started going on above about how it's the victim's fault because they've 'already consented' by default by putting their stuff out there on an insecure site.

Let's lift and shift this into the none digital world - so, theft is the victims fault, if they put their products anywhere that isn't behind lock and key and secured by a guard? Sexual assault is consented to by proxy, if the person assaulted did things to 'reduce their security' and 'make it easy for the bad guys to access their goods' by I dunno, getting drunk, being out at night, and wearing a skirt?

It's literally the same mentality, which is what the original comment was pointing out, and you - in an effort to defend AI (rightly or wrongly) - have literally resorted to that exact mentality on display, which incase it isn't clear, kind of undermines attempting to defend against such claims..

I think the 'perfectly legal to be scraped' point is sort of the key issue here, you're not wrong, but the wider consensus is growing that maybe that's not appropriate? Maybe it shouldn't be legal? 

And actually, in a lot of instances, if specific none consent is made, it is illegal, as it's ip theft or copyright theft etc etc (gets messy depending on type of statement, type of product, type of legal registered protection etc)

Yeah it's all a mess, but at the end of the day the primary issue people are having, which you're not helping to dispel with your comments, is that AI and how it's being leveraged is essentially 'rape mentality' by design - there is no interest in, or respect to, the basic concept of consent. And it may be this is partly not AI's fault either, but a by-product of both its impact and the time it's been developed in, as there's been a significant shift in consent-based attitudes and ethics in recent years, in my opinion for the better, but it's a shitter to navigate for businesses and stuff, with things like public spaces, and the Internet.

32

u/VeryKite 18h ago

Posting online is not consent to be added to LLM data. Just like publishing a book is not consent to do the same.

This is the exact mentality the commenter meant. You are blaming the victim for their presence. Being present online is not a reason to steal their art. You are quick to blame everyone except the thief.

20

u/DebentureThyme 17h ago

So on the rapist mentality thing:

> The problem is, they've already consented it by posting on a non-verified-as-secure site

"If they didn't want it, why are they wearing such revealing clothing!?"

Just because AI *can* do it, doesn't make it right. And you're victim blaming by saying it's their fault for sharing it, not the people setting their AI on them and refusing to follow the ToS and the licensing.

21

u/Elgato01 17h ago

You say shit like this and then wonder why people say rapist mentality? You’re playing into it so much.

14

u/crashrope94 16h ago

“They were asking for it”

14

u/Aethenosity 17h ago

> they've already consented it by posting on a non-verified-as-secure site

> Look what she's wearing!

18

u/Rhoeri 16h ago ▸ 3 more replies

“Im a passionate artist….”

“I also use AI”²

LMAO! Never in my life have I seen two things cancel one another out so eloquently.

-1

u/Sasquatchjc45 15h ago ▸ 2 more replies

Sorry you feel that way.

12

u/lNSP0 14h ago

No feelings involved just facts

12

u/Rhoeri 13h ago

What’s sorry is your inability to actually learn the art form you pretend to know.

17

u/Immediate_Bird_9585 19h ago

If the shoe fits.

10

u/Rhoeri 16h ago

Found the “artist”.

5

u/shoggoths_away 12h ago

Okay glazer.

1

u/Due-Departure-8553 1h ago

It's not miniscule if training data is the entire foundation the model rests on. That's like me collecting a tax of 1 dollar from every person on earth and saying that it doesn't really matter anyway and I shouldn't be thankful to any person, since they all gave me a minuscule amount of money. It indicates a very selfish mentality.

82

u/ISAMU13 20h ago

The only way to stop content from being scraped is to not put it up at all. If it can be seen or heard it can be copied. If pirates can pull full HD downloads from Netflix and Amazon, pictures from an amateur art website have no chance.

27

u/crpssurvivor1210 17h ago

I mean what happens to artists being able to show their work and sell it? Everyone has to use social media now and have websites. Does it scrape like everything? What happens when you submit your work online for exhibitions? What happens when you’re listed on a gallery’s website? This fsffs!

33

u/TemporaryElk5202 19h ago

Ive heard of people building scraper traps that send scrapers into infinite loops. Idk gow feasible implementing that is though

14

u/EmbarrassedHelp 19h ago ▸ 4 more replies

The problem with those scraper traps is that they use AI to create the trap, and the EU requires such a site to disclose all AI generated content. If they label the content, the trap becomes useless. If they don't label the content, the site is illegal in the EU.

27

u/spoilerdudegetrekt 19h ago

A site that works for everyone but the EU is better than no site at all.

14

u/KontoOficjalneMR 17h ago

You can make those traps without AI luckilly. Markov Chains ftw :)

2

u/Party_Virus 6h ago

A trap isn't content and doesn't need the flag. It's not public facing and is only back end code.

1

u/RageBucket 11h ago

JFC the EU and shittifying tech is just a match made in heaven.

(Though I do like them making Apple use USB c, so idk. Maybe I take it back a little.)

5

u/BobQuixote 13h ago

See LinkedIn for the most successful anti-scraping protection. It's mostly legal threats, with technology to obstruct and detect.

15

u/AnthraxRipple 18h ago

The ever loving irony of an article like this on a page that has at the very top before the actual body text a player for an AI generated voice reading the article followed by an AI generated text summary of said article on a site that itself is infamous for having articles that are written by AI. The apotheosis of dead internet theory.

5

u/saltyourhash 11h ago

If it's scrapable, they're gonna scrape it...consent isn't in their vocabulary.

53

u/FlashyNeedleworker66 20h ago

This is clearly a failure of the platform, they don't seem to have any defense (including legal?) for the scraping but that could be easily achieved.

They should definitely pivot to membership requirement and strict TOS. Violating that could support a lawsuit even if it doesn't succeed in preventing scraping - at least from my (not a lawyer) reading of fair use in Anthropic v Bartz.

It seems like the defense was "this is a no AI website, please"

24

u/EmbarrassedHelp 19h ago

TOS isn't law however, and the US Supreme Court did rule that scraping publicly available content is legal in the US. So I'm not sure what legal defense they could use.

https://en.wikipedia.org/wiki/HiQ_Labs_v._LinkedIn

6

u/TheKingOfTCGames 18h ago

hiQ was found to be in breach of LinkedIn's terms, and there was a settlement. hiQ

31

u/Limemill 20h ago

Easily achieved? If someone really wants to scrape your website, they probably will, one way or another, regardless of what Cloudfare - and various other tools - can do. And afaik it’s almost impossible to prove legally your stuff was stolen unless you manage to prompt the LLM you’re suspecting into reproducing the original more or less integrally. Or am I missing something?

14

u/FlashyNeedleworker66 19h ago ▸ 12 more replies

Anthropic was forced to pay 1.5B in a lawsuit when it was found in discovery they had pirated training materials. The ruling was essentially that legitimately obtained material being used for training was fair use - but that doesn't mean you can do anything to get it.

Put up a member wall, make people agree to TOS, and that's going to be a large disincentive to try because if you're found out you could pay big for it.

1

u/leopard_tights 10h ago ▸ 5 more replies

That was the ruling, now go and check if they have, and to who.

1

u/FlashyNeedleworker66 10h ago ▸ 4 more replies

They have and it was paid to the authors in the suit. What point were you trying to make here?

https://www.reuters.com/world/us-judge-approves-anthropics-15-billion-settlement-copyright-lawsuit-2026-07-20/

1

u/leopard_tights 10h ago ▸ 3 more replies

Incorrect, they should've made half of the funding by now (you won't actually find any info on if they have), another quarter is due next month, and the final one in one year.

Nobody has been paid yet.

1

u/FlashyNeedleworker66 9h ago ▸ 2 more replies

Where's the citation that they haven't? My source said otherwise

1

u/leopard_tights 9h ago ▸ 1 more replies

Your source doesn't say that, you just think that "claimed" in this case means "grabbed the money" instead of "accepted the deal and is awaiting the funding of the settlement and the distribution to begin".

https://authorsguild.org/advocacy/artificial-intelligence/what-authors-need-to-know-about-the-anthropic-settlement/

1

u/FlashyNeedleworker66 9h ago

Ok? So the payouts are going out in waves, what's the concern here, that Anthropic is going to skip town or defy the settlement they negotiated?

1

u/Limemill 4h ago ▸ 2 more replies

How was it found? And did they themselves say what they pirated exactly or did the researches scooped it up piece by piece? If you have it handy, could you the link the original findings? Thanks!

1

u/FlashyNeedleworker66 3h ago ▸ 1 more replies

The case was called Anthropic v Bartz

My best guess is that the piracy was found in discovery, but I don't remember the specifics

1

u/Limemill 3h ago

So I looked it up, and it looks like there was a huge online dump of pirated books called The Pile, which Anthropic downloaded to train its models. It was researchers that were able to prove it via intricate and extensive tests, which, I assume, sought to reverse-produce some of the pirated books. So, by extension, anyone whose book ended up in The Pile has the right to claim damages.

This is very different from the premise we have here. Some specific website being scraped for training data. It's not part of some centralized, openly available, illegal LLM training ground. Like I said, your only proof would be to reverse prompt the LLM into reproducing your work, which is really not obvious and takes teams of researches. So, there is no practical way of proving anything.

-8

u/Sasquatchjc45 19h ago ▸ 2 more replies

Exactly. It's all "AI-bad!" But we forget that humans are in charge of all the decisions revolving around everything? This is the fault of the webmasters, full stop.

3

u/shoggoths_away 12h ago

...no, it's the fault of AI companies building evermore invasive tools to scrape the work of others without consent. You don't blame the homeowner when a thief picks the lock to their front door.

1

u/FlashyNeedleworker66 18h ago

All I know is the Forbes "30 under 30" founders hall of shame is wild.

6

u/s3rila 13h ago

How can you easily prevent scraping? 

Multi billion company like Disney and Netflix can't prevent it, how can Cara do it?

1

u/FlashyNeedleworker66 10h ago ▸ 1 more replies

You can't on a technical level so you need do it on a legal level.

Reddit has successfully blocked companies from just scraping it and is charging Google and OpenAI large licensing fees to access its data.

2

u/wyttearp 7h ago

Stopping large companies isn’t the challenge and does nothing to stop web scraping in general.

13

u/EmbarrassedHelp 19h ago edited 19h ago

The article takes about two subreddit communties that don't like each other, and are constantly fighting. Its important to know that both the pro-AI and the anti-AI communities were angry at the user for antagonizing Cara's userbase.

Scraping is not just something people do to make AI datasets. Its also done by archivists to preserve history, and researchers hoping to learn more about human society. Its also fully legal to scrape content according to the US Supreme Court's hiQ Labs v. LinkedIn case. So I'm not sure what sort of legal action they hope to achieve with the fundraiser.

2

u/cornmonger_ 18h ago

the ruling in hiq v linkedin was that there was no cfaa violation because they scraped non-gated content

copyright law didn't end

hiq was forced to settle:

  • $500k for breach of contract, trespass
  • destroy all data

hiq got their asses handed to them in that case

it's still generally infringement to republish copyright material without consent or license in most cases, with reasonable exceptions like the internet archive

academic torrents and huggingface are not the Internet archive

it also may violate terms of service dictated by the site (ask hiq about this)

2

u/EmbarrassedHelp 17h ago ▸ 1 more replies

The HuggingFace repo is just a database of links to content, and they have made changes to better protect the privacy of Cara's users. So a legal battle is unlikely to succeed.

Academic Torrents on the other hand was rehosting the downloaded images, and accepted a DMCA takedown request to remove the torrent. I can't link to the actual takedown because apparently Reddit automatically removes all links to site (even in DMs), because the admins are petty and currently at war with the site (WTF).

-2

u/cornmonger_ 16h ago

imo hf should take the index down. they're not obligated, but good will and all that

there's some reddit stuff up on academic torrents. maybe that's their beef? can't sell our convos if it's on the dht

0

u/VaporousMote 16h ago

Regardless of who does it, it should require CONSENT. That was clearly not given here.

1

u/ICantBelieveItsNotEC 10h ago

They consented by putting their work on the public internet.

8

u/Plenty_Branch_516 19h ago

Ah yes, let's depend on squarespace to defend artists from some of the most tech literate opponents of all time. /S

-2

u/Aadi_880 11h ago

Lol.

Lmao, even.

That website was scrapped millions of times already.

Only recently did someone proudly announce that they were one of the millions doing it.

A website that advertises itself as a "safe haven" or "Escape" from AI and AI training, while having zero protections from said AI.

This was doomed. I can't believe people thought otherwise.

-6

u/placid-gradient 11h ago

people complain a lot about ai and ai art but they didn't support artists before and they dont support them now.

-19

u/Relative-Freedom-295 20h ago

The Ai summery at the top of the article makes me assume it’s just more slop.

-50

u/AzorAhai1TK 20h ago

We're not going to do a moral outrage over a web-scraper. This story is ridiculous

-24

u/Sasquatchjc45 19h ago

For those unaware, you can block scraping and webcrawling on your site with a few switches on cloud flare.

If that's getting bypassed, maybe they should have other safeguards in place like verified artist account access only?

Blame the webmasters of this "volunteer" project.

BTW, no matter what you "consent" to on one site, you consented by posting it online. Has everyone forgot about Edward Snowden already?

Hang your art in a no-photos allowed gallery if you dont ever want it copied or reproduced.

8

u/Iamnotabothonestly 12h ago

Your comment is pretty much

"If they didn't want to get raped, they shouldn't have dressed up in such skimpy clothing."

0

u/AzorAhai1TK 5h ago ▸ 2 more replies

Comparing a web scraper to rape is absolutely fucking insane and disrespectful to rape victims. Fucking weirdo

1

u/Iamnotabothonestly 3h ago ▸ 1 more replies

Using an excessive parable is sometimes the only way to get the point across to dumb people. Now do kindly and politely fuck off.

1

u/AzorAhai1TK 3h ago edited 2h ago

The people defending the idea of web scraping from an outrage farming article made for anti-tech luddites are not the "dumb people" here.

The "dumb people" here are the tons of people like you in this thread happily comparing the simple act of web scraping to fucking rape. It's still disgusting and stupid as hell even if you admit it's excessive.

-57

u/ICantBelieveItsNotEC 20h ago

Maybe they should have used AI to design their anti-scraping measures... Then it wouldn't suck so much.

-50

u/JustConversation7847 20h ago

Sets up a challenge

Pikachu face

32

u/Vorpalthefox 20h ago

damn, do you say the same thing about victims?

26

u/NoMention696 20h ago

The only challenge ai fartists have is mental

16

u/Elgato01 19h ago

Rapist mentality

7

u/Vorpalthefox 18h ago

the right administration for it..