this post was submitted on 21 May 2025
462 points (99.1% liked)

Technology

70199 readers
4098 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 2 years ago
MODERATORS
 

Researchers published a massive database of more than 2 billion Discord messages that they say they scraped using Discord’s public API. The data was pulled from 3,167 servers and covers posts made between 2015 and 2024, the entire time Discord has been active.

Though the researchers claim they’ve anonymized the data, it’s hard to imagine anyone is comfortable with almost a decade of their Discord messages sitting in a public JSON file online. Separately, a different programmer released a Discord tool called "Searchcord" based on a different data set that shows non-anonymized chat histories.

you are viewing a single comment's thread
view the rest of the comments
[–] [email protected] 44 points 12 hours ago (2 children)

I see a lot of drama here in the thread, people decrying data leaks, how Discord is very very bad, and a number of people wanting the "good old days" of forums.

Yes. I like forums too, but, uh...

These researchers scraped publicly posted messages. Keyword here being "public". How would anything similarly public, like a forum, be better?

I actually remember the times when forums were at their peak. I hung out on BZPower for Bionicle things, and the Relic News Forum for Homeworld modding. You know what they had? Google bots that scraped messages, looked for certain words, and populated websites with advertisements based on what it could scrape from forums.

Pretty sure Lemmy doesn't do encryption either, unless there's some very special, private Lemmy server that nobody has access to. So the researchers could've just as well scraped the fediverse.

[–] [email protected] 5 points 7 hours ago

How would anything similarly public, like a forum, be better?

Forums were the primary way that groups would talk with one another pre-global scale social media.

They could contain public subforums, but the majority of all of the forums that I've been a part of were not viewable without an account, which was manually approved or required a small payment (to make bans have a chance to actually stick).

[–] [email protected] 14 points 11 hours ago (1 children)

Yeah this being just as easy on bb forums or literally any webpage with a public comment section was my first thought as well..

Isn't most of the internet scraped anyways, by the internet archive? The concerning part is that this is 100% going to be used to train some coomer brained AI. Scraping, botting, scamming: all those things are going to happen on large public communities.

[–] [email protected] 4 points 9 hours ago* (last edited 9 hours ago)

Yeah, a lot of this push is about ushering in new laws to prevent data scraping.

Propaganda spreads easily through fake accounts—but how do we detect large-scale operations if they’re constantly creating and deleting accounts or trying to blend in with the rest of us? We’d need access to massive data sets to mine for patterns and expose coordinated behavior.

But the powers that benefit from shaping the narrative are the same ones pushing the idea that all scraping is bad. They want people to hate it, so they can justify laws that lock down access. That’s the end game.