example.com/path/to/article
000 points · username · 0 hours ago
example.com610 points · 645 comments · 2 months ago · montroser
benjiro29
felooboolooomba
In 1-2 years, you won't be able to browse the internet without scanning a body part and providing a ID for your ass.
https://www.gadgetreview.com/reddit-user-uncovers-who-is-beh...
tdeck
I guess credit to whoever fought internally to keep old.reddit running all these years. I'm sure it's been a candidate for the axe many times.
galleywest200
Example: https://www.reddit.com/r/whatisit/comments/1v3mzl6/what_is_t...
aucisson_masque
And the discussions quality on Reddit are abysmal, its filled with bots.
account42
This is because appending site: reddit.com to a search query is basically a surefire way to find results written by genuine humans.
This hasn't been true for a while now. A lot of what you find on reddit these days feels incredibly fake. It's like there should be a variant of Goodhart's law: Once a forum becomes known for genuine human writing it stops being a forum for genuine human writing.
khurs
https://mediaandthemachine.substack.com/p/reddits-new-ai-lic...
msny
DivingForGold
It is javascript that allows corporate interests to pop up and nag screen you, and otherwise control your web browsing experience to their favor.
I have an app in my browser that allows me to toggle off javascript and vanquish most pop ups and nag screens.
jdw64
The business model where a company owns a public community and monetizes through ads and search traffic seems to have just been a failed model.
The value of a community comes from the conversations and knowledge that users accumulate for free. But when the content company lowers accessibility under the pretense of protecting that content, new members stop coming in.
Platforms try to protect their assets, but in doing so, they end up damaging the ecosystem that created those assets in the first place.
tassadarforaiur
Klonoar
Behavioral tagging of bot traffic is actually easier when you're not serving plain HTML pages, because the number of signals thrown off by bad bots that don't exactly replicate browsing habits goes through the roof. The larger anti-bot implementations do a lot of this type of thing nowadays.
Bot authors no doubt have their own incentives for using old.reddit (lower bandwidth, conceptually simpler to crawl, etc) but it would not surprise me to learn that Reddit wants exactly one door for bots to have to shuffle through. They also don't work on old.reddit much at all these days, so any effort they have to put in to handling bot traffic is kept to one route.
Pretty dumb era of the internet we've wound up in.
corbet
superkuh
LaurensBER
It's pretty clear that their focus is no longer on providing the best communities (see the many examples of power hungry mods, see below) or the best UX (see the deprecation of the API or the recent changes to old.reddit.com described in the post).
The "product" is now selling data to LLM companies, everything else is clearly a secondary concern. It's a shame because at some point I really enjoyed participating on Reddit and it was a good source of valuable information.
to give some examples of mod abuse:
- /r/energy used to (or still does?) ban anyone in favour of nuclear power
- /r/ubisoft and /r/assasinscreed do not allow any posts criticizing the writing or characters, these are clearly run by the company
- /r/unitedkingdom is famous for shadow banning everyone who's critical of some government program (I got a ban for criticizing motability cars)
- /r/europe banned me 1 minute after submitting a chat control post that got 1.5k upvotes before it was removed
- some vaccination subs ban anyone who posted in anti-vaccination subs (irrespective of what they post)
p2detar
All of this work done by people for free for a commercial platform, so that it can then build a walled-garden around all that content and charge others for API usage.
Hats off.
pwndByDeath
I'm curious what HN minds would consider a useful fix to this cyborg Internet problem. It seems having humans and bots coexisting is problematic.
If you had to start over knowing what could happened and what we should avoid, how would you do it?
Right now the reticulum network stack has my vote, it enables and cripples just the right things to make me think it might help.
But I don't earn my living from the commercial form of the Internet, I find I blame the commercial interest for the bulk of the problems (enabled by human fragility)
But I may just be a crusty old mind.
phreack
plqbfbv
dredmorbius
It's ... deader now.
I'd noticed this change myself a few days ago, suspected something was up.
(Occasionally visit for information that's not readily available elsewhere. Log in only to pull content from my now-private subs, perhaps a few times/year.)
doodlebugging
I did my normal morning routine of checking 8-10 subs that still have interesting content and did not see any nag screen for logins at all. Everything worked on my end just as it always has worked. I read comments and post titles with no issues and no login prompts.
I wonder if the login nags only happen on mobile devices. I surf entirely from a desktop using Firefox or Librewolf and a couple of extensions. I never log in to reddit unless there is some issue that I can help with. I avoid most of the popular subs and focus on things that I have direct experience doing or using. If I do post I leave it up for a day or two and then edit it a bit like Yossarian-style censoring - death to all modifiers!
decimalenough
This is because appending site: reddit.com to a search query is basically a surefire way to find results written by genuine humans.
This is, sadly, not even remotely the case anymore.
y-c-o-m-b
https://lemmy.today/?dataType=Post&listingType=All&sort=TopD...
OR
https://old.lemmy.world/?listingType=All&sort=TopDay
https://piefed.social/ (more or less the same as above)
bo1024
Thanks for making sure I’m secure enough to read your three-paragraph news article…
junon
They're killing themselves the same way Tumblr did. Oh well.
zahlman
This is because appending site: reddit.com to a search query is basically a surefire way to find results written by genuine humans.
This take seems... very out of date.
Anyway, old.reddit.com is still loading fine for me in a Private Browsing tab. Firefox + NoScript on desktop, just like OP.
jwrallie
It started dying for me when they started censoring posts and comments to appeal to advertisers, followed by killing non-official clients, then the light mobile website. Killing old.reddit feels like the last nail in the coffin.
Unfortunately, we have yet to see an alternative that can sustain mass migration.
LoganDark
ggm
Money and legal compliance pressures are far higher in my view.
sco1
I tried looking into how to do things the "correct" way nowadays and they seem to have completely changed their app platform since the last time I looked. They've completely locked out the old method for requesting API access in favor of whatever this Devvit thing is: https://developers.reddit.com/docs/
If you want just plain old data access you have to submit a help ticket, which as far as I can tell just goes completely ignored if it doesn't get rejected outright by some automation.
egorfine
I vaguely remember some problems with new reddit initially. But the "new" frontend looks nice to me in all aspects. What am I missing?
knorker
When they eventually remove the old reddit, they'll lose this 18 year reddit veteran. No way in hell am I using that infinite scroll abomination.
Gareth321
storus
gblargg
programmertote
I was a little annoyed by being asked to log in. But I don't want to browse 'reddit.com' because I don't like the way the new UI looks. It's hard to read and scan the first few pages quickly in the new UI (also looking at you Vanguard; you destroyed your old UI which was compact and informative).
Worse, now that I'm logged in and can use "old.reddit.com", Reddit serves me what it "thinks" I like, so now I'm trapped in the subreddit/topic bubble that Reddit selects for me (btw, it selects things poorly). I want to go to the old front page, so I click on 'Hot' or 'Top'. But Reddit still interleaves the posts with what it thinks I'd like to see.
Madness and sadness ensue.
moritzwarhier
The "error tolerance" introduces algorithmic complexity into what could be a parsing problem.
Root cause: DOM parsing can lead to a different result when you take the first-pass output, serialize it to HTML again, and use it as input for the second parsing pass. And if it sounds stupid: browser do this, AFAIK.
Side note, not security-related: CSS has introduced two-pass rendering since Flexbox, AFAIK. Not the same thing regarding what "rendering" means (there, it's about boxes on screen, not about transforming HTML->HTML, but...).
Lehacy inline attributes that allow JS execution are a favorite example, but disabling them via CSP is not a silver bullet.
Even if you disallow all attributes like "onclick", "onload", etc, modern document parsing is not context-free, and allows for arbitrarily complex exploits, AFAIK (a href, src or style payload can be just as bad as arbitrary JS execution).
The author of the ancient library DOMPurify made a decent amount of money from explaining this thoroughly and providing solutions.
In short, detecting XSS is a variant of the Halting Problem in some especially hard cases, as far as I remember. Because HTML parsing is not context-feee.
ChrisArchitect
Tell HN: Old Reddit Now Requires Login
jokoon
reddit is probably the worst echo chamber of the internet
ChrisArchitect
Reddit will require you to log in to use old.reddit.com
https://smackernews.com/item/48740217 HN
(Source: https://old.reddit.com/r/modnews/comments/1ujtebf/logging_in...)
wktmeow
TZubiri
To be specific, the defensive measures probably rely on javascript, captchas and captchaless bot detection measures (probably javascript), like measuring mouse movement and stuff like that. So sending a client with javascript that loads content if some security checks pass, is harder to abuse than a plain html site that sends all content in the first request.
Besides scraping, this probably has implications for posting as well. Do you like reading comments from humans and not from other bots? Well, in that case a system that uses plain html will be easier to automate with bots than a javascript clusterfuck.
You can't have the cake of web 1.0 and the eat-it of a botless experience too.
modinfo
The article is right about one thing: JavaScript and frontend bloat do not secure public content. They just make scraping more "expensive"
But frontend design does affect how easily a site can detect and restrict automated traffic. Reddit may have a real abuse problem. Calling the change necessary 'to keep Reddit safe' though, is slippery wording. this is about controlling access to content and automated traffic, not protecting anyone from unsafe HTML.
1970-01-01
darepublic
Then this wonderful benefactor starts adding more things.. fencing.. signage, guardrails. Then one day the benefactor says to everyone I've made all this for you. And yet some of you still don't use this place in the intended way. I will now put up a security perimeter. You must show your papers to enter, walk along the narrow lanes we've set up and obey the signage and pay a toll to justify all I've done. Some of you are so ungrateful
I wonder if we can find a new patch of greenery
moralestapia
I did that several years ago, after encountering a very abusive mod.
I was a heavy user, visited the site several times a day for like a decade, and yet I was surprised to realize that I did not miss it one bit.
Just leave Reddit for good. :).
busymom0
This is because appending site: reddit.com to a search query is basically a surefire way to find results written by genuine humans.
I don't think this has been true for a while now. Reddit has been filled with bots and shill accounts for a while.
zmmmmm
shevy-java
The reason was censorship by insane moderators. I am also hardly the only one - there is so much criticism about reddit on reddit that it is quite interesting to see reddit has not yet died already. On the one hand, some info on reddit is useful, but on the other hand, I would not be in the slightest concerned if reddit were to be removed - nobody needs censordit.
Thinking about it, I actually should have also deleted all my contributions; feels a bit unfair that they can still benefit from what I wrote even though they (many moderators) censored me.
TZubiri
To be specific, the defensive measures probably rely on javascript, captchas and captchaless bot detection measures (probably javascript), like measuring mouse movement and stuff like that. So sending a client with javascript that loads content if some security checks pass, is harder to abuse than a plain html site that sends all content in the first request.
BryantD
ChrisArchitect
Mountain_Skies
ChrisArchitect
What I'm getting is a temporary/false "You've been blocked by network security" page but refreshing then brings up the reddit page as per usual.
https://old.reddit.com/r/modnews/comments/1ujtebf/logging_in...
Melatonic
dyauspitr
luxuryballs
dncornholio
Also you can scrape Reddit with JSON now, which is 10x better than any HTML, safe or unsafe, whatever that means.
throw7
Hitton
giancarlostoro
tempfile
nottorp
is basically a surefire way to find results written by genuine humans
Was a surefire way?
janalsncm
I stopped using Reddit when they blocked 3rd party apps and forced people to use their first party ad-ridden and painful UX. Every time I go back I find some other annoying unilateral change like deleting r/all.
In any case, Reddit was so filled with politics and bots and low-effort drive-by comments that I don’t miss it too much.
wwalexander
exabrial
Pxtl
darepublic
imzadi
[deleted]
silon42
lisp2240
jLaForest
GaryBluto
A recent example: if you use a VPN and are signed out, YouTube will occasionally claim you need to sign in to "Protect our community", despite signed out users being unable to affect anything in the first place.
traceroute66
I avoid YouTube like the plague, but sadly there are lots of people out there who don't know anything other than YouTube exists and so they post their videos there (e.g. conference presentation recordings etc.).
latexr
This is because appending `site: reddit.com` to a search query is basically a surefire way to find results written by genuine humans.
People share that tip as if it’s some kind of cheat code for the internet, and it baffles me how they can’t see they’re being manipulated. Of course advertisers took notice and there are Reddit posts which are ads disguised as content. That has been happening since before LLMs. After LLMs, of course plenty of posts there are written by them. Not that the human posts are great, either: for sharing opinions it’s of course subjective, but for factual stuff they are just as wrong as anywhere else on the internet, but with the bad results even more upvoted to the top. Reddit is full of propaganda and confidently ignorant people.
rose-knuckle17
impish9208
dclaw
[deleted]
alex1138
You should become a pariah if you were anywhere close to companies like Reddit. Or at least Steve Huffman. There are some companies that clearly decided at some point "yeah, no. fuck the user"
holoduke
pestat0m
russellbeattie
I find I spend a lot less time on the site as a result, which really is a blessing in retrospect. I'm pretty convinced it's at least half AI bots now. Given how much AI sites use reddit as a source of "truth", it's not a good sign.
Still, I was among the first couple thousand people to create an account, so it still annoys me that the automated process just happily nuked me. Oh well.
dust-jacket
Old reddit is going to be MUCH, MUCH easier to scrape. Scraping has the potential to bring down the site. Gating old reddit to logged in users will massively mitigate that.
I don't like it either. But it's not hard to see a good reason why doing so helps "keep reddit safe"
Even if it also coincidentally makes it easier for them to make more money. Grr.
analog8374
Havoc
inigyou
inigyou
Best to treat Reddit archives as read-only and modern Reddit as a marketing slopfarm. It mostly consists of bots trying to get their brands into AI training data.
m0llusk
While yes, you can not simply scrap new reddit as easily as a pure html scraper is cheaper to run. And while the new reddit its js "slow down" scraping as it need to run over a headless browser. Do you slow scraping down a lot?
No. Because you can simply spin up more instances and route over more proxies. Given the fact that you can scrap millions of pages per day, on a single cheap 1 CPU node...
The limiting factor for scraping a website is not html vs js / client vs headless browser. Its avoiding detection by running from different ips, faking browser information and masking your traffic to not give off the smell of automatization (and avoiding tls fingerprinting).
I wrote this type of stuff before we had AI to help write it, in a few days time. Now with a LLM at your fingertips in a few hours, you will have a complete working version with cloaking build in.
And the larger size difference does not matter, you can block irrelevant files via faked cache responds, to give the server the impression that your a repeat client, and no need to flood your browser instance with images, fonts etc. Some think that this is a defense, aka throttling you with increased bandwidth usage lol
So in my personal opinion, its not about dealing with scrapers, but trying to push people to new reddit, so they can finally phase out old reddit.
PS: So much fun reading new reddit on your phone browser, to have it pop-up a "use our app" banner, that prevents you from interacting with the website anymore. And even blocks firefox its scroll ability (and thus accessing to the browser its nav bar). Tip for people: delete the sites cookies and cache regularly, to prevent this usage tracking, so the banner newer shows.