example.com/path/to/article
000 points · username · 0 hours ago
example.com676 points · 511 comments · 3 years ago · costco
Here's Paul Graham:
https://stylometry.net/user?username=pg
Here are some frequent HN commenters: (EDIT: Removed due to privacy concerns)
sillysaurusx
jll29
Also, the cosine of the vectors of word frequencies conflates author-specific vocabulary and topics; in other words, my account is grouped (with >51% similarity, according to the demo) with someone probably because we wrote about similar things. A strong stylometric matcher ought to be robust against topic shifts (our personal writing style is what stays constant when we move from writing about one topic to writing about another topic, just like our personality is what stays constant about our behavior over time - of course styles do change, but the premise then has to be that such changes happen very slowly).
Stylometrics/authorship identification is interesting and has led to some surprising findings, e.g. in forensic linguistics (Malcolm Coulthard wrote several good books about the topic).
This paper lists some other features that could be used and compares a bunch of techniques: https://research.ijcaonline.org/volume86/number12/pxc3893384...
sillysaurusx
This is a fascinating way to find similar HN users who aren’t the same person. It’s a surprisingly great recommendation engine. “If you like pg, you might also like…”
Sure, the privacy concerns are valid, but the cat’s out of the boot. Might as well enjoy the benefits.
montrose is almost definitely pg. Someone who talks about ancient history, Occam’s razor, VCs and startups, uses the phrase “YC cos” (relatively uncommon), etc. https://smackernews.com/item/17112567 HN
Nicely done. One of the best hacks I’ve seen in a long time.
rcarr
Excerpts from wiki:
Before the publication of Industrial Society and Its Future, Kaczynski's brother, David, was encouraged by his wife to follow up on suspicions that Ted was the Unabomber.[91] David was dismissive at first, but he took the likelihood more seriously after reading the manifesto a week after it was published in September 1995. He searched through old family papers and found letters dating to the 1970s that Ted had sent to newspapers to protest the abuses of technology using phrasing similar to that in the manifesto.[92]
In early 1996, an investigator working with Bisceglie contacted former FBI hostage negotiator and criminal profiler Clinton R. Van Zandt. Bisceglie asked him to compare the manifesto to typewritten copies of handwritten letters David had received from his brother. Van Zandt's initial analysis determined that there was better than a 60 percent chance that the same person had written the manifesto, which had been in public circulation for half a year. Van Zandt's second analytical team determined a higher likelihood. He recommended Bisceglie's client contact the FBI immediately.[96]
In February 1996, Bisceglie gave a copy of the 1971 essay written by Ted Kaczynski to Molly Flynn at the FBI.[87] She forwarded the essay to the San Francisco-based task force. FBI profiler James R. Fitzgerald[98][99] recognized similarities in the writings using linguistic analysis and determined that the author of the essays and the manifesto was almost certainly the same person. Combined with facts gleaned from the bombings and Kaczynski's life, the analysis provided the basis for an affidavit signed by Terry Turchie, the head of the entire investigation, in support of the application for a search warrant.[87]
drc500free
I appear to be a well-educated, over-confident know-it-all.
jsnell
And yeah, there's a bunch of high confidence (.6-.8) hits for that account, and from a quick browse of the comments of the recently active ones, they look really likely to be alts. Like, all three that I looked at had comments that made it very clear it was this person writing pseudonymously. (E.g. writing on their signature issue, and saying they couldn't go into more detail due to fear of self-doxxing; or somebody literally saying that the alt's claims reminded them of the public writings of the notorious guy years ago).
Obviously I'm not naming the account, but this functionality turned out way creepier than I thought the moment I tried it on the account of somebody who has a reason to disassociate from an existing public persona, but still wants to participate here.
gus_massa
I don't think the list of pg alternate account is accurate. I checked a few. They have many oneliners that is typical of pg, but the topics and style don't look similar.
I searched a few more and got better results. :)
I searched myself (that I know that I have no alternate accounts). I recognize a few users that are interested in similar topics, and I discuss/upvote them many times. But I didn't recognize most of the user of the list.
Fnoord
I don't do throwaway. I either post or STFU. I also STFU on darknet. Its why I found it fun to read/lurk on things like I2P back when it was new. And I know that on a pseudonymous account it is only a matter of time until it can be linked to another pseudonymous account. It would not surprise me if stylometry was used on Dread Pirate Roberts or the people behind The Pirate Bay or the people behind Wikileaks (Assange's sockpuppet accounts). Such can also have been used to verify afterwards instead of beforehand. Though with TPB since it was on clearweb an advanced adversary could have used correlation/timing attack to figure who wrote what.
I'm having fun times recognizing other Dutch people though their usage of English language. For example, a distinctive word I see Dutch people use a lot is 'oke' instead of 'OK' or 'okay'. Its a red flag the person is native Dutch. I wonder if there are stylometry tools available for figuring if someone used physical vs touchscreen keyboard (I used Glider to write this post, spellchecker unavailable).
And yes, organizations like secret service and police should use such tools as well. It is a known tool, why not use it for good? As with any tool, it can be used for good and evil. On HN this could be useful for the mod team (AFAIK nowadays only dang) to find banned people's sockpuppets. Cross-community could also be a fun project: find a HN user's Twitter or Reddit account. And I hope this method is also used to find Russian trolls on social media.
dlkf
I wonder if this explains our similarity. And if so, could we tweak the algo by e.g. Removing text that is prepended with ”>”
bscphil
Also, since this uses only word frequency, there are probably relatively easy improvements to make that would make it even more powerful, like looking at particular runs of words that are unique. Some expressions or figurative language only show up in combinations of words, and tend to be highly style specific.
setr
saurik
dsr_
I'm 0.566 correlated with logfromblammo -- and while we are definitely not the same person, I could easily imagine writing a sentence such as:
"For some bizarre reason, management has not yet assigned a task to their programmer underlings to automated themselves out of existence. I can't imagine why."
which is theirs, not mine, from about a year ago. I like that.
On the other hand, I'm nearly as correlated with peterwwillis: 0.5485 -- who has no comments and no submissions.
DenisM
As you're typing out a comment the software gives you a list of accounts you're becoming similar to. That way you can adjust your writing as you type.
davebillyhock
serhack_
super256
They were always communicating in some kind of meme-russian, and their texts were funny to read. [1]
I believe their writing mostly defeated this kind of analysis, at the cost of looking like idiots (which was probably the reason no one sent them crypto-dollars to buy that stuff exclusively).
Here's an excerpt:
"Attention government sponsors of cyber warfare and those who profit from it !!!!
How much you pay for enemies cyber weapons? Not malware you find in networks. Both sides, RAT + LP, full state sponsor tool set? We find cyber weapons made by creators of stuxnet, duqu, flame. Kaspersky calls Equation Group. We follow Equation Group traffic. We find Equation Group source range. We hack Equation Group. We find many many Equation Group cyber weapons. You see pictures. We give you some Equation Group files free, you see. This is good proof no? You enjoy!!! You break many things. You find many intrusions. You write many words. But not all, we are auction the best files."
[1] https://archive.ph/20160815133924/http://pastebin.com/NDTU5k...
spdustin
(FWIW, it didn’t find my throwaways; my own model didn’t, either, because I knew that word choice wasn’t enough to avoid being outed by stylometry)
Edit: by bigrams and trigrams, I mean reducing word to their parts of speech labels and using THOSE as word tokens. You’ll find that native English speakers have higher weights on some phrase construction patterns than, say, folks from Romania. TF-IDF is useful for these POS-grams (just made that word up) as well.
zxcvbn4038
“Find hot single women who write just like you”
interroboink
It's so easy for something like this to be turned into a tool for a witch hunt, targeting innocents.
psychphysic
ufmace
stavros
https://stylometry.net/user?username=stavros
The next person is 30% less certain, that's huge! This would basically identify any alt I might have with near certainty.
4qz
schappim
macintux
chriskanan
Yeahsureok
Also think it's probably poor form to list users as examples without their permission.
Arathorn
SevenNation
... This site works primarily by analyzing for each user the frequencies of the most common words and phrases in the English language. Accordingly, the easiest way to avoid being identified is to simply use different words than you ordinarily would when writing. More sophisticated models than the one I made can use punctuation, comma usage, and capitalization to identify you so try alternating those as well. Services like Quillbot can help with you this but depending on your circmstances you may not want to send your writings to a third party service.
HN offers many other threads which could be tied together, including:
- time of posting
- ratio of replies to top-level comments
- comments being mainly upvoted or downvoted
- sentiment (mostly angry, dismissive, questioning, etc.)
- most common topics (keyword analysis of post being replied to)
- ratio of new posting to post replies
- first-to-comment on a post
- lone comment on a post
- etc...
It seems very likely that sooner or later every pseudonym for posting content will get discovered and linked. The lesson here is don't post anything that would cause you undue shame or harm if linked directly to your legal name.
bhaney
operator-name
mygentys
The second that I found out that requesting deletion of an account and its posts needed a MANUAL request to a single user (dang) I noped out so fast
But happy that the rest of you are still happy to contribute :)
ggerganov
Edit: From the "How to avoid .." page, there is the following sentence:
Also, most authorship identification algorithms have poor accuracy when working with small amounts of words. This means the optimal strategy would be discarding an account either after every comment or after a small number of comments. Unfortunately, this is against HN rules and may result in a ban.
Can you clarify what this means and why it would result in a ban?
Bhurn00985
Unless the author would run this against all HN user accounts, no need to flag the ones "of interest".
StrangeDoctor
Very nice clean site, great work.
wizofaus
bumble_bee900
I guess it's difficult to evade it as the word frequency certainly catches all about the countries I frequently refer, programming languages, interests etc.
culi
.. [0] https://hackaday.com/2022/10/20/render-yourself-invisible-to...
nostylometry
lettergram
https://twitter.com/austingwalters/status/104189476543920128...
Made both Metacortex.me and insideropinion.com
The idea being you don’t actually need an active directory. It would drop in, figure out all the users (provided one account was on the AD) and would monitor everyone’s skill sets, morale, schedule, etc. Worked super well for what it was / is.
woodruffw
Out of curiosity: do you filter sentences than begin with ‘>’, indicating a block quote from another user? That might improve the accuracy a little here, if you don’t already.
pkos98
nwiswell
DrStrangeLoop
I wouldn't rely on these results
antirez
[deleted]
jefftk
jl6
chronogram
nvr219
nr2x
;-)
CrypticShift
How long before large commercial indexers start offering an efficient (AI based ?) stylometry to agencies and states ?
wait... do you think the NSA is already doing this?
garbagetime
dibt
Does this ignore stop words? Or do all words have the same weighting? I wonder if only focusing on stop words would give a more accurate measure. Maybe we are more comfortable with certain stop words more than others?
https://en.wikipedia.org/wiki/Stop_words
"Stop words are the words in a stop list (or stoplist or negative dictionary) which are filtered out (i.e. stopped) before or after processing of natural language data (text) because they are insignificant."
MikePlacid
2. There is a Russian mnemonic verse, which can’t be properly translated to English, at least it’s beyond my humble capabilities. It goes:
“Это я знаю и помню прекрасно:
Пи многие знаки мне лишни, напрасны”
The number of letters in the words give you the pi number: 3,1415… The meaning is: “I know and remember perfectly: too many signs (positions) of pi are useless and impractical”. Sometimes it’s nice to remember both things.
Trouble_007
Edit to add;
Would be nice to have the https://news.ycombinator.com/user?id=username links included.
jimhi
I am afraid to combine all these methods
elteto
WaitWaitWha
- Why is the author costco[0] not in this lookup?
weinzierl
dibt
I ran it on Brian Armstrong's temp account from here, and it said it didn't write 10,000 characters:
https://smackernews.com/item/3754664 HN
EDIT: Or maybe it's something else because Brian only wrote less than 6k characters. But then why can my account be looked up?
Also, I would guess quoted replies are included, which muddies the analysis. Seems to be a very naive implementation. Much more can be done, but this was probably just a quick project.
[deleted]
lifeisstillgood
A lot of discussion on the thread are over "how can we prevent this". I would like to know why should we not embrace this and similar technologies?
The benefits in my view are large - online behaviour tracks back to real life - and epidemiology speaking the value of millions of test subjects across every question are invaluable - from traditional medicine to "mass psychology recommendations"
I can guess some downsides (hiding from abusive exes) but am interested in studies, surveys, reports etc - any HN thoughts welcome
jaredsohn
Wistar
saurik
nickstinemates
https://stylometry.net/user?username=nickstinemates
Number 2 for me is someone I worked closely with for a few years, and then putting his name into this results in all of the people we worked with for a few years. So it seems content>style, or, we are all more alike than we thought.
lostmyacctoops
Also, talk about a chilling effect. I was already vaguely aware of this, and now I'm overthinking every word I'm thinking/typing.
agumonkey
silasdavis
godisdad
Ros2
I know their blog, which is their HN username, and this tool found their other account.
Perhaps ironically, this person stood out a lot because of this and I didn't forget them.
p4bl0
samwillis
My guess is that as they commonly mention the project and I have on a number of occasions, that has formed the link. Plus maybe usage of common British terms, but that seems far less significant.
It's super interesting!
It would be good if there were more controls to filter the type of words and language that are used for the matching algorithm. So you could say exclude words not in the dictionary. I wander how that would effect my link with this other person.
throwboi123
Long live throwaways.
hobobaggins
This means the optimal strategy would be discarding an account either after every comment or after a small number of comments. Unfortunately, this is against HN rules and may result in a ban.
Is this? I thought that it was ok to have throwaway accounts, as long as they're not specifically to avoid a ban or something like that.
ThrowSet
A question for the author (costco): You created that account in 2019 but you didn't post or submit a single thing until 4 hours ago. Why did you create an account almost 3 years ago for no purpose?
chucksmash
Does a low correlation with other users imply higher susceptibility to de-anonymization if I were using alts regularly?
kevmo314
Maybe we could be friends :)
[deleted]
anpat
seydor
throwaway5434q
gavinray
Needs sentiment analysis IMO, otherwise you'll get "Here's a bunch of people who are JUST LIKE YOU", except they use a similar grammar style but hold opposite opinions on the same nouns.
soneca
When I searched the other one, “soneca” was the first guess, with 0.4.
But when I searched “soneca”, the other one was not in the top 20.
laurex
iHateStylo
notacoward
joshstrange
SnowHill9902
Would translating to other language and back defend against this algorithm?
andreareina
Lichtso
rand_user_100
On the other hand, can we agree that this product is unethical?
In many cases, when a person uses an alt, it is a direct and strong signal that they do not wish their other posts to be associated.
So this product is circumventing the explicit will of the person, and making it available to anyone with zero effort i.e. there is no barrier to getting this info.
I met someone about 10 years ago who said they built this at a university. And their argument also was "actually this enhances privacy because it lets you know something something something". And yet their research grants were coming from one source only.
It can be used for good, but most often it won't.
yyt554
All those scared folks who naively think that it's not too late yet. Busted.
abhaynayar
JKCalhoun
Running then the most similar person to my account did not put me in their top 20.
hamburglar
oblib
rcarr
notduncansmith
mancerayder
rglover
[deleted]
harryvederci
C'mon guys, work harder. That's not even close! :-D
Btw, I myself am only at 0.9999999999999999 so I guess I need to work harder at being myself.
srean
Good job.
SkyMarshal
> Most likely candidates:
skymarshal: 0.9999999999999997
The other few usernames I tested (pg, dang, some random ones from this thread) all matched themselves at 1.0.
timeon
Retr0id
account-5
kfichter
CobaltFire
I’m a native English speaker as well, so I’m unsure how to feel about that.
[deleted]
musicale
I made this site mostly to show how easy this is and how it can erode online privacy
looks like it can indeed
Here are some frequent HN commenters: (EDIT: Removed due to privacy concerns)
How surprising that someone might object to being included in a demonstration of the erosion of privacy!
Is the site opt-in or opt-out?
tomrod
stephc_int13
I ever had only one account here and the closest match is at 0.47.
00F_
anyway, this app was able to identify a lot of my accounts. but a lot of the matches werent me. bold matches were almost all me. but i know there are many more matches than those that were listed. it mainly showed my most recent accounts.
i think most people would get a sick feeling in their stomach if they tried this app. i dont think people are prepared for a world where you can type someones name into an app like this and produce everything ever recorded online that was created by that person. not only this but everything highlighted and summarized to answer any question about that person. this is what advanced ai will bring us. an information implosion where the planet-sized ocean of data that is just floating all around us suddenly and violently coalesces into the objects of our new societal calculus. violent is a good word. and this is just the change that one can see coming with ai.
the_cat_kittles
dvh
robertlagrant
paulpauper
ed25519FUUU
sdsd
thot_experiment
neodypsis
aryc19
[deleted]
scarface74
balls187
And that accounts last several comments were flagged as dead.
I'm a native speaker, but my english succcccks.
scotty79
Which user has lowest best match?
Mine is 0.58 so I'm really not that unique.
a-dub
also maybe a tf-idf vector of top n words per user.
also could maybe do a same phrase analysis across the corpus to find some hand picked features.
timestamps could be interesting.
or, of course, let the machine do it with comment2vec.
mysterydip
iambateman
If an account returns a high score for many accounts, does that also mean they’re relatively less original in style?
msla
oliwary
peacelilly
Semaphor
jonnycomputer
el_dev_hell
Is there a common open source library (Python, JS, whatever) that implements something like this?
silasdavis
imagine what a company with millions of dollars and a couple dozen PhD linguists could do.
Could they do much better?
medellin
serf
vxNsr
I'd love to have the experience and or apparent wealth my "alts" have
SanjayMehta
One funny thing though, while your example says 1.0, for my own account it says 0.99lotsof9s4
dsr_
Perhaps 6 or 7 digits is enough?
2OEH8eoCRo0
sedatk
ChrisMarshallNY
elorant
[deleted]
pugworthy
uberduper
uberduper: 0.9999999999999991
jonnycomputer
[deleted]
[deleted]
kiernanmcgowan
McDyver
[deleted]
throwawayhghcj
This is breaking anonymity that people incorrectly thought would not be revealed.
For some it might be awkward, others it might be quite problematic.
F_r_k
ThrowawayTestr
[deleted]
[deleted]
[deleted]
afarviral
theGnuMe
atum47
hk1337
Ikatza
b800h
And I'm no lawyer, but it seems like there's also an outside chance of a breach of section 171 here as well, which is a criminal offence committed by a person who reidentifies de-identified data.
Plus - the laws have extraterritoriality. Vanishingly unlikely that you'd actually be pursued for it, but it's worth bearing in mind when you munge people's personal data.
canadiantim
kuramitropolis
AtlasBarfed
karol
AviationAtom
zem
thr0v_awway
Holy shit, it works really, really good. It found all of my older accounts.
moneywoes
spaniard89277
cbracketdash
[deleted]
user-
julienreszka
seydor
t0bia_s
ruined
ALittleLight
rmelhem
andsoitis
[deleted]
joxel
Woah
franze
my current and my old account
[deleted]
ecec
xwolfi
[deleted]
WalterBright
This is to protect high profile users who are secretly enjoying programming in D rather than the language they are supposed to use.
And, of course, to protect users who feel they might be discriminated against if their background was known.
RepAgent
j_s,password4321,carolinew,colinwright,kuharich etc.
https://stylometry.net/user?username=j_s https://stylometry.net/user?username=carolinew https://stylometry.net/user?username=colinwright https://stylometry.net/user?username=password4321
Lowest match for j_s is 0.80 and all but one is black.
jallasprit
pg: 1.0
montrose: 0.604073065373204
mattmaroon: 0.5900372458160795
natsu: 0.5519832271289953
rauljara: 0.5418566694533273
waterlesscloud: 0.5378996309342633
damoncali: 0.5292014150349463
gruseom: 0.5290151637991445
kemiller2002: 0.5254174524920762
jfengel: 0.5231938496089998
jamesaguilar: 0.5229081613163672
houseabsolute: 0.5219738531025365
danssig: 0.5195368367601849
austenallred: 0.519343009683366
loewenskind: 0.5177030083877397
baguasquirrel: 0.5153841099708854
asdfasgasdgasdg: 0.5146704002447524
aptwebapps: 0.5144149629369845
allenbrunson: 0.512802806408646
danielweber: 0.5123620795710832honkler
You fail, I win.
The most interesting thing is that my writing style changed pretty drastically since a decade ago. Searching for my oldest account matches my earliest usernames, whereas searching this account matched the rest.
The details of the algorithm are fascinating: https://stylometry.net/about Mostly because of how simple it is. I assumed it would measure word embeddings against a trained ML model, but nothing so fancy.