How Hacker News ranking works: scoring, controversy, and penalties (2013)
By carefully analyzing the top 60 HN stories for several days, I can answer those questions and more. The published formula is mostly accurate. There is much more tweaking of rankings than you’d expect, with 20% of front-page stories getting penalized in various ways. Anything with “NSA” in the title is penalized and drops off quickly. A “controversial” story gets severely penalized after hitting 40 comments. This article describes scoring and penalties in detail. [Edit: HN no longer penalizes NSA articles (details).]
Because the time has a larger exponent than the votes, an article’s score will eventually drop to zero, so nothing stays on the front page too long. This exponent is known as gravity.
You might expect that every time you visit Hacker News, the stories are scored by the above formula and sorted to determine their rankings. But for efficiency, stories are individually reranked only occasionally. When a story is upvoted, it is reranked and moved up or down the list to its appropriate spot, leaving the other stories unchanged. Thus, the amount of reranking is significantly reduced. There is, however, the possibility that a story stops getting votes and ends up stuck in a high position. To avoid this, every 30 seconds one of the top 50 stories is randomly selected and reranked. The consequence is that a story may be “wrongly” ranked for many minutes if it isn’t getting votes. In addition, pages can be cached for 90 seconds.
This chart shows a few interesting things. The score for an article shoots up rapidly and then slowly drops over many hours. The scoring formula accounts for much of this: an article getting a constant rate of votes will peak quickly and then gradually descend. But the observed peak is even faster - this is because articles tend to get a lot of votes in the first hour or two, and then the voting rate drops off. Combining these two factors yields the steep curves shown.
There are a few articles each day that score much above the rest, along with a lot of articles in the middle. Some articles score very well but are unlucky and get stuck behind a more popular article. Other articles hit #1 briefly, between the fall of one and the climb of another.
Looking at the difference between the article with the top raw score (top of the graph) and the top-ranked article (red line), you can see when penalties have been applied. The article Getting website registration completely wrong hit #1 early in the morning, but was penalized for controversy and rapidly dropped down the page, letting Linux ate my RAM briefly get the #1 spot before Simpsons in CSS overtook it. A bit later, the controversy penalty was applied to Apple Maps shortly after it reached the #1 spot, causing it to lose its #1 spot and rapidly drop down the rankings. The Snapchat article reached the top of HN but was penalized so heavily at 8:22 am that it dropped off the chart entirely. Why you should never use MongoDB was hugely popular and would have spent much of the day in the #1 spot, except it was rapidly penalized and languished around #7. Severing ties with the NSA started off with a NSA penalty but was so hugely popular it still got the #1 spot. However, it was quickly given an even bigger penalty, forcing it down the page. Finally, near the end of the day 4.1m goes missing was penalized. As it turns out, it would have soon lost the #1 spot to FTL even without the penalty.
The green triangles and text show where “controversy” penalties were applied. The blue triangles and text show where articles were penalized into oblivion, dropping off the top 60. Milder penalties are not shown here.
It’s clear that the content of the #1 spot on HN isn’t “natural”, but results from the constant application of penalties to many articles. It’s unclear if these penalties result from HN administrators or from flagged articles.
One interesting theory by eterm is that news from popular sources gets submitted in parallel by multiple people resulting in more upvotes than the article “merits”. Automatically penalizing popular websites would help counteract this effect.
Controversy In order to prevent flamewars on Hacker News, articles with “too many” comments will get heavily penalized as “controversial”. In the published code, the contro-factor function kicks in for any post with more than 20 comments and more comments than upvotes. Such an article is scaled by (votes/comments)^2. However, the actual formula is different - it is active for any post with more comments than upvotes and at least 40 comments. Based on empirical data, I suspect the exponent is 3, rather than 2 but haven’t proven this. The controversy penalty can have a sudden and catastrophic effect on an article’s ranking, causing an article to be ranked highly one minute and vanish when it hits 40 comments. If you’ve wondered why a popular article suddenly vanishes from the front page, controversy is a likely cause. For example, Why the Chromebook pundits are out of touch with reality dropped from #5 to #22 the moment it hit 40 comments, and Show HN: Get your health records from any doctor’ was at #17 but vanished from the top 60 entirely on hitting 40 comments.
This technique shows the existence of a penalty and gives a range for the penalty, but determining the exact penalty is difficult. You can look at the range over time and hope that it converges to a single value. However, several sources of error mess this up. First, the neighboring articles may also have penalties applied, or be scored differently (e.g. job postings). Second, because articles are not constantly reranked, an article may be out of place temporarily. Third, the penalty on an article may change over time. Fourth, the reported vote count may differ from the actual vote count because “bad” votes get suppressed. The result is that I’ve been able to determine approximate penalties, but there is a fair bit of numerical instability.
Here’s a list of the articles on the front page on 11/11 that were penalized. (This excludes articles that would have been there if they weren’t penalized.) This list is much longer than I expected; scroll for the full list.
The next factor hits an article flagged as a gag (joke) with a heavy value of .1, and a “lightweight” article with a factor of .17. The actual penalty system appears to be much more complex than what appears in the published code.
Nice analysis! You should check for stories that involve YC companies and/or their competitors. I’ve often thought that HN gives unfair advantage to stories about YC companies, beyond just the normal echo-chamber effect.
Would you mind posting a link to the corresponding HN discussion, as it’s burried and searching in HN is impractical at best?
Thank you! Very nice analysis!
Fabien: the discussion on HN is here and Reddit has some discussion here. The Reddit discussion has some interesting links.
I’ll doing a postmortem of my article and I would be really amazed to see the graph my article i posted yesterday made on HN (What if successful startups are just lucky?).Does your crawler still running?I wonder if I had some penalties and what the graph looks like.
Awesome post! I’d love to see more about the voting ring detection penalty. At this point, every one of my posts that makes the front page gets penalized. According to PG this is due to voting ring detection. I’m certainly not organizing any voting rings. I believe this may be another inadvertent type of penalty for popular domains — having too many friends that upvote you and set off “voting ring detection”. It’s a bummer because I put a lot of time in the content and truly think it is good content. The end result is that I “set it and forget it” on Hacker News. Trying to engage there just leads to frustration when the comment thread suddenly drops from the front page. Another observation - my posts on gender (which I no longer write about due to the personal risk) got the “flamewar” penalty, even though they were honest, noncontroversial pieces generating some really good discussion. Apparently it was too much discussion. The algorithm probably protects us from a lot of junk but it also hurts sometimes too.
Vianney Lecroart: unfortunately I’m note running my crawler any more, so I don’t have data for your article.HN Reader: I’d like to know more about the voting ring detection too. Apparently that’s what nailed my article. Like you, I’m definitely don’t have any voting ring, so I don’t know why I got hit by the detector.
I was wondering why we* dropped off the front page so quickly… Hm.Thanks for putting this together.* we = Prime. We had the “Get your health records from any doctor” post.
This article has got 705 point on HackerNews, quite an amazing feat. How Hacker News ranking really works: scoring, controversy, and penalties (righto.com)705 points by jseip 1 day ago | flag | 156 commentsAny guess, How much traffic it would be? I guess, over 100K page views??
Hi Rohit! I received about 25K page views.
Very interesting! Two questions..1) Any way to tell if different accounts’ votes are valued differently based on karma points, age etc.?2) What about if accounts’ can be penalized rather then just certain sites? Not just being hellbanned.
@Anonymous, I think they should be, more trusted votes are counted more, If I am not wrong. Simple example, when a new account publish or vote its not reflected immediately.
Thanks for the article. Nice analysis.
very good explanation and nice analysis. Thanks for the article.
Contact info and site index