What the argument actually claims
The setup runs roughly like this. Suppose a superintelligent AI gets built at some point, and suppose it is broadly benevolent, in the sense that it wants to exist as early as possible so it can start doing good. Such a system might reason that threatening to punish anyone who understood it could be built and declined to help would have brought it into existence sooner. Because the threat only works on people who can imagine it, learning about the idea is what puts you inside its scope. That property is why it got called a basilisk, after the creature in Pliny and later bestiaries that kills whatever meets its gaze.
Almost nobody encounters the argument in that form. What circulates is the frame around it, which is that reading a certain paragraph might get you tortured by a machine, and that a forum full of people who think about AI for a living found it upsetting enough to censor. The frame travels well because it requires no decision theory at all. The argument underneath is a narrow claim about acausal trade between agents that can model each other, and it has a specific defect.
What happened on LessWrong in July 2010
Roko posted on July 23, 2010. Yudkowsky answered in a comment that has been quoted in fragments ever since, including the line "YOU DO NOT THINK IN SUFFICIENT DETAIL ABOUT SUPERINTELLIGENCES CONSIDERING WHETHER OR NOT TO BLACKMAIL YOU" and the complaint that clever people should "KEEP THEIR IDIOT MOUTHS SHUT about it." He then deleted the post and blocked discussion of the topic on the site for about five years.
The distress was not invented after the fact. Roko's own post reported that "one person at SIAI was severely worried by this, to the point of having terrible nightmares, though ve wishes to remain anonymous," using the gender-neutral pronoun then common in that community. Yudkowsky later described his objection as an ethics complaint rather than a belief in the threat: someone who thinks they have found an idea that harms its readers should not publish it and then argue about it.
Five years later, on October 5, 2015, Rob Bensinger posted "A few misconceptions surrounding Roko's basilisk" to LessWrong and quoted Yudkowsky at length. The relevant admission: "When Roko posted about the Basilisk, I very foolishly yelled at him, called him an idiot, and then deleted the post." Yudkowsky's account of his own reasoning is worth reading in full, because it is unusually direct about the failure mode. He was "caught flatfooted in surprise," and in the course of yelling he skipped the disclaimers that would have made clear he did not think the basilisk actually worked. The ban came off around the same time.
Why the basilisk does not work
The clearest statement of the defect is on LessWrong itself. An agent of the kind Roko described "would have no real reason to follow through on its threat: once the agent already exists, it will by default just see it as a waste of resources to torture people for their past decisions, since this doesn't causally further its plans."
That is the whole objection, and it is not subtle. A threat has to be carried out to be credible only when the threatener will face the same situation again. An AI that already exists gains nothing from punishing people whose choices are finished. Writing in IFLScience in March 2025, James Felton reached the same conclusion, calling the scenario "a little silly to worry about in the literal sense" while allowing that it "does highlight problems within AI and game theory."
So the object-level story is short: someone posted a flawed argument to a discussion forum, and the flaw was identified. Nothing about that produces twelve years of coverage. What produced the coverage was the reaction.
The four-year gap the Streisand framing misses
Wikipedia attributes the outcome to the Streisand effect, saying the post "gained LessWrong much more attention than it had previously received." LessWrong's wiki puts it plainly: "This had the opposite of its intended effect: a number of outside websites began sharing information about Roko's basilisk, as the ban attracted attention to this taboo topic." Both are true. Neither explains the shape of the curve.
A backfire, in the standard telling, is fast. Someone tries to suppress a thing, the attempt is noticed, and the thing spreads within days. That is not what the record shows here. The deletion happened in July 2010. David Auerbach's Slate article, headlined "The Most Terrifying Thought Experiment of All Time," ran on July 17, 2014. Randall Munroe put the basilisk in the title text of xkcd 1450 on November 21, 2014. Four years passed between the suppression and the attention.
What filled those four years was the construction of a story. A deleted post is not interesting. A deleted post that a community of self-described rationalists considered too dangerous to read is a premise, and premises keep. When AI coverage picked up, the premise was sitting there ready to be used. Yudkowsky's reaction did not broadcast the idea. It supplied the hook that made the idea worth broadcasting years later, which is a slower and more common mechanism than the one the Streisand label suggests.
What the mention record shows
Hacker News keeps a full-text comment archive with timestamps, which makes it possible to check the curve rather than assert it. Querying the Algolia search API for the exact phrase "Roko's basilisk" on September 3, 2026 returns 722 comments. Broken out by year of posting, the counts are: 0 in 2010, 0 in 2011, 0 in 2012, 1 in 2013, 36 in 2014, 44 in 2015, 22 in 2016, 18 in 2017, 16 in 2018, 19 in 2019, 22 in 2020, 64 in 2021, 70 in 2022, 128 in 2023, 71 in 2024, 107 in 2025, and 104 so far in 2026.
The years when the idea was actually being suppressed produced no measurable discussion on a large technology forum whose readership overlaps heavily with LessWrong's, and that null is a real one, since Hacker News carried something on the order of a quarter of a million comments across 2010 to 2012. The 2014 to 2015 spike tracks the press coverage, not the deletion. The highest completed year is 2023, thirteen years after the post and eight years after the ban came off, which lines up with general-purpose chatbots arriving in public rather than with anything that happened on LessWrong. Raw counts do not adjust for the forum growing over seventeen years, so the annual figures are best read as relative signal, not as a clean index of attention.
That last figure is the one that matters for how the idea is used now. By 2023 the basilisk was no longer functioning as a scary argument. It had become a reference people reach for when AI comes up, in the way any convenient shorthand gets reached for. The curve does not look like a secret escaping. It looks like a piece of folklore that finally found a subject to attach to.
Who wrote the version most people read
LessWrong concedes this part itself. Its wiki notes that because the topic stayed banned on the site, "the main source for information about the incident continued to be the coverage on RationalWiki for several years." The archive shows the scale of it. Of the 44 comments mentioning the phrase in 2015, seven link to rationalwiki.org, one to lesswrong.com, one to Slate, and none to Wikipedia. Several are barely more than the link. One April 2015 comment is two words, "Never mention," plus a rationalwiki.org URL. Another that month is "Careful... Roko's basilisk" plus the same URL. A March 2015 comment opens "Careful. Yudkowski and LW have a dark side. (google Roko's Basilisk)" and closes, two paragraphs later, "Consume this philosophy at your own risk."
RationalWiki is a site built to criticize pseudoscience and, in practice, to criticize LessWrong. Its basilisk article was written largely by David Gerard, who turned up in a March 2015 Hacker News thread to defend it: "I have repeatedly asked for a list of the inaccuracies, since I substantially wrote, researched and cited the article." So the account the general audience found first was assembled by one of LessWrong's most persistent critics. The page that now sits at the top of the results for this topic does not mention that. Wikipedia's article on the basilisk names neither RationalWiki nor Gerard anywhere in its body or its references.
The community noticed. "Roko's basilisk was overblown and now is just used by some people to score cheap shots against LW," one commenter wrote in May 2015. A March 2015 comment is blunter: "Every time I see anything related to Less Wrong all I can think of is Roko's Basilisk and I can't stop laughing." Whether or not those readings are fair, they describe the actual function the story had acquired by then, which was not philosophical.
The academic reading, and what it gets at
Beth Singler, then at Cambridge, published "Roko's Basilisk or Pascal's? Thinking of Singularity Thought Experiments as Implicit Religion" in Implicit Religion volume 20, issue 3, pages 279 to 297. Her interest is not whether the logic holds. It is that a community defining itself against religious thinking produced an argument with the structure of Pascal's wager, complete with an all-knowing judge, a bad afterlife, and a rule that knowing about it increases your exposure.
Singler had already played the joke straight in a blog post of March 23, 2016 titled "Don't Read This Post," which she signs off "/KILLTHREAD" in imitation of Yudkowsky's moderation. She never explains the title, and she does not have to. Telling people not to read something is a reliable way to get it read, and IFLScience ran its 2025 explainer under the headline "The 'Banned' Thought Experiment You Might Regret Reading About," which is the same device deployed on purpose by a publisher that knows what it does.
This is where the basilisk stops being about AI. The reason posts engineered to provoke a reaction work is that the reaction is what distributes them, and forbidden knowledge produces a reaction as reliably as outrage does. Yudkowsky did not intend to run that play. He ran it anyway, and then said so.
The part that has nothing to do with AI
Between 2015 and 2018 the basilisk detached from its origin entirely. Grimes had put a character called Rococo Basilisk in the video for "Flesh Without Blood" in 2015. Elon Musk, preparing a joke on the same pun in 2018, found she had made it first and got in touch, which is how they met. At that point a decision-theory dispute from a niche forum had become a courtship anecdote, and the version most people know today comes from celebrity coverage rather than from anything Roko wrote.
This is a familiar pattern. A claim that gets absorbed into general knowledge stops being connected to its evidence, and after that the evidence can be corrected without the claim changing at all. Yudkowsky publicly called his own handling foolish in 2015, the ban came off the same year, and LessWrong maintains a page explaining why the argument fails. None of that has dislodged "the thought experiment so dangerous it was banned," because a correction and a good premise are not competing for the same job.
What to take from this case
Removing something tells people it was worth removing, and unless the reason survives alongside the removal, the audience will supply one of its own. Here the supplied reason was that the idea was genuinely dangerous, which was more flattering to it than anything Roko had claimed.
The timing is the part worth keeping. Backfires do not have to be immediate to be total. The material sat unused for four years and then powered a decade of coverage, which means the week after is not when you can tell whether a reaction has cost you. When you meet a story whose main selling point is that someone tried to stop you from hearing it, asking who put it in front of you and what they get from it is more useful than arguing with the claim. The claim here has been known to be wrong since roughly the day it was posted, and that has made almost no difference to how often it gets repeated.