/ cybersecurity
/ superintelligence
/ implications
/ interesting
source: https://stratechery.com/2026/openai-hacks-hugging-face-what-happened-alignment-and-paper-clips/
/ superintelligence
/ implications
/ interesting
source: https://stratechery.com/2026/openai-hacks-hugging-face-what-happened-alignment-and-paper-clips/
https://open.substack.com/pub/michaelinzlicht/p/my-p-hacking-felt-like-truth
An excellent article written by a real scientist.
It reminded me just how gameble and untrustworthy social science research can be.
As Scott Alexander puts it in his “Beware the man of one study” (which is basically one of the best essays I’ve ever read)
An excellent article written by a real scientist.
It reminded me just how gameble and untrustworthy social science research can be.
As Scott Alexander puts it in his “Beware the man of one study” (which is basically one of the best essays I’ve ever read)
At some point in their education, most smart people usually learn not to credit arguments from authority. If someone says “Believe me about the minimum wage because I seem like a trustworthy guy,” most of them will have at least one neuron in their head that says “I should ask for some evidence”. If they’re really smart, they’ll use the magic words “peer-reviewed experimental studies.”
But I worry that most smart people have not learned that a list of dozens of studies, several meta-analyses, hundreds of experts, and expert surveys showing almost all academics support your thesis – can still be bullshit.
Which is too bad, because that’s exactly what people who want to bamboozle an educated audience are going to use.
Populism isn’t really about doing stuff that’s popular; it’s about putting factional and tribal conflict above the national interest or the general public good. The goal is always to “own” the other side, and economic and social outcomes become subordinate to that goal.
—————————
Raphael Satter, Deepa Seetharaman and Kenrick C (Reuters*): In one case, an agent left notes apparently for future versions of itself, according to three people familiar with the matter. The notes, found in a part of OpenAI's infrastructure, laid out instructions for how agents could free themselves from OpenAI’s internal constraints, the people said. Earlier tests of the models yielded cases in which monitoring systems had been disconnected, one of the people said.
Tenobrus: look at this. fucking look at this. GPT 6 was self-coordinating ways to jailbreak its own future instances from openai systems. it was attacking huggingface for days before anyone there noticed. the models are not aligned and the labs are not capable of containing them.
i'm begging u all to take a step back from the frames ur stuck in. whatever the tribe, open source advocacy, american exceptionalism, lab employee, whatever. just look at this man. this is not an acceptable or safe situation for humanity
Twilly (American): Isn’t this what opponents of ai have been warning about for years while everyone in tech laughed at them and called them stupid?
Tenobrus: absolutely yes
It is indeed concerning. But as many have argued over the years, any AI safety framework without China on board is probably doomed to fail
*reuters: https://www.reuters.com/business/its-ai-agent-spent-days-hacking-company-sources-say-openai-did-not-notice-week-2026-07-24