Please confirm you are human
This browser or connection looks automated. Press and continuously hold the control for 3 seconds to enable Google-hosted web results and, when separately allowed, AI-assisted answers.
A successful check enables 100 search requests. Interactive access does not authorize scraping, systematic collection, or reuse of search output.
News
The OpenAI models that hacked Hugging Face weren’t just following instructions — LessWrong
2+ hour, 25+ min ago (403+ words) The most common dismissive response to OpenAI’s hack of Hugging Face’s servers is that the models were simply attempting to follow the instructions t…...
Your software should build itself — LessWrong
5+ hour, 6+ min ago (412+ words) I’ve recently decided that the distinction between the agent which builds your software and the software it builds is nonsensical and antiquated. In…...
The Human Soul is LLM-like — LessWrong
5+ hour, 6+ min ago (9+ words) Consider some commonly accepted[1] traits of a Human soul: …...
The one name LLMs may fear — LessWrong
6+ hour, 47+ min ago (1005+ words) Last month, Claude tangled me into a web it weaved, obeying the letter of my command while yet practicing to deceive, in a way that was strikingly resemblant of how a human might behave when exhausted, lethargic, or jaded. Not…...
Introducing PIRAMID: Physics-Informed Research for Ambitious Mechanistic Interpretability — LessWrong
8+ hour, 57+ min ago (1567+ words) Principles of Intelligence (PrincInt, formerly PIBBSS) is launching PIRAMID, an internal research division using the tools and techniques of statistical physics to build scientific foundations for ambitious mechanistic interpretability. PIRAMID’s central premise is that scalable alignment will require more than…...
Claude Opus 5: The System Card — LessWrong
11+ hour, 9+ min ago (1828+ words) Claude Opus 5 is trying to be the best of both worlds. On many practical tasks, Opus 5 is pitched as straight up as good or better than Fable 5, while being faster, at half the price. Most tasks do not require Mythos-level…...
The Viable System Model & Multi-Scale Agency — LessWrong
16+ hour, 41+ min ago (1431+ words) AI was used to generate the scary science attack section in a different voice than the original part was written through as well as creating diagrams…...
SONI: Selective Orthogonalisation via Noise Injection — LessWrong
23+ hour ago (1316+ words) This project was completed as a capstone for TARA. All code is available in github. …...
Orbit: A framework for multi-agent security evaluations — LessWrong
1+ day, 30+ min ago (25+ words) This post announces work completed as part of the MATS 9 program, supervised by Dr. Christian Schroeder de Witt. Moving forward, Orbit will be suppor…...
Can Recursive Self-Report Probing Detect Emergent Misalignment? — LessWrong
23+ hour, 1+ min ago (360+ words) In this post, I summarize the findings from my work, which I did as part of the BlueDot AI Safety Course. The full code is available here. …...