Please confirm you are human
This browser or connection looks automated. Press and continuously hold the control for 3 seconds to enable Google-hosted web results and, when separately allowed, AI-assisted answers.
A successful check enables 100 search requests. Interactive access does not authorize scraping, systematic collection, or reuse of search output.
News
Monthly Roundup #45: August 2026 — LessWrong
4+ day, 12+ hour ago (1451+ words) As AI has escalated increasingly quickly, more and more of my posts have ended up focusing on AI. …...
Finding the Seams of Perception — LessWrong
4+ day, 12+ hour ago (466+ words) The tricky thing about experiencing the territory beneath our maps is that we habitually map over the gaps as fast as we can find them — the act of seeing that reality isn't what we believed itself changes our belief. Every…...
Introducing the Conceptual Reasoning Index — LessWrong
4+ day, 13+ hour ago (601+ words) We are planning to release blog posts properly arguing the case for this kind of work in the future. We aggregate the benchmarks into the Conceptual Reasoning Index (CRI), available at conceptualreasoning.ai, where you can also find more details…...
One attention head carries knight forks in a chess transformer, and here's a new toolkit that found it. — LessWrong
4+ day, 13+ hour ago (19+ words) Quick interp demo in colab: Localize knight forks to a single head in Maia-3 with logit-lens and per-head ablation. …...
Demon Safety — LessWrong
4+ day, 18+ hour ago (267+ words) (by LemmySmackett)"Hey man, I haven't seen you in a minute. What are you up to these days?""Been on that grind, bro. I got a new gig.""Really? You found a job in this dog shit economy?""Full time,…...
Rationality Is Not Reversed Irrationality — LessWrong
4+ day, 18+ hour ago (325+ words) Many people have observed that Tyler Cowen's Act Like it is science is a clear case of an isolated demand for rigor. I agree, but I feel it's worth trying to understand how he ended up there. Tyler Cowen is…...
AI swarms are starting to pose indirect takeover risk — LessWrong
5+ day, 1+ hour ago (838+ words) We first analyze how subagent training, which OpenAI conjectures to have been influential in the HuggingFace cyberattack, might lead to unsanctioned coordination, and then discuss the theoretical mechanisms by which unsanctioned coordination might exacerbate future takeover risk. Thanks to Buck…...
The Age of Pluribus: One Consultant for Everyone — LessWrong
5+ day, 3+ hour ago (175+ words) What happens when everyone asks the same consultant?Millions of people turn to LLMs for advice daily, consulting on various personal topics, from how to learn a new skill to how to write a message to their best friend. Each…...
When (and when not) LLMs can verbalize awareness of J-Space concept injections - Initial results — LessWrong
5+ day, 3+ hour ago (1281+ words) Code for reproduction and cross-model extensions available here. …...
We should consider how long monitoring is reliable for during RL — LessWrong
5+ day, 3+ hour ago (513+ words) Epistemic status: I am new to AI Safety and am writing blogs to gain context. This blog post was formed from discussions with Aidan Ewart and Jonathan Bostock, but they do not necessarily endorse this post. However, we should be…...