News

lesswrong.com
lesswrong.com > posts > PG2RupAvCNqZ5PAvu > a-solution-to-cryptographic-boxes-for-unfriendly-ai

A Solution to Cryptographic Boxes for Unfriendly AI — LessWrong

3+ hour, 6+ min ago  (1437+ words) Summary In 2010, Paul Christiano wrote Cryptographic Boxes for Unfriendly AI, in which he asks how we can sandbox arbitrarily dangerous AIs and recom…...

lesswrong.com
lesswrong.com > posts > CCt9oy8axcdvaGcPE > a-red-line-and-oversight-framework-for-government-ai

A Red Line and Oversight Framework for Government AI Contracts — LessWrong

8+ hour, 41+ min ago  (1067+ words) Principled contract language for selling AI to governments, with transparency that preserves principles against pressure....

lesswrong.com
lesswrong.com > posts > 2JQzDZXjoG2opnAjk > my-payorian-fairbot-was-just-the-original-fairbot

My "Payorian FairBot" was just the original FairBot — LessWrong

10+ hour, 42+ min ago  (618+ words) MIRI's proof-based prisoner's dilemma tournament defined agents encoded as formulas of Peano arithmetic (PA) with one free variable. means the agent…...

lesswrong.com
lesswrong.com > posts > xWZpwPPp5nR9rZq3x > endogenous-alignment

Endogenous Alignment — LessWrong

13+ hour, 29+ min ago  (223+ words) When we think about building aligned AIs, we often imagine an AI that’s in harmony with something like the coherent extrapolation of our highest values. But in humans, we don’t get there directly. Instead, we typically progress from exogenous to…...

lesswrong.com
lesswrong.com > posts > iwKT6rYQsBwLbMXZA > the-coming-of-the-global-brain-a-review-of-the-god-test

The Coming of the Global Brain: A Review of "The God Test" — LessWrong

10+ hour, 32+ min ago  (77+ words) I have published a brief review of Robert Wright's new book on AI and took the opportunity to discuss Teilhard de Chardin and the directionality of evolution:https://theanticompletionist.substack.com/p/the-coming-of-the-global-brainDiscuss The Coming of the Global Brain: A Review of "The God Test…...

lesswrong.com
lesswrong.com > posts > zvpocS7CswgT8bw7E > nuances-in-the-workings-of-the-eye-and-retina

Nuances in the Workings of the Eye and Retina — LessWrong

9+ hour, 30+ min ago  (1434+ words) Image edited from source at Britannica • The eye (any eye) is a miracle of evolution, which has filled its structure at every scale with an apparent…...

lesswrong.com
lesswrong.com > posts > Pnk4kcQ95qLdyLCwR > map-and-territory-predictably-wrong

Map and Territory, Predictably Wrong — LessWrong

19+ hour, 52+ min ago  (162+ words) Crossposted (with small tweaks) from my Substack. What Do I Mean By “Rationality”? … What’s a Bias, Again? Some Thoughts of mine Note: I used ChatGPT for minimal proofreading/copyediting - fixing typos, grammar mistakes and suggestions of awkward or unidiomatic phrasing....

lesswrong.com
lesswrong.com > posts > tEFD2bgNWZ6XcurKA > the-most-forbidden-technique-is-not-always-forbidden-1

The Most Forbidden Technique is not always forbidden — LessWrong

1+ day, 2+ hour ago  (21+ words) A few days ago, Goodfire announced a private beta of Silico, their LLM training platform. As part of the announcement, they made a post describing Si…...

lesswrong.com
lesswrong.com > posts > 5kEezahvkEiQuWEfx > should-we-benchmark-conceptual-capabilities-using-judgment

Should we benchmark conceptual capabilities using judgment prediction tasks? — LessWrong

1+ day, 3+ hour ago  (322+ words) A bunch of conceptual reasoning tasks involve very subjective judgments, which makes them poorly suited for benchmarking AI capabilities. For example, it seems unreasonable to benchmark how well AIs can predict the probability of misaligned AI takeover. Perhaps instead we…...

lesswrong.com
lesswrong.com > posts > F3FqoM45j8sHdo9bL > longtermism-is-very-intuitive

Longtermism is very intuitive. — LessWrong

1+ day, 4+ hour ago  (972+ words) As a relatively new species, humanity could have time to settle on distant planets, create artificial biospheres, upload digital minds, and achieve unfathomably high amounts of total utility. It seems intuitive from any consequentialist standpoint that we have a moral…...