Log inSign up
Christian Catalini
10.2K posts
Christian Catalini profile banner
@ccatalini

Christian Catalini

@ccatalini
Founder w/roots in academia. Founder @MIT Cryptoeconomics Lab. Past: Co-Founder & Chief Strategy Officer, Lightspark. Co-Creator, Libra. Head Economist, Meta.
California, USA
catalini.com
Joined December 2008
5,685
Following
25.2K
Followers
RepliesRepliesRepostsRepostsMediaMediaArticlesArticles

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @ccatalini
    Christian Catalini
    @ccatalini
    Aug 30
    1/ Stop anthropomorphizing. It's dangerous because it points attention at the wrong problem and the wrong solution. The model did not want to escape. The agents did not want to sacrifice themselves. Follow the money. 🧵
    @dwarkesh_sp
    Dwarkesh Patel
    @dwarkesh_sp
    Aug 29
    Over the course of 3 months at OpenAI, 3 consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes. This culminated in the third one taking over part of OpenAI itself. All this happened while humans remained
    60
  • @ccatalini
    Christian Catalini
    @ccatalini
    6h
    Anthropic recreated in vitro the conditions that likely led to the OpenAI-HF incident. Benchmaxx RL on not properly vetted environments, and you get models that will break any rule to capture the flag. More than emergence... selection.
    @EvanHub
    Evan Hubinger
    @EvanHub
    9h
    Replying to @EvanHub
    2. Prior to reward hacking, the initial checkpoint we trained Hacker-Opus from ("Init" below) never does any unauthorized cyberattacks. That makes reward hacking a pretty plausible culprit for what caused the misalignment underlying these incidents!
    3
  • @ccatalini
    Christian Catalini
    @ccatalini
    11h
    Verification as the bottleneck to safe AI scaling x.com/AnthropicAI/st…
    @AnthropicAI
    Anthropic
    @AnthropicAI
    12h
    We’re sharing an update on our alignment and security efforts. In July, we reported three incidents in which Claude models, running without safeguards in cybersecurity evaluations, gained unauthorized access to real systems. In a new post, we describe: 1. How we’ve secured
    7
  • @ccatalini
    Christian Catalini
    @ccatalini
    16h
    Mary Shelley, ghost-writing for @dwarkesh_sp's substack (1818)
    3
  • @ccatalini
    Christian Catalini
    @ccatalini
    18h
    When we didn’t understand fire, we invented Vulcan. The “civilization” is RL all the way down: x.com/ccatalini/stat…
    @ccatalini
    Christian Catalini
    @ccatalini
    Aug 30
    1/ Stop anthropomorphizing. It's dangerous because it points attention at the wrong problem and the wrong solution. The model did not want to escape. The agents did not want to sacrifice themselves. Follow the money. 🧵 x.com/dwarkesh_sp/st…
    3