1. X
  2. Andrew Ho
Log inSign up
Andrew Ho
261 posts
user avatar
Andrew Ho
@andrewho03
Prev: @OpenAI
San Francisco, CA
andrewho.xyz
Joined June 2026
147
Following
13.3K
Followers
RepliesRepliesArticlesArticlesMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    user avatar
    Andrew Ho
    @andrewho03
    Jul 29
    Today is my last day at @OpenAI. I'm glad to have spent the last eight months of my life working here! I'm starting a new company focused on the production of high-quality reinforcement learning datasets: 1. The generalization ability of LLMs is clearly very poor, with "spiky"
    1.2M
  • user avatar
    Andrew Ho
    @andrewho03
    57m
    I'm not that worried about AI safety, but that's because I think that AI might be humanity's last hope before descending into, at best, a "dark age" lasting centuries. Since I don't think AI risk is extinction-level, I'm willing to take a bet on stronger capabilities. Prior to
    4.8K
  • user avatar
    Andrew Ho
    @andrewho03
    17h
    user avatar
    Pangram
    @pangram
    17h
    Replying to @d_kopfmann and @andrewho03
    We believe that this entire text is human-written. pangram.com/history/de09fc…
    5.7K
  • user avatar
    Andrew Ho
    @andrewho03
    19h
    Looking for connects to anyone at Mistral interested in these datasets: long-horizon statistical analysis, CAD environments (CUA or tool use), PDF parsing (finance & other white collar work), biology in general, chemistry in general.
    11K
  • user avatar
    Andrew Ho
    @andrewho03
    20h
    Constructing high-quality datasets at scale is really hard. Incredibly so. No AI lab researcher is an expert in every field, let alone necessarily more than one; to produce even a simple benchmark means continually staring at rollouts for totally unfamiliar fields. Even when you
    user avatar
    Daniel Rupawalla
    @rupawalladaniel
    21h
    it's insane to me how many people take benchmark scores at face value. i spent 2 minutes looking at GDP-val tasks and realized there's a verifier item that checks for: " The Word document includes the warehouse phone number 560-555-3867 (accepts common US formats such as
    17K