Skip to content
View klieret's full-sized avatar
💭
🦊
💭
🦊

Highlights

  • Pro

Block or report klieret

Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
klieret/README.md

📣 NEWS: Check out ProgramBench, a 0% software from scratch benchmark
📣 NEWS: Check out mini-swe-agent, a 100 line AI agent that scores 77% on SWE-bench verified!
📣 NEWS: Check out CodeClash, the first goal-oriented SWE benchmark!

Hi 👋 I'm an AI research scientist at Meta Superintelligence focusing on agentic AI for software development. Some of my smaller open source projects are described here (but it's an incomplete list).

Pinned Loading

  1. SWE-agent/SWE-agent SWE-agent/SWE-agent Public

    SWE-agent takes a GitHub issue and tries to automatically fix it, using your LM of choice. It can also be employed for offensive cybersecurity or competitive coding challenges. [NeurIPS 2024]

    Python 20.2k 2.2k

  2. SWE-agent/mini-swe-agent SWE-agent/mini-swe-agent Public

    The 100 line AI agent that solves GitHub issues or helps you in your command line. Radically simple, no huge configs, no giant monorepo—but scores >74% on SWE-bench verified!

    Python 6.9k 958

  3. facebookresearch/ProgramBench facebookresearch/ProgramBench Public

    Can Language Models Rebuild Programs From Scratch?

    Python 914 65

  4. SWE-agent/SWE-ReX SWE-agent/SWE-ReX Public

    Sandboxed code execution for AI agents, locally or on the cloud. Massively parallel, easy to extend. Powering SWE-agent and more.

    Python 581 119

  5. SWE-bench/SWE-smith SWE-bench/SWE-smith Public

    [NeurIPS 2025 D&B Spotlight] Scaling Data for SWE-agents

    Python 756 127

  6. CodeClash-ai/CodeClash CodeClash-ai/CodeClash Public

    Benchmarking Goal-Oriented Software Engineering

    Python 205 20