SWE-bench (Lite, Verified, Multimodal, Multilingual) all in one place!
SWE-bench
community
AI & ML interests
None defined yet.
Recent Activity
View all activity
Papers
SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Organization Card
SWE-bench
We are a team of researchers across Stanford University and Princeton University working on LMs and AI systems for software engineering.
In this organization, you will find the assets for several projects in the SWE-* research ecosystem, notably:
datasets 17
SWE-bench/SWE-bench_Multilingual
Benchmark • Updated • 300 • 40.5k • 22
SWE-bench/SWE-bench_Multimodal
Viewer • Updated • 612 • 3.7k • 11
SWE-bench/SWE-prime
Viewer • Updated • 1.36k • 1.63k • 1
SWE-bench/SWE-smith-cpp
Viewer • Updated • 5.12k • 1.56k
SWE-bench/SWE-smith-ts
Viewer • Updated • 5.03k • 1.47k
SWE-bench/SWE-bench_Verified
Benchmark • Updated • 500 • 77.3k • 135
SWE-bench/SWE-smith-java
Viewer • Updated • 7.47k • 2.28k
SWE-bench/SWE-smith-rs
Viewer • Updated • 5.31k • 1.31k • 2
SWE-bench/SWE-smith-js
Viewer • Updated • 6.07k • 2.03k
SWE-bench/SWE-smith-py
Viewer • Updated • 50.9k • 4.26k • 5