{"benchmark": {"slug": "swe-bench", "name": "SWE-bench", "description": "SWE-bench is the original benchmark suite (2,294 real GitHub issues) introduced in Jimenez et al., ICLR 2024. Verified (500 tasks) is the human-curated subset used for most modern leaderboards.\n\nFull swe-bench remains relevant for training and historical comparisons. OpenHands Index publishes swe-bench component scores under `openhands-swe-bench` (OpenHands SDK runs) \u2014 not identical to swebench.com Verified submissions.", "category": "eval", "split": null, "language": "python", "task_count": 2294, "homepage_url": "https://www.swebench.com/", "paper_url": "https://arxiv.org/abs/2310.06770", "hf_dataset_id": null, "official_leaderboard_url": "https://www.swebench.com/", "is_public_leaderboard": true, "harness_count": 0, "model_count": 0, "score_count": 0, "best_success_rate": null}, "protocol": null, "rankings": {"benchmark": {"slug": "swe-bench", "name": "SWE-bench", "description": "SWE-bench is the original benchmark suite (2,294 real GitHub issues) introduced in Jimenez et al., ICLR 2024. Verified (500 tasks) is the human-curated subset used for most modern leaderboards.\n\nFull swe-bench remains relevant for training and historical comparisons. OpenHands Index publishes swe-bench component scores under `openhands-swe-bench` (OpenHands SDK runs) \u2014 not identical to swebench.com Verified submissions.", "category": "eval", "split": null, "language": "python", "task_count": 2294, "homepage_url": "https://www.swebench.com/", "paper_url": "https://arxiv.org/abs/2310.06770", "hf_dataset_id": null, "official_leaderboard_url": "https://www.swebench.com/", "is_public_leaderboard": true, "harness_count": null, "model_count": null, "score_count": null, "best_success_rate": null}, "sort": "success_rate", "rows": []}}