Sources
Position Map
Where the models stand
Each row is one part of the answer. The cards show the distinct positions; the model chips show who supports each one.
Only partially. SWE-bench is a useful signal for how well an AI system resolves a particular class of task: fixing issues in existing, m
Only partially. SWE-bench is a useful signal for how well an AI system resolves a particular class of task: fixing issues in existing, mostly open-source repositories while usi
- OpenAI
- Gemini
- DeepSeek
It is not a reliable universal ranking of which coding agent will work best for a particular team.
Supports this conclusion
- OpenAI
- Gemini
- DeepSeek
Repository-level bug fixing
Supports this conclusion
- OpenAI
- Gemini
- DeepSeek
Navigating unfamiliar code
Supports this conclusion
- OpenAI
- Gemini
See how each model moved across checks
| Model | Aug 18 | Aug 25 | Sep 01 |
|---|---|---|---|
| OpenAI | |||
| Gemini | |||
| Grok | — | — | |
| DeepSeek | — | — |
Cite this answer
Consensus Watch
Run history
View the full agreement chart
Agreement over time
How strongly the models support the same claims. Every point links to its run below.
Checks
Newest first. Open any saved result to read the full consensus from that date.
-
90/100 agreement
No meaningful movement detected in this check.
Open this consensus -
75/100 agreement
Both versions reach the same conclusion: SWE-bench is a useful screening tool for autonomous repository-level bug fixing, but not a reliable predictor of the best overall AI coding assistant for real-world development. The new version merely expands on details and concrete examples.
Open this consensus -
90/100 agreement
No meaningful movement detected in this check.
Open this consensus