AI is ahead
AI beats the best humans in the world.
IBM's Deep Blue beat world champion Garry Kasparov. Today's engines are far beyond any human.
IBM's Watson beat the show's two greatest champions.
DeepMind's AlphaGo beat Lee Sedol, a feat experts expected to be a decade away.
Pluribus beat elite professionals at six-player no-limit Texas hold'em, a game of hidden information and bluffing.
OpenAI Five beat OG, the reigning world champion team, in a complex real-time team strategy game.
DeepMind's AlphaFold predicted 3D protein shapes with near-lab accuracy, earning a share of the 2024 Nobel Prize in Chemistry.
At the ICPC World Finals, an OpenAI system solved all 12 problems, a result no human team matched.
AI is level
AI performs at expert or skilled-human level, but not clearly above the best.
Microsoft's system beat the estimated human error rate on the ImageNet photo-labeling benchmark.
Microsoft reported transcription accuracy matching professional human transcribers on a standard benchmark.
DeepMind's AlphaStar reached Grandmaster, ranking above 99.8% of active players.
GPT-4 earned a passing score on the US Uniform Bar Exam.
Systems from Google DeepMind and OpenAI scored at gold-medal level at the International Mathematical Olympiad.
Top AI systems now beat the average human tester on puzzles designed to resist memorization.
Humans are ahead
Areas where people still clearly outperform AI.
Launched in March 2026; humans solve it readily while the best AI systems fall far short.
On Humanity's Last Exam, built from questions that stump specialists, top AI gets roughly half to about 60% right.
AI agents handle tasks lasting hours, but still struggle with projects that take a person weeks.
Robotaxis run without drivers in selected cities, but not on every road and in every condition a human can handle.
Folding laundry or tidying a strange kitchen remains hard for robots. See sirobot.us for where they stand.
The year is when the milestone was reached or, for "Humans are ahead", when the status was last checked. Benchmark scores are reported inconsistently, so we use these three categories rather than precise percentages. Last reviewed October 2026.