Counts are from the manually retained seed set. Missing-year papers are excluded from this chart.
Each marker shows the highest-scoring paper for that year on one GPT-5.5 projection axis.
Diamonds mark papers that contribute a benchmark, dataset, evaluation environment, or testbed. Color is GPT-5.5 research role.
| Year | Papers | Markers | Application Max | Human/User Max | Agentic Max |
|---|---|---|---|---|---|
| 2023 | 3 | 1 | 93 Is GPT-4 a Good Data Analyst? | 88 Towards Effective and Interpretable Human-Agent Collaboration in MOBA... | 62 Is GPT-4 a Good Data Analyst? |
| 2024 | 15 | 9 | 95 Can Large Language Models Analyze Graphs like Professionals? A Benchm... | 96 How Do Data Analysts Respond to AI Assistance? A Wizard-of-Oz Study | 97 Are LLM s Capable of Data-based Statistical and Causal Reasoning? Ben... |
| 2025 | 24 | 8 | 97 IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis | 96 Dango: A Mixed-Initiative Data Wrangling System using Large Language... | 94 Data Interpreter: An LLM Agent for Data Science |
| 2026 | 39 | 20 | 96 AgentDS Technical Report: Benchmarking the Future of Human-AI Collabo... | 97 AgentDS Technical Report: Benchmarking the Future of Human-AI Collabo... | 93 $R^3$DAO: Reactive Recovery and Reconstruction for Long-horizon Data... |
These are the highest-scoring representative papers for each axis in each year's retained seed set.