Back to Portal

3D Research Space

Experimental Three.js projection of Stage 5 papers onto problem-space axes. Treat this as a framework-audit tool: clusters and outliers are clues, not ground truth.

Trend Report

Structured readout over 81 manually retained Stage 5 papers using the same GPT-5.5 projection scores as the 3D browser. This is seed-set analysis, not corpus-wide prevalence.

Overview

81Stage 5 papers
38testbed / benchmark markers
35collaboration core
29autonomy baselines
7graph related
81papers with year metadata
Collaboration CoreAutonomy BaselineGraph RelatedFramework / EvaluationBoundary Reference

Interpretive Findings

Benchmark/testbed-heavy, collaboration-light: 38/81 retained papers carry a benchmark/testbed marker (19 primary benchmark/dataset, 14 evaluation testbed, 5 method-with-benchmark). Papers with both high human involvement and high application grounding remain much rarer.
Two literatures are adjacent but not merged: autonomy/data-agent papers have strong application and agentic scores, while HAI/collaboration papers have stronger human scores but weaker data/graph grounding.
Graph is a focused side line: the graph-related set is small (7) and mostly evaluates autonomous graph reasoning or graph-problem solving rather than human-agent graph analysis.
Likely thesis opening: turn benchmark/task infrastructure into selective human-agent collaboration settings: clarification, verification burden, handoff, and recovery are the missing connective tissue.

Role Trend by Year

2023 Collaboration Core: 12023 Autonomy Baseline: 12023 Boundary Reference: 12024 Collaboration Core: 62024 Autonomy Baseline: 52024 Graph Related: 32024 Boundary Reference: 12025 Collaboration Core: 122025 Autonomy Baseline: 52025 Graph Related: 22025 Framework / Evaluation: 12025 Boundary Reference: 42026 Collaboration Core: 162026 Autonomy Baseline: 182026 Graph Related: 22026 Framework / Evaluation: 12026 Boundary Reference: 220233202415202524202639

Counts are from the manually retained seed set. Missing-year papers are excluded from this chart.

Axis Max by Year

Application Human/User Agentic Depth02550751002023 Application max: 93 (Is GPT-4 a Good Data Analyst?)2023 Application max: 93 (Is GPT-4 a Good Data Analyst?)2024 Application max: 95 (Can Large Language Models Analyze Graphs like Professionals? A Benchmark, Datasets and Models)2024 Application max: 95 (Can Large Language Models Analyze Graphs like Professionals? A Benchmark, Datasets and Models)2025 Application max: 97 (IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis)2025 Application max: 97 (IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis)2026 Application max: 96 (AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science,)2026 Application max: 96 (AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science,)2023 Human/User max: 88 (Towards Effective and Interpretable Human-Agent Collaboration in MOBA Games: A Communication Perspective)2023 Human/User max: 88 (Towards Effective and Interpretable Human-Agent Collaboration in MOBA Games: A Communication Perspective)2024 Human/User max: 96 (How Do Data Analysts Respond to AI Assistance? A Wizard-of-Oz Study)2024 Human/User max: 96 (How Do Data Analysts Respond to AI Assistance? A Wizard-of-Oz Study)2025 Human/User max: 96 (Dango: A Mixed-Initiative Data Wrangling System using Large Language Model)2025 Human/User max: 96 (Dango: A Mixed-Initiative Data Wrangling System using Large Language Model)2026 Human/User max: 97 (AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science,)2026 Human/User max: 97 (AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science,)2023 Agentic Depth max: 62 (Is GPT-4 a Good Data Analyst?)2023 Agentic Depth max: 62 (Is GPT-4 a Good Data Analyst?)2024 Agentic Depth max: 97 (Are LLM s Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data)2024 Agentic Depth max: 97 (Are LLM s Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data)2025 Agentic Depth max: 94 (Data Interpreter: An LLM Agent for Data Science)2025 Agentic Depth max: 94 (Data Interpreter: An LLM Agent for Data Science)2026 Agentic Depth max: 93 ($R^3$DAO: Reactive Recovery and Reconstruction for Long-horizon Data Agent Orchestration)2026 Agentic Depth max: 93 ($R^3$DAO: Reactive Recovery and Reconstruction for Long-horizon Data Agent Orchestration)2023202420252026

Each marker shows the highest-scoring paper for that year on one GPT-5.5 projection axis.

Human vs Agentic Mechanism Space

"I Need to Find That One Chart": How Data Workers Navigate, Summarize and Communicate Analytical Conversations$R^3$DAO: Reactive Recovery and Reconstruction for Long-horizon Data Agent OrchestrationA call for collaborative intelligence: Why human-agent systems should precede ai autonomyA task-driven human-AI collaboration: When to automate, when to collaborate, when to challengeAdaptive Collaboration with Humans: Metacognitive Policy Optimization for Multi-Agent LLMs with Continual LearningAgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science,AIRepr: An Analyst-Inspector Framework for Evaluating Reproducibility of LLM s in Data ScienceAmbig-DS: A Benchmark for Task-Framing Ambiguity in Data-Science AgentsAre Large Language Models Ready for Multi-Turn Tabular Data Analysis?Are LLM s Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with DataAutodata: An agentic data scientist to create high quality synthetic dataBenchmarking Data Science AgentsCan Large Language Models Analyze Graphs like Professionals? A Benchmark, Datasets and ModelsCan LLM s Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise EnvironmentsCentaurEval: Benchmarking Human-in-the-Loop Value in Agentic CodingCocoa: Co-Planning and Co-Execution with AI AgentsCoDA-Bench: Can Code Agents Handle Data-Intensive Tasks?CoDA: Agentic Systems for Collaborative Data VisualizationCollaborative Gym: A Framework for Enabling and Evaluating Human-Agent CollaborationCoMind: Towards Community-Driven Agents for Machine Learning EngineeringConDABench: Interactive Evaluation of Language Models for Data AnalysisCowPilot: A Framework for Autonomous and Human-Agent Collaborative Web NavigationDA-Code: Agent Data Science Code Generation Benchmark for Large Language ModelsDAComp: Benchmarking Data Agents across the Full Data Intelligence LifecycleDango: A Mixed-Initiative Data Wrangling System using Large Language ModelDARE-bench: Evaluating Modeling and Instruction Fidelity of LLMs in Data ScienceData Formulator 2: Iterative Creation of Data Visualizations, with AI Transforming Data Along the WayData Interpreter: An LLM Agent for Data ScienceDatawise Agent: A Notebook-Centric LLM Agent Framework for Adaptive and Robust Data Science AutomationDeepAnalyze: Agentic Large Language Models for Autonomous Data ScienceDS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based ReasoningDSBench: How Far Are Data Science Agents from Becoming Data Science Experts?DSGym: A Holistic Framework for Evaluating and Training Data Science AgentsDuetUI: A Bidirectional Context Loop for Human-Agent Co-Generation of Task-Oriented InterfacesEnhancing Human Experience in Human-Agent Collaboration: A Human-Centered Modeling Approach Based on Positive Human GainEnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise SettingsExplanations are a Means to an End: Decision Theoretic Explanation EvaluationExposing Weaknesses of Large Reasoning Models through Graph Algorithm ProblemsFacilitating Proactive and Reactive Guidance for Decision Making on the Web: A Design Probe with WebSeekFeedback by Design: Understanding and Overcoming User Feedback Barriers in Conversational AgentsFlow-of-Options: Diversified and Improved LLM Reasoning by Thinking Through OptionsFrom Knowing to Teaching: Scaffolding Pedagogical Decisions for LLM AgentFrom Prompting to Verification: How Experience Shapes Vibe Coding PracticesGraph Counselor: Adaptive Graph Exploration via Multi-Agent Synergy to Enhance LLM ReasoningGraphArena: Evaluating and Exploring LLMs on Graph ComputationGraphOmni: A Comprehensive and Extensible Benchmark Framework for Large Language Models on Graph-theoretic TasksHow Do Analysts Understand and Verify AI-Assisted Data Analyses?How Do Data Analysts Respond to AI Assistance? A Wizard-of-Oz StudyHunt Instead of Wait: Evaluating Deep Data Research on Large Language ModelsIDA-Bench: Evaluating LLMs on Interactive Guided Data AnalysisIEvoAgent: Evolving Conversational Agent based on User Implicit FeedbackInfiAgent-DABench: Evaluating Agents on Data Analysis TasksInsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight GenerationInteractive Speculative Planning: Enhance Agent Efficiency through Co-design of System and User InterfaceIs GPT-4 a Good Data Analyst?KRAMABENCH: A Benchmark for AI Systems on Data-to-Insight Pipelines over Data LakesLAPS: Automating Hypothesis-Driven Statistical Analysis of Public Survey Using Large Language ModelsLearning from Active Human Involvement through Proxy Value PropagationLeveraging Dual Process Theory in Language Agent Framework for Real-time Simultaneous Human- AI CollaborationLongDS-Bench: On the Failure of Long-Horizon Agentic Data AnalysisMA - GTS : A Multi-Agent Framework for Solving Complex Graph Problems in Real-World ApplicationsMapping the Design Space of Teachable Social Media Feed ExperiencesMedAgentGym: A Scalable Agentic Training Environment for Code-Centric Reasoning in Biomedical Data ScienceOpenAgents: An Open Platform for Language Agents in the WildPlanTogether: Facilitating AI Application Planning Using Information Graphs and Large Language ModelsPosition paper: Towards open complex human-AI agents collaboration systems for problem solving and knowledge managementPreemptive Detection and Correction of Misaligned Actions in LLM AgentsReinforcing Real-world Service Agents: Balancing Utility and Cost in Task-oriented DialogueScaling Generalist Data-Analytic AgentsSPIO : Ensemble and Selective Strategies via LLM -Based Multi-Agent Planning in Automated Data ScienceStructVizor: Interactive Profiling of Semi-Structured Textual DataSubstance over Style: Evaluating Proactive Conversational Coaching AgentsTalk, Evaluate, Diagnose: User-aware Agent Evaluation with Automated Error AnalysisTell Me More! Towards Implicit User Intention Understanding of Language Model Driven AgentsThe Value of Information in Human-AI Decision-makingTowards Effective and Interpretable Human-Agent Collaboration in MOBA Games: A Communication PerspectiveTowards Human-centered Proactive Conversational AIUniDataBench: Evaluating Data Analytics Agents Across Structured and Unstructured DataVisionGraph: Leveraging Large Multimodal Models for Graph Theory Problems in Visual ContextWebDS: An End-to-End Benchmark for Web-based Data ScienceWho Does What? Archetypes of Roles Assigned to LLMs During Human-AI Decision-Making Human/User Involvement -> Agentic Depth -> autonomy-heavy collaboration mechanisms

Diamonds mark papers that contribute a benchmark, dataset, evaluation environment, or testbed. Color is GPT-5.5 research role.

Year Max Papers

YearPapersMarkersApplication MaxHuman/User MaxAgentic Max
2023 3 1 93 Is GPT-4 a Good Data Analyst? 88 Towards Effective and Interpretable Human-Agent Collaboration in MOBA... 62 Is GPT-4 a Good Data Analyst?
2024 15 9 95 Can Large Language Models Analyze Graphs like Professionals? A Benchm... 96 How Do Data Analysts Respond to AI Assistance? A Wizard-of-Oz Study 97 Are LLM s Capable of Data-based Statistical and Causal Reasoning? Ben...
2025 24 8 97 IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis 96 Dango: A Mixed-Initiative Data Wrangling System using Large Language... 94 Data Interpreter: An LLM Agent for Data Science
2026 39 20 96 AgentDS Technical Report: Benchmarking the Future of Human-AI Collabo... 97 AgentDS Technical Report: Benchmarking the Future of Human-AI Collabo... 93 $R^3$DAO: Reactive Recovery and Reconstruction for Long-horizon Data...

These are the highest-scoring representative papers for each axis in each year's retained seed set.

High Application, Low Human: Autonomy Gap

  1. Data Interpreter: An LLM Agent for Data Science
    The paper is best treated as an autonomous data-science workflow system that provides a strong substrate and contrast case for human-agent collaboration research. It is not primarily about collaboration, and although graph modeling is important, the main claim is autonomous DS agent execution rather than graph analysis itself.
  2. MA - GTS : A Multi-Agent Framework for Solving Complex Graph Problems in Real-World Applications
    Graph problems, graph extraction, graph algorithm selection, and a graph-problem benchmark are central to the paper’s claim. Although it is also an autonomous agent method, its strongest role in the research map is as a graph-specific method/benchmark reference.
  3. Graph Counselor: Adaptive Graph Exploration via Multi-Agent Synergy to Enhance LLM Reasoning
    Graph reasoning, GraphRAG, KGQA, graph operators, and GRBENCH evaluation are central to the paper’s claim. Although it is also an autonomy baseline, the graph-specific contribution is the dominant reason to include it.
  4. DA-Code: Agent Data Science Code Generation Benchmark for Large Language Models
    Although it is user-labeled as a benchmark and highly relevant to data-science workflows, its main role for the user's problem is as an autonomous-agent baseline: it shows what agents can and cannot do alone in complex analytical tasks, providing contrast and substrate for future human-agent oversight, handoff, or recovery mechanisms.
  5. $R^3$DAO: Reactive Recovery and Reconstruction for Long-horizon Data Agent Orchestration
    The paper is best treated as an autonomous data-science agent baseline or substrate. It clarifies what the agent can plan, execute, diagnose, and repair without human help, making it useful contrast material for later human-agent division-of-labor designs.
  6. Are LLM s Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data
    A useful near-core benchmark paper showing that autonomous LLM agents can execute some data-based statistical/causal analyses, but still fail badly at integrating domain knowledge with data—especially for causal reasoning.
  7. InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks
    The main value for the user's research space is as a strong autonomous-agent baseline for realistic data-analysis workflows. It defines what agents can attempt alone and where mechanisms like execution and self-debugging matter, but it is not centered on human-agent collaboration.
  8. CoDA-Bench: Can Code Agents Handle Data-Intensive Tasks?
    The paper is best treated as an autonomy baseline for data-science workflows: it evaluates what autonomous code agents can and cannot do in realistic data-intensive analysis settings. Although benchmark construction uses a dataset co-occurrence graph, graph reasoning is not central enough to classify it as graph_related.

High Human, Lower Application: Transferable HAI

  1. Leveraging Dual Process Theory in Language Agent Framework for Real-time Simultaneous Human- AI Collaboration
    The central contribution concerns real-time human-agent collaboration, mixed-initiative coordination, human-intent modeling, and autonomous execution under latency constraints. It is not graph-related and not primarily an evaluation framework or benchmark.
  2. Mapping the Design Space of Teachable Social Media Feed Experiences
    Best treated as a boundary reference: it is useful HCI background on teachable AI, user agency, and human-as-teacher interaction patterns, but it is outside the core domains of data science, graph analysis, and complex analytical workflow collaboration.
  3. Substance over Style: Evaluating Proactive Conversational Coaching Agents
    Useful as a boundary reference for mixed-initiative proactivity, user-centered evaluation, and the risk of auto-rater/user misalignment, but not central evidence for data-science, analytical workflow, or graph-agent systems.
  4. Towards Effective and Interpretable Human-Agent Collaboration in MOBA Games: A Communication Perspective
    Despite being a boundary-domain paper for data-analysis research, its main contribution is directly about human-agent role allocation, bidirectional communication, mixed initiative, and command-following decisions. It is best colored as collaboration_core rather than boundary_reference because the collaboration mechanism is the core claim.
  5. Learning from Active Human Involvement through Proxy Value Propagation
    Useful as a boundary reference for oversight, interruption, takeover, corrective feedback, and handoff concepts, but not central evidence for data-science or graph-analysis human-agent collaboration.
  6. Who Does What? Archetypes of Roles Assigned to LLMs During Human-AI Decision-Making
    The primary value is as conceptual evidence for human-agent role allocation, mixed initiative, oversight, disagreement, and division of labor. Although it is background rather than domain-specific, its central claim directly concerns collaboration structure.
  7. A call for collaborative intelligence: Why human-agent systems should precede ai autonomy
    This paper is best categorized as collaboration_core because it centrally addresses human-agent collaboration, autonomy versus mixed initiative, role specification, oversight, verification, handoff-like control, and evaluation of collaborative systems. Although it is also a framework paper, its main value for this research space is as a core conceptual source on human-agent division of labor.
  8. DuetUI: A Bidirectional Context Loop for Human-Agent Co-Generation of Task-Oriented Interfaces
    Despite being a boundary/background paper for the user’s domain, its main value is as central evidence for human-agent collaboration patterns: mixed initiative, bidirectional feedback, shared artifacts, delegation, human refinement, and handoff. It is not graph-related, not a benchmark, and not mainly an evaluation framework.

Closest to Target Intersection

  1. Dango: A Mixed-Initiative Data Wrangling System using Large Language Model
    The main contribution is a mixed-initiative human-AI data wrangling system with explicit role allocation, clarification, oversight, verification, and repair. It is therefore best treated as collaboration-core rather than an autonomy baseline or benchmark paper.
  2. IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis
    Although it is also clearly a benchmark, its strongest role in this research space is as core evidence for human-agent collaboration and division of labor in data-analysis workflows: the human provides evolving intent and domain judgment while the agent executes technical analysis. It should be visually marked as benchmark-shaped but colored as collaboration_core.
  3. Ambig-DS: A Benchmark for Task-Framing Ambiguity in Data-Science Agents
    Although it is also clearly a benchmark, its primary value for this research map is as a central human-agent division-of-labor paper: it directly studies when an agent should proceed autonomously versus ask a human for missing framing information in data-science work.
  4. Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration
    This is a core collaboration paper because its main contribution is a framework and benchmark for mixed-initiative human-agent collaboration, dual control, shared environments, process metrics, and division-of-labor analysis. Although it is also a framework/evaluation paper, its primary value for the user's map is as collaboration-core evidence.
  5. AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science,
    Although it is also a benchmark paper, its most important role for this research space is as core evidence about human-agent role allocation in realistic data-science workflows, including oversight, handoff, and mixed-initiative collaboration.
  6. Cocoa: Co-Planning and Co-Execution with AI Agents
    The main contribution is a concrete human-agent collaboration method for role allocation across planning, execution, verification, and replanning. It is not primarily a graph paper, benchmark paper, or autonomy baseline.
  7. LAPS: Automating Hypothesis-Driven Statistical Analysis of Public Survey Using Large Language Models
    The main value for this research space is its explicit treatment of human-agent division of labor in a real analytical workflow: user agency, oversight, approval, refinement, transparency, trust, and bounded automation are central to the paper's claim.
  8. Are Large Language Models Ready for Multi-Turn Tabular Data Analysis?
    Although it is also a benchmark, its value for this map is as a core example of mixed-initiative data-analysis collaboration: users clarify intent and provide feedback while agents generate, revise, and interpret analyses. It is not graph-related, and it is more than a generic evaluation framework.

Graph-Related Side Line

  1. MA - GTS : A Multi-Agent Framework for Solving Complex Graph Problems in Real-World Applications
    Graph problems, graph extraction, graph algorithm selection, and a graph-problem benchmark are central to the paper’s claim. Although it is also an autonomous agent method, its strongest role in the research map is as a graph-specific method/benchmark reference.
  2. Can Large Language Models Analyze Graphs like Professionals? A Benchmark, Datasets and Models
    Graph analysis and graph-problem benchmarking are central to the paper's claim. Although it is also benchmark evidence and an autonomy baseline, the most distinctive role in the user's space is as a graph-specific benchmark for tool-using LLM agents.
  3. Exposing Weaknesses of Large Reasoning Models through Graph Algorithm Problems
    Graph algorithm problems and graph-specific benchmarking are central to the paper’s claim. Although it is also a benchmark and autonomy baseline, the most distinctive role for the visualization is graph-related evidence.
  4. Graph Counselor: Adaptive Graph Exploration via Multi-Agent Synergy to Enhance LLM Reasoning
    Graph reasoning, GraphRAG, KGQA, graph operators, and GRBENCH evaluation are central to the paper’s claim. Although it is also an autonomy baseline, the graph-specific contribution is the dominant reason to include it.
  5. VisionGraph: Leveraging Large Multimodal Models for Graph Theory Problems in Visual Context
    Graph-theory reasoning and graph-problem benchmarking are central to the paper's claim. Although it also contributes benchmark infrastructure and an agentic method, its strongest role in the research map is as a graph-specific benchmark/reference paper.
  6. GraphArena: Evaluating and Exploring LLMs on Graph Computation
    Graph computation is central to the paper's claim and contribution. Although it is also a benchmark, the most distinctive role in this research map is as a graph-problem benchmark anchor rather than a general evaluation-framework paper.
  7. GraphOmni: A Comprehensive and Extensible Benchmark Framework for Large Language Models on Graph-theoretic Tasks
    Graph reasoning is central to the claim: the paper contributes a benchmark specifically for graph-theoretic tasks and LLM graph-problem reasoning. Although it is also a benchmark paper, the graph-related role is the most distinctive placement for the research-space color scheme.

Method Papers with Benchmark Contribution

  1. MA - GTS : A Multi-Agent Framework for Solving Complex Graph Problems in Real-World Applications
    Graph problems, graph extraction, graph algorithm selection, and a graph-problem benchmark are central to the paper’s claim. Although it is also an autonomous agent method, its strongest role in the research map is as a graph-specific method/benchmark reference.
  2. VisionGraph: Leveraging Large Multimodal Models for Graph Theory Problems in Visual Context
    Graph-theory reasoning and graph-problem benchmarking are central to the paper's claim. Although it also contributes benchmark infrastructure and an agentic method, its strongest role in the research map is as a graph-specific benchmark/reference paper.
  3. OpenAgents: An Open Platform for Language Agents in the Wild
    The best role is collaboration_core because the paper is centrally useful for studying in-the-wild human-agent interaction, oversight, intervention, and feedback in deployed agent workflows. It is also an autonomy substrate, but the user-facing human-in-the-loop platform framing is more central than a purely autonomous baseline role.
  4. Is GPT-4 a Good Data Analyst?
    The paper is best treated as an autonomy baseline: GPT-4 performs the full data-analysis workflow with humans mainly serving as comparison baselines and judges. It helps motivate later collaboration and oversight work by showing both autonomous capabilities and verification gaps.
  5. From Knowing to Teaching: Scaffolding Pedagogical Decisions for LLM Agent
    Best treated as a boundary/reference paper: useful for ideas about multi-agent decomposition, provenance, coherence constraints, and benchmark design, but not central evidence for human-AI division of labor in data science, graph analysis, or real-world analytical workflows.

Framework Stress Cases

  1. No papers.

Primary Benchmark / Dataset Papers

  1. Are LLM s Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data
    A useful near-core benchmark paper showing that autonomous LLM agents can execute some data-based statistical/causal analyses, but still fail badly at integrating domain knowledge with data—especially for causal reasoning.
  2. UniDataBench: Evaluating Data Analytics Agents Across Structured and Unstructured Data
    Although it is a benchmark paper, its most useful role for the user's research space is as a core autonomous data-analytics-agent baseline and substrate for comparison against human-in-the-loop or mixed-initiative systems. It is not classified as graph_related because graph structures are used for metadata/link discovery rather than being the central benchmark target.
  3. KRAMABENCH: A Benchmark for AI Systems on Data-to-Insight Pipelines over Data Lakes
    Although it is also benchmark evidence, its main role in this research space is as an autonomous data-science-agent baseline/substrate. It defines realistic analytical tasks and failure modes against which human-agent collaboration or oversight methods could later be compared. It is not graph-related and not a collaboration-core paper.
  4. LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis
    The paper is best categorized as an autonomy baseline because it evaluates mostly autonomous data-analysis agents and exposes their limits in long-horizon workflows. Although it is also a benchmark, its role for the user’s research is primarily to provide contrast and motivation for human oversight and mixed-initiative recovery.
  5. Ambig-DS: A Benchmark for Task-Framing Ambiguity in Data-Science Agents
    Although it is also clearly a benchmark, its primary value for this research map is as a central human-agent division-of-labor paper: it directly studies when an agent should proceed autonomously versus ask a human for missing framing information in data-science work.
  6. InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks
    The main value for the user's research space is as a strong autonomous-agent baseline for realistic data-analysis workflows. It defines what agents can attempt alone and where mechanisms like execution and self-debugging matter, but it is not centered on human-agent collaboration.
  7. Are Large Language Models Ready for Multi-Turn Tabular Data Analysis?
    Although it is also a benchmark, its value for this map is as a core example of mixed-initiative data-analysis collaboration: users clarify intent and provide feedback while agents generate, revise, and interpret analyses. It is not graph-related, and it is more than a generic evaluation framework.
  8. Benchmarking Data Science Agents
    Although the user label is benchmark and the paper is benchmark evidence, its most useful role in this research map is as an autonomy baseline: it characterizes what autonomous data-science agents can and cannot do in realistic workflows, against which human-agent division-of-labor systems can be compared.

Evaluation Environment / Testbed Papers

  1. IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis
    Although it is also clearly a benchmark, its strongest role in this research space is as core evidence for human-agent collaboration and division of labor in data-analysis workflows: the human provides evolving intent and domain judgment while the agent executes technical analysis. It should be visually marked as benchmark-shaped but colored as collaboration_core.
  2. Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration
    This is a core collaboration paper because its main contribution is a framework and benchmark for mixed-initiative human-agent collaboration, dual control, shared environments, process metrics, and division-of-labor analysis. Although it is also a framework/evaluation paper, its primary value for the user's map is as collaboration-core evidence.
  3. DSGym: A Holistic Framework for Evaluating and Training Data Science Agents
    The paper is best treated as an autonomy baseline: it studies mostly autonomous data-science agents on realistic workflows and provides infrastructure and failure evidence useful for contrasting with human-in-the-loop or division-of-labor systems. It is not graph-central and not primarily a collaboration paper.
  4. WebDS: An End-to-End Benchmark for Web-based Data Science
    Best treated as an autonomy baseline for realistic data science workflows. It is highly useful for showing where autonomous agents fail and where human oversight might help, but it is not itself a collaboration or division-of-labor study.
  5. Hunt Instead of Wait: Evaluating Deep Data Research on Large Language Models
    The paper is best used as an autonomy baseline for complex data-analysis agents: it shows what LLM agents can and cannot do alone in exploratory data science. Although it is a benchmark paper, its main relevance to the user's problem is as a substrate and contrast point for adding human oversight, guidance, handoff, and recovery.
  6. DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle
    The paper is best treated as an autonomy baseline and benchmark substrate for complex data work. It defines realistic data-agent tasks and exposes autonomy failures, but does not center mixed-initiative collaboration or human-agent division of labor. Although it is a benchmark, its role in this research map is primarily to characterize autonomous data-agent capability in the target application domain.
  7. InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation
    Although it is a benchmark, its main role for the user's research space is to define the autonomous end of data-analysis work: agents perform the full analytics pipeline without interactive human guidance. It provides a strong baseline/substrate for later comparisons with mixed-initiative, oversight, or verification-centered systems.
  8. CoMind: Towards Community-Driven Agents for Machine Learning Engineering
    Best categorized as an autonomous data-science-agent baseline/substrate. Although it includes community knowledge and a benchmark, the main relevance to the user's problem is as a contrast case: agents exploiting human-produced artifacts autonomously rather than collaborating with humans in the loop.