Draft:DABstep Benchmarks
Submission declined on 3 August 2026 by EaglesFan37 (talk).
Where to get help
How to improve a draft
You can also browse Wikipedia:Featured articles and Wikipedia:Good articles to find examples of Wikipedia's best writing on topics similar to your proposed article. Improving your odds of a speedy review To improve your odds of a faster review, tag your draft with relevant WikiProject tags using the button below. This will let reviewers know a new draft has been submitted in their area of interest. For instance, if you wrote about a female astronomer, you would want to add the Biography, Astronomy, and Women scientists tags. Editor resources
|
DABstep (Data Agent Benchmark for Multi-step Reasoning) is a benchmark for evaluating artificial intelligence agents on complex, multi-step data analysis tasks. It was developed jointly by Adyen and Hugging Face and first published in 2025.[1]
Background
Existing benchmarks for data analysis AI evaluated models on isolated code generation tasks or synthetic question-answering problems.[1] DABstep was designed to address limitations in these earlier evaluations by grounding tasks in actual operational workloads from a financial analytics environment.[1]
Design and structure
DABstep comprises over 450 real-world challenges derived from a financial analytics platform, requiring models to combine code-based data processing with contextual reasoning over heterogeneous documentation.[1] Tasks are derived directly from operational workloads at Adyen and reflect the complex, iterative problem-solving scenarios faced by professional data analysts.[1]
The benchmark integrates both structured data, including CSV and JSON files representing transaction telemetry and business metadata, and unstructured documentation such as Markdown files defining domain-specific formulas and rules.[1] Tasks require technical proficiency in data manipulation, including filtering, aggregation, and joins, as well as the ability to extract and apply domain-specific rules from documentation.[1]
Tasks are divided into two difficulty splits. The easy split covers tasks resolvable by reasoning over a single file or data source. The hard split requires multi-step reasoning across multiple files and source documents simultaneously and accounts for the majority of the benchmark's tasks.[1]
DABstep uses factoid-style evaluation in which each task output maps to a binary outcome — correct or incorrect — enabling objective scoring at scale without requiring human interpretation.[1] Unlike benchmarks such as SWE-bench or MLE-bench, DABstep is designed for low-barrier usage; generating answers requires only access to a code execution environment, and participants can submit answers directly to a leaderboard for automatic evaluation.[1] Results
Results from the benchmark's original evaluation revealed a substantial performance gap between AI agent capability and real-world data analysis requirements.[1] Even the best agent at the time of publication achieved only 14.55% accuracy on the hard tasks, while performance on the easy split was considerably higher, with the top model reaching 76.39% accuracy.[1]
Subsequent systems improved substantially on the original baseline results. DS-STAR, a role-decomposed data science agent developed at Google, raised accuracy on DABstep to 45.2% and secured the top rank on the public leaderboard as of September 2025.[2]
Publication and availability
The benchmark paper was submitted to the NeurIPS 2025 Datasets and Benchmarks Track.[3] The benchmark is released with a public leaderboard on Hugging Face Spaces and an open dataset at huggingface.co/datasets/adyen/dabstep.[1]
References
[edit]- ^ a b c d e f g h i j k l m Egg, Alex; Iglesias Goyanes, Martin; Kingma, Friso; Mora, Andreu; von Werra, Leandro; Wolf, Thomas (2025). "DABstep: Data Agent Benchmark for Multi-step Reasoning". arXiv:2506.23719.
{{cite arXiv}}: CS1 maint: missing class (link) A bot will complete this citation soon. Click here to jump the queue - ^ "DS-STAR: A state-of-the-art versatile data science agent". Google Research. 2025.
- ^ "DABstep: Data Agent Benchmark for Multi-step Reasoning". OpenReview. April 2025.
References Egg, Alex; Iglesias, Martin; Kingma, Friso; Mora, Andreu; Von Werra, Leandro; Wolf, Thomas (2025). "DABstep: Data Agent Benchmark for Multi-step Reasoning". arXiv:2506.23719. Submitted to NeurIPS 2025 Datasets and Benchmarks Track. Adyen Tech (February 2025). "Data Agent Benchmark for Multi-step Reasoning (DABstep)". Medium / Adyen Tech Blog. Hugging Face (February 2025). "DABstep: Data Agent Benchmark for Multi-step Reasoning". Hugging Face Blog. Google Research (2025). "DS-STAR: A state-of-the-art versatile data science agent". Google Research Blog. DABstep Leaderboard. Hugging Face Spaces. huggingface.co/spaces/adyen/DABstep. DABstep Dataset Card. Hugging Face Datasets. huggingface.co/datasets/adyen/dabstep.

The draft requires multiple published secondary sources that:
Please add references that meet all three of these criteria. If none exist, the subject is not yet suitable for Wikipedia.
You must place an inline citation directly after:
Learn how to create inline citations in the: