Draft:DABstep Benchmarks
This draft is not written from a neutral point of view. Wikipedia articles must be written neutrally in a formal, impersonal, and dispassionate way. They should not read like a blog post, advertisement, or fan page. Rewrite the draft to remove:
Where to get help
How to improve a draft
You can also browse Wikipedia:Featured articles and Wikipedia:Good articles to find examples of Wikipedia's best writing on topics similar to your proposed article. Improving your odds of a speedy review To improve your odds of a faster review, tag your draft with relevant WikiProject tags using the button below. This will let reviewers know a new draft has been submitted in their area of interest. For instance, if you wrote about a female astronomer, you would want to add the Biography, Astronomy, and Women scientists tags. Editor resources
|
Submission declined on 3 August 2026 by EaglesFan37 (talk). This draft does not include any sources or inline citations. Wikipedia's verifiability policy requires that all content be supported by reliable sources. You should also use inline citations (footnotes) to show which source supports which specific statement.
Declined by EaglesFan37 43 hours ago.The draft requires multiple published secondary sources that:
Please add references that meet all three of these criteria. If none exist, the subject is not yet suitable for Wikipedia. You must place an inline citation directly after:
Learn how to create inline citations in the:
|
DABstep (Data Agent Benchmark for Multi-step Reasoning) is a benchmark for evaluating artificial intelligence agents on complex, multi-step data analysis tasks. It was developed jointly by Adyen and Hugging Face and first published in 2025.[1]
Background
Existing benchmarks for data analysis AI evaluated models on isolated code generation tasks or synthetic question-answering problems.[1] DABstep was designed to address limitations in these earlier evaluations by grounding tasks in actual operational workloads from a financial analytics environment.[1]
Design and structure
DABstep comprises over 450 real-world challenges derived from a financial analytics platform, requiring models to combine code-based data processing with contextual reasoning over heterogeneous documentation.[1] Tasks are derived directly from operational workloads at Adyen and reflect the complex, iterative problem-solving scenarios faced by professional data analysts.[1]
The benchmark integrates both structured data, including CSV and JSON files representing transaction telemetry and business metadata, and unstructured documentation such as Markdown files defining domain-specific formulas and rules.[1] Tasks require technical proficiency in data manipulation, including filtering, aggregation, and joins, as well as the ability to extract and apply domain-specific rules from documentation.[1]
Tasks are divided into two difficulty splits. The easy split covers tasks resolvable by reasoning over a single file or data source. The hard split requires multi-step reasoning across multiple files and source documents simultaneously and accounts for the majority of the benchmark's tasks.[1]
DABstep uses factoid-style evaluation in which each task output maps to a binary outcome — correct or incorrect — enabling objective scoring at scale without requiring human interpretation.[1] Unlike benchmarks such as SWE-bench or MLE-bench, DABstep is designed for low-barrier usage; generating answers requires only access to a code execution environment, and participants can submit answers directly to a leaderboard for automatic evaluation.[1] Results
Results from the benchmark's original evaluation revealed a substantial performance gap between AI agent capability and real-world data analysis requirements.[1] Even the best agent at the time of publication achieved only 14.55% accuracy on the hard tasks, while performance on the easy split was considerably higher, with the top model reaching 76.39% accuracy.[1]
Subsequent systems improved substantially on the original baseline results. DS-STAR, a role-decomposed data science agent developed at Google, raised accuracy on DABstep to 45.2% and secured the top rank on the public leaderboard as of September 2025.[2]
Publication and availability
The benchmark paper was submitted to the NeurIPS 2025 Datasets and Benchmarks Track.[3] The benchmark is released with a public leaderboard on Hugging Face Spaces and an open dataset at huggingface.co/datasets/adyen/dabstep.[1]
References
[edit]- 1 2 3 4 5 6 7 8 9 10 11 12 13 Egg, Alex; Iglesias Goyanes, Martin; Kingma, Friso; Mora, Andreu; von Werra, Leandro; Wolf, Thomas (2025). "DABstep: Data Agent Benchmark for Multi-step Reasoning". arXiv:2506.23719.
{{cite arXiv}}: CS1 maint: missing class (link) A bot will complete this citation soon. Click here to jump the queue - ↑ "DS-STAR: A state-of-the-art versatile data science agent". Google Research. 2025.
- ↑ "DABstep: Data Agent Benchmark for Multi-step Reasoning". OpenReview. April 2025.
References Egg, Alex; Iglesias, Martin; Kingma, Friso; Mora, Andreu; Von Werra, Leandro; Wolf, Thomas (2025). "DABstep: Data Agent Benchmark for Multi-step Reasoning". arXiv:2506.23719. Submitted to NeurIPS 2025 Datasets and Benchmarks Track. Adyen Tech (February 2025). "Data Agent Benchmark for Multi-step Reasoning (DABstep)". Medium / Adyen Tech Blog. Hugging Face (February 2025). "DABstep: Data Agent Benchmark for Multi-step Reasoning". Hugging Face Blog. Google Research (2025). "DS-STAR: A state-of-the-art versatile data science agent". Google Research Blog. DABstep Leaderboard. Hugging Face Spaces. huggingface.co/spaces/adyen/DABstep. DABstep Dataset Card. Hugging Face Datasets. huggingface.co/datasets/adyen/dabstep.

- provide significant coverage: discuss the subject in detail, not just brief mentions or routine announcements;
- are reliable: from reputable outlets with editorial oversight;
- are independent: not connected to the subject, such as interviews, press releases, the subject's own website, or sponsored content.
Please add references that meet all three of these criteria. If none exist, the subject is not yet suitable for Wikipedia.