The document is a technical report presenting DRACO, a benchmark designed for evaluating complex deep research tasks. It outlines the methodology for creating 100 tasks that span 10 domains and utilize information from 40 countries, derived from real-world usage patterns of a large-scale deep research system. The tasks are anonymized and designed to be open-ended, complex, and objectively evaluable, graded against specific rubrics that assess factual accuracy, breadth and depth of analysis, presentation quality, and citation quality. The report discusses the challenges of evaluating deep research systems and the importance of having a comprehensive dataset that reflects realistic use cases. It details the task construction process, which includes sampling, pre-processing, augmentation, filtering, and curation, ensuring that the tasks are representative of actual user needs. The document also compares DRACO with existing benchmarks and highlights its unique contributions to the field of deep research evaluation.