This report is a benchmark of modern speech recognition systems evaluated across various real-world conditions, including conversational, multilingual, noisy, and long-form audio. The evaluation methodology ensures that all models are tested under consistent conditions using a shared evaluation pipeline. The findings indicate that performance varies significantly depending on the domain, with systems that perform well on clean speech often struggling in conversational and noisy environments. The report identifies two notable patterns: conversational speech presents the greatest challenges due to interruptions and speaker variability, while performance on multilingual and noisy audio is inconsistent among different models. To facilitate fair comparisons, all transcripts are normalized before scoring, enhancing the accuracy of the Word Error Rate (WER) as a measure of model performance. The benchmark is designed for reproducibility, allowing results to be verified or extended, with all datasets and evaluation steps clearly defined.