Perhimpunan Mahasiswa SUTD Indonesia (PADI
Leveraging Hallucinations in LLMs for Drug Discovery
Pages
17
Time to read
46 mins
Publication
Language
English
Pages
17
Time to read
46 mins
Publication
Language
English
This technical report investigates the potential benefits of hallucinations generated by large language models (LLMs) in the context of drug discovery, specifically focusing on molecule property prediction. The study evaluates seven instruction-tuned LLMs across five datasets to determine whether hallucinated outputs can enhance predictive accuracy. The findings indicate that hallucinations can significantly improve performance for certain models, with Falcon3-Mamba-7B showing notable gains when such text is included. The report categorizes over 18,000 beneficial hallucinations, identifying structural misdescriptions as the most impactful type. Additionally, ablation studies reveal that larger models tend to benefit more from hallucinations, while the temperature of generation has a limited effect on performance. This research challenges the conventional view of hallucinations as purely problematic, suggesting instead that they may serve as useful signals in scientific modeling tasks. The study contributes to the understanding of how LLMs can be utilized creatively in early-stage drug discovery processes.