DECIMER—hand-drawn molecule images dataset
- The translation of images of chemical structures into machine-readable representations of the depicted molecules is known as optical chemical structure recognition (OCSR). There has been a lot of progress over the last three decades in this field, but the development of systems for the recognition of complex hand-drawn structure depictions is still at the beginning. Currently, there is no data for the systematic evaluation of OCSR methods on hand-drawn structures available. Here we present DECIMER — Hand-drawn molecule images, a standardised, openly available benchmark dataset of 5088 hand-drawn depictions of diversely picked chemical structures. Every structure depiction in the dataset is mapped to a machine-readable representation of the underlying molecule. The dataset is openly available and published under the CC-BY 4.0 licence which applies very few limitations. We hope that it will contribute to the further development of the field.
Verfasserangaben: | Henning Otto Brinkhaus, Achim Zielesny, Christoph Steinbeck, Kohulan Rajan |
---|---|
ISSN: | 1758-2946 |
Titel des übergeordneten Werkes (Englisch): | Journal of Cheminformatics |
Verlag: | BioMed Central |
Verlagsort: | London |
Dokumentart: | Wissenschaftlicher Artikel |
Sprache: | Englisch |
Datum der Veröffentlichung (online): | 09.06.2022 |
Datum der Erstveröffentlichung: | 09.06.2022 |
Veröffentlichende Institution: | Westfälische Hochschule Gelsenkirchen Bocholt Recklinghausen |
Datum der Freischaltung: | 21.12.2022 |
Jahrgang: | 14.2022 |
Seitenzahl: | 5 |
Erste Seite: | 1 |
Letzte Seite: | 5 |
Fachbereiche / Institute: | Institute / Institut für biologische und chemische Informatik |
Lizenz (Deutsch): | Es gilt das Urheberrechtsgesetz |