• Treffer 1 von 1
Zurück zur Trefferliste

RanDepict: Random chemical structure depiction generator

  • The development of deep learning-based optical chemical structure recognition (OCSR) systems has led to a need for datasets of chemical structure depictions. The diversity of the features in the training data is an important factor for the generation of deep learning systems that generalise well and are not overfit to a specific type of input. In the case of chemical structure depictions, these features are defined by the depiction parameters such as bond length, line thickness, label font style and many others. Here we present RanDepict, a toolkit for the creation of diverse sets of chemical structure depictions. The diversity of the image features is generated by making use of all available depiction parameters in the depiction functionalities of the CDK, RDKit, and Indigo. Furthermore, there is the option to enhance and augment the image with features such as curved arrows, chemical labels around the structure, or other kinds of distortions. Using depiction feature fingerprints, RanDepict ensures diversely picked image features. Here, the depiction and augmentation features are summarised in binary vectors and the MaxMin algorithm is used to pick diverse samples out of all valid options. By making all resources described herein publicly available, we hope to contribute to the development of deep learning-based OCSR systems.

Metadaten exportieren

Weitere Dienste

Teilen auf Twitter Suche bei Google Scholar
Metadaten
Verfasserangaben:Henning Otto Brinkhaus, Kohulan Rajan, Achim Zielesny, Christoph Steinbeck
ISSN:1758-2946
Titel des übergeordneten Werkes (Englisch):Journal of Cheminformatics
Verlag:BioMed Central
Verlagsort:London
Dokumentart:Wissenschaftlicher Artikel
Sprache:Englisch
Datum der Veröffentlichung (online):06.06.2022
Datum der Erstveröffentlichung:06.06.2022
Veröffentlichende Institution:Westfälische Hochschule Gelsenkirchen Bocholt Recklinghausen
Datum der Freischaltung:23.12.2022
Freies Schlagwort / Tag:CDK; Chemical image depiction; Depiction generator image augmentation; Indigo; OCSR; RDKit
Jahrgang:14.2022
Ausgabe / Heft:31
Seitenzahl:7
Erste Seite:1
Letzte Seite:7
Lizenz (Deutsch):License LogoEs gilt das Urheberrechtsgesetz

$Rev: 13159 $