Draft:Chemical Data Processing Toolkit
Submission declined on 2 June 2026 by Stuartyeates (talk).
Where to get help
How to improve a draft
You can also browse Wikipedia:Featured articles and Wikipedia:Good articles to find examples of Wikipedia's best writing on topics similar to your proposed article. Improving your odds of a speedy review To improve your odds of a faster review, tag your draft with relevant WikiProject tags using the button below. This will let reviewers know a new draft has been submitted in their area of interest. For instance, if you wrote about a female astronomer, you would want to add the Biography, Astronomy, and Women scientists tags. Editor resources
|
Submission declined on 29 May 2026 by A.Classical-Futurist (talk). This draft appears to contain text generated by a large language model (such as ChatGPT). You cannot use LLMs to generate article content.
LLM-generated pages with certain obvious signs of being machine generated may be deleted without notice. Instead, only summarize in your own words a range of independent, reliable, published sources that discuss the subject. See the advice page on large language models for more information.This draft reads like an advertisement. Wikipedia is an encyclopedia, not a platform for promotion or marketing. Drafts that are exclusively promotional may be deleted without notice.
Declined by A.Classical-Futurist 2 months ago.Wikipedia articles must be written neutrally in a formal, impersonal, and dispassionate way. They should not read like a blog post, advertisement, or fan page. Rewrite the draft to remove:
Instead, only summarize in your own words a range of independent, reliable, published sources that discuss the subject. If you have a conflict of interest (e.g. you are the subject, an employee, or a relative) or are being paid to edit, you must disclose this to comply with Wikipedia's Terms of Use. |
This draft appears to contain copyrighted material from https://cdpkit.org/introduction.html, which has been removed. Wikipedia strictly prohibits copyright violations. Assume all text is copyrighted unless released under a compatible license.
This draft reads like an advertisement. Wikipedia is an encyclopedia, not a platform for promotion or marketing. Drafts that are exclusively promotional may be deleted without notice.
Declined by NeoGaze 2 months ago.Wikipedia articles must be written neutrally in a formal, impersonal, and dispassionate way. They should not read like a blog post, advertisement, or fan page. Rewrite the draft to remove:
Instead, only summarize in your own words a range of independent, reliable, published sources that discuss the subject. If you have a conflict of interest (e.g. you are the subject, an employee, or a relative) or are being paid to edit, you must disclose this to comply with Wikipedia's Terms of Use. |
Comment: Lack of secondary sources, clearly still based on copyright source. Stuartyeates (talk) 23:55, 2 June 2026 (UTC)
Comment: In accordance with Wikipedia's Conflict of interest guideline, I disclose that I have a conflict of interest regarding the subject of this article. SeidelT (talk) 13:05, 11 May 2026 (UTC)
| Chemical Data Processing Toolkit | |
|---|---|
| Developer | Thomas Seidel |
| Release | 2023 |
| Stable release | 1.3.0
/ April 29, 2026 |
| Written in | C++ and Python |
| Operating system | Linux, macOS, and Microsoft Windows |
| Platform | Many |
| Available in | English |
| Type | Chemoinformatics |
| License | LGPL-2.1-or-later |
| Website | cdpkit |
The Chemical Data Processing Toolkit (CDPKit) is an open-source cheminformatics toolkit implemented in C++. CDPKit comprises a suite of command line and GUI tools as well as a programming library called the Chemical Data Processing Library (CDPL) which provides a modular implementation of basic functionality typically required by any higher-level software application in the field of cheminformatics. In addition to the CDPL C++ API, an equivalent Python-interfacing layer is provided that allows to harness all of CDPL’s functionality from Python code.[1]
CDPKit is developed at the Department of Pharmaceutical Sciences/University of Vienna on behalf of the Christian Doppler Laboratory for Molecular Informatics in the Biosciences (CD-Lab MIB) and receives funding from the Federal Ministry of Economy, Energy and Tourism of the Republic of Austria (BMWET), the Christian Doppler Forschungsgesellschaft, BASF SE and Boehringer Ingelheim RCV.
CDPKit seamlessly integrates with machine learning (ML) libraries like scikit-learn, PyTorch, and TensorFlow. The utility of CDPKit in the context of ML is showcased by several published scientific software tools that predict attributes of potential drug candidates such as lipophilicity and solubility,[2] biological activity,[3][4] and site of metabolism.[5] Apo2Ph4[6], PharmacoMatch[7] and CHA[8] represent further examples of computer-aided drug design software projects that rely on CDPKit's functionality.
Key Features (excerpt)
[edit]- Data structures for the representation and processing of molecules, chemical reactions and pharmacophores
- Routines for all typical cheminformatics pre-processing tasks (e.g. ring and aromaticity perception, stereochemistry processing, …)
- Methods for molecule and reaction substructure searching
- Readers/writers for various file formats (Mol, SDF, Rxn, RDF, Mol2, PDB, mmCIF, MMTF, SMILES, SMARTS, etc.) allowing the I/O of small molecule, macromolecular, reaction and pharmacophore data
- Molecule fragmentation algorithms such as RECAP,[9] and BRICS[10]
- Generation of molecule and pharmacophore fingerprints (e.g. ECFP[11])
- Large collection of implemented chemical structure descriptors
- 2D structure layout and rendering of molecules and reactions
- Gaussian shape-based molecule alignment and descriptor calculation[12]
- Pharmacophore generation, alignment and screening
- 3D structure and conformer generation[13]
- Prediction of a wide panel of physicochemical properties
- Test-suite compliant implementation of the MMFF94 force field
- C++ implementation follows best practices for robustness and computational efficiency
References
[edit]- ^ "Introduction — CDPKit 1.3.0 documentation". cdpkit.org. Retrieved 29 May 2026.
This article incorporates text from this source, which is available under the CC BY-SA 4.0 license. Also licensed under the GNU Free Documentation License (unversioned, with no invariant sections, front-cover texts, or back-cover texts.
- ^ Wieder, Oliver; Kuenemann, Mélaine; Seidel, Thomas; Meyer, Christophe; Bryant, Sharon D.; Langer, Thierry (2021). "Improved lipophilicity and aqueous solubility prediction with composite graph neural networks". Molecules. 26 (20): 6185. doi:10.3390/molecules26206185. PMC 8539502. PMID 34684766.
- ^ Fellinger, Christian; Seidel, Thomas; Merget, Benjamin; Schleifer, Klaus-Jürgen; Langer, Thierry (2025). "GRADE and X-GRADE: unveiling novel protein–ligand interaction fingerprints based on GRAIL scores". Journal of Chemical Information and Modeling. 65 (5): 2456–2475. doi:10.1021/acs.jcim.4c01902. PMC 11898076. PMID 39980202.
- ^ Kohlbacher, Stefan M.; Langer, Thierry; Seidel, Thomas (2021). "QPhAR: quantitative pharmacophore activity relationship: method and validation". Journal of Cheminformatics. 13 (1): 57. doi:10.1186/s13321-021-00537-9. PMC 8351372. PMID 34372940.
- ^ Jacob, Roxane Axel; Gaskin, Leo; Seidel, Thomas; Chen, Ya; Mazzolari, Angelica; Kirchmair, Johannes (2026). "FAME3R: an efficient, practical and reliable open-source tool for predicting phase 1 and phase 2 sites of metabolism". Journal of Cheminformatics. 18 (1): 37. doi:10.1186/s13321-026-01161-1. PMC 13011438. PMID 41691256.
- ^ Heider, Jörg; Kilian, Jonas; Garifulina, Aleksandra; Hering, Steffen; Langer, Thierry; Seidel, Thomas (2023). "Apo2ph4: a versatile workflow for the generation of receptor-based pharmacophore models for virtual screening". Journal of Chemical Information and Modeling. 63 (1): 101–110. doi:10.1021/acs.jcim.2c00814. PMC 9832483. PMID 36526584.
- ^ Rose, Daniel; Wieder, Oliver; Seidel, Thomas; Langer, Thierry (2025). "PharmacoMatch: efficient 3D pharmacophore screening via neural subgraph matching" (PDF). International Conference on Learning Representation. 2025: 85726–85749.
- ^ Wieder, Marcus; Garon, Arthur; Perricone, Ugo; Boresch, Stefan; Seidel, Thomas; Almerico, Anna Maria; Langer, Thierry (2017). "Common hits approach: combining pharmacophore modeling and molecular dynamics simulations". Journal of Chemical Information and Modeling. 57 (2): 365–385. doi:10.1021/acs.jcim.6b00674. PMID 28072524.
- ^ Lewell, Xiao Qing; Judd, Duncan B.; Watson, Stephen P.; Hann, Michael M. (1998). "RECAP - retrosynthetic combinatorial analysis procedure: a powerful new technique for identifying privileged molecular fragments with useful applications in combinatorial chemistry". Journal of Chemical Information and Computer Sciences. 38 (3): 511–522. doi:10.1021/ci970429i. PMID 9611787.
- ^ Degen, Jörg; Wegscheid-Gerlach, Christof; Zaliani, Andrea; Rarey, Matthias (2008). "On the art of compiling and using 'drug-like' chemical fragment spaces". ChemMedChem. 3 (10): 1503–150. doi:10.1002/cmdc.200800178. PMID 18792903.
- ^ Rogers, David; Hahn, Mathew (2010). "Extended-connectivity fingerprints". Journal of Chemical Information and Modeling. 50 (5): 742–754. doi:10.1021/ci100050t. PMID 20426451.
- ^ Grant, J. A.; Gallardo, M. A.; Pickup, B. T. (1996). "A fast method of molecular shape comparison: a simple application of a gaussian description of molecular shape". Journal of Computational Chemistry. 17 (14): 1653–1666. doi:10.1002/(SICI)1096-987X(19961115)17:14<1653::AID-JCC7>3.0.CO;2-K.
- ^ Seidel, Thomas; Permann, Christian; Wieder, Oliver; Kohlbacher, Stefan M.; Langer, Thierry (2023). "High-quality conformer generation with conforge: algorithm and performance assessment". Journal of Chemical Information and Modeling. 63 (17): 5549–5570. doi:10.1021/acs.jcim.3c00563. PMC 10498443. PMID 37624145.

- provide significant coverage: discuss the subject in detail, not just brief mentions or routine announcements;
- are reliable: from reputable outlets with editorial oversight;
- are independent: not connected to the subject, such as interviews, press releases, the subject's own website, or sponsored content.
Please add references that meet all three of these criteria. If none exist, the subject is not yet suitable for Wikipedia.