Draft:Digital chemistry
Draft article not currently submitted for review.
This is a draft Articles for creation (AfC) submission. It is not currently pending review. While there are no deadlines, abandoned drafts may be deleted after six months. To edit or make changes to this draft, simply click on the "Edit" tab at the top of the window. To be accepted, a draft should:
It is strongly discouraged to write about either yourself or your business or employer. If you do so, you must declare it. Where to get help
How to improve a draft
You can also browse Wikipedia:Featured articles and Wikipedia:Good articles to find examples of Wikipedia's best writing on topics similar to your proposed article. Improving your odds of a speedy review To improve your odds of a faster review, tag your draft with relevant WikiProject tags using the button below. This will let reviewers know a new draft has been submitted in their area of interest. For instance, if you wrote about a female astronomer, you would want to add the Biography, Astronomy, and Women scientists tags. Editor resources
Last edited by Citation bot (talk | contribs) 16 days ago. (Update) |

Digital chemistry refers to the emerging area of focus in the chemical sciences formed at the intersection of chemistry with chemical engineering, computer science, and digital technologies. Typically, digital chemistry makes use of the knowledge and technologies from these fields to enhance or expedite the discovery of new chemicals.[1][2][3][4] Key areas of application include the discovery of new pharmaceuticals,[5] agrochemicals,[6][7] and materials.[8][9]
In a 2024 article[1] in the academic journal Digital Discovery, chemist Stefan Bräse highlighted seven key digital technologies that can enable chemical discovery: chemical robotics, computational chemistry, high-throughput screening, cheminformatics, dedicated software tools and toolkits, machine learning/artificial intelligence, and quantum computing. Work that develops these technologies without explicitly attempting to discover new substances may also be considered digital chemistry. For example, the use of artificial intelligence to interpret and aggregate data from chemical literature,[10][11][12] or the creation and validation of new machine-readable formats for communicating synthetic instructions to chemical robots.[13][14] These technologies are strongly connected: robotics and simulation enable high-throughput screening, which allows for the generation of large amounts of experimental or simulated data that may be analyzed by cheminformaticians and used to create dedicated software or train machine learning models.
Chemical Robots and Self-Driving Laboratories
[edit]With regard to digital chemistry, the term "chemical robot" is usually taken to mean platforms for automated synthesis for use in at the same scale as a typical chemical laboratory (i.e. affording micrograms to grams of product in one experiment).[15][16] There is significant overlap with areas of chemical engineering, particularly reactor engineering. However, chemical engineering has industrial manufacture as its chief concern, and robotics designed for chemical discovery tend to operate on this smaller microgram-to-gram scale and with a higher degree of parallelization.
Benefits of Automation
[edit]There are a number of benefits to automated synthesis over manual synthesis, including:
- Safety - Removing the human experimenter from contact with hazardous substances or environmental conditions, reduces the risks associated with these. In the case of a laboratory accident, damage to a chemical robot, while potentially expensive, is preferrable to harm befalling a human experimenter. Additionally, automatic monitoring systems (based on e.g. analytical chemical detection or computer vision) can also be used in tandem with robotic synthesis platforms to shut down the platform or alert human experimenters if dangerous conditions are detected.[8][17][18] However, additional safety concerns may be introduced by using chemical robots, such as the possibility for the system being hacked for disruption or malicious misuse, highlighting the important of cybersecurity.[18] In the future, robotic systems may also be linked to databases of chemical hazard information and machine learning tools for predicting hazards of unknown substances or for suggesting alternative, safer routes to target chemicals.[19][20]
- Reproducibility - In theory, robotic systems are immune from human error, and so should be better suited to precisely replicating and recording actions, leading to more reproducible chemical procedures. Encoding chemical syntheses as machine-readable instructions has also been presented as allowing for a given procedure to be replicated easily and precisely on multiple machines.[8][21]
- Reliability - Automated synthesisers are generally capable of a high degree of accuracy and precision (particularly when manipulating liquids) and reduce the risk of human error. However, systematic errors may still be present.[17] The reduction in human error is of particular note for highly complex or sensitive processes with many opportunities for, or little toleration of, anthropogenic error.
- Speed - Robotic systems do not need to rest, and may be able to conduct certain manual operations more rapidly than a human experimenter (although the opposite may also be true).[17]
- Productivity - Automating synthetic procedures can allow human researchers to do different types of work (for example, interpreting analytical data or reading chemical literature) or monitor multiple robotic systems simultaneously, increasing individual output.[17]
- Laboratory Expenses - Opportunities for savings and efficiencies may take a variety of forms. For example, robots may require less space than a human researcher would to manually complete a task, or may lead to reduction in laboratory waste through increased accuracy of measurement.[17] This benefit is closely related to the potential for increased individual productivity, where human resources are deployed more efficiently.
- Parallelization - Closely related to the benefits of speed and productivity, parallelization offers the potential for dramatically increased experimental throughput.[22] This is particularly valuable when attempting to optimise synthetic conditions, conduct substrate screening to determine the scope of a reaction, or explore chemical process space, for example.
- Miniaturization - Reliable precision robotics enable reactions to be conducted on smaller scales. Choosing to work on smaller scale reduces material costs and may also decrease the severity of any hazards.[23]
Many of these benefits contribute directly to the suitability of automation for chemical discovery tasks. Automation can reduce the time and resource costs of performing chemical reactions, while also moving the burden of running manual procedures from the human experimenter to allow for a more efficient division of labor. Parallelization and miniaturization further allow for the implementation of high-throughput screening techniques to rapidly sample many experimental conditions and to generate datasets that can be used for data science tasks and the training of deep learning models.
Hardware
[edit]Automated chemical synthesis first gained popularity in the latter half of the 20th century, particularly following the development of a set of highly parallelizable robots for screening ammonia fixation catalysts by BASF in the 1940s,[24] and the Nobel prize-winning work of Robert Merrifield on the iterative solid phase synthesis of oligopeptides published in 1965.[25][26][27] Practical chemistry in the laboratory consists of a fundamental core of operations that are common to many different procedures. This makes task-specific automation quite attractive, particularly for labor-intensive or high-precision steps. Merrifield's batch reactor is a key example of this, as it was built for the iterative workflow of growing pepetide chains by adding amino acids in sequence, which required the same operations be performed in the same sequence for each step. Sample preparation and chemical analysis were also early beneficiaries of task-specific automation.[8][26]
While task-specific automation of these core operations or analytical techniques has had a huge impact on chemistry, synthesis usually requires multi-step workflows, requiring more complex platforms. The most straightforward approach for these is to simply link task-specific hardware together into a static combination of individual modules. Flow chemistry uses this approach to great effect, often making use of a continuous flow of material, but discretized batch systems also exist. For example, 3D printing has been used to generate bespoke multistage reactors for specific chemical syntheses.[28]
General-Purpose Chemical Robots
[edit]Linking task-specific hardware together into a workflow may be a very efficient way of automating a particular workflow, but lacks flexibility. General-purpose or universal robots that can interface with task-specific hardware in an adaptive manner more reminiscent of human researchers are increasingly becoming the target of chemical researchers. The flexibility required for this universality may be achieved in different ways: through use of robotic arms, gantry systems, or other modes of manipulation that can move experiments physically between specialized workstations,[16][29][30] or through selection valves linking task-specific hardware, through which liquid can be dynamically routed between any two locations.[31][32][33] Such systems offer the potential to conduct a vast number of potential chemical syntheses, but the flexibility comes at a cost; reconfigurable systems are less efficient and harder to parallelize than workflow-specific systems.[8][34]
In a 2024 review[8] of this area, researchers primarily from the University of Toronto suggested that robotic arms are the superior approach to general-purpose robots, as these can easily interface with equipment designed for human use already in a chemical laboratory, minimizing the reconfiguration or adaptation of task-specific proprietary equipement needed. Although the researchers highlight that perception of the robot's own environment and adaptive decision-making remain large obstacles to widespread adoption.
General-purpose robots can be further augmented by digital twins: virtual environments (i.e. physics-based simulations) corresponding to a physical chemical robot and its real environment, which can be used to allow the robot to learn manipulation skills and test workflows. This has particular value for robots which will interact with and manipulate a variety of real-world objects (such as transparent rigid laboratory glassware, hinged cupboard doors, or soft and deformable rubber tubing), as the robot can learn through trial-and-error without risk. Even the best physics simulators are not perfect recreations of reality, however, and so this process has limitations.[8][35][36]
Integrated Analysis
[edit]A notable example of task-specific hardware are machines for chemical analysis. Analytical modules can be integrated into chemical robots to afford the capacity to monitor and assess experiments as they run.[37] Depending on the level of autonomy, analytical interpretation may be simple (e.g. presence/absence or intensity of a particular signal) or may be sufficiently complex to require machine learning to process the output of the analytical measurement and use it in the experimental assessment and choice of future experimental conditions, if applicable.[1][8]
A number of terms can be used to describe the integration of analytical techniques. Note that the exact terms used to describe analytical integration may vary based on the field of chemistry and the specific use case. Some common examples include:[38]
- Inline measurement involves sensors integrated directly into the flow of material in a flow system or the reactor in a batch system.
- Online measurement involves automated sampling from the material flow or reactor and analysis of this sample rather than the live reactor or material flow.
- Offline measurements are those where the sensors are not integrated into the robot at all, and usually require human intervention to take samples and analyse.
- Atline measurement may be used to describe cases between online and offline measurements, where a human is required to take samples from the platform and transport these to nearby analysis equipment, but where this is conducted as close to real-time as possible.
Inline and online systems allow for far more frequent monitoring than atline, and may be used for real-time measurements of the chemical system and automatic adjustment of the robot to control the process. This is possible with atline, but is limited by the more time-consuming sampling and transport processes. Inline and online measurements have several other notable benefits over offline or atline measurements, including reproducibility, cost efficiency, and safety, all deriving from less human exposure. However, online (and, to a lesser extent, atline) measurements also provide more flexibility than inline measurements as external instruments are easier to adjust and maintain.[38]
Hardware Challenges
[edit]A key chemistry-specific challenge for any automated synthesis machine is that of handling hazardous substances with a wide variety of physical properties. Hazards may negatively affect human researchers in the sames space (e.g. toxicity) or the robot itself (e.g. corrosion). Chemical handling is typically more straightforward for liquids, where pumps can be calibrated to deliver specific volumes of liquid and thus easily control and record the amounts used in a particular experiment. Relatively low-cost, robust technologies (such as syringe or peristatltic pumps and positive displacement pipettes) have already been developed for these purposes, and largely unreactive materials can be chosen for tubing and pump construction. Highly viscous or reactive liquid chemicals may still pose difficulties.[8]
However, dispensing solid materials are more difficult as it calls for real-time measurement of the mass of material dispensed, which requires expensive specialized equipment. Some such systems have been developed, including ChemSpeed's POWDERDOSE ranges, which use an overhead gantry system with an inbuilt scale to measure the amount of powder dispensed, with up to milligram precision.[39][30] However, these systems still have limitations - for example, dispensing soft matter or powders with strong electrostatic repulsion between grains. A lower-cost solution to the problems presented by solid handling is to make up solutions of known concentration and handle these as any other liquid. This circuvents the issues around solid handling, but is labor intensive for human researchers, places constraints on the solvents that can be used during experiments, and may lead to degradation of the solid reagent, particularly if the stock solutions are prepared far in advance of their use.[8][40]
Another notable challenge to the widespread adoption of chemical robots is the large expense that buying or building such machinery can entail. Given the high precision required for many applications but the relatively small market of users, this presents a puzzle for commercialisation of the technology. Some companies have chosen to address this by centralising and scaling the technology in their own purpose-built facilities, to then work independently or with clients or collaborators on a project-by-project basis.[41][42] Others have chosen to address the issue by pursuing low-cost open-source designs for chemical robots. In theory, these can be constructed from widely-available materials or 3D-printed for costs in the range of hundreds, rather than thousands, of American dollars. The open-source aspect also allows for crowd-sourcing of development, testing, and customization of the hardware. However, there is still a requirement for technical know-how to set up the platforms and a need to grow user communities before these open-source designs are likely to receive widespread adoption.[8][43]
Self-Driving Laboratories
[edit]Self-driving laboratories, SDLs (also known as: self-directing laboratories, autonomous laboratories, or materials acceleration platforms, MAPs) take laboratory automation to a higher level by combining it with data-driven decision-making, usually implying the use of machine learning to determine future experiments using data already collected. This allows for an autonomous machine that is capable of closed-loop discovery (i.e. discovering new molecules, or at the extreme, new scientific laws, without requiring a human researcher's intervention).[8][44][45]
Autonomy
[edit]To be truly 'self-driving', the gold standard for an SDL is to be able to make discoveries independently of a human researcher. This requires both independent hardware and software. In determining the level of autonomy of an SDL there are a number of dimensions that can be evaluated, including: the specificity of the goal, the level of freedom in the machine agent's choice of search space, the machine agent's independence in performing and analysing experiments, the machine agent's independence in the choice of experiments, the efficiency of the exploration of the search space (in comparison to a brute-force or random search), and the machine agent's independence in the interpretation of results.[8][44][45]
Automated hardware can range from machines capable of a single automated task or experiment, through to those capable of full workflows comprising multiple tasks strung together, and ultimately to fully-automated machines that are able to string components together into reconfigurable workflows and are thus capable of a diverse range of experiments. The experiment planner can range from methods to select experiments within a human-chosen search space, starting with a single iteration of experiment selection and execution and moving to multiple iterations with more autonomous machines. True autonomy can be envisioned as the computer taking responsibility, not just for the experiment selection, but the search space selection as well. At present, a truly autonomous self-driving laboratory has not yet been realized.[8]
There is variation in the outputs that an autonomous platform may deliver - specifically in their human intelligibility. Where human interpretation is required 'after the fact' to evaluate what was learned, the level of autonomy of the system can be considered lower than a machine agent that is capable of transfering the understanding it generates to another entity. The ultimate realisation of autonomy, then, is not just an SDL that can discover new molecules, but one that can contribute to broader scientific understanding. For example, through the discovery of new insights or scientific laws (which may be broad or narrow in scope) and the explanation of these to a human.[44][45][46]
It is of note that some research has shown that collaborative human-robot teams have the potential to outperform purely human or purely robotic researchers.[47] The machine learning for experiment planning used has a limited level of autonomy. However, the level of autonomy does not necessarily correlate with performance and human-robot teamwork may still provide beneficial synergy with truly-autonomous SDLs.
Orchestration
[edit]SDLs, particularly those that make use of a general-purpose chemical robots to convey substances between different workstations, require software capable of a higher degree of orchestration than task-specific or single workflow hardware. Orchestration includes tasks such as: queuing jobs and assigning resources to them, logging actions taken, handling the data generated in an experiment. Queue management is particularly important as a mitigating factor for the loss in parallelization capabilities that general-purpose robots often suffer from, as multiple experiments could theoretically be run asynchronously using the same set of workstations. Orchestrators could also be used to coordinate multiple SDLs, potentially even across disparate physical locations by leaning on cloud technologies.[8] Several orchestrators, largely targeted at in-house use for the developers, have been created, including ChemOS,[48] AresOS,[49] and HELAO-async.[50] Although translation from manual operation to automated orchestration is still complicated by issues such as a lack of standardized APIs provided by instrument manufacturers, or limited exposure to the programming skills required in current chemical and materials science education.[8]
Protocol Management
[edit]SDL technology is in its early stages of development for chemistry, and there is as yet no strong consensus on standardisation. The representation of experimental protocols is a key concern. To prevent data fragmentation and allow interoperability and experimental reproducibility between various SDLs, it is important that experimental procedures can be recorded in a fashion that accurately records the procedure while not assuming a specific design of SDL. Specialized programming languages have been suggested as a solution, such as Chemical Description Language (χDL).[51] This is an XML-based language that is used to describe chemical procedures, with the added benefit of also being understandable by human researchers. The barrier to entry with χDL has been further softened by CLAIRify, an LLM that is used as an interface to allow human users to easily generate χDL descriptions of experimental procedures.[52] While standardization of protocols is a chemistry-specific problem, general good practice in data management is also important for SDLs (see Challenges in the Curation of Chemical Data). This should be further augmented by including integration of SDL-generated data with complimentary human-generated data, combining records of integrated analysis with human observations.[53]
"Cloud" Laboratories
[edit]*** SDLs and the democratisation of science potentially permits folk with bad juju to produce nasty stuff
Computational Simulation for Discovery
[edit]Computational simulation of individual molecules or dynamic ensembles can be used to determine the structure and behaviour of molecules at the atomic level.[1] Computational chemistry often has an explanatory role that is used to support observations of chemical properties or behaviour. For example, by determining the mechanism of a chemical transformation or calculating the expected results of an analytical test that can then be compared to experimentally measured analytics. However, it is the power of computational chemistry to calculate properties or simulate behaviour for chemicals not previously prepared or assayed for a particular property that is chiefly of interest to digital chemistry.
Level of theory
Structure and Property Calculation
[edit]Simulation accelerated by Machine Learning
[edit]*** Also want to talk about how automation of chemical simulation is democratizing it
Machine-Learned Interatomic Potentials (MLIPs)
[edit]Protein Structure Simulation and AlphaFold
[edit]Quantum Computing
[edit]Quantum computing may provide possible routes to perform faster and more efficient calculations required for the simulation of chemical species, particularly in cases where high accuracy is required or large numbers of particles are being simulated.[57]
The major advantage of quantum computing for digital chemistry is in the development of hardware capable of optimizing molecular structures, reaction paths, or experimental parameters that are beyond the capabilities of existing hardware.[1][57] For example, high-fidelity simulation techniques (ab initio calculations such as CCSD(T) or full CI methods) are only capable of simulating very simple molecules using current computing systems.[57]
High-Throughput Screening (HTS)
[edit]High-throughput screening (HTS) is the usually-automated testing of many possible candidate molecules or systems against an assay or measure of a desired property. HTS combines many of the benefits of automation (particularly speed, parallelization, and minaturization) to enable rapid performance, quantification, and iteration of chemical synthesis or assays.[58][59][60] HTS systems may be used for the optimization of synthetic conditions[61] and the discovery of new materials and reactions.[58][62] They may also be used in conducting rapid tests of a particular property of interest, notably in drug discovery where HTS is used to assay the activity of many candidate chemicals against a specific biological target. For example, the candidate molecule's inhibition of a protein's function might be measured.[63] HTS allows for gathering a large amount of data efficiently and this can then be used to uncover trends in the data, often employing machine learning to interpret data and make predictions based on it.
Screening may be conducted virtually, by the computational simulation of molecules, or experimentally, usually via chemical robots or more rarely human labor. Experimental HTS requires that the candidate molecules are physically available - either commercially, by isolation from a natural source, or through chemical synthesis. As such, virtual HTS may have advantages in terms of time and the cost of materials and human labor.[60][64]
Virtual screening hits require experimental validation to confirm the accuracy of the simulated properties, but the virtual HTS process can drastically shrink the number of candidate molecules that need to be prepared and manually assayed through sequential virtual and experimental screening. This sequential screening approach can also be applied to entirely virtual screens, where more expensive calculations or simulations are reserved only for those candidate chemicals that show promise in initial, relatively inexpensive calculations. Such an approach is called a "computational funnel".[60][65]
Cheminformatics and Data Curation
[edit]Cheminformatics is defined by the IUPAC as "the science of handling, indexing, archiving, searching, and evaluating information that is specific to chemical structures and is used in data mining, information retrieval, information extraction, and machine learning."[66] Within digitial chemistry, large databases of chemical compounds are frequently used to identify patterns (such as quantitative structure-activity relationships) and train machine learning models for descriptive or predictive tasks, such as directing the discovery of chemicals with specific properties. As such, the availability and quality of these data sources is a key concern and bottleneck.[1][67]
Databases and Software
[edit]A large variety of widely-used chemical databases are available, in many cases on a free-to-use basis.[68] Some types of chemical data stored and examples of databases include:
- Databases of organic molecules and their biological properties for medicinal chemistry applications, such as ChEMBL,[69] PubChem,[70] and ZINC[71]
- Databases of materials and their structural or material properties, such as the Materials Project[72]
- Databases of spectroscopic or crystallographic data, such as the NIST Chemical Webbook,[73] AIST's Spectral Database for Organic Compounds (SDBS),[74] and the Cambridge Structural Database[75]
- Databases of synthetic transformations, such as the Organic Chemistry Portal[76]
- Multi-dimensional databases that combine substance, property, reaction, and bibliographic data, such as Elsevier's Reaxys[77] or the American Chemical Society's SciFinder.[78]
There are many examples of software aimed at the chemical sciences, particularly to enable analysis, however, for digital chemistry, a chief concern is in the collection and sharing of structural, synthetic, or chemical property data. This can be achieved in multiple ways, for example: by uploading data to a pre-existing database, by explicitly generating or curating a new database, by using an electronic lab notebook to digitally capture first-hand experimental data and observations or by standardising the reporting of a given type of chemical data even when this is not centrally collected.[1]
Software that allows the manipulation of this data is also vital to allow for trends to be recognised and predictions to be made. Often this involves using chemical data to train a machine learning model and so integration with coding languages is key. For example, a number of packages have been developed for the Python programming language that enable the machine-readable encoding or interpretation of molecular representations,[79][80] comparison of structural similarity,[79] calculation of structural or material properties,[81] or generation of chemistry-specific visualizations,[79][82] among many other functions.[83][84][85]
Challenges in the Curation of Chemical Data
[edit]Data Availability and Standardization
[edit]Science often suffers from data fragmentation, the disparate spread of available data across formats, locations, and behind paywalls. Data fragmentation results in a landscape where data is difficult to access and data from different sources may be incompatible. While many chemistry-focused datasets exist, these often have a high dimensionality but a relatively low sample size, making it difficult to train machine learning models that can make the general predictions or inferences required for true scientific understanding. Additionally, research institutions often do not promote the curation of data when considering career rewards.[86]
The variety of data sources also leads to difficulties combining multiple datasets. Even where similar data is available from two different sources, a lack of standardization in what and how data is stored can prevent these data from being readily combined. Different funding agencies and journal groups rarely coalesce around the same standard for reporting data, although there are some notable exceptions, such as the Cambridge Structural Database (CSD) as a repository for small molecule crystallographic data.[86]
Historically, the recording of chemical data has also been optimised for human interpretation, rather than AI. This also presents challenges for cheminformaticians or digital chemists in scraping this data to allow for machine learning training. In particular, a large amount of chemical data is stored in PDFs available behind journal paywalls. The huge wealth of scientific data contained solely in this format requires that machine learning (for example, natural language processing, large language models) be employed to parse it all, as the scale makes manual curation of this data infeasible. Machine learning also faces challenges, as key data in chemical publications is often contained within images, captions, or tables, which can be challenging to parse.[86][87][88]
The extraction of experimental procedures is of particular interest. Many synthetic procedures follow a similar, formulaic pattern, but there is no true single standardized way of reporting each operation. Novel synthetic operations are also always being invented, notably in the subdiscipline of chemical automation. These often do not have a large existing corpus of similar transformations to guide the style they are written in. Furthermore, procedures can often make implicit assumptions about the techniques used that must be reverse-engineered and explicitly presented to the model being trained.[86][87] A number of researchers have proposed consistent machine-readable representations for synthesis, however these have yet to be adopted as a standard by funding or data collection agencies, and are not commonly used by chemists working outside digital chemistry.[13]
Lack of Negative Data
[edit]The trial-and-error process demanded by chemical research on the path to a final functional molecule or material results in many failed experiments. These unsuccessful conditions are rarely reported, as chemical literature is focused on reporting research with positive results. However, a number of studies have shown that the ability of machine learching models to predict synthetic success or likely synthetic conditions is markedly improved by the inclusion of negative data.[89][90]
There is a suggestion that very small positive datasets augumented with larger amounts of negative data can still be used to train high quality models. Using a large language model (fine-tuned through reinforcement learning), a team from IBM was able to show that an asymmetric dataset of as few as 20 positive datapoints and a negative dataset at least 40 times larger still gave high quality performance when predicting reaction success. This reinforces the importance of negative data in digitally-enabled chemical discovery.[90]
Whilst recording and inclusion of negative data is a clear solution to this challenge, some machine learning techniques may also provide ways to mitigate the lack of negative data. Positive-unlabelled learning is a family of semi-supervised machine learning techniques that don't require negative data to train, and have already begun showing promise for chemical problems.[91]
Just as positive results in chemistry can prove difficult to reproduce or may be a false positive, it is important to note that reported negative results may also suffer from the same issues (i.e. irreproducibility or false negatives). Whether a result is positive or negative is also sensitive to context. A reaction that produces a new chemical, but one that is not of interest to the current study may be classified as 'unsuccessful'. However, a future study with different aims and interests may view the previously 'unsuccessful' result as a 'success' when measured against their own criteria. How these cases can best be reported has not yet been addressed.
Anthropogenic Bias
[edit]Anthropogenic bias refers to a prejudice or preference in scientific data that arises through the cognitive biases, heuristics, and social influences of human researchers. This leads to unrepresentative datasets that may hinder the predictive or generative power of machine learning models trained on them. It can be highly challenging to identify and quantify such biases, and there is limited work specific to chemistry in this area.[92]
In a seminal study from 2019, researchers from Haverford College (Pennsylvania, USA) were able to show that the success of a hybrid perovskite synthesis (defined as providing crystals of suitable quality for single-crystal x-ray diffraction characterisation) had little to no correlation with the popularity of the amine used in all hybrid perovskites recorded in the Cambridge Structural Database (CSD). This was despite 17% of the amines present in hybrid perovskite materials in the CSD being responsible for 79% of all hybrid perovskite entries. Hybrid perovskites were chosen for the experiment as they are well-explored, functionally-interesting, provide a sizeable dataset, have well-understood and robust syntheses, and use amines that were all commercially-available in similar quantities and for similar prices. Taken together, this implies that the popularity of an amine in perovskite materials recorded in the CSD is the result of anthropogenic bias and not a consequence of inate chemical reasons.[92]
This study goes on to evidence several additional claims using the same chemical system as a basis:[92]
- Machine learning models to predict reaction success are improved by training on data without an anthropogenic bias.
- The anthropogenic data implicitly obscures the contribution of certain variables through the biased selection of reaction conditions.
- Models trained on the anthropogenic data were more likely to incorrectly predict a reaction would fail, likely due to the presence of human loss aversion bias in their training data.
- The anthropogenic dataset was less able to predict successful reaction conditions for a diverse range of amines.
Given that anthropogenic bias is likely more widespread in chemistry and may have negative effects on analogue and digital chemical discovery for the reasons detailed, combatting this is of interest to researchers in digital chemistry. More sophisticated strategies for choosing experiments, new databases, and high-throughput experimentation to screen a greater number of reaction conditions have all been suggested as possible mitigation for the issues around anthropogenic bias.[93]
Machine Learning and Artificial Intelligence (AI) in Chemistry
[edit]Machine learning and AI have already attracted great interest from chemists, and have been applied to the prediction of molecular properties, optimization of chemical reactions, and design of new molecules and materials.[1]
Predictive Models
[edit]x
Predictive models run the gamut from simple linear regression through to deep learning methods for the prediction of chemical properties or profitable synthetic conditions.
Prediction as a proxy for true simulation/calculation
Trajectory-Based Methods
- Monte Carlo Search
- Simulated Annealing
- Gradient Descent
- Design of Experiments
Population-Based Methods
- Genetic Algorithms
- Covariance Matrix Adaptation Evolution Strategy (CMA-ES)
- Particle Swarm Optimization
Surrogate Models
- Bayesian Optimisation (and Gaussian Process Regression)
- Random Forests and Tree-Based Ensembles
- Surrogate-Based Global Optimization (such as Radial Basis Functions, Kriging, Support Vector Regression)
Generative Models
[edit]x[94]
VAE/Self-organising Maps, GAN, Diffusion Models, RL
Foundation Models
[edit]The term "foundation model" indicates a broad class of deep neural networks that are trained on data that is broad in scope and vast in scale, producing general-purpose models that may be adapted into more specialized models for different downstream tasks. Most current models that fulfill this definition are large language models (LLMs), which are designed principally for tasks related to natural language processing (NLP) or language generation.[95][96] These are directly applicable to chemistry, as the vast majority of chemical data is stored as written text in journal articles, particularly historic data before the advent of modern online databases.[97] The text-to-image model, Stable Diffusion, may also be considered a foundation model, but is less broadly applied to chemistry.[98]

Large Language Models (LLMs)
[edit]The dominant architecture underpinning major LLMs is the transformer, and this architecture is widely explored for chemical applications.[97]
*** Brief description of the transfomer - try and make relevant to the illustration
*** Key uses in chemistry, maybe picking out some particular types of LLM that have been used effectively
AI Scientists
[edit]AI Scientists, Chemical oracles
Interpretability of Machine Learning Models
[edit]Where a neural network (or other machine learning model of similar or greater complexity) is used to find a pattern in a chemical dataset, it would be beneficial if that pattern could be interrogated and transformed into a form that human chemists can understand. However, many deep learning models often function as a "black box", where***
While patterns in chemical properties or behaviour may be identified by a particular AI model, they are not easily generalised or used to inform future experiments beyond those that the model itself might design as part of an iterative discovery process.[99]
To combat this...**
Digital Tools for Chemical Synthesis
[edit]Optimization of Synthetic Conditions
[edit]One of the most well-established areas of digital chemistry is in the optimization of chemical synthesis.[100][101]
Digitization of Chemical Synthesis
[edit]Computer-Aided Synthesis Planning (CASP)
[edit]Digital Strategies for Chemical Discovery
[edit]Digital discovery strategies fall into two main camps: data-driven discovery and curiousity-driven discovery.
Exploration vs. Exploitation
[edit]Data-Driven Discovery
[edit]Inverse design
Synthesizability
Curiosity-Driven Discovery
[edit]Random or Quasi-Random Methods
- Brute-Force Search
- Random Search
- Grid Search
- Orthogonal Sampling - LHS, Sobol, etc.
ADD TO AUTOMATED SYNTHESIS: Chemical Cobotics (Collaborative Robotics)
[edit]Whilst there is much focus in this research area on autonomous discovery, where the human investigator is out of the iterative loop of experiment, analysis, and selection of subsequent experimental conditions, some researchers have explicitly considered the potential of synergistic approaches, where human and robotic or AI "co-investigators" collaborate on a particular discovery aim.
Cobots
NEW PAGE: Data-Driven Discovery in Chemistry
[edit]Inverse Design
[edit]*** Exscientia -> First digitally designed drug candidate[104][105][5][106] Exscientia now a part of Recursion[107]
Synthesizability
[edit]Synthesizability refers to the challenge of ensuring that candidate materials or molecules proposed by a generative algorithm during a data-driven discovery process are capable of being made in the laboratory (i.e. that they are "synthetically accessible"). Without explicit constraints, generative algorithms will generate chemical structures that may have a high likelihood of having desired material properties, but which cannot be physically manufactured.[108][109]
Synthesizability cannot be estimated by considering the structure of a proposed chemical alone, but must be assessed with respect to a specific proposed route or recipe for that chemical. The conditions of synthesis are important as they dictate the path across the potential energy surface that is taken by the reagents during a given transformation (i.e. the potential intermediates and transition states that are encountered en route to the final structure). Practical limitations (such as material cost, time, available equipment, or route safety) may also be considered as part of an assessment of synthesizability.[108][109]
Assessing Synthesisability in Solid State Chemistry
[edit]For crystalline inorganic species, an estimate of synthesizability can be gained by computing the thermodynamic stability of the crystal, where higher relative stability is equated to a higher likelihood of formation. This allows for comparison of relative stability between phases with the same composition at given values of pressure and temperature. However, this purely thermodynamic picture is complicated by the range of equilibrium and non-equilibrium growth techniques that may be used. A variety of factors, including precursor identity and purity, reaction kinetics, phase purity, the use of solvents or fluxes, and the applied thermal/pressure profile can all influence the observed product of a chemical reaction. In particular, kinetic factors may influence crystal nucleation and growth of phases other than the most thermodynamically stable.[108]
A number of methods have been suggested to aid this complicated assessment of synthetic accessibility for sold-state chemistry. They include:
- The calculation of thermodynamic stability outlined above can be enhanced by the inclusion of additional variables (such as basing calculations on Gibbs free energy rather than internal energy) to extend their utility for amorphous, metastable, or disordered crystalline structures.[110]
- Chemical heuristic models (such as Pauling electronegativity, the Hume-Rothery rules, or electron counting) can be used to filter chemically implausible candidates from large sets of generated molecules.[111][112]
- Machine learning models may be used in a similar manner to chemical heuristics,[9] potentially identifying patterns in sparse datasets or those unidentified by human chemists.[108]
- Measures of crystal likeness (CL) provide a numerical measure of a given structure's similarity to already-synthesised materials, which can be used to estimate their synthetic feasibility. Note that crystal likeness is not explicitly a measure of synthesizability as defined above, because it is divorced from any specific synthetic route.[108]
Assessing Synthesisability in Organic Chemistry
[edit]For molecular organic species, thermodynamic qualities cannot be readily used as a proxy for synthesizability. Instead, consideration of synthesizability can be incorporated at various points in a generative AI workflow:[109]
- AI training data may be curated for synthetic accessibility. Several of these already exist, such as the MOSES subset of the ZINC database,[113] or the "make-on-demand" libraries of chemical supply companies,[114] and such datasets using only easy-to-synthesize molecules have shown potential in testing.[115][116] Alternatively, datasets can be labelled with scores indicating the synthetic accessibility of each datum and conditionally train the generative model on this value.[117]
- AI generation of molecules can be constrained to make use of only commericially available building blocks and predefined reaction templates.
- During the iterative process of optimizing the generated molcules towards a particular chemical property, synthesizability can be included as an additional optimization objective (i.e. multi-objective optimization). For the highest possible confidence in the synthesizability prediction, the best option is retrosynthetic route prediction (i.e. computationally generating possible recipes for the synthesis of the AI-generated candidate molecule). However, this is computationally expensive, and so heuristic-based proxy scores are commonly used as an alternative.
See also
[edit]References
[edit]- 1 2 3 4 5 6 7 8 9 Bräse, Stefan (2024-10-09). "Digital chemistry: navigating the confluence of computation and experimentation – definition, status quo, and future perspective". Digital Discovery. 3 (10): 1923–1932. doi:10.1039/D4DD00130C. ISSN 2635-098X.
- ↑ "Digital Discovery". Royal Society of Chemistry. Archived from the original on 2026-01-13. Retrieved 2026-02-17.
- ↑ "Digital Chemistry | Study | Imperial College London". www.imperial.ac.uk. Retrieved 2024-03-13.
- ↑ "MSC Digital Chemistry | University of Southampton". www.southampton.ac.uk. Retrieved 2024-03-13.
- 1 2 "Exscientia claims world first as AI-created drug enters clinic". pharmaphorum. 30 January 2020. Retrieved 2024-03-13.
- ↑ Agarwal, Aditi; Shankar, Ganesh; Thukral, Sujata; Gupta, Manoj Kumar (2025). "AI: The Catalyst in Agrochemical Discovery" (PDF). www.infosys.com. Archived from the original (PDF) on 2025-11-28. Retrieved 2026-02-17.
- ↑ Djoumbou-Feunang, Yannick; Wilmot, Jeremy; Kinney, John; Chanda, Pritam; Yu, Pulan; Sader, Avery; Sharifi, Max; Smith, Scott; Ou, Junjun; Hu, Jie; Shipp, Elizabeth; Tomandl, Dirk; Kumpatla, Siva P. (2023-11-29). "Cheminformatics and artificial intelligence for accelerating agrochemical discovery". Frontiers in Chemistry. 11 1292027. Bibcode:2023FrCh...1192027D. doi:10.3389/fchem.2023.1292027. ISSN 2296-2646. PMC 10716421. PMID 38093816.
- 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 Tom, Gary; Schmid, Stefan P.; Baird, Sterling G.; Cao, Yang; Darvish, Kourosh; Hao, Han; Lo, Stanley; Pablo-García, Sergio; Rajaonson, Ella M.; Skreta, Marta; Yoshikawa, Naruki; Corapi, Samantha; Akkoc, Gun Deniz; Strieth-Kalthoff, Felix; Seifrid, Martin (2024-08-28). "Self-Driving Laboratories for Chemistry and Materials Science". Chemical Reviews. 124 (16): 9633–9732. doi:10.1021/acs.chemrev.4c00055. ISSN 0009-2665. PMC 11363023. PMID 39137296.
- 1 2 Merchant, Amil; Batzner, Simon; Schoenholz, Samuel S.; Aykol, Muratahan; Cheon, Gowoon; Cubuk, Ekin Dogus (2023-12-07). "Scaling deep learning for materials discovery". Nature. 624 (7990): 80–85. Bibcode:2023Natur.624...80M. doi:10.1038/s41586-023-06735-9. ISSN 1476-4687. PMC 10700131. PMID 38030720.
- ↑ Mavračić, Juraj; Court, Callum J.; Isazawa, Taketomo; Elliott, Stephen R.; Cole, Jacqueline M. (2021-09-27). "ChemDataExtractor 2.0: Autopopulated Ontologies for Materials Science". Journal of Chemical Information and Modeling. 61 (9): 4280–4289. doi:10.1021/acs.jcim.1c00446. ISSN 1549-9596. PMID 34529432.
- ↑ Isazawa, Taketomo; Cole, Jacqueline M. (2022-03-14). "Single Model for Organic and Inorganic Chemical Named Entity Recognition in ChemDataExtractor". Journal of Chemical Information and Modeling. 62 (5): 1207–1213. doi:10.1021/acs.jcim.1c01199. ISSN 1549-9596. PMC 9049593. PMID 35199519.
- ↑ "ChemDataExtractor v2". chemdataextractor2.org. Retrieved 2026-02-17.
- 1 2 Mehr, S. Hessam M.; Craven, Matthew; Leonov, Artem I.; Keenan, Graham; Cronin, Leroy (2020-10-02). "A universal system for digitization and automatic execution of the chemical synthesis literature". Science. 370 (6512): 101–108. Bibcode:2020Sci...370..101M. doi:10.1126/science.abc2986. PMID 33004517.
- ↑ Roch, Loïc M.; Häse, Florian; Kreisbeck, Christoph; Tamayo-Mendoza, Teresa; Yunker, Lars P. E.; Hein, Jason E.; Aspuru-Guzik, Alán (2018-06-20). "ChemOS: Orchestrating autonomous experimentation". Science Robotics. 3 (19) eaat5559. doi:10.1126/scirobotics.aat5559. PMID 33141686.
- ↑ Crow, James Mitchell. "The robots revolutionising chemistry". Chemistry World. Retrieved 2026-02-09.
- 1 2 Burger, Benjamin; Maffettone, Phillip M.; Gusev, Vladimir V.; Aitchison, Catherine M.; Bai, Yang; Wang, Xiaoyan; Li, Xiaobo; Alston, Ben M.; Li, Buyi; Clowes, Rob; Rankin, Nicola; Harris, Brandon; Sprick, Reiner Sebastian; Cooper, Andrew I. (2020-07-09). "A mobile robotic chemist". Nature. 583 (7815): 237–241. Bibcode:2020Natur.583..237B. doi:10.1038/s41586-020-2442-2. ISSN 0028-0836. PMID 32641813.
- 1 2 3 4 5 "5 benefits of automated chemistry systems". Syrris. Retrieved 2026-04-14.
- 1 2 Leong, Shi Xuan; Griesbach, Caleb E.; Zhang, Rui; Darvish, Kourosh; Zhao, Yuchi; Mandal, Abhijoy; Zou, Yunheng; Hao, Han; Bernales, Varinia; Aspuru-Guzik, Alán (October 2025). "Steering towards safe self-driving laboratories". Nature Reviews Chemistry. 9 (10): 707–722. doi:10.1038/s41570-025-00747-x. ISSN 2397-3358. PMID 40826175.
- ↑ Harada, Takaaki; Hayashi, Rumiko; Tomita, Kengo (2025-07-28). "Direct Prediction of Chemical Hazards in GHS Classification Using SMILES Representation for Health and Safety Applications". ACS Chemical Health & Safety. 32 (4): 440–448. doi:10.1021/acs.chas.5c00035.
- ↑ Johnes, Austin; Khan, Faisal; Hasan, M. M. Faruque (2026-01-01). "Safety and risks informed economic process synthesis". Computers & Chemical Engineering. 204 109410. doi:10.1016/j.compchemeng.2025.109410. ISSN 0098-1354.
- ↑ Rauschen, Robert; Guy, Mason; Hein, Jason E.; Cronin, Leroy (April 2024). "Universal chemical programming language for robotic synthesis repeatability". Nature Synthesis. 3 (4): 488–496. Bibcode:2024NatSy...3..488R. doi:10.1038/s44160-023-00473-6. ISSN 2731-0582.
- ↑ Furka, Árpád (2025-10-01). "The parallel and combinatorial synthesis and screening in drug discovery". Structural Chemistry. 36 (5): 1925–1929. Bibcode:2025StrCh..36.1925F. doi:10.1007/s11224-025-02595-3. ISSN 1572-9001.
- ↑ Gesmundo, Nathan; Dykstra, Kevin; Douthwaite, James L.; Kao, Yu-Ting; Zhao, Ruheng; Mahjour, Babak; Ferguson, Ron; Dreher, Spencer; Sauvagnat, Bérengère; Saurí, Josep; Cernak, Tim (November 2023). "Miniaturization of popular reactions from the medicinal chemists' toolbox for ultrahigh-throughput experimentation". Nature Synthesis. 2 (11): 1082–1091. Bibcode:2023NatSy...2.1082G. doi:10.1038/s44160-023-00351-1. ISSN 2731-0582.
- ↑ Topham, Susan A. (1985), "The History of the Catalytic Synthesis of Ammonia", in Anderson, John R.; Boudart, Michel (eds.), Catalysis: Science and Technology, Berlin, Heidelberg: Springer, pp. 1–50, doi:10.1007/978-3-642-93281-6_1, ISBN 978-3-642-93281-6, retrieved 2026-04-21
- ↑ Merrifield, R. B. (July 1963). "Solid Phase Peptide Synthesis. I. The Synthesis of a Tetrapeptide". Journal of the American Chemical Society. 85 (14): 2149–2154. Bibcode:1963JAChS..85.2149M. doi:10.1021/ja00897a025. ISSN 0002-7863.
- 1 2 Merrifield, R. B. (1965-10-08). "Automated Synthesis of Peptides: Solid-phase peptide synthesis, a simple and rapid synthetic method, has now been automated". Science. 150 (3693): 178–185. doi:10.1126/science.150.3693.178. ISSN 0036-8075. PMID 5319951.
- ↑ "Nobel Prize in Chemistry 1984". NobelPrize.org. Retrieved 2026-02-09.
- ↑ Hou, Wenduan; Bubliauskas, Andrius; Kitson, Philip J.; Francoia, Jean-Patrick; Powell-Davies, Henry; Gutierrez, Juan Manuel Parrilla; Frei, Przemyslaw; Manzano, J. Sebastián; Cronin, Leroy (2021-02-24). "Automatic Generation of 3D-Printed Reactionware for Chemical Synthesis Digitization using ChemSCAD". ACS Central Science. 7 (2): 212–218. doi:10.1021/acscentsci.0c01354. ISSN 2374-7943. PMC 7908023. PMID 33655058.
- ↑ Godfrey, Alexander G.; Masquelin, Thierry; Hemmerle, Horst (2013-09-01). "A remote-controlled adaptive medchem lab: an innovative approach to enable drug discovery in the 21st Century". Drug Discovery Today. 18 (17): 795–802. doi:10.1016/j.drudis.2013.03.001. ISSN 1359-6446. PMID 23523957.
- 1 2 "Automated Gravimetric Solid Dispensing | Chemspeed FLEX POWDERDOSE". www.chemspeed.com. Retrieved 2026-04-21.
- ↑ Steiner, Sebastian; Wolf, Jakob; Glatzel, Stefan; Andreou, Anna; Granda, Jarosław M.; Keenan, Graham; Hinkley, Trevor; Aragon-Camarasa, Gerardo; Kitson, Philip J.; Angelone, Davide; Cronin, Leroy (2019-01-11). "Organic synthesis in a modular robotic system driven by a chemical programming language". Science. 363 (6423) eaav2211. Bibcode:2019Sci...363v2211S. doi:10.1126/science.aav2211. ISSN 0036-8075. PMID 30498165.
- 1 2 Porwol, Luzian; Kowalski, Daniel J.; Henson, Alon; Long, De-Liang; Bell, Nicola L.; Cronin, Leroy (2020). "An Autonomous Chemical Robot Discovers the Rules of Inorganic Coordination Chemistry without Prior Knowledge". Angewandte Chemie International Edition. 59 (28): 11256–11261. Bibcode:2020ACIE...5911256P. doi:10.1002/anie.202000329. ISSN 1521-3773. PMC 7384156. PMID 32419277.
- ↑ Chatterjee, Sourav; Guidi, Mara; Seeberger, Peter H.; Gilmore, Kerry (2020-03-18). "Automated radial synthesis of organic molecules". Nature. 579 (7799): 379–384. Bibcode:2020Natur.579..379C. doi:10.1038/s41586-020-2083-5. ISSN 1476-4687. PMID 32188949.
- ↑ Mullin, Rick (2025-01-21). "The lab of the future is now". Chemical & Engineering News. Retrieved 2026-04-21.
- ↑ Klami, Arto; Damoulas, Theodoros; Engkvist, Ola; Rinke, Patrick; Kaski, Samuel (2022-08-08), Virtual Laboratories: Transforming research with AI, doi:10.36227/techrxiv.20412540.v1, retrieved 2026-04-21
- ↑ Beeler, Chris; Subramanian, Sriram Ganapathi; Sprague, Kyle; Baula, Mark; Chatti, Nouha; Dawit, Amanuel; Li, Xinkai; Paquin, Nicholas; Shahen, Mitchell; Yang, Zihan; Bellinger, Colin; Crowley, Mark; Tamblyn, Isaac (2024-04-17). "ChemGymRL: A customizable interactive framework for reinforcement learning for digital chemistry". Digital Discovery. 3 (4): 742–758. doi:10.1039/D3DD00183K. ISSN 2635-098X.
- ↑ Ashikari, Yosuke; Tamaki, Takashi; Tomite, Kyosuke; Yonekura, Yuya; Nagaki, Aiichiro (2025-09-30). "Real-time inline-IR-analysis via linear-combination strategy and machine learning for automated reaction optimization". Communications Chemistry. 8 (1) 287. Bibcode:2025CmChe...8..287A. doi:10.1038/s42004-025-01676-y. ISSN 2399-3669. PMC 12484558. PMID 41028417.
- 1 2 "Optimizing process control: Inline, Online, Atline, and Offline". Advanced Microfluidics. Retrieved 2026-02-09.
- ↑ "Crystal Powderdose | Benchtop Gravimetric Solid Dispensing". www.chemspeed.com. Retrieved 2026-04-21.
- 1 2 Kowalski, Daniel J.; MacGregor, Catriona M.; Long, De-Liang; Bell, Nicola L.; Cronin, Leroy (2023-02-01). "Automated Library Generation and Serendipity Quantification Enables Diverse Discovery in Coordination Chemistry". Journal of the American Chemical Society. 145 (4): 2332–2341. Bibcode:2023JAChS.145.2332K. doi:10.1021/jacs.2c11066. ISSN 0002-7863. PMC 9896557. PMID 36649125.
- ↑ "New Chemifarm Revolutionizes Molecular Design and Manufacture". chemify.reportablenews.com. Retrieved 2026-04-19.
- ↑ Philippidis, Alex (2025-07-17). "Chemify "Farm" Cultivates Molecules by Marrying Chemistry, AI, and Robotics". GEN - Genetic Engineering and Biotechnology News. Retrieved 2026-04-19.
- ↑ Lo, Stanley; Baird, Sterling G.; Schrier, Joshua; Blaiszik, Ben; Carson, Nessa; Foster, Ian; Aguilar-Granda, Andrés; Kalinin, Sergei V.; Maruyama, Benji; Politi, Maria; Tran, Helen; Sparks, Taylor D.; Aspuru-Guzik, Alán (2024-05-15). "Review of low-cost self-driving laboratories in chemistry and materials science: the "frugal twin" concept". Digital Discovery. 3 (5): 842–868. doi:10.1039/D3DD00223C. ISSN 2635-098X.
- 1 2 3 Coley, Connor W.; Eyke, Natalie S.; Jensen, Klavs F. (2020-12-14). "Autonomous Discovery in the Chemical Sciences Part I: Progress". Angewandte Chemie International Edition. 59 (51): 22858–22893. arXiv:2003.13754. Bibcode:2020ACIE...5922858C. doi:10.1002/anie.201909987. ISSN 1433-7851. PMID 31553511.
- 1 2 3 Coley, Connor W.; Eyke, Natalie S.; Jensen, Klavs F. (2020-06-11). "Autonomous Discovery in the Chemical Sciences Part II: Outlook". Angewandte Chemie International Edition. 59 (52): 23414–23436. arXiv:2003.13755. Bibcode:2020ACIE...5923414C. doi:10.1002/anie.201909989. ISSN 1433-7851. PMID 31553509. Archived from the original on 2021-07-18.
- ↑ Krenn, Mario; Pollice, Robert; Guo, Si Yue; Aldeghi, Matteo; Cervera-Lierta, Alba; Friederich, Pascal; dos Passos Gomes, Gabriel; Häse, Florian; Jinich, Adrian; Nigam, AkshatKumar; Yao, Zhenpeng; Aspuru-Guzik, Alán (2022-10-11). "On scientific understanding with artificial intelligence". Nature Reviews Physics. 4 (12): 761–769. arXiv:2204.01467. Bibcode:2022NatRP...4..761K. doi:10.1038/s42254-022-00518-3. ISSN 2522-5820. PMC 9552145. PMID 36247217.
- ↑ Duros, Vasilios; Grizou, Jonathan; Sharma, Abhishek; Mehr, S. Hessam M.; Bubliauskas, Andrius; Frei, Przemysław; Miras, Haralampos N.; Cronin, Leroy (2019-06-24). "Intuition-Enabled Machine Learning Beats the Competition When Joint Human-Robot Teams Perform Inorganic Chemical Experiments". Journal of Chemical Information and Modeling. 59 (6): 2664–2671. doi:10.1021/acs.jcim.9b00304. ISSN 1549-9596. PMC 6593393. PMID 31025861.
- ↑ Roch, Loïc M.; Häse, Florian; Kreisbeck, Christoph; Tamayo-Mendoza, Teresa; Yunker, Lars P. E.; Hein, Jason E.; Aspuru-Guzik, Alán (2020-04-16). "ChemOS: An orchestration software to democratize autonomous discovery". PLOS ONE. 15 (4) e0229862. Bibcode:2020PLoSO..1529862R. doi:10.1371/journal.pone.0229862. ISSN 1932-6203. PMC 7161969. PMID 32298284.
- ↑ Sloan, Arthur W. N.; Waelder, Robert W.; Smith, Morgen L.; Kleiner, Nicholas; Babeckis, Arnas; Wheeler, Jason; Hooper, Daylond; Maruyama, Benji (2026). "ARES OS 2.0: An Orchestration Software Suite for Autonomous Experimentation Systems and Self-Driving Labs". arXiv:2604.03440 [cs.CE].
- ↑ Guevarra, Dan; Kan, Kevin; Lai, Yungchieh; Jones, Ryan J. R.; Zhou, Lan; Donnelly, Phillip; Richter, Matthias; Stein, Helge S.; Gregoire, John M. (2023-12-04). "Orchestrating nimble experiments across interconnected labs". Digital Discovery. 2 (6): 1806–1812. doi:10.1039/D3DD00166K. ISSN 2635-098X.
- ↑ Mehr, S. Hessam M.; Craven, Matthew; Leonov, Artem I.; Keenan, Graham; Cronin, Leroy (2020-10-02). "A universal system for digitization and automatic execution of the chemical synthesis literature". Science. 370 (6512): 101–108. Bibcode:2020Sci...370..101M. doi:10.1126/science.abc2986. PMID 33004517.
- ↑ Yoshikawa, Naruki; Skreta, Marta; Darvish, Kourosh; Arellano-Rubach, Sebastian; Ji, Zhi; Bjørn Kristensen, Lasse; Li, Andrew Zou; Zhao, Yuchi; Xu, Haoping; Kuramshin, Artur; Aspuru-Guzik, Alán; Shkurti, Florian; Garg, Animesh (2023-12-01). "Large language models for chemistry robotics". Autonomous Robots. 47 (8): 1057–1086. doi:10.1007/s10514-023-10136-2. ISSN 1573-7527.
- ↑ Pendleton, Ian M.; Cattabriga, Gary; Li, Zhi; Najeeb, Mansoor Ani; Friedler, Sorelle A.; Norquist, Alexander J.; Chan, Emory M.; Schrier, Joshua (September 2019). "Experiment Specification, Capture and Laboratory Automation Technology (ESCALATE): a software pipeline for automated chemical experimentation and data management". MRS Communications. 9 (3): 846–859. Bibcode:2019MRSCo...9..846P. doi:10.1557/mrc.2019.72. ISSN 2159-6859.
- ↑ Jumper, John; Evans, Richard; Pritzel, Alexander; Green, Tim; Figurnov, Michael; Ronneberger, Olaf; Tunyasuvunakool, Kathryn; Bates, Russ; Žídek, Augustin; Potapenko, Anna; Bridgland, Alex; Meyer, Clemens; Kohl, Simon A. A.; Ballard, Andrew J.; Cowie, Andrew (2021-08-26). "Highly accurate protein structure prediction with AlphaFold". Nature. 596 (7873): 583–589. Bibcode:2021Natur.596..583J. doi:10.1038/s41586-021-03819-2. ISSN 1476-4687. PMC 8371605. PMID 34265844.
- ↑ Yang, Zhenyu; Zeng, Xiaoxi; Zhao, Yi; Chen, Runsheng (2023-03-14). "AlphaFold2 and its applications in the fields of biology and medicine". Signal Transduction and Targeted Therapy. 8 (1): 115. doi:10.1038/s41392-023-01381-z. ISSN 2059-3635. PMC 10011802. PMID 36918529.
- ↑ Akdel, Mehmet; Pires, Douglas E. V.; Pardo, Eduard Porta; Jänes, Jürgen; Zalevsky, Arthur O.; Mészáros, Bálint; Bryant, Patrick; Good, Lydia L.; Laskowski, Roman A.; Pozzati, Gabriele; Shenoy, Aditi; Zhu, Wensi; Kundrotas, Petras; Serra, Victoria Ruiz; Rodrigues, Carlos H. M. (2022-11-07). "A structural biology community assessment of AlphaFold2 applications". Nature Structural & Molecular Biology. 29 (11): 1056–1067. doi:10.1038/s41594-022-00849-w. ISSN 1545-9985. PMC 9663297. PMID 36344848.
- 1 2 3 Cao, Yudong; Romero, Jonathan; Olson, Jonathan P.; Degroote, Matthias; Johnson, Peter D.; Kieferová, Mária; Kivlichan, Ian D.; Menke, Tim; Peropadre, Borja; Sawaya, Nicolas P. D.; Sim, Sukin; Veis, Libor; Aspuru-Guzik, Alán (2019-10-09). "Quantum Chemistry in the Age of Quantum Computing". Chemical Reviews. 119 (19): 10856–10915. arXiv:1812.09976. Bibcode:2019ChRv..11910856C. doi:10.1021/acs.chemrev.8b00803. ISSN 0009-2665. PMID 31469277.
- 1 2 Buitrago Santanilla, Alexander; Regalado, Erik L.; Pereira, Tony; Shevlin, Michael; Bateman, Kevin; Campeau, Louis-Charles; Schneeweis, Jonathan; Berritt, Simon; Shi, Zhi-Cai; Nantermet, Philippe; Liu, Yong; Helmy, Roy; Welch, Christopher J.; Vachal, Petr; Davies, Ian W. (2015-01-02). "Nanomole-scale high-throughput chemistry for the synthesis of complex molecules". Science. 347 (6217): 49–53. Bibcode:2015Sci...347...49B. doi:10.1126/science.1259203. ISSN 0036-8075. PMID 25554781.
- ↑ Battersby, Bronwyn J.; Trau, Matt (2002-04-01). "Novel miniaturized systems in high-throughput screening". Trends in Biotechnology. 20 (4): 167–173. doi:10.1016/S0167-7799(01)01898-4. ISSN 0167-7799. PMID 11906749.
- 1 2 3 Pyzer-Knapp, Edward O.; Suh, Changwon; Gómez-Bombarelli, Rafael; Aguilera-Iparraguirre, Jorge; Aspuru-Guzik, Alán (2015-07-01). "What Is High-Throughput Virtual Screening? A Perspective from Organic Materials Discovery". Annual Review of Materials Research. 45 (1): 195–216. Bibcode:2015AnRMS..45..195P. doi:10.1146/annurev-matsci-070214-020823. ISSN 1531-7331.
- ↑ Krska, Shane W.; DiRocco, Daniel A.; Dreher, Spencer D.; Shevlin, Michael (2017-12-19). "The Evolution of Chemical High-Throughput Experimentation To Address Challenging Problems in Pharmaceutical Synthesis". Accounts of Chemical Research. 50 (12): 2976–2985. doi:10.1021/acs.accounts.7b00428. ISSN 0001-4842. PMID 29172435.
- 1 2 McNally, Andrew; Prier, Christopher K.; MacMillan, David W. C. (2011-11-25). "Discovery of an α-Amino C–H Arylation Reaction Using the Strategy of Accelerated Serendipity". Science. 334 (6059): 1114–1117. Bibcode:2011Sci...334.1114M. doi:10.1126/science.1213920. PMC 3266580. PMID 22116882.
- ↑ Broach, James R.; Thorner, Jeremy (1996-11-07). "High-throughput screening for drug discovery" (PDF). Nature. 384 (6604): Supplement 'Intelligent Drug Design' 14-16. PMID 8895594. Retrieved 2026-03-11.
- ↑ Shoichet, Brian K. (December 2004). "Virtual screening of chemical libraries". Nature. 432 (7019): 862–865. Bibcode:2004Natur.432..862S. doi:10.1038/nature03197. ISSN 0028-0836. PMC 1360234. PMID 15602552.
- ↑ Bajorath, Jürgen (November 2002). "Integration of virtual and high-throughput screening". Nature Reviews Drug Discovery. 1 (11): 882–894. Bibcode:2002NRvDD...1..882B. doi:10.1038/nrd941. ISSN 1474-1776. PMID 12415248.
- ↑ Chemistry (IUPAC), The International Union of Pure and Applied. "IUPAC - cheminformatics (11421)". goldbook.iupac.org. doi:10.1351/goldbook.11421. Retrieved 2026-02-09.
- ↑ Team, Neovarsity (2023-07-01). "Cheminformatics: an in-depth guide for beginners". neovarsity.org. Retrieved 2025-11-19.
- ↑ "List of useful databases | Chemistry Library". library.ch.cam.ac.uk. Retrieved 2025-11-10.
- ↑ "ChEMBL - ChEMBL". www.ebi.ac.uk. Retrieved 2025-11-10.
- ↑ "PubChem". pubchem.ncbi.nlm.nih.gov. Retrieved 2025-11-10.
- ↑ "ZINC". zinc.docking.org. Retrieved 2025-11-10.
- ↑ "Materials Project". next-gen.materialsproject.org/. Retrieved 2025-11-10.
- ↑ Informatics, NIST Office of Data and. "Welcome to the NIST WebBook". webbook.nist.gov. Retrieved 2025-11-10.
- ↑ "AIST:Spectral Database for Organic Compounds,SDBS". sdbs.db.aist.go.jp. Retrieved 2025-11-10.
- ↑ "The Largest Curated Crystal Structure Database | CCDC". www.ccdc.cam.ac.uk. Retrieved 2025-11-10.
- ↑ "Organic Chemistry Portal". www.organic-chemistry.org. Retrieved 2025-11-19.
- ↑ "Reaxys | An expert-curated chemistry database | Elsevier". www.elsevier.com. Retrieved 2025-11-10.
- ↑ "CAS SciFinder® Training Overview". www.cas.org. Retrieved 2025-11-10.
- 1 2 3 "The RDKit Documentation — The RDKit 2025.09.2 documentation". www.rdkit.org. Retrieved 2025-11-19.
- ↑ rxn4chemistry (2026-03-05), rxnmapper GitHub Repository, rxn4chemistry, retrieved 2026-03-11
{{citation}}: CS1 maint: numeric names: authors list (link) - ↑ "Welcome to Chemprop's documentation! — Chemprop 2.2.2 documentation". chemprop.readthedocs.io. Retrieved 2026-03-11.
- ↑ "How to use ChemPlot — ChemPlot 1.3.1 documentation". chemplot.readthedocs.io. Retrieved 2026-03-11.
- ↑ "Welcome to Computational Problem Solving in the Chemical Sciences - Computational Problem Solving in the Chemical Sciences". wexlergroup.github.io. Retrieved 2026-03-11.
- ↑ "Top 18 Python Chemistry Projects | LibHunt". www.libhunt.com. Retrieved 2026-03-11.
- ↑ Ryzhkov, Fedor V.; Ryzhkova, Yuliya E.; Elinson, Michail N. (2023-09-30). "Python in Chemistry: Physicochemical Tools". Processes. 11 (10): 2897. doi:10.3390/pr11102897. ISSN 2227-9717.
- 1 2 3 4 Channing, Georgia; Ghosh, Avijit (2026), "AI for scientific discovery is a social problem", Patterns, 7 (3) 101497, arXiv:2509.06580, doi:10.1016/j.patter.2026.101497, PMC 13100680, PMID 42028402
- 1 2 Swain, Matthew C.; Cole, Jacqueline M. (2016-10-24). "ChemDataExtractor: A Toolkit for Automated Extraction of Chemical Information from the Scientific Literature". Journal of Chemical Information and Modeling. 56 (10): 1894–1904. doi:10.1021/acs.jcim.6b00207. ISSN 1549-9596. PMID 27669338.
- ↑ Gurulingappa, Harsha; Mudi, Anirban; Toldo, Luca; Hofmann-Apitius, Martin; Bhate, Jignesh (2013-08-28). "Challenges in mining the literature for chemical information". RSC Advances. 3 (37): 16194–16211. Bibcode:2013RSCAd...316194G. doi:10.1039/C3RA40787J. ISSN 2046-2069.
- ↑ Raccuglia, Paul; Elbert, Katherine C.; Adler, Philip D. F.; Falk, Casey; Wenny, Malia B.; Mollo, Aurelio; Zeller, Matthias; Friedler, Sorelle A.; Schrier, Joshua; Norquist, Alexander J. (May 2016). "Machine-learning-assisted materials discovery using failed experiments". Nature. 533 (7601): 73–76. Bibcode:2016Natur.533...73R. doi:10.1038/nature17439. ISSN 1476-4687. PMID 27147027.
- 1 2 Toniato, Alessandra; Vaucher, Alain C.; Laino, Teodoro; Graziani, Mara (2025-06-13). "Negative chemical data boosts language models in reaction outcome prediction". Science Advances. 11 (24) eadt5578. Bibcode:2025SciA...11.5578T. doi:10.1126/sciadv.adt5578. PMC 12164950. PMID 40512839.
- ↑ Nishii, Takafumi; Ichizawa, Kaname; Nagano, Haruka; Mukai, Hiroya; Sakaguchi, Daimon; Gotoh, Hiroaki (2025-10-28). "Predicting Substrate Reactivity in Oxidative Homocoupling of Phenols Using Positive and Unlabeled Machine Learning". ACS Omega. 10 (42): 49805–49815. doi:10.1021/acsomega.5c05523. PMC 12572977. PMID 41179213.
- 1 2 3 Jia, Xiwen; Lynch, Allyson; Huang, Yuheng; Danielson, Matthew; Lang'at, Immaculate; Milder, Alexander; Ruby, Aaron E.; Wang, Hao; Friedler, Sorelle A.; Norquist, Alexander J.; Schrier, Joshua (September 2019). "Anthropogenic biases in chemical reaction data hinder exploratory inorganic synthesis". Nature. 573 (7773): 251–255. Bibcode:2019Natur.573..251J. doi:10.1038/s41586-019-1540-5. ISSN 1476-4687. PMID 31511682.
- ↑ Nature, Research Communities by Springer (2020-11-12). "After the Paper | Rolling the dice again… a look back at 'Anthropogenic biases in chemical reactions hinder exploratory inorganic synthesis'". Research Communities by Springer Nature. Retrieved 2026-04-14.
- ↑ "What is Generative AI? - Generative Artificial Intelligence Explained - AWS". Amazon Web Services, Inc. Retrieved 2024-03-20.
- ↑ Bommasani, Rishi; Hudson, Drew A.; Adeli, Ehsan; Altman, Russ; Arora, Simran; Arx, Sydney von; Bernstein, Michael S.; Bohg, Jeannette; Bosselut, Antoine (2022-07-12), On the Opportunities and Risks of Foundation Models, arXiv:2108.07258
- ↑ Friedland, Alex (2023-05-12). "What Are Generative AI, Large Language Models, and Foundation Models?". Center for Security and Emerging Technology. Retrieved 2026-03-10.
- 1 2 Ramos, Mayk Caldas; Collison, Christopher J.; White, Andrew D. (2025-02-05). "A review of large language models and autonomous agents in chemistry". Chemical Science. 16 (6): 2514–2572. doi:10.1039/D4SC03921A. ISSN 2041-6539. PMC 11739813. PMID 39829984.
- ↑ "What are Foundation Models? - Foundation Models in Generative AI Explained - AWS". Amazon Web Services, Inc. Retrieved 2024-03-20.
- ↑ Oviedo, Felipe; Ferres, Juan Lavista; Buonassisi, Tonio; Butler, Keith T. (2022-06-24). "Interpretable and Explainable Machine Learning for Materials Science and Chemistry". Accounts of Materials Research. 3 (6): 597–607. arXiv:2111.01037. Bibcode:2022AMatR...3..597O. doi:10.1021/accountsmr.1c00244.
- ↑ Clayton, Adam D.; Manson, Jamie A.; Taylor, Connor J.; Chamberlain, Thomas W.; Taylor, Brian A.; Clemens, Graeme; Bourne, Richard A. (2019). "Algorithms for the self-optimisation of chemical reactions". Reaction Chemistry & Engineering. 4 (9): 1545–1554. doi:10.1039/C9RE00209J. ISSN 2058-9883.
- ↑ Taylor, Connor J.; Pomberger, Alexander; Felton, Kobi C.; Grainger, Rachel; Barecka, Magda; Chamberlain, Thomas W.; Bourne, Richard A.; Johnson, Christopher N.; Lapkin, Alexei A. (2023-03-22). "A Brief Introduction to Chemical Reaction Optimization". Chemical Reviews. 123 (6): 3089–3126. Bibcode:2023ChRv..123.3089T. doi:10.1021/acs.chemrev.2c00798. ISSN 0009-2665. PMC 10037254. PMID 36820880.
- ↑ Reymond, Jean-Louis; Deursen, Ruud van; Blum, Lorenz C.; Ruddigkeit, Lars (2010-07-01). "Chemical space as a source for new drugs". MedChemComm. 1 (1): 30–38. doi:10.1039/C0MD00020E. ISSN 2040-2511.
- ↑ Reymond, Jean-Louis (2015-03-17). "The Chemical Space Project". Accounts of Chemical Research. 48 (3): 722–730. doi:10.1021/ar500432k. ISSN 0001-4842. PMID 25687211.
- ↑ "AI drug discovery: assessing the first AI-designed drug candidates to go into human clinical trials | CAS". www.cas.org. 2022-09-23. Retrieved 2024-03-13.
- ↑ Burki, Talha (May 2020). "A new paradigm for drug development". The Lancet Digital Health. 2 (5): e226–e227. doi:10.1016/S2589-7500(20)30088-1. PMC 7194950. PMID 32373787.
- ↑ "Exscientia Case Study". Amazon Web Services, Inc. Retrieved 2026-02-10.
- ↑ Nasdaq.com (2024-11-20). "Recursion and Exscientia, two leaders in the AI drug discovery space, have officially combined to advance the industrialization of drug discovery". Nasdaq. Retrieved 2026-02-10.
- 1 2 3 4 5 Park, Hyunsoo; Mastej, Kinga O.; Detrattanawichai, Panyalak; Nduma, Ryan; Walsh, Aron (2025-07-22), Closing the synthesis gap in computational materials design, ChemRxiv, doi:10.26434/chemrxiv-2025-sbc0c, retrieved 2025-09-29
- 1 2 3 Papidocha, Sven Michael; Burger, Andreas; Bernales, Varinia; Aspuru-Guzik, Alán (2025-09-02), The elephant in the lab: Synthesizability in generative small-molecule design, ChemRxiv, doi:10.26434/chemrxiv-2025-1lcpq, retrieved 2025-10-01
- ↑ Bartel, Christopher J. (2022-06-01). "Review of computational approaches to predict the thermodynamic stability of inorganic solids". Journal of Materials Science. 57 (23): 10475–10498. Bibcode:2022JMatS..5710475B. doi:10.1007/s10853-022-06915-4. ISSN 1573-4803. OSTI 1865532.
- ↑ Davies, Daniel W.; Butler, Keith T.; Jackson, Adam J.; Morris, Andrew; Frost, Jarvist M.; Skelton, Jonathan M.; Walsh, Aron (2016-10-13). "Computational Screening of All Stoichiometric Inorganic Materials". Chem. 1 (4): 617–627. Bibcode:2016Chem....1..617D. doi:10.1016/j.chempr.2016.09.010. ISSN 2451-9294. PMC 5074417. PMID 27790643.
- ↑ "Introduction — smact". smact.readthedocs.io. Retrieved 2025-09-29.
- ↑ molecularsets/moses, MOSES, 2025-09-25, retrieved 2025-10-01
- ↑ "Enamine REAL Space - Enamine". enamine.net. Retrieved 2025-10-01.
- ↑ Gao, Wenhao; Coley, Connor W. (2020-12-28). "The Synthesizability of Molecules Proposed by Generative Models". Journal of Chemical Information and Modeling. 60 (12): 5714–5723. doi:10.1021/acs.jcim.0c00174. ISSN 1549-9596. PMID 32250616.
- ↑ Ghazi Vakili, Mohammad; Gorgulla, Christoph; Snider, Jamie; Nigam, AkshatKumar; Bezrukov, Dmitry; Varoli, Daniel; Aliper, Alex; Polykovsky, Daniil; Padmanabha Das, Krishna M.; Cox III, Huel; Lyakisheva, Anna; Hosseini Mansob, Ardalan; Yao, Zhong; Bitar, Lela; Tahoulas, Danielle (2025-01-22). "Quantum-computing-enhanced algorithm unveils potential KRAS inhibitors". Nature Biotechnology. 43 (12): 1954–1959. doi:10.1038/s41587-024-02526-3. ISSN 1087-0156. PMC 12700792. PMID 39843581.
- ↑ Vost, Lucy; Chenthamarakshan, Vijil; Das, Payel; Deane, Charlotte M. (2025-04-09). "Improving structural plausibility in diffusion-based 3D molecule generation via property-conditioned training with distorted molecules". Digital Discovery. 4 (4): 1092–1099. doi:10.1039/D4DD00331D. ISSN 2635-098X.
