🔬

Protein Structure Prediction

Biochemistry / AI

Computational protein biology. AlphaFold, protein design, structure prediction, drug discovery through molecular simulation, and de novo protein engineering.

27 Indexed Papers
3 API Sources
Sep 8 Last Updated

Top Publications

Ranked by citation impact across Semantic Scholar, OpenAlex & arXiv

#1
Semantic Scholar Open Access 38,861 citations

Highly accurate protein structure prediction with AlphaFold

This work validated an entirely redesigned version of the neural network-based model, AlphaFold, in the challenging 14th Critical Assessment of protein Structure Prediction (CASP14)15, demonstrating accuracy competitive with experimental structures in a majority of cases and greatly outperforming other methods.

Abstract

Proteins are essential to life, and understanding their structure can facilitate a mechanistic understanding of their function. Through an enormous experimental effort1–4, the structures of around 100,000 unique proteins have been determined5, but this represents a small fraction of the billions of known protein sequences6,7. Structural coverage is bottlenecked by the months to years of painstaking effort required to determine a single protein structure. Accurate computational approaches are needed to address this gap and to enable large-scale structural bioinformatics. Predicting the three-dimensional structure that a protein will adopt based solely on its amino acid sequence—the structure prediction component of the ‘protein folding problem’8—has been an important open research problem for more than 50 years9. Despite recent progress10–14, existing methods fall far short of atomic accuracy, especially when no homologous structure is available. Here we provide the first computational method that can regularly predict protein structures with atomic accuracy even in cases in which no similar structure is known. We validated an entirely redesigned version of our neural network-based model, AlphaFold, in the challenging 14th Critical Assessment of protein Structure Prediction (CASP14)15, demonstrating accuracy competitive with experimental structures in a majority of cases and greatly outperforming other methods. Underpinning the latest version of AlphaFold is a novel machine learning approach that incorporates physical and biological knowledge about protein structure, leveraging multi-sequence alignments, into the design of the deep learning algorithm. AlphaFold predicts protein structures with an accuracy competitive with experimental structures in the majority of cases using a novel deep learning architecture.

Source DOI PDF llms.txt
#2
Semantic Scholar Open Access 13,574 citations

Accurate structure prediction of biomolecular interactions with AlphaFold 3

The new AlphaFold model demonstrates substantially improved accuracy over many previous specialized tools: far greater accuracy for protein–ligand interactions compared with state-of-the-art docking tools, much higher accuracy for protein–nucleic acid interactions compared with nucleic-acid-specific predictors and substantially higher antibody–antigen prediction accuracy.

Abstract

The introduction of AlphaFold 21 has spurred a revolution in modelling the structure of proteins and their interactions, enabling a huge range of applications in protein modelling and design2, 3, 4, 5–6. Here we describe our AlphaFold 3 model with a substantially updated diffusion-based architecture that is capable of predicting the joint structure of complexes including proteins, nucleic acids, small molecules, ions and modified residues. The new AlphaFold model demonstrates substantially improved accuracy over many previous specialized tools: far greater accuracy for protein–ligand interactions compared with state-of-the-art docking tools, much higher accuracy for protein–nucleic acid interactions compared with nucleic-acid-specific predictors and substantially higher antibody–antigen prediction accuracy compared with AlphaFold-Multimer v.2.37,8. Together, these results show that high-accuracy modelling across biomolecular space is possible within a single unified deep-learning framework. AlphaFold 3 has a substantially updated architecture that is capable of predicting the joint structure of complexes including proteins, nucleic acids, small molecules, ions and modified residues with greatly improved accuracy over many previous specialized tools.

Source DOI PDF llms.txt
#3
OpenAlex Open Access 10,045 citations

ColabFold: making protein folding accessible to all

Abstract

ColabFold offers accelerated prediction of protein structures and complexes by combining the fast homology search of MMseqs2 with AlphaFold2 or RoseTTAFold. ColabFold's 40-60-fold faster search and optimized model utilization enables prediction of close to 1,000 structures per day on a server with one graphics processing unit. Coupled with Google Colaboratory, ColabFold becomes a free and accessible platform for protein folding. ColabFold is open-source software available at https://github.com/sokrypton/ColabFold and its novel environmental databases are available at https://colabfold.mmseqs.com .

Source DOI PDF llms.txt
#4
OpenAlex Open Access 9,568 citations

The STRING database in 2023: protein–protein association networks and functional enrichment analyses for any sequenced genome of interest

Abstract

Much of the complexity within cells arises from functional and regulatory interactions among proteins. The core of these interactions is increasingly known, but novel interactions continue to be discovered, and the information remains scattered across different database resources, experimental modalities and levels of mechanistic detail. The STRING database (https://string-db.org/) systematically collects and integrates protein-protein interactions-both physical interactions as well as functional associations. The data originate from a number of sources: automated text mining of the scientific literature, computational interaction predictions from co-expression, conserved genomic context, databases of interaction experiments and known complexes/pathways from curated sources. All of these interactions are critically assessed, scored, and subsequently automatically transferred to less well-studied organisms using hierarchical orthology information. The data can be accessed via the website, but also programmatically and via bulk downloads. The most recent developments in STRING (version 12.0) are: (i) it is now possible to create, browse and analyze a full interaction network for any novel genome of interest, by submitting its complement of encoded proteins, (ii) the co-expression channel now uses variational auto-encoders to predict interactions, and it covers two new sources, single-cell RNA-seq and experimental proteomics data and (iii) the confidence in each experimentally derived interaction is now estimated based on the detection method used, and communicated to the user in the web-interface. Furthermore, STRING continues to enhance its facilities for functional enrichment analysis, which are now fully available also for user-submitted genomes.

Source DOI PDF llms.txt
#5
OpenAlex Open Access 8,528 citations

AlphaFold Protein Structure Database: massively expanding the structural coverage of protein-sequence space with high-accuracy models

Abstract

The AlphaFold Protein Structure Database (AlphaFold DB, https://alphafold.ebi.ac.uk) is an openly accessible, extensive database of high-accuracy protein-structure predictions. Powered by AlphaFold v2.0 of DeepMind, it has enabled an unprecedented expansion of the structural coverage of the known protein-sequence space. AlphaFold DB provides programmatic access to and interactive visualization of predicted atomic coordinates, per-residue and pairwise model-confidence estimates and predicted aligned errors. The initial release of AlphaFold DB contains over 360,000 predicted structures across 21 model-organism proteomes, which will soon be expanded to cover most of the (over 100 million) representative sequences from the UniRef90 data set.

Source DOI PDF llms.txt
#6
OpenAlex Open Access 5,913 citations

Accurate prediction of protein structures and interactions using a three-track neural network

Abstract

DeepMind presented notably accurate predictions at the recent 14th Critical Assessment of Structure Prediction (CASP14) conference. We explored network architectures that incorporate related ideas and obtained the best performance with a three-track network in which information at the one-dimensional (1D) sequence level, the 2D distance map level, and the 3D coordinate level is successively transformed and integrated. The three-track network produces structure predictions with accuracies approaching those of DeepMind in CASP14, enables the rapid solution of challenging x-ray crystallography and cryo-electron microscopy structure modeling problems, and provides insights into the functions of proteins of currently unknown structure. The network also enables rapid generation of accurate protein-protein complex models from sequence information alone, short-circuiting traditional approaches that require modeling of individual subunits followed by docking. We make the method available to the scientific community to speed biological research.

Source DOI PDF llms.txt
#7
OpenAlex 4,167 citations

Protein complex prediction with AlphaFold-Multimer

Abstract

While the vast majority of well-structured single protein chains can now be predicted to high accuracy due to the recent AlphaFold [1] model, the prediction of multi-chain protein complexes remains a challenge in many cases. In this work, we demonstrate that an AlphaFold model trained specifically for multimeric inputs of known stoichiometry, which we call AlphaFold-Multimer, significantly increases accuracy of predicted multimeric interfaces over input-adapted single-chain AlphaFold while maintaining high intra-chain accuracy. On a benchmark dataset of 17 heterodimer proteins without templates (introduced in [2]) we achieve at least medium accuracy (DockQ [3] ≥ 0.49) on 13 targets and high accuracy (DockQ ≥ 0.8) on 7 targets, compared to 9 targets of at least medium accuracy and 4 of high accuracy for the previous state of the art system (an AlphaFold-based system from [2]). We also predict structures for a large dataset of 4,446 recent protein complexes, from which we score all non-redundant interfaces with low template identity. For heteromeric interfaces we successfully predict the interface (DockQ ≥ 0.23) in 70% of cases, and produce high accuracy predictions (DockQ ≥ 0.8) in 26% of cases, an improvement of +27 and +14 percentage points over the flexible linker modification of AlphaFold [4] respectively. For homomeric inter-faces we successfully predict the interface in 72% of cases, and produce high accuracy predictions in 36% of cases, an improvement of +8 and +7 percentage points respectively.

Source DOI llms.txt
#8
OpenAlex Open Access 3,601 citations

Improved protein structure prediction using potentials from deep learning

Source DOI PDF llms.txt
#9
Semantic Scholar Open Access 2,584 citations

Highly accurate protein structure prediction for the human proteome

The state-of-the-art machine learning method, AlphaFold, is applied at a scale that covers almost the entire human proteome, and the resulting dataset covers 58% of residues with a confident prediction, of which a subset (36% of all residues) have very high confidence.

Abstract

Protein structures can provide invaluable information, both for reasoning about biological processes and for enabling interventions such as structure-based drug development or targeted mutagenesis. After decades of effort, 17% of the total residues in human protein sequences are covered by an experimentally determined structure1. Here we markedly expand the structural coverage of the proteome by applying the state-of-the-art machine learning method, AlphaFold2, at a scale that covers almost the entire human proteome (98.5% of human proteins). The resulting dataset covers 58% of residues with a confident prediction, of which a subset (36% of all residues) have very high confidence. We introduce several metrics developed by building on the AlphaFold model and use them to interpret the dataset, identifying strong multi-domain predictions as well as regions that are likely to be disordered. Finally, we provide some case studies to illustrate how high-quality predictions could be used to generate biological hypotheses. We are making our predictions freely available to the community and anticipate that routine large-scale and high-accuracy structure prediction will become an important tool that will allow new questions to be addressed from a structural perspective. AlphaFold is used to predict the structures of almost all of the proteins in the human proteome—the availability of high-confidence predicted structures could enable new avenues of investigation from a structural perspective.

Source DOI PDF llms.txt
#10
OpenAlex Open Access 2,096 citations

AlphaFold Protein Structure Database in 2024: providing structure coverage for over 214 million protein sequences

Abstract

The AlphaFold Database Protein Structure Database (AlphaFold DB, https://alphafold.ebi.ac.uk) has significantly impacted structural biology by amassing over 214 million predicted protein structures, expanding from the initial 300k structures released in 2021. Enabled by the groundbreaking AlphaFold2 artificial intelligence (AI) system, the predictions archived in AlphaFold DB have been integrated into primary data resources such as PDB, UniProt, Ensembl, InterPro and MobiDB. Our manuscript details subsequent enhancements in data archiving, covering successive releases encompassing model organisms, global health proteomes, Swiss-Prot integration, and a host of curated protein datasets. We detail the data access mechanisms of AlphaFold DB, from direct file access via FTP to advanced queries using Google Cloud Public Datasets and the programmatic access endpoints of the database. We also discuss the improvements and services added since its initial release, including enhancements to the Predicted Aligned Error viewer, customisation options for the 3D viewer, and improvements in the search engine of AlphaFold DB.

Source DOI PDF llms.txt
#11
OpenAlex Open Access 2,057 citations

Robust deep learning–based protein sequence design using ProteinMPNN

Abstract

Although deep learning has revolutionized protein structure prediction, almost all experimentally characterized de novo protein designs have been generated using physically based approaches such as Rosetta. Here, we describe a deep learning-based protein sequence design method, ProteinMPNN, that has outstanding performance in both in silico and experimental tests. On native protein backbones, ProteinMPNN has a sequence recovery of 52.4% compared with 32.9% for Rosetta. The amino acid sequence at different positions can be coupled between single or multiple chains, enabling application to a wide range of current protein design challenges. We demonstrate the broad utility and high accuracy of ProteinMPNN using x-ray crystallography, cryo-electron microscopy, and functional studies by rescuing previously failed designs, which were made using Rosetta or AlphaFold, of protein monomers, cyclic homo-oligomers, tetrahedral nanoparticles, and target-binding proteins.

Source DOI PDF llms.txt
#12
OpenAlex Open Access 686 citations

AlphaFold and Implications for Intrinsically Disordered Proteins

Abstract

Accurate predictions of the three-dimensional structures of proteins from their amino acid sequences have come of age. AlphaFold, a deep learning-based approach to protein structure prediction, shows remarkable success in independent assessments of prediction accuracy. A significant epoch in structural bioinformatics was the structural annotation of over 98% of protein sequences in the human proteome. Interestingly, many predictions feature regions of very low confidence, and these regions largely overlap with intrinsically disordered regions (IDRs). That over 30% of regions within the proteome are disordered is congruent with estimates that have been made over the past two decades, as intense efforts have been undertaken to generalize the structure-function paradigm to include the importance of conformational heterogeneity and dynamics. With structural annotations from AlphaFold in hand, there is the temptation to draw inferences regarding the "structures" of IDRs and their interactomes. Here, we offer a cautionary note regarding the misinterpretations that might ensue and highlight efforts that provide concrete understanding of sequence-ensemble-function relationships of IDRs. This perspective is intended to emphasize the importance of IDRs in sequence-function relationships (SERs) and to highlight how one might go about extracting quantitative SERs to make sense of how IDRs function.

Source DOI PDF llms.txt
#13
OpenAlex 294 citations

Protein structure predictions to atomic accuracy with AlphaFold

Source DOI llms.txt
#14
OpenAlex Open Access 97 citations

De novo protein design by inversion of the AlphaFold structure prediction network

Abstract

De novo protein design enhances our understanding of the principles that govern protein folding and interactions, and has the potential to revolutionize biotechnology through the engineering of novel protein functionalities. Despite recent progress in computational design strategies, de novo design of protein structures remains challenging, given the vast size of the sequence-structure space. AlphaFold2 (AF2), a state-of-the-art neural network architecture, achieved remarkable accuracy in predicting protein structures from amino acid sequences. This raises the question whether AF2 has learned the principles of protein folding sufficiently for de novo design. Here, we sought to answer this question by inverting the AF2 network, using the prediction weight set and a loss function to bias the generated sequences to adopt a target fold. Initial design trials resulted in de novo designs with an overrepresentation of hydrophobic residues on the protein surface compared to their natural protein family, requiring additional surface optimization. In silico validation of the designs showed protein structures with the correct fold, a hydrophilic surface and a densely packed hydrophobic core. In vitro validation showed that 7 out of 39 designs were folded and stable in solution with high melting temperatures. In summary, our design workflow solely based on AF2 does not seem to fully capture basic principles of de novo protein design, as observed in the protein surface's hydrophobic vs. hydrophilic patterning. However, with minimal post-design intervention, these pipelines generated viable sequences as assessed experimental characterization. Thus, such pipelines show the potential to contribute to solving outstanding challenges in de novo protein design.

Source DOI PDF llms.txt
#15
OpenAlex Open Access 84 citations

Protein structure prediction beyond AlphaFold

Source DOI PDF llms.txt
#16
Semantic Scholar Open Access 42 citations

Proteins with alternative folds reveal blind spots in AlphaFold-based protein structure prediction

Three blind spots that alternative conformations reveal about AF-based protein structure prediction are reviewed and suggested approaches to predict alternative folds more reliably are suggested.

Abstract

In recent years, advances in artificial intelligence (AI) have transformed structural biology, particularly protein structure prediction. Though AI-based methods, such as AlphaFold (AF), often predict single conformations of proteins with high accuracy and confidence, predictions of alternative folds are often inaccurate, low-confidence, or simply not predicted at all. Here, we review three blind spots that alternative conformations reveal about AF-based protein structure prediction. First, proteins that assume conformations distinct from their training-set homologs can be mispredicted. Second, AF overrelies on its training set to predict alternative conformations. Third, degeneracies in pairwise representations can lead to high-confidence predictions inconsistent with experiment. These weaknesses suggest approaches to predict alternative folds more reliably.

Source DOI PDF llms.txt
#17
Semantic Scholar 33 citations

Advancements in protein structure prediction: A comparative overview of AlphaFold and its derivatives

This review provides a comprehensive analysis of AlphaFold and its derivatives (AF2 and AF3) in protein structure prediction, which enables groundbreaking advancements in protein design, disease research and discusses future integration with experimental techniques.

Abstract

This review provides a comprehensive analysis of AlphaFold (AF) and its derivatives (AF2 and AF3) in protein structure prediction. These tools have revolutionized structural biology with their highly accurate predictions, driving progress in protein modeling, drug discovery, and the study of protein dynamics. Its exceptional accuracy has redefined our understanding of protein folding, which enables groundbreaking advancements in protein design, disease research and discusses future integration with experimental techniques. In addition, their achievement features, architectures, important case studies, and noteworthy effects in the field of biology and medicine were evaluated. In consideration of the fact that AF2 is a relatively recent innovation, it has already been taken into account in many studies that highlight its applications in many ways. Moreover, the limitations of AF2 that directed to the introduction of AF3 are also reported, which is a great improvement as it provides precise predictions of the structures and interactions of proteins, DNA, RNA, and ligands, thereby aiding in the understanding of the molecular level. Addressing current challenges and forecasting future developments, this work underscores the lasting significance of AF in reshaping the scientific landscape of protein research.

Source DOI llms.txt
#18
Semantic Scholar Open Access 32 citations

Protein Design Using Structure-Prediction Networks: AlphaFold and RoseTTAFold as Protein Structure Foundation Models.

This work reviews recent studies that use structure-prediction neural networks to design proteins, via approaches such as activation maximization, inpainting, or denoising diffusion, and suggests that continued improvement of their accuracy and generality will be key to unlocking the full potential of protein design.

Abstract

Designing proteins with tailored structures and functions is a long-standing goal in bioengineering. Recently, deep learning advances have enabled protein structure prediction at near-experimental accuracy, which has catalyzed progress in protein design as well. We review recent studies that use structure-prediction neural networks to design proteins, via approaches such as activation maximization, inpainting, or denoising diffusion. These methods have led to major improvements over previous methods in wet-lab success rates for designing protein binders, metalloproteins, enzymes, and oligomeric assemblies. These results show that structure-prediction models are a powerful foundation for developing protein-design tools and suggest that continued improvement of their accuracy and generality will be key to unlocking the full potential of protein design.

Source DOI PDF llms.txt
#19
Semantic Scholar 31 citations

Proteins with alternative folds reveal blind spots in AlphaFold-based protein structure prediction

Three blind spots that alternative conformations reveal about AF-based protein structure prediction are reviewed and suggested approaches to predict alternative folds more reliably are suggested.

Abstract

In recent years, advances in artificial intelligence (AI) have transformed structural biology, particularly protein structure prediction. Though AI-based methods, such as AlphaFold (AF), often predict single conformations of proteins with high accuracy and confidence, predictions of alternative folds are often inaccurate, low-confidence, or simply not predicted at all. Here, we review three blind spots that alternative conformations reveal about AF-based protein structure prediction. First, proteins that assume conformations distinct from their training-set homologs can be mispredicted. Second, AF overrelies on its training set to predict alternative conformations. Third, degeneracies in pairwise representations can lead to high-confidence predictions inconsistent with experiment. These weaknesses suggest approaches to predict alternative folds more reliably.

Source llms.txt
#20
Semantic Scholar 21 citations

Emerging frontiers in protein structure prediction following the AlphaFold revolution

This review focuses on the application of state-of-the-art protein structure prediction to these advanced applications of deep learning, and suggests a set of guidelines for reporting AlphaFold predictions.

Abstract

Models of protein structures enable molecular understanding of biological processes. Current protein structure prediction tools lie at the interface of biology, chemistry and computer science. Millions of protein structure models have been generated in a very short space of time through a revolution in protein structure prediction driven by deep learning, led by AlphaFold. This has provided a wealth of new structural information. Interpreting these predictions is critical to determining where and when this information is useful. But proteins are not static nor do they act alone, and structures of proteins interacting with other proteins and other biomolecules are critical to a complete understanding of their biological function at the molecular level. This review focuses on the application of state-of-the-art protein structure prediction to these advanced applications. We also suggest a set of guidelines for reporting AlphaFold predictions.

Source DOI llms.txt
#21
Semantic Scholar Open Access 18 citations

AlphaFold-latest: revolutionizing protein structure prediction for comprehensive biomolecular insights and therapeutic advancements

This AlphaFold framework has the ability to yield atomically-accurate structural predictions for a variety of biomolecular interactions, hence facilitating advancements in drug discovery.

Abstract

Breakthrough achievements in protein structure prediction have occurred recently, mostly due to the advent of sophisticated machine learning methods and significant advancements in algorithmic approaches. The most recent version of the AlphaFold model, known as “AlphaFold-latest,” which expands the functionalities of the groundbreaking AlphaFold2, is the subject of this article. The goal of this novel model is to predict the three-dimensional structures of various biomolecules, such as ions, proteins, nucleic acids, small molecules, and non-standard residues. We demonstrate notable gains in precision, surpassing specialized tools across multiple domains, including protein–ligand interactions, protein–nucleic acid interactions, and antibody–antigen predictions. In conclusion, this AlphaFold framework has the ability to yield atomically-accurate structural predictions for a variety of biomolecular interactions, hence facilitating advancements in drug discovery.

Source DOI PDF llms.txt
#22
Semantic Scholar Open Access 9 citations

Boosting AlphaFold Protein Tertiary Structure Prediction through MSA Engineering and Extensive Model Sampling and Ranking in CASP16

The results show that MSA engineering through the use of different protein sequence databases, alignment tools, and domain segmentation as well as extensive model sampling are the key to generate accurate and correct structural models.

Abstract

AlphaFold2 and AlphaFold3 have revolutionized protein structure prediction by enabling high-accuracy tertiary structure predictions for most single-chain proteins. However, obtaining high-quality predictions for some hard protein targets with shallow or noisy multiple sequence alignments (MSAs) and complicated multi-domain architectures remains challenging. Here, we present MULTICOM4, an integrative protein structure prediction system that uses diverse MSA generation, large-scale model sampling, and an ensemble model quality assessment (QA) strategy of combining individual QA methods to improve model generation and ranking of AlphaFold2 and AlphaFold3. In the 16th Critical Assessment of Techniques for Protein Structure Prediction (CASP16), our predictors built on MULTICOM4 ranked among the top performers out of 120 predictors in tertiary structure prediction and outperformed a standard AlphaFold3 predictor. The average TM-score of our best performing predictor MULTCOM’s top-1 prediction for 84 CASP16 domain is 0.902. It achieved high accuracy (TM-score > 0.9) for 73.8% of the 84 domains and correct fold predictions (TM-score > 0.5) for 97.6% domains in terms of top-1 prediction. In terms of best-of-top-5 prediction, it predicted correct folds for all the domains. The results show that MSA engineering through the use of different protein sequence databases, alignment tools, and domain segmentation as well as extensive model sampling are the key to generate accurate and correct structural models. Additionally, using multiple complementary QA methods and model clustering can improve the robustness and reliability of model ranking.

Source DOI PDF llms.txt
#23
Semantic Scholar 8 citations

From CASP13 to the Nobel Prize: DeepMind's AlphaFold Journey in Revolutionizing Protein Structure Prediction and Beyond.

This review will begin by revisiting DeepMind's early efforts in CASP13, detailing the architecture and the remarkable progress that led to their breakthrough of AlphaFold2 in CASP14 (2020), and delve into two main areas: (1) AlphaFold's contributions to the scientific community across various fields over the past four years, and (2) the latest improvements, enhancements, and achievements by DeepMind.

Abstract

Four years ago, at the 14th Critical Assessment of Structure Prediction (CASP14), John Moult made a historic announcement that the long-standing challenge of Protein Structure Prediction- a problem that had confounded scientists for over five decades-had been "solved" for single protein chains. Supporting this groundbreaking statement was a plot depicting the median Global Distance Test (GDT) across 87 out of 92 domains, where AlphaFold2, developed by DeepMind, achieved an unprecedented score of 92.4. The bar chart not only underscored AlphaFold2' s remarkable performance-standing out prominently among other methods-but also revealed a level of accuracy that exceeded all prior expectations. In the years since this breakthrough, DeepMind's team has made significant strides. The AlphaFold Database now hosts approximately 214 million structures for various model organisms, covering nearly the entire genome. Research continues to explore multiple facets of protein science, including the prediction of multi-chain protein complex structures and the impact of missense mutations on protein function. The open availability of this extensive database and the suite of AlphaFold2 algorithms has catalysed remarkable advancements in protein biology and bioinformatics. This review will begin by revisiting DeepMind's early efforts in CASP13, detailing the architecture and the remarkable progress that led to their breakthrough of AlphaFold2 in CASP14 (2020). It will then delve into two main areas: (1) AlphaFold's contributions to the scientific community across various fields over the past four years, and (2) the latest improvements, enhancements, and achievements by DeepMind, including AlphaFold3 and the Nobel Prize in Chemistry.

Source DOI llms.txt
#24
Semantic Scholar Open Access 5 citations

Receptor- ligand interactions in plant inmate immunity revealed by AlphaFold protein structure prediction

Detecting interaction of both Ptr and Pi39(t) with AVR-Pita, and Pi-9 with both AVR-Pi9 and AVR-Pik, revealed a new insight into recognition of pathogen signaling molecules by these host R genes in triggering plant innate immunity.

Source DOI PDF llms.txt
#25
Semantic Scholar 4 citations

Beyond Current Boundaries: Integrating Deep Learning and AlphaFold for Enhanced Protein Structure Prediction from Low-Resolution Cryo-EM Maps

This study introduces DeepTracer-LowResEnhance, an innovative computational framework that uniquely integrates structural predictions from AlphaFold with a deep-learning-based map refinement strategy specifically tailored to enhance low-resolution maps.

Abstract

Constructing atomic models from cryo-electron microscopy (cryo-EM) maps is a crucial yet intricate task in structural biology. While advancements in deep learning, such as convolutional neural networks (CNNs) and graph neural networks (GNNs), have spurred the development of sophisticated map-to-model tools like DeepTracer and ModelAngelo, their efficacy notably diminishes with low-resolution maps beyond 4 Å. To address this critical gap, this study introduces DeepTracer-LowResEnhance, an innovative computational framework that uniquely integrates structural predictions from AlphaFold with a deep-learning-based map refinement strategy specifically tailored to enhance low-resolution maps. Unlike existing techniques, our approach leverages the strengths of AlphaFold's sequence-based predictions combined with advanced neural network-driven refinement processes to significantly improve map interpretability and modeling accuracy. DeepTracer-LowResEnhance demonstrates substantial and consistent improvements on an extensive dataset comprising 37 diverse protein cryo-EM maps, covering resolutions from 2.5 to 8.4 Å and including 22 challenging cases below 4 Å resolution. DeepTracer-LowResEnhance achieves an average TM-score improvement of 3.53x compared to baseline DeepTracer predictions. Notably, our enhanced methodology showed performance gains across 95.5% of the tested low-resolution datasets. A comparative analysis alongside traditional sharpening methods such as Phenix's auto-sharpening illustrates DeepTracer-LowResEnhance's superior capability in rendering more detailed and precise atomic models, thereby pushing the boundaries of current computational structural biology methodologies.

Source DOI llms.txt
#26
Semantic Scholar 3 citations

Revolutionizing structural biology: AI-driven protein structure prediction from AlphaFold to next-generation innovations.

This research emphasizes AI's importance in structural biology and envisions a future in which predictive tools will provide comprehensive insights into protein function, dynamics, and therapeutic potential.

Source DOI llms.txt
#27
Semantic Scholar Open Access 2 citations

AlphaFold as a Prior: Guiding Protein Structure Prediction Using Experimental Data with ROCKET

An augmentation of AlphaFold2, ROCKET, is presented that refines its predictions using cryo-EM, cryo-ET, and X-ray crystallography data, and it is demonstrated that this approach captures biologically important structural variation that AlphaFold2 does not.

Abstract

Advances in machine learning have transformed structural biology, enabling swift and accurate prediction of protein structure from sequence. However, challenges persist in capturing sidechain packing, condition-dependent conformational dynamics, and biomolecular interactions, primarily due to scarcity of high-quality training data. Emerging techniques, including cryo-electron tomography (cryo-ET) and high-throughput crystallography, promise vast new sources of structural data, but translating experimental observations into mechanistically interpretable atomic models remains a key bottleneck. Here, we address these challenges by improving the efficiency of structural analysis through combining experimental measurements with a landmark protein structure prediction method – AlphaFold2. We present an augmentation of AlphaFold2, ROCKET, that refines its predictions using cryo-EM, cryo-ET, and X-ray crystallography data, and demonstrate that this approach captures biologically important structural variation that AlphaFold2 does not. By performing structure optimization in the space of coevolutionary embeddings, rather than Cartesian coordinates, ROCKET automates difficult modeling tasks, such as flips of functional loops and domain rearrangements at low resolution. ROCKET does not require retraining of AlphaFold2 and is readily adaptable to other data modalities. This new type of structure refinement that optimizes latent representations in evolutionary space could unlock possibilities for high-throughput ligand screening, assemblies solved at low resolution, and conformational landscapes.

Source DOI PDF llms.txt