Research starter report
Deep Learning Approaches for Protein Structure Prediction from Sequence
- Topic
- Deep learning methods for predicting protein structure from amino acid sequence, including AlphaFold and related approaches
- Field
- Biology / Life Sciences
- Papers
- 15, every DOI checked against Crossref
- Sources
- Semantic Scholar, OpenAlex, Crossref
- Generated
- 2026-10-04
Field overview
Determining the three-dimensional structure of a protein from its one-dimensional amino acid sequence, known as the protein folding problem, has been one of the central challenges in molecular biology for over half a century. The structure of a protein determines its function, and understanding that structure is essential for drug discovery, enzyme engineering, understanding disease mechanisms, and interpreting genomic data. Historically, experimental methods such as X-ray crystallography, cryo-electron microscopy, and NMR spectroscopy have been the gold standards for structure determination, but these approaches are time-consuming, expensive, and not always feasible. Computational prediction methods have therefore been pursued in parallel, and the advent of deep learning has fundamentally transformed what is computationally achievable. The field has moved from predicting secondary structure elements and contact maps to producing near-atomic-accuracy three-dimensional models of entire protein chains and complexes, representing one of the most dramatic demonstrations of artificial intelligence applied to a core scientific problem.
The historical arc of computational protein structure prediction spans decades of progress through comparative modeling, threading, and ab initio approaches, with progress benchmarked through the biennial Critical Assessment of Protein Structure Prediction (CASP) competitions. Early machine learning methods improved contact prediction by exploiting co-evolutionary signals extracted from multiple sequence alignments (MSAs), recognizing that pairs of residues that co-vary across evolutionary time are likely to be in spatial proximity. The introduction of deep residual networks and attention mechanisms progressively improved the accuracy of these predictions throughout the late 2010s. The watershed moment came with the system described in “Highly accurate protein structure prediction with AlphaFold” (2021), which demonstrated that a deep learning architecture combining MSA processing, pairwise residue representations, and an equivariant structure module could achieve accuracy competitive with experimental methods on many targets. This result, validated at CASP14, effectively solved the protein folding problem for single-chain proteins under favorable conditions, triggering a rapid reorganization of the field. Complementary work described in “Accurate prediction of protein structures and interactions using a 3-track neural network” (2021) introduced RoseTTAFold, which similarly leveraged a multi-track architecture processing sequence, distance, and coordinate information simultaneously, establishing that the core architectural innovations were not unique to a single implementation.
The current state of the art has expanded well beyond single-chain structure prediction. “Highly accurate protein structure prediction for the human proteome” (2021) demonstrated that AlphaFold2 could be applied at proteome scale, producing structure predictions for virtually every human protein and catalyzing the creation of databases covering hundreds of millions of predicted structures. “ColabFold: making protein folding accessible to all” (2022) democratized access by coupling faster MSA generation tools with the AlphaFold2 inference pipeline, dramatically reducing computational requirements and enabling researchers without large-scale infrastructure to use these methods. The field has since pushed toward predicting biomolecular interactions, including protein-protein, protein-DNA, protein-RNA, and protein-ligand complexes, with “Accurate structure prediction of biomolecular interactions with AlphaFold 3” (2024) representing a major extension of the framework using a diffusion-based generative architecture capable of jointly modeling all classes of biomolecules. Hybrid and hierarchical approaches such as those described in “Deep-learning-based single-domain and multidomain protein structure prediction with D-I-TASSER” (2025) continue to integrate deep learning components with physics-based refinement pipelines, particularly for challenging multidomain and novel fold targets.
Despite remarkable progress, substantial open challenges remain. Intrinsically disordered proteins and regions that lack a stable folded structure are not well handled by current methods, which are trained to produce single structural states. Capturing conformational dynamics and predicting how protein structures change upon ligand binding, post-translational modification, or interaction with partners remains an active frontier. The accuracy of predicted interaction interfaces in protein complexes is considerably lower than for monomers, and the treatment of membrane proteins, glycoproteins, and large assemblies continues to be challenging. Questions also persist about the extent to which predicted structures reflect biologically relevant conformations versus low-energy ground states that may not dominate under physiological conditions. Reviews such as “Protein structure prediction via deep learning: an in-depth review” (2025), “AlphaFold2 and its applications in the fields of biology and medicine” (2023), and “Before and after AlphaFold2: An overview of protein structure prediction” (2023) collectively underscore that while deep learning has transformed the field, integration with experimental validation, molecular dynamics simulation, and functional annotation remains essential for translating structural predictions into biological insight. For a new PhD student, this landscape offers rich opportunities at the intersection of machine learning methodology, structural biology, and biomedical application.
Reading order
Read them in this order. Each entry says what the paper is, why it sits at this position, and what to take from it.
Before and after AlphaFold2: An overview of protein structure prediction
Letícia Machado Favery Bertoline et al., 2023. Frontiers in Bioinformatics. doi:10.3389/fbinf.2023.1120370 312 citations
This mini-review surveys the history of protein structure prediction methods before and after AlphaFold2, covering template-based and free modeling approaches, the impact of AlphaFold2, and the emergence of protein language model-based methods that can predict structures from sequence alone without multiple sequence alignments. It also discusses AlphaFold2's limitations and how newer approaches address some of them.
Start here: this historical overview of protein structure prediction before and after AlphaFold2 gives essential context for understanding where the field came from and why the deep learning revolution mattered.
Why read it. This review provides useful historical context and a comparative perspective on how AlphaFold2 transformed the field, while also introducing protein language model alternatives, helping situate deep learning approaches within the broader evolution of structure prediction methodology.
Protein structure prediction via deep learning: an in-depth review
Yajie Meng et al., 2025. Frontiers in Pharmacology. doi:10.3389/fphar.2025.1498662 39 citations
This 2025 review comprehensively covers deep learning methodologies applied to protein structure prediction, surveying relevant databases, large language models, and state-of-the-art deep learning tools. It also discusses future challenges and opportunities in the field.
Read after the historical overview, as this in-depth review of deep learning methods for protein structure prediction consolidates the broader technical landscape and prepares you for diving into specific architectures.
Why read it. As a recent survey, this review provides a current snapshot of the deep learning landscape for protein structure prediction, including coverage of protein language models alongside AlphaFold-style methods, which is useful for situating new work within the broader field.
AI-Driven Deep Learning Techniques in Protein Structure Prediction
Lingtao Chen et al., 2024. International Journal of Molecular Sciences. doi:10.3390/ijms25158426 118 citations
This 2024 survey comprehensively reviews computational methods for protein structure prediction, tracing the evolution from homology modeling and ab initio methods to state-of-the-art deep learning models such as AlphaFold (versions 1-3), RoseTTAFold, and ProteinBERT. It also compares model performance using CASP14, CASP15, and CAMEO benchmarks and discusses evaluation metrics including TM-score, GDT_TS, and lDDT.
Read after the deep learning review, as this survey of AI-driven techniques provides a focused bridge between general deep learning concepts and the specific neural network designs used in modern protein structure prediction.
Why read it. With 118 citations, this survey offers a well-structured comparative overview of the full landscape of deep learning approaches for protein structure prediction, including benchmark performance data that is useful for situating AlphaFold and related models within the broader field.
Highly accurate protein structure prediction with AlphaFold
J. Jumper et al., 2021. Nature. doi:10.1038/s41586-021-03819-2 39,594 citations
AlphaFold2 is a neural network-based computational method that predicts protein three-dimensional structures from amino acid sequences with atomic accuracy, even when no homologous structure is available. Validated in CASP14, it dramatically outperformed all previous methods and achieved accuracy competitive with experimental structures in the majority of cases.
Read after the surveys: this is the landmark AlphaFold2 paper that transformed the field, and the earlier surveys give you the vocabulary and context needed to fully appreciate its technical contributions.
Why read it. This is the foundational paper for the research topic, introducing the AlphaFold2 architecture and demonstrating the breakthrough in deep learning-based protein structure prediction. Any study of deep learning methods for protein structure prediction must begin here.
Accurate prediction of protein structures and interactions using a 3-track neural network
M. Baek et al., 2021. Science. doi:10.1126/science.abj8754 3,512 citations
This paper introduces RoseTTAFold, a three-track neural network that simultaneously processes protein sequence information at the 1D sequence level, 2D distance map level, and 3D coordinate level to predict protein structures. The approach achieves accuracy approaching AlphaFold2 and demonstrates utility for solving X-ray crystallography and cryo-EM modeling problems as well as predicting protein-protein complexes.
Read directly after the AlphaFold2 paper, as this RoseTTAFold 3-track neural network paper represents the major concurrent alternative approach and offers an illuminating architectural contrast.
Why read it. RoseTTAFold is one of the two landmark deep learning systems that transformed protein structure prediction in 2021, making it a foundational reference for understanding the architectural innovations, particularly multi-track information integration, that define the current state of the field.
Protein complex prediction with AlphaFold-Multimer
Richard Evans et al., 2021. bioRxiv. doi:10.1101/2021.10.04.463034 4,226 citations
AlphaFold-Multimer extends the original AlphaFold2 model to predict structures of multi-chain protein complexes of known stoichiometry, significantly improving interface prediction accuracy over prior methods. On benchmark datasets it achieves high-accuracy interface predictions in 26% of heteromeric cases and 36% of homomeric cases, representing improvements of 14 and 7 percentage points over the best prior approaches.
Read after the core AlphaFold2 paper, since AlphaFold-Multimer extends that foundational framework to protein complexes and is a natural next step in understanding the ecosystem of AlphaFold methods.
Why read it. As one of the most highly cited papers in this set (4226 citations) and a direct extension of AlphaFold2 to protein complexes, this paper is essential for understanding how deep learning structure prediction has been scaled to multi-chain systems, a major frontier in the field.
Highly accurate protein structure prediction for the human proteome
Kathryn Tunyasuvunakool et al., 2021. Nature. doi:10.1038/s41586-021-03828-1 3,347 citations
This paper applies AlphaFold2 at proteome scale to predict structures for 98.5% of human proteins, expanding confident structural coverage from 17% to 58% of residues. It also introduces metrics for interpreting prediction confidence and identifying disordered regions, with all predictions made freely available.
Read after AlphaFold-Multimer, as this paper on proteome-scale structure prediction demonstrates how AlphaFold2 was applied at massive scale and introduces key practical and methodological considerations.
Why read it. This work demonstrates the practical deployment of AlphaFold2 at large scale and provides the methodology and metrics used to assess prediction quality across diverse protein families, offering essential context for understanding how AlphaFold2 performs across the full range of protein sequences.
ColabFold: making protein folding accessible to all
Milot Mirdita et al., 2022. Nature Methods. doi:10.1038/s41592-022-01488-1 10,204 citations
ColabFold accelerates protein structure prediction by pairing the fast MMseqs2 homology search with AlphaFold2 or RoseTTAFold, achieving 40-60-fold faster searches and enabling prediction of nearly 1,000 structures per day on a single GPU. It is released as open-source software and integrated with Google Colaboratory to make structure prediction freely accessible.
Read after the proteome-scale prediction paper, since ColabFold addresses practical accessibility and computational efficiency, building directly on the AlphaFold2 ecosystem you have now studied in depth.
Why read it. ColabFold is a widely adopted practical implementation that makes AlphaFold2-based prediction accessible at scale, and understanding its design choices, including the use of MMseqs2 for MSA generation, is important for evaluating how deep learning methods are deployed in practice.
AlphaFold2 and its applications in the fields of biology and medicine
Zhenyu Yang et al., 2023. Signal Transduction and Targeted Therapy. doi:10.1038/s41392-023-01381-z 606 citations
This review covers the principles, system architecture, and key innovations behind AlphaFold2, explaining why it achieved atomic-level accuracy in protein structure prediction from amino acid sequences. It also surveys AlphaFold2 applications across biology and medicine, including drug discovery, protein design, and protein function prediction, while discussing current limitations.
Read after ColabFold, as this applications-focused review of AlphaFold2 in biology and medicine broadens your perspective on real-world impact now that you have a solid grasp of the underlying methods.
Why read it. This review provides a thorough explanation of AlphaFold2's architecture and the reasons for its success, making it a useful resource for understanding the technical foundations of the system and surveying the breadth of its downstream applications.
Recent Progress of Protein Tertiary Structure Prediction
Qiqige Wuyun et al., 2024. Molecules. doi:10.3390/molecules29040832 35 citations
This 2024 review surveys methodologies for predicting protein tertiary structure from amino acid sequences, covering traditional approaches such as template-based and template-free modeling as well as deep learning methods including contact/distance-guided models, end-to-end folding methods like AlphaFold2, and protein language model-based approaches. It also discusses CASP assessments, relevant databases including the AlphaFold Protein Structure Database, and multi-domain prediction challenges.
Read after the applications review, as this recent survey of protein tertiary structure prediction progress synthesizes developments across multiple methods and helps consolidate your knowledge before moving to the newest research.
Why read it. This review provides a thorough taxonomy of both classical and modern deep learning methods for protein structure prediction, with particular attention to how end-to-end methods like AlphaFold2 compare to earlier approaches, making it a useful reference for contextualizing methodological advances.
Accurate structure prediction of biomolecular interactions with AlphaFold 3
Josh Abramson et al., 2024. Nature. doi:10.1038/s41586-024-07487-w 14,251 citations
AlphaFold 3 extends the AlphaFold framework with a diffusion-based architecture capable of jointly predicting the structures of complexes involving proteins, nucleic acids, small molecules, ions, and modified residues. It substantially outperforms specialized tools for protein-ligand docking, protein-nucleic acid interactions, and antibody-antigen prediction within a single unified deep-learning framework.
Read after the tertiary structure survey, as AlphaFold 3 represents a major architectural evolution toward predicting biomolecular interactions beyond proteins alone, and prior reading makes its innovations fully accessible.
Why read it. This paper represents the most recent major evolution of the AlphaFold system, showing how the core deep learning approach has been generalized beyond single-chain protein structure prediction to full biomolecular complex modeling, which is directly relevant for understanding the trajectory of the field.
Boltz-1 Democratizing Biomolecular Interaction Modeling
Jeremy Wohlwend et al., 2024. bioRxiv. doi:10.1101/2024.11.19.624167 497 citations
Boltz-1 is an open-source deep learning model that achieves AlphaFold3-level accuracy in predicting 3D structures of biomolecular complexes, including protein-ligand and protein-protein interactions. The work introduces architectural innovations, speed optimizations, and a novel inference-time steering technique called Boltz-steering to correct hallucinations, and releases all code, weights, and datasets under an MIT license.
Read after AlphaFold 3, since Boltz-1 is a direct open-source competitor targeting similar biomolecular interaction modeling, making comparison between the two papers highly instructive.
Why read it. Boltz-1 is a highly cited (497 citations) recent model that directly extends the AlphaFold lineage to complex biomolecular interactions while being fully open-source, making it an important reference for understanding the current frontier of deep learning-based structure prediction.
Deep-learning-based single-domain and multidomain protein structure prediction with D-I-TASSER
Wei Zheng et al., 2025. Nature Biotechnology. doi:10.1038/s41587-025-02654-4 87 citations
D-I-TASSER is a hybrid protein structure prediction method that combines multisource deep learning potentials with iterative threading fragment assembly and physics-based simulations, along with a domain splitting and assembly protocol for large multidomain proteins. Benchmarks including CASP15 show it outperforms AlphaFold2 and AlphaFold3 on both single-domain and multidomain proteins.
Read after Boltz-1 to see how classical iterative assembly methods like I-TASSER have been modernized with deep learning, offering a perspective on hybrid approaches distinct from the pure transformer-based systems.
Why read it. This paper is directly relevant as a recent alternative to pure deep learning approaches, illustrating how integrating classical force field simulations with deep learning potentials can surpass AlphaFold on challenging multidomain targets and offering a complementary perspective on the design space for structure prediction methods.
From sequence to protein structure and conformational dynamics with artificial intelligence/machine learning
Alexander M. Ille et al., 2025. Structural Dynamics. doi:10.1063/4.0000765 9 citations
This 2025 review covers AI/ML models for protein structure prediction, including AlphaFold2, RoseTTAFold, and ESMFold, explaining their attention-based neural network architectures and the foundational sequence-structure hypothesis. It extends the discussion to conformational dynamics, proposing that amino acid sequence also encodes dynamic behavior and sketching a conceptual model architecture using NMR-derived conformational data to predict protein conformational ensembles.
Read after the D-I-TASSER paper, as this review of AI and ML methods for capturing conformational dynamics extends the field from static structure prediction toward understanding protein motion and flexibility.
Why read it. This paper provides a current (2025) synthesis of how deep learning methods connect sequence to structure and, uniquely, to conformational dynamics, offering forward-looking perspective on extending AlphaFold-style approaches beyond static structure prediction.
Deep Learning for Protein Structure Prediction: Advancements in Structural Bioinformatics
Daniel James Szelogowski, 2025. bioRxiv. doi:10.1101/2023.04.26.538026 8 citations
This paper presents a literature review of deep learning advances in protein structure prediction and introduces ProteiNN, a Transformer-based model for end-to-end single-sequence prediction of protein secondary and tertiary structure from integer-encoded amino acid sequences. The work benchmarks Transformer architectures for this task and provides a system for user-input sequence prediction and visualization.
Read last: this 2025 advancements review synthesizes the most recent progress across structural bioinformatics and is best appreciated once you have a comprehensive foundation in all preceding methods and papers.
Why read it. This recent paper offers a current survey of the deep learning landscape for structure prediction and provides a concrete example of applying Transformer architectures directly to single sequences, which is useful for understanding alternative approaches to multi-sequence methods like AlphaFold2.
Current trends
- Expansion from Single Proteins to Biomolecular Complexes and Interactions
- Research has rapidly moved beyond predicting individual protein structures toward modeling full biomolecular assemblies, including protein-protein, protein-nucleic acid, and protein-ligand interactions. Works such as Evans et al. (2021), Abramson et al. (2024), and Wohlwend et al. (2024) exemplify this trajectory, with AlphaFold-Multimer, AlphaFold 3, and Boltz-1 each extending predictive scope to increasingly complex molecular systems.
- Democratization and Accessibility of Structure Prediction Tools
- A strong trend exists toward making state-of-the-art prediction methods computationally accessible to researchers without large infrastructure, as demonstrated by Mirdita et al. (2022) with ColabFold and by Wohlwend et al. (2024) with Boltz-1. These efforts lower barriers by optimizing multiple sequence alignment pipelines and enabling predictions on consumer-grade hardware or cloud platforms.
- Integration of Conformational Dynamics and Ensemble Prediction
- There is growing recognition that static structure prediction is insufficient and that capturing conformational flexibility, disorder, and dynamic transitions is essential for biological relevance. Ille et al. (2025) and Meng et al. (2025) highlight emerging approaches that combine deep learning with molecular dynamics or generative models to represent protein conformational landscapes.
- Proteome-Scale and High-Throughput Structure Determination
- Deep learning methods are now being applied at proteome scale, enabling systematic structural coverage of entire organisms. Tunyasuvunakool et al. (2021) demonstrated this for the human proteome, and subsequent efforts catalogued in reviews by Yang et al. (2023) and Bertoline et al. (2023) show accelerating application to diverse species and functional annotation pipelines.
- Multi-Track and Hybrid Architectural Innovations
- Novel neural network architectures that jointly process sequence, multiple sequence alignment, and pairwise residue information have become a defining trend. Baek et al. (2021) introduced a 3-track architecture with RoseTTAFold, while Zheng et al. (2025) extended iterative prediction with D-I-TASSER, reflecting broad experimentation with how information from different biological representations can be fused for improved accuracy.
Research gaps
- Reliable Prediction of Intrinsically Disordered Regions and Proteins
- Current top-performing models such as AlphaFold (Jumper et al., 2021) were primarily benchmarked on well-folded globular proteins, and their confidence metrics do not straightforwardly translate to intrinsically disordered proteins or regions. Despite acknowledgment of this limitation in reviews by Meng et al. (2025) and Szelogowski et al. (2025), systematic methods for predicting the ensemble behavior of disordered regions remain underexplored.
- Modeling Protein Conformational Plasticity Beyond a Single State
- Most deployed tools output a single or small set of representative structures, which inadequately captures the multiple functional states, allostery, and induced-fit dynamics that govern biological activity. While Ille et al. (2025) identifies this as a frontier, robust deep learning frameworks that natively predict multi-state ensembles or transition pathways have not yet achieved the accuracy level of single-state prediction.
- Prediction Accuracy for Orphan and Low-Homology Sequences
- Performance of sequence-based deep learning methods degrades substantially when multiple sequence alignments are sparse or unavailable, a scenario common for novel or rapidly evolving proteins. This dependence on evolutionary co-variation signals, noted across Mirdita et al. (2022) and Wuyun et al. (2024), represents a gap for single-sequence or few-homolog prediction without sacrificing accuracy.
- Interpretability and Physical Validity of Predicted Structures
- Deep learning models produce structures with high predicted confidence scores that can nonetheless contain physically implausible local geometries or fail experimental validation, yet interpretability of model reasoning remains limited. Chen et al. (2024) and Bertoline et al. (2023) note this tension, and there is a clear gap in methods that link model outputs to mechanistic understanding or provide uncertainty estimates grounded in physical chemistry.
- Incorporation of Post-Translational Modifications and Non-Standard Chemistry
- The vast majority of structure prediction pipelines model canonical amino acid sequences and do not account for phosphorylation, glycosylation, methylation, or other post-translational modifications that substantially alter structure and function. AlphaFold 3 (Abramson et al., 2024) begins to incorporate some non-standard entities, but coverage remains incomplete, and systematic benchmarking of modified-protein prediction is lacking.
Open questions
- 1
Can a deep learning model be trained or fine-tuned to predict multiple distinct conformational states of a protein from sequence alone, and how should training data be curated to capture functionally relevant structural diversity rather than crystallographic noise?
- 2
To what extent do confidence metrics produced by models such as AlphaFold (Jumper et al., 2021) correlate with experimentally measurable structural uncertainty, and can these metrics be recalibrated to provide reliable estimates for intrinsically disordered regions?
- 3
How can multiple sequence alignment-free or low-homology prediction strategies be developed that maintain accuracy comparable to alignment-dependent approaches for proteins lacking known homologs, building on insights from ColabFold (Mirdita et al., 2022) regarding alignment bottlenecks?
- 4
What architectural modifications or training objectives would allow a 3-track or Evoformer-style network to natively represent post-translational modifications, and how would such a model perform on benchmark sets of experimentally characterized modified proteins?
- 5
Given that proteome-scale prediction has been demonstrated for individual organisms (Tunyasuvunakool et al., 2021), can systematic comparison of predicted structural repertoires across evolutionarily divergent organisms reveal novel functional relationships or previously uncharacterized protein families?
- 6
How can open, democratized tools such as ColabFold (Mirdita et al., 2022) and Boltz-1 (Wohlwend et al., 2024) be benchmarked against each other and against laboratory-scale resources on diverse biomolecular complex types, and what practical guidance can be derived for researchers choosing among accessible platforms?
Methods
End-to-end deep learning architectures for structure prediction, exemplified by AlphaFold2 using Evoformer and structure modules with attention mechanisms, Multiple sequence alignment (MSA) based coevolutionary analysis to extract residue-residue contact and distance information, Three-track and multi-track neural networks integrating sequence, MSA, and template information simultaneously (e.g., RoseTTAFold), Diffusion-based generative modeling for biomolecular structure prediction, as used in AlphaFold3 and Boltz-1, Homology search and template-based modeling using fast sequence search tools such as MMseqs2, integrated in ColabFold, Hybrid approaches combining deep learning with physics-based force field simulations and fragment assembly (e.g., D-I-TASSER), Transfer learning and large-scale pretraining on protein sequence and structural databases to improve generalization, Geometric deep learning and equivariant neural networks for modeling 3D atomic coordinates and molecular interactions
Datasets
Protein Data Bank (PDB), the primary repository of experimentally determined protein structures used for training and validation, UniRef and UniProt sequence databases, used for constructing multiple sequence alignments and training sequence-based models, CASP (Critical Assessment of Structure Prediction) benchmarks, the gold-standard competition for evaluating prediction accuracy across CASP13, CASP14, and CASP15, AlphaFold Protein Structure Database, providing predicted structures for the human proteome and millions of other proteins for downstream evaluation, Template Modeling Score (TM-score) and GDT-TS (Global Distance Test Total Score) as primary evaluation metrics for structural accuracy, Local Distance Difference Test (lDDT) and predicted lDDT (pLDDT) confidence scores used to assess per-residue prediction quality
Keywords
protein structure prediction, deep learning protein folding, AlphaFold neural network, amino acid sequence analysis, protein tertiary structure, end-to-end structure prediction, multiple sequence alignment, attention-based protein models, protein complex prediction, biomolecular interaction modeling, co-evolutionary information extraction, residue contact prediction, human proteome structure, multidomain protein structure, structural bioinformatics deep learning, protein folding accuracy assessment, transformer architecture proteins, three-track neural network, protein conformational dynamics prediction, ab initio structure prediction
Next steps
- 1
Reproduce a small-scale AlphaFold inference run using ColabFold (Mirdita et al., 2022), which provides free GPU access via Google Colab. Pick a well-characterized protein with a known crystal structure from the PDB, predict its structure, and quantitatively compare your prediction to the experimental structure using TM-score and RMSD metrics to build intuition for what 'good' predictions look like in practice.
- 2
Work through the architectural details of the original AlphaFold system (Jumper et al., 2021) and the RoseTTAFold 3-track network (Baek et al., 2021) side by side. Focus specifically on how each model encodes multiple sequence alignments, represents pairwise residue relationships, and iteratively refines structure. Sketch out the data flow diagrams for both to identify where their core design philosophies differ.
- 3
Design a small benchmark experiment comparing AlphaFold-Multimer (Evans et al., 2021) and AlphaFold 3 (Abramson et al., 2024) on a set of protein-protein complexes with known experimental structures. Use metrics like interface RMSD and DockQ score to assess complex-level accuracy, and document cases where one model clearly outperforms the other to develop a concrete sense of current capability gaps.
- 4
Identify a specific biological or therapeutic problem where structure prediction is a limiting step, such as predicting a drug target structure in an understudied protein family or modeling a disease-associated variant. Use the application landscape described in Yang et al. (2023) and Tunyasuvunakool et al. (2021) as a guide to scope the problem, then formulate a testable hypothesis your PhD project could address.
- 5
Study the treatment of conformational dynamics and disordered regions across the methods surveyed in Ille et al. (2025) and Meng et al. (2025). These are known weaknesses of static structure predictors. Identify one concrete limitation (for example, predicting alternative conformations or intrinsically disordered protein behavior) and find at least two recent preprints or publications outside your current list that propose solutions, to begin mapping the frontier of open problems.
- 6
Implement a systematic literature tracking routine: set up keyword alerts for terms like 'protein structure prediction', 'language model folding', and 'biomolecular interaction prediction' in Google Scholar and bioRxiv. Given the pace of the field illustrated by the progression from AlphaFold 2 (Jumper et al., 2021) through D-I-TASSER (Zheng et al., 2025) and Boltz-1 (Wohlwend et al., 2024), new architectures and benchmarks appear frequently, and staying current from the first year of your PhD will prevent costly gaps in your literature review.