← Blog

How to Map Protein Aggregation Propensity with Aggrescan3D and a PDB Structure

M
MindCell Research
2026-07-27
Share
Aggrescan3Dprotein engineeringstructural bioinformaticsaggregation propensityRCSB PDBLinux CPU

Table of contents

Static Aggrescan3D 1.0.2 analysis of the official RCSB Protein Data Bank structure 2GB1 produced finite scores for all 56 residues of chain A. The minimum score was −3.5629, the maximum was 1.1983, and the arithmetic mean was −1.4135875. The highest-scoring residue was valine A:21 at 1.1983; methionine A:1 ranked second at 0.5750. The validated outputs include the complete residue table, a structure annotated with scores, and native PNG and SVG profiles. This result identifies positions prioritized by the selected structure-based scoring method. It does not demonstrate that 2GB1 aggregates experimentally, predict an aggregation rate, or validate a mutation. Only static mode ran on Linux AMD64 CPU; FoldX- or CABS-dependent dynamic modes were not installed or tested.

Scientific introduction

Protein aggregation is a context-dependent process

Protein aggregation describes association into assemblies that may range from reversible oligomers to amorphous deposits or ordered amyloid fibrils. It matters in protein production, formulation, storage, disease biology, and therapeutic development. Aggregation can reduce soluble yield, alter activity, trigger particles, complicate purification, and create safety or immunogenicity concerns. Yet “aggregation propensity” is not one universal property measured by one universal number. Temperature, concentration, pH, ionic strength, agitation, interfaces, oxidation, proteolysis, cofactors, and molecular crowding can all change the pathway.

A protein sequence contains local physicochemical tendencies, but a folded structure hides some segments and exposes others. Hydrophobic residues buried in a stable core may have little opportunity to interact intermolecularly until partial unfolding occurs. Conversely, an exposed patch can create association risk even if the full sequence appears generally soluble. Structure-aware methods try to combine intrinsic residue tendencies with the spatial neighborhood and exposure represented by a three-dimensional model.

Computational screening is most useful as a prioritization layer. It can highlight positions for inspection, suggest mutations to model, and help compare variants under a consistent representation. It cannot replace experimental assays that measure the behavior of the actual construct under its intended conditions. A score should therefore travel with the input structure, algorithm version, mode, and assumptions.

What Aggrescan3D estimates

Aggrescan3D extends sequence-based aggregation concepts into three dimensions. The method assigns residue-level structural aggregation-propensity values using intrinsic residue tendencies and the local spatial environment. The original work introduced a structure-based approach that accounts for residues near one another in the folded protein, while later Aggrescan3D 2.0 work expanded capabilities such as dynamic analyses and stability-informed redesign. The retained package in this dossier is specifically Aggrescan3D 1.0.2, not the current web service or every feature described by later publications.

Positive scores are commonly interpreted as relatively aggregation-prone regions under the model, while more negative values indicate lower modeled propensity in that structural context. The sign is not a probability, energy, concentration, or kinetic rate. It should not be compared numerically with a value from an unrelated method. Even within the same package, a changed structure, chain selection, protonation, missing residues, or mode can change the profile.

The command used here was aggrescan -i data/2gb1.pdb -w outputs/a3d -O -v 1. It contains no dynamic option. The validator explicitly rejects a command containing --dynamic or -d. This matters because a static coordinate model represents one structural state, whereas dynamic workflows may explore flexibility or optimize variants using additional programs. Static validation cannot be promoted into dynamic-mode validation.

The 2GB1 structure and its provenance

The input is the official PDB-format file for RCSB entry 2GB1, downloaded from the RCSB file service with a PDBe fallback recorded in sources.json. The header identifies an engineered immunoglobulin-binding domain of streptococcal protein G, chain A. The experiment is solution NMR. The deposited coordinates are a restrained minimized structure derived from an ensemble; the file remarks describe the experimental restraints and point to entry 1GB1 for the set of 60 structures.

The protein G B1 domain is a compact 56-residue model protein widely used in folding and stability research. Its small size makes it practical for a transparent E2E: every residue identifier can be checked, and the complete score table remains readable. Small size does not mean the biological or physical interpretation is trivial. The deposited coordinate set is one processed representation from NMR evidence, not a snapshot of every conformation in solution.

Official provenance is essential. A PDB identifier alone does not specify which assembly, model, chain, altloc, or preprocessing operation was analyzed. This test keeps the downloaded file, validates chain/residue identifiers, and preserves the copied input beside the output. For a production study, investigators should also record retrieval date, revision, biological assembly decision, missing atoms, mutations, protonation, and any repair.

Reading per-residue scores

The output contains one row per residue, with protein label, chain, residue number, one-letter residue name, and score. All 56 chain-A residue identifiers are unique and match the annotated PDB. The range spans −3.5629 to 1.1983. The mean of all rows is −1.4135875. A:21 V is the maximum at 1.1983, followed by A:1 M at 0.5750. Several residues have exactly 0.0000 in the generated profile.

Ranking is useful for triage, but the top residue is not automatically the correct mutation site. A residue can contribute to a core, binding surface, secondary-structure pattern, or functional epitope. Replacing a hydrophobic side chain may lower a computed aggregation score while destabilizing the fold or destroying function. Candidate engineering therefore needs structural inspection, conservation, functional annotations, and stability analysis.

The arithmetic mean summarizes the profile but can hide a localized positive patch. Conversely, a positive single residue does not prove an extended aggregation-prone surface. Researchers should inspect neighboring residues in sequence and space, compare variants with an unchanged pipeline, and examine whether the input conformation plausibly exposes the region under formulation conditions.

Static structures and conformational ensembles

Proteins fluctuate. Side chains change rotamers, loops move, domains breathe, and partially unfolded states can expose surfaces absent from a deposited model. A static score evaluates the supplied coordinates. It does not sample those fluctuations or account for temperature-dependent populations. NMR-derived structures especially remind us that a coordinate file may summarize an ensemble.

Aggrescan3D documentation and later literature discuss dynamic approaches. Some modes may invoke FoldX for stability-related calculations or CABS-flex for flexibility simulations. Those dependencies add separate install, licensing, parameter, runtime, and validation questions. They can also produce different structures and therefore different aggregation profiles. None of those paths ran in this CPU-lightweight case.

The correct compatibility statement is narrow: Aggrescan3D 1.0.2 static mode passed in a retained legacy Python 2.7 environment on Linux AMD64 CPU. This result says nothing about FoldX installation, CABS execution, mutation scanning, dynamic relaxation, Windows, macOS, ARM, CUDA, or newer Aggrescan3D releases.

Aggregation prediction and experimental validation

Experimental aggregation assays observe different endpoints. Size-exclusion chromatography can resolve soluble oligomers, dynamic light scattering reports size distributions, turbidity captures light scattering, analytical ultracentrifugation examines sedimentation, microscopy visualizes particles, and dyes such as thioflavin T can report amyloid-like structures under suitable controls. Stress studies probe agitation, heat, freeze–thaw, light, or interfaces. No single assay captures every pathway.

A computational residue profile can help design a panel of variants, but the experimental plan should test the intended construct and formulation. Controls should distinguish expression or folding failure from reduced aggregation. Measurements need replicates, time points, concentration, buffer, temperature, and preanalytical handling. If a mutation lowers aggregation but also lowers stability, activity, or yield, it may not be an improvement.

For biologics, aggregation risk assessment sits within a broader developability and quality framework. A model can prioritize, but regulatory or manufacturing conclusions require validated analytical methods and product-specific evidence. This article reports software execution, not product quality.

Why legacy environments require explicit containment

Aggrescan3D 1.0.2 in this transaction depends on Python 2.7, which is end-of-life. Running it inside a scoped retained environment avoids changing the system Python and makes the package set inspectable. It does not remove security risk. Legacy environments should not be exposed as network services, should process trusted inputs, and should be separated from credentials and sensitive data.

The solved micromamba transaction linked 97 packages totaling 262,639,912 bytes. That measured Linux AMD64 transaction stayed below the local ceiling of 500,000,000 bytes. The number is not an upstream universal package size; it is the resolved transaction for the recorded channels and date. Environments and packages were retained after testing to support repair, not uninstalled.

A future production integration should consider containerization, restricted execution, checksummed packages, and migration to a maintained implementation. Reproducibility does not justify ignoring unmaintained dependencies. It means the legacy boundary is declared rather than hidden.

Semantic validation of structural outputs

The semantic validator does more than check filenames. It opens the residue CSV, parses every score as finite, verifies unique chain/residue identifiers, and compares reported minimum, maximum, and mean against the actual rows within tolerance. It checks that the package version contains 1.0.2, confirms static command morphology, requires at least two ranked residues, and ensures the annotated PDB, PNG, and SVG are nontrivial files.

The validator accepts explicit schema aliases where meaning is preserved. A version may appear under version, package_version, or a nested runtime map. Score statistics may be top-level fields or a score object. The ranked list may use one of two clear names. This flexibility became necessary because multiple chat runs produced correct science with different structured naming.

Alias handling must remain constrained. Accepting any numeric field would hide missing evidence; accepting a prose claim would allow fabricated summaries. The final validator recognizes named equivalents but always recomputes statistics from A3D.csv.

Test progress

GateAttempt-4 statusRetained evidence
Skill installationPassedPackaged Aggrescan3D instructions loaded into isolated context
Package preflightPassed97-package solved transaction, 262,639,912 bytes
Package installationPassedAggrescan3D 1.0.2 in isolated Python 2.7 environment
Demo structureReadyOfficial RCSB 2GB1 PDB with recorded source/fallback
Native executionPassedStatic per-residue calculation on Linux AMD64 CPU
Chat executionPassedAgent-directed request created all canonical outputs
Artifact validationPassedScores recomputed; identifiers, PDB, PNG, and SVG checked
Publication evidencePassedFocused result capture and manifested native/generated visuals

Demo user request

Run Aggrescan3D 1.0.2 in static mode on the supplied official RCSB structure at data/2gb1.pdb. Save the complete residue score CSV, annotated PDB, native PNG and SVG profile, and a JSON summary. Report residue count, minimum, maximum, mean, and the most aggregation-prone residues. Validate residue identifiers and finite scores. Do not run or claim FoldX, CABS, mutation, or dynamic modes.

The request is phrased as a scientific objective with explicit scope. It fixes static mode and canonical outputs while leaving the installed instructions to choose the legacy runtime safely.

Demo data

The input 2gb1.pdb is the official RCSB PDB entry for the engineered protein G immunoglobulin-binding domain, chain A. Its header states solution NMR and 56 residues in the analyzed chain. sources.json records the RCSB download URL and PDBe fallback.

FieldRetained valueMeaning
PDB entry2GB1Official structure identifier
ChainAOnly analyzed chain
Residues scored56One finite row per unique residue identifier
Minimum score−3.5629Lowest static A3D score in the table
Maximum score1.1983Highest static A3D score
Mean score−1.4135875Arithmetic mean over 56 rows
Highest positionA:21 VValine 21, score 1.1983
Second positionA:1 MMethionine 1, score 0.5750
ModeStaticNo dynamic flag, FoldX, or CABS execution

Validated workflow

The installer performed a micromamba dry-run with the lcbio and conda-forge channels before installation. It recorded a solved 97-package Linux AMD64 transaction of 262,639,912 bytes for Python 2.7 and Aggrescan3D. The package probe then confirmed the command in the retained environment.

The run copied the official input into its isolated project, invoked static mode, and preserved the native work directory. The wrapper reopened A3D.csv and the annotated output.pdb, calculated summary statistics from the scores, ranked residues, and recorded package/runtime details. Chat execution repeated the intent and the semantic validator compared the JSON claims with the raw CSV.

official RCSB 2GB1 coordinate file
  → isolated Aggrescan3D 1.0.2 / Python 2.7 static runtime
  → one finite score per chain-A residue
  → annotated PDB + native PNG/SVG profile
  → recomputed range, mean, and ranked positions
  → identifier and artifact validation

Transparent developer reproduction uses:

python test/scientific-skills/run_skill_cycle.py aggrescan3d
python test/scientific-skills/skills/aggrescan3d/chat_e2e.py
python test/scientific-skills/validate_how_to.py \
  test/scientific-skills/skills/aggrescan3d

These commands expose the validation chain; scientists do not need to write them to formulate the analysis request.

Results and artifacts

The validated static profile covers 56 residues. Scores range from −3.5629 to 1.1983, with mean −1.4135875. Valine A:21 is the top-ranked residue at 1.1983, followed by methionine A:1 at 0.5750. These values have no statistical uncertainty because they are deterministic outputs of the supplied coordinates and software version, not repeated experimental measurements.

Focused static Aggrescan3D results report for the 2GB1 protein structure

The focused application result presents the retained values and static-mode limitation. It is not a prompt screenshot, file list, raw JSON editor, or generic terminal card.

Native per-residue Aggrescan3D profile for chain A of RCSB 2GB1

The native plot visualizes the per-residue profile generated by Aggrescan3D. The editable A.svg, complete A3D.csv, and annotated output.pdb are retained alongside it.

Validated Aggrescan3D summary fields and observed static-mode values

The derived table image provides a compact review, while summary.json and the raw residue table remain authoritative.

What attempts 1–4 repaired

Attempt 1 executed real static analysis and produced the correct 56-residue range, mean, top residue, CSV, and annotated PDB. It failed because the validator required summary["version"], while the output used a different explicit package-version field. Correct scientific artifacts were retained, but the run received no feature credit.

Attempt 2 again produced correct outputs, including native PNG and SVG, but used a different score-summary schema. The validator attempted the missing top-level minimum_score key and failed. Attempt 3 produced the correct results yet supplied version information in a form that did not satisfy the then-current assertion.

The repair was to recognize narrow semantic aliases while preserving recomputation from raw evidence. Attempt 4 passed with package_version, top-level minimum_score, maximum_score, and mean_score, plus most_aggregation_prone_residues. Every score was checked against the CSV, and the command was checked for absence of dynamic flags. Publication therefore reflects attempt 4 only.

Designing a responsible protein-engineering follow-up

Begin with structural inspection of A:21 and neighboring residues. Determine solvent exposure, secondary structure, packing, binding interfaces, and conservation. Examine whether residue numbering corresponds to the expressed construct and whether tags, linkers, missing residues, or oligomeric partners change the surface.

Generate a small mutation panel rather than optimizing one score blindly. Exclude substitutions likely to disrupt core packing or function. Predict stability with an independently validated method, then express and purify variants under matched conditions. Measure soluble yield, monomer fraction, thermal stability, activity, and aggregation under intended formulation stresses.

Compare computational profiles only under an identical pipeline. If structures are modeled or relaxed, document the modeling method and uncertainty. Consider multiple conformations where flexibility may expose hidden patches. A score improvement should be treated as a hypothesis that earns experimental testing, not as proof of developability.

Reproducibility

Validation was completed on 2026-07-27 using Linux AMD64 CPU, legacy Python 2.7, and Aggrescan3D 1.0.2. The solved transaction was 262,639,912 bytes across 97 packages. CUDA was neither required nor tested, and the generated artifacts passed independent semantic checks.

Reproduction requires the exact official 2GB1 file and its revision, static command, package environment, chain/residue mapping, and raw CSV. A new structure preparation, package release, operating system, or dynamic option is a new validation condition. Visual and screenshot manifests connect every publication image to source evidence with SHA-256 hashes.

Limitations

Static scoring uses one supplied structural representation. It does not sample solution dynamics, partial unfolding, interfaces, concentration effects, formulation, post-translational modifications, or degradation. The 2GB1 coordinate file is an NMR-derived restrained minimized structure and does not represent every ensemble member.

Aggrescan3D 1.0.2 uses a legacy Python 2.7 environment. FoldX, CABS-flex, dynamic mode, mutation design, and newer Aggrescan3D functionality were not tested. No experimental aggregation assay, stability measurement, confidence interval, or variant comparison was performed.

The top score at A:21 V is a prioritization result, not a mutation recommendation. Protein function and stability require independent review. Compatibility is established only for the retained Linux AMD64 CPU transaction.

References

Try this workflow

MindPlot has built-in support for the demonstrated Aggrescan3D skill. A user can provide a protein structure and request a qualified static aggregation-propensity analysis in ordinary language; the MindPlot agent runs the isolated legacy tool, preserves residue tables and structural artifacts, and presents the validated result. Users do not need to write the reproduction commands above. Try it at https://mindplot.ai, or download the desktop version for a more integrated experience and stronger local-data privacy.