跳转到主要内容
搜索

2026 年 10 月 9 日

Peptide therapeutics are expanding quickly. Approvals have climbed at an accelerating rate in recent years, and the drivers behind that growth are converging rather than singular: the clinical success of GLP-1 receptor agonists, maturing delivery technologies, and the ability of peptides to reach targets — particularly protein-protein interactions — that remain largely undruggable by small molecules, often with a favorable safety profile. Peptide pipelines are growing across biotech and large pharma as a result.

What gets discussed less is the informatics gap underneath that growth — specifically, the tooling needed for peptide sequence alignment and peptide sequence analysis. Small molecule discovery has a long history of purpose-built tooling — R-group decomposition, matched molecular series analysis, clustering, prediction — built around a well-understood structural representation. Peptides don’t have the same inheritance, and analyzing a peptide sequence as though it were “a small molecule with extra atoms” tends to break down at exactly the point where the interesting science starts: understanding how a single monomer substitution shifts potency, stability, or selectivity.

Representation is the first problem

A peptide sequence sounds simple until it needs to be represented in a way that supports reliable peptide sequence alignment downstream. Is it linear, cyclic, or branched? Does it contain non-natural monomers that are not handled by standard bioinformatics techniques? Does the representation need to be legible to a bench scientist, machine-parseable for downstream computation, or both?

Two notations have emerged, and they make different trade-offs. HELM (Hierarchical Editing Language for Macromolecules) offers strong computational completeness for complex biomolecule structures, including inline SMILES for non-natural monomers — but a raw HELM string for a branched peptide is dense and genuinely hard to read at a glance. BILN (Boehringer Ingelheim Line Notation) takes the opposite emphasis: a directly human-readable line notation that still handles branched and cyclized sequences, and encoding sequence alignment information in a way a scientist can scan without a specialized rendering step. Neither replaces the other. A platform supporting real peptide programs needs to render HELM clearly and accept BILN for fast human review, rather than forcing a choice between the two.

Table comparing BILN and HELM notation with HELM renderings of a cyclic peptide and its linear counterpart.

D360 dataset containing cyclic and linear peptide sequences as BILN and HELM sequences, with HELM rendering column

Underneath either notation sits a further requirement: monomer-level structure. Knowing a sequence contains a non-natural residue isn’t enough on its own — the actual atomistic structure of that residue affects how alignments should be scored and how that residue is likely to influence potency, stability, or selectivity. This is why standardized monomer libraries — public/standard and proprietary — function as core infrastructure for peptide sequence analysis, not an add-on layered in afterward.

Alignment is where the analysis actually happens

Once sequences are represented consistently, sequence alignment becomes the analytical backbone for everything downstream. Peptide alignment has requirements that small-molecule tools don’t anticipate:

  • Alignment scoring matrices need to account for non-natural monomers, using identity- or similarity-based approaches
  • Alignment needs to respect sequence cyclization and branching, not just linear sequences
  • Gap penalties, branch bonuses, and grouping by structural class need to be tunable, not fixed to generate alignments indicative of sequence-activity relationships
Alignment dialog with Helm Rendering aligned to reference CHEMBL508411 using the Needleman-Wunsch-Gotoh algorithm.

Alignment Setup Options

Alignment also needs to remain a working surface, not a static output. Manual editing — dragging a selected monomer to a new position, opening or closing a gap — is essential to understanding the relationship of sequence to bioprofile, and those edits shouldn’t disappear the next time a query runs. Capturing manual alignment adjustments themselves, so they’re reproduced automatically when new sequence and assay data is available, is a vital detail with a real effect on how much a scientist can trust the tool: their judgment persists as the underlying dataset grows.

Once aligned, the payoff is visual and immediate: coloring sequences by a monomer property such as AlogP, ordering by activity or similarity to a reference, and — most usefully — asking a positional question. Which monomer, at which position, correlates with higher potency or better stability? That’s difficult to answer from a spreadsheet of sequence strings, which is exactly where basic peptide sequence analysis stops short. It becomes tractable once alignment, monomer properties, and activity data sit in the same view, colored and ordered by the variable that matters.

Sequence alignment viewer with monomers colored by AlogP from blue (low) to magenta (high), aligned to reference CHEMBL508411.

Sequence Alignment Viewer with monomer coloring based on AlogP values

From analysis to design

The natural next step is using what the alignment reveals to design better sequences — and design works best when it isn’t disconnected from the analysis behind it. An integrated design workspace lets scientists build new sequences directly within the alignment context, combining favorable features from multiple known sequences, and promote those designs into a dataset for further analysis or synthesis prioritization. The design decision and the evidence supporting it stay in the same view.

Alignment viewer with a virtual sequence sandbox below, holding three edited peptide sequences for design.

Sequence Alignment Viewer with sandbox enabled for virtual peptide design

Where this is heading

The same approach extends naturally: head-to-tail cyclized peptides, matched sequence pairs and series, activity cliff detection, and systematic enumeration methods such as alanine scanning are all natural extensions of alignment-driven analysis rather than departures from it.

Enumerate Sequences dialog with base sequence CHEMBL508411 and monomer A set as the replacement.

Setup of an Alanine scan in Sequence Alignment Viewer

The takeaway

Peptide discovery is growing faster than the informatics supporting it in many organizations. Representation that serves both computational and human needs, monomer-level structural understanding, alignment that preserves a scientist’s manual judgment, and design integrated into the same analytical view — these aren’t extras added to a small-molecule platform. They’re a distinct analytical requirement that peptide programs need met directly.

Learn more about this subject heading

Certara’s D360 platform provides modality-specific analytics for peptides alongside small molecules, antibodies, ADCs, and other modalities — federating data across source systems into a single environment for analysis and design.

了解有关 D360 的更多信息

Author

Anikó Horváth-Farkas

Product Marketing Manager

Anikó is a Product Marketing Manager for Certara’s Discovery portfolio, connecting drug discovery professionals with technologies that accelerate their path from compound to candidate.

沪ICP备2022021526号

Powered by Translations.com GlobalLink Web Software