A peptide sequence is a string of letters, and almost everything you need to predict about a compound's behaviour is in it. How it dissolves, how many counter-ions the vial carries, which residues will oxidise, whether it can form a disulphide bridge, which protease will cut it. Reading a sequence is a small skill that pays back constantly, and it takes about ten minutes to learn.
Two notations, one molecule
Amino acids are written either as three-letter abbreviations or as single letters. Three-letter is readable; single-letter is compact and is what databases use.
| Group | Residues |
|---|---|
| Basic (positive at pH 7) | Lys K, Arg R, His H |
| Acidic (negative at pH 7) | Asp D, Glu E |
| Polar | Ser S, Thr T, Asn N, Gln Q, Tyr Y, Cys C |
| Hydrophobic | Ala A, Val V, Leu L, Ile I, Met M, Phe F, Trp W, Pro P |
| Special | Gly G (no side chain), Pro P (ring), Cys C (thiol) |
Sequences are written N-terminus to C-terminus, always. The first letter is the amino end. This is not optional formatting — the reverse sequence is a different molecule with an identical mass, which is why a mass measurement confirms composition without proving order.
How do I work out the net charge?
Count K, R and H as positive; count D and E as negative; add one positive for the free N-terminus and one negative for the free C-terminus, then take the difference. A peptide with three lysines, one glutamate and free termini comes out around +3 at neutral pH.
That single number predicts two practical things. It tells you which solvent the peptide wants — clearly charged peptides dissolve in water, near-neutral ones resist it. And it tells you roughly how many counter-ions the vial carries, since every basic site pairs with one, which is a direct contributor to how much of the powder is not peptide.
What do the modifications written around a sequence mean?
They are as important as the letters and are easy to skim past. Ac- at the front is N-terminal acetylation. -NH2 at the end is C-terminal amidation. A residue in brackets or lower case — (D-Phe) or d-Phe — is the mirror-image form, used to defeat stereospecific proteases. Aib is 2-aminoisobutyric acid, a non-natural residue that blocks DPP-4. Nle is norleucine, substituted for methionine to remove an oxidation site.
Each of these changes the mass, which is why a value calculated from the bare letters will not match a modified peptide's certificate. The modifications are usually there to extend half-life, and their presence is a reliable signal that the molecule was engineered rather than copied.
What to look for on a first read
Four things, in about the order they cause problems.
Cysteines. Count them. One free cysteine means the peptide can form intermolecular disulphide bridges — dimers, at concentration — and it is the first thing that makes a compound a poor candidate for co-formulation with another peptide. Two cysteines may mean an intended internal bridge, which is a structural feature rather than a hazard.
Methionine and tryptophan. Both oxidise. Methionine gains 16 mass units, tryptophan a similar small increment. In a short sequence a single oxidised residue is a large proportional change, and it is a common impurity on a certificate.
Asparagine and glutamine, especially followed by glycine. These deamidate, converting to aspartate and glutamate and shifting the mass by one unit. An Asn-Gly pair is the classic fast-deamidating motif. A one-dalton discrepancy is easy to dismiss and is often this.
Alanine or proline at position two. This is the DPP-4 cleavage site. A native sequence with it is short-lived; an engineered one usually has Aib there instead, and spotting the substitution tells you what the designer was defending against.
Why does proline matter so much for stability?
Because its side chain loops back and bonds to the backbone nitrogen, forming a ring that constrains the local geometry. Most peptidases need a backbone they can position in an active site, and a proline-constrained stretch does not adopt that geometry. This is why proline-rich sequences resist degradation generally, and why a Pro-Gly-Pro tail was bolted onto two unrelated fragments to buy them time. Proline also disrupts helices and introduces kinks, so in a longer peptide its position carries structural meaning as well.
Can I tell how soluble a peptide will be just from the letters?
Approximately, and well enough to choose a first solvent. Count the hydrophobic residues as a fraction of the total: if more than about half are drawn from A, V, L, I, M, F, W and P with little compensating charge, expect difficulty in water and plan for an organic co-solvent. If the sequence carries clear net charge in either direction, water will usually work. The prediction fails for sequences that aggregate — alternating hydrophobic patterns that stack into sheets can be insoluble despite reasonable composition — so it is a starting point rather than a guarantee.
What does the sequence not tell you?
How much peptide is in the vial, which is the question people most often try to answer from it. The sequence gives molecular weight; the vial's contents depend on net peptide content, counter-ion burden and residual water, none of which is predictable from the letters. It also does not give you purity, identity confirmation for the actual batch, or endotoxin. The sequence is a specification of what the molecule should be. The certificate is the evidence about what arrived.
All products referenced here are supplied for laboratory and research use only. They are not drugs, foods, supplements or cosmetics, and are not for human or veterinary use.




