Peptide LexiconAll compounds

Amino Acid Chart: The Twenty, With the Masses Derived Rather Than Copied

A reference table of the twenty standard amino acids with codes, formulas and residue masses — every number computed from the formula, and the method checked against a real peptide's published mass.

Caroline S · Published 2026-09-16

Illustration: Twenty identical clear glass vials, each with a distinct colorless amino acid powder.
Illustration

Most amino acid charts are copied from other amino acid charts. This one was computed, and then the method was checked against a molecule whose mass somebody else published — which is a slower way to make a table and the only way to know it is right.

The twenty

Codes follow the IUPAC-IUB conventions. Masses are average masses in daltons, computed from each molecular formula using standard atomic weights. The residue mass is the free amino acid minus one water — what the amino acid contributes once it is inside a chain.

Name 3-letter 1-letter Formula Free (Da) Residue (Da)
Glycine Gly G C₂H₅NO₂ 75.07 57.05
Alanine Ala A C₃H₇NO₂ 89.09 71.08
Serine Ser S C₃H₇NO₃ 105.09 87.08
Proline Pro P C₅H₉NO₂ 115.13 97.12
Valine Val V C₅H₁₁NO₂ 117.15 99.13
Threonine Thr T C₄H₉NO₃ 119.12 101.11
Cysteine Cys C C₃H₇NO₂S 121.15 103.14
Leucine Leu L C₆H₁₃NO₂ 131.18 113.16
Isoleucine Ile I C₆H₁₃NO₂ 131.18 113.16
Asparagine Asn N C₄H₈N₂O₃ 132.12 114.10
Aspartate Asp D C₄H₇NO₄ 133.10 115.09
Glutamine Gln Q C₅H₁₀N₂O₃ 146.15 128.13
Lysine Lys K C₆H₁₄N₂O₂ 146.19 128.18
Glutamate Glu E C₅H₉NO₄ 147.13 129.12
Methionine Met M C₅H₁₁NO₂S 149.21 131.19
Histidine His H C₆H₉N₃O₂ 155.16 137.14
Phenylalanine Phe F C₉H₁₁NO₂ 165.19 147.18
Arginine Arg R C₆H₁₄N₄O₂ 174.20 156.19
Tyrosine Tyr Y C₉H₁₁NO₃ 181.19 163.18
Tryptophan Trp W C₁₁H₁₂N₂O₂ 204.23 186.21

Three things fall straight out of the table. The range is narrower than people expect — tryptophan is only about 3.3 times heavier than glycine, so peptides of similar length have similar masses regardless of composition. Leucine and isoleucine are identical in formula and in mass; they differ only in where the side chain branches, which is why no mass measurement can separate them. And only two residues contain sulphur, cysteine and methionine — cysteine's being the one that matters structurally, because it forms the disulphide bridges that hold folded molecules in shape.

The water, and the arithmetic everyone gets wrong

The most common mistake in checking a peptide's molecular weight is adding up its amino acids.

That answer is always too high, because forming a peptide bond releases a molecule of water. A chain of n residues has n − 1 bonds and has therefore shed n − 1 waters. At 18.015 daltons each, this is not a rounding error: a 30-residue peptide is about 522 daltons lighter than the sum of its parts.

So the working formula is:

peptide mass = Σ(residue masses) + 18.015

— the residue masses from the right-hand column, plus one water back for the two free ends of the finished chain.

Checking the method against something real

A table is only as good as its arithmetic, so here is the arithmetic run against a molecule whose mass was published by somebody else.

BPC-157, sequence GEPPPGKPADDAGLV — 15 residues, therefore 14 bonds.

Step Value
Sum of the 15 free amino acids 1,671.77 Da
Less 14 × water (14 × 18.015) − 252.21 Da
Computed mass 1,419.56 Da
PubChem's stated mass (CID 9941957) 1,419.5 Da

They agree. The table and the method behind it reproduce an independently published figure to within the rounding, which is the only reason to trust either.

The same arithmetic run on C-peptide — 31 residues, EAEDLQVGQVELGGGPGAGSLQPLALEGSLQ — gives about 3,020 daltons, and anybody can repeat it with the column above.

The codes that are not amino acids

Four one-letter codes appear routinely in sequences and stand for no particular residue:

Code Means
B Aspartate or asparagine (Asx)
Z Glutamate or glutamine (Glx)
J Leucine or isoleucine (Xle)
X Any residue

B and Z are historical honesty. A classical way of analysing composition converts asparagine to aspartate and glutamine to glutamate in the process of measuring them, so the method cannot distinguish the pairs — and rather than guessing, the notation records exactly what was established. J exists for the same reason at the mass-spectrometry end: leucine and isoleucine weigh the same, so an instrument reporting mass alone cannot choose between them.

A sequence containing B, Z, J or X is not a sequence with unusual chemistry in it. It is a sequence with a stated limit on what was measured.

Twenty, or twenty-two

The count depends on what is being counted, and the honest answer names the criterion.

Twenty is the standard set specified by the genetic code and present in essentially every protein — the table above.

Twenty-two adds two more that are genuinely incorporated during synthesis rather than attached afterwards, by machinery that reads what would otherwise be a stop codon:

  • Selenocysteine (U) — cysteine with selenium in place of sulphur. Not a curiosity in humans: glutathione peroxidase 1 (UniProt P07203, 203 residues) carries one at position 49, written as a U in its own sequence.
  • Pyrrolysine (O) — found in certain archaea and bacteria, not in humans.

Anything beyond that — hydroxyproline in collagen, phosphoserine in a signalling protein, the acetylated end of a synthetic peptide — is a modification, made after the chain is built. Modifications are extremely common and they change a molecule's mass, which is why a computed mass from this table sometimes disagrees with a published one. When it does, a modification is usually the reason, and the modification is usually the point of the molecule.

Why a library of compounds needs this page

Because nearly every claim about a peptide reduces to something in the table above. A sequence tells you its length; the residue masses tell you its weight; the weight tells you a great deal about whether it can survive a stomach or cross skin. The 500-dalton rule that governs what can be absorbed through skin is a statement about this column, and almost every peptide in this library is on the wrong side of it.

It also gives a reader something to check with. A vendor's certificate of analysis states a molecular weight. The sequence is usually published. Twenty numbers and one subtraction are enough to find out whether the two agree.

What a peptide actually is · C-peptide, a 31-residue chain worked through end to end · BPC-157, the molecule used to check this table.

Frequently asked questions

How many amino acids are there?

Twenty is the answer that is almost always meant: the twenty specified directly by the genetic code and found in essentially every protein. Twenty-two is also defensible, because selenocysteine and pyrrolysine are genuinely built into proteins during synthesis rather than added afterwards. Hundreds is true as well if you count every amino acid that exists in nature, most of which never appear in a protein. The number is not in dispute; what is being counted is.

Why is a peptide lighter than its amino acids added up?

Because making a bond costs a water molecule. Joining two amino acids releases one H₂O, so a chain of n residues has n−1 bonds and has shed n−1 waters. At 18.015 daltons each, that adds up fast: a 30-residue peptide is about 522 daltons lighter than the sum of its parts. This is the single most common arithmetic error people make when checking a peptide's molecular weight.

What is the difference between an amino acid mass and a residue mass?

The free amino acid is the molecule on its own. The residue is what is left of it once it is inside a chain, after the water has been released — so every residue mass in the table below is exactly 18.015 daltons lower than the free acid above it. To compute a peptide's mass, add the residue masses and then add one water back for the two free ends of the finished chain.

Why can mass spectrometry not distinguish leucine from isoleucine?

Because they weigh exactly the same. Both are C₆H₁₃NO₂, both 131.18 daltons free and 113.16 as residues; the only difference is where the side chain branches. An instrument that measures mass measures the same number for both, which is why the ambiguity code J exists for cases where a sequencing method cannot tell which one is present.

What do B, Z, J and X mean in a sequence?

They are ambiguity codes, not amino acids. B means the residue is aspartate or asparagine, Z means glutamate or glutamine, J means leucine or isoleucine, and X means any residue at all. The first two exist because a classical analysis method converts asparagine to aspartate and glutamine to glutamate, destroying the distinction — so the code records what was actually established rather than guessing.

What are the 21st and 22nd amino acids?

Selenocysteine (U) and pyrrolysine (O). Both are inserted during protein synthesis by special machinery that reads what would otherwise be a stop codon, which is what makes them different from the many amino acids that are modified after a protein is built. Selenocysteine is present in humans — glutathione peroxidase 1 carries one at position 49 — while pyrrolysine occurs in certain archaea and bacteria.

Are these masses monoisotopic or average?

Average, computed from standard atomic weights. Average masses are what molecular-weight figures on product pages and in chemical databases normally report, which makes them the right choice for checking such a figure. Mass spectrometry usually works in monoisotopic masses instead, which use the mass of the most common isotope of each element and give slightly lower numbers; a chart mixing the two silently is a chart that will not reconcile.