A sequence is an ordered description. Each residue symbol occupies a position, and changing either the symbol or its position changes what is being described. You do not need to memorise the entire amino-acid alphabet to start reading a specification, but you do need to preserve its direction and annotations.
Start with the notation system
IUPAC–IUB recommendations define both one-letter and three-letter amino-acid symbols. In the one-letter system, G corresponds to glycine, H to histidine and K to lysine. The letters are codes, not a rule that the first letter of each amino-acid name is always used.IUPAC sequence notation (opens in a new tab)
| Position | One-letter code | Three-letter code | Residue name |
|---|---|---|---|
| 1 | G | Gly | Glycine |
| 2 | H | His | Histidine |
| 3 | K | Lys | Lysine |
Thus GHK and Gly–His–Lys describe the same residue order at this level of notation. The expansion helps you read the string; it does not add a copper ion, specify a salt form or prove that a supplied sample has that sequence. Those details belong elsewhere in the molecular description and analytical record.
Read from the N-terminus towards the C-terminus
The conventional written direction is from the amino end, or N-terminus, towards the carboxyl end, or C-terminus. EMBL-EBI explains this directionality as part of primary structure. Modified ends and cyclic structures need their own annotations; a simple linear example is the starting convention, not every possible structure.EMBL-EBI structure guide (opens in a new tab)
In GHK, glycine occupies position 1, histidine position 2 and lysine position 3. KHG is a different sequence. Reversing the display does not preserve the same connectivity just because the residue counts stay unchanged. When copying from a diagram, confirm which end is labelled before transcribing left to right.
Count residues, not punctuation
A one-letter sequence without annotations generally gives one residue per symbol. In a three-letter representation, Gly–His–Lys still contains three residues, not nine. Spaces, line breaks and conventional separators usually organise the display; they are not additional amino acids.
As an original comparison exercise, write GHK on one line and GHKGHK on another. The second string repeats the three-symbol pattern and contains six positions. A repeated motif may help you recognise a pattern, but it does not justify dropping one copy when recording the sequence.
If a document uses brackets, a prefix or a special symbol, look for its key before counting. A terminal modification may add chemical information without adding an amino-acid residue. A range of residue numbers may describe a fragment of a larger sequence rather than a free-standing count beginning at one.
Keep uncertainty and modifications visible
The nomenclature includes symbols for ambiguity. For example, B can represent an unresolved aspartic-acid/asparagine choice and Z an unresolved glutamic-acid/glutamine choice. Do not replace an ambiguous symbol with one definite residue unless the source provides that resolution.IUPAC sequence notation (opens in a new tab)
A plain sequence string can also be incomplete as a chemical specification. It may not state terminal chemistry, side-chain modifications, stereochemical qualifications or additional connectivity. Copying only the letters from a more detailed description can therefore lose information even when every letter is correct.
A useful transcription keeps two fields: the residue string and the accompanying structural annotations. Then retain the source identifier or document version so that an unfamiliar mark can be checked later instead of being silently normalised away.
Check a transcription before using it
- Confirm one-letter versus three-letter notation.
- Check the stated N-to-C direction and any residue numbering.
- Compare the beginning, end and total residue count with the source.
- Preserve modifications, ambiguity codes and the source reference.
Sequence reading is not sequence verification. The result of this process is an accurate account of what the document says. Establishing whether a physical sample matches that account requires the relevant analytical evidence.
Sources and further detail
- IUPAC–IUB — One-letter amino-acid notation (opens in a new tab)
Nomenclature recommendations, sections 3AA-20 and 3AA-21; symbols, direction and ambiguous residues.
- EMBL-EBI — The peptide bond and primary structure (opens in a new tab)
Background on residue order, covalent linkage and conventional sequence direction. No claim about a supplier's synthesis process.
Sources checked 19 September 2026. Worked examples are illustrative unless a supplied report is explicitly identified. This article has not undergone independent scientific peer review.