Reading an HPLC Chromatogram: What the Trace Shows That the Percentage Doesn't
A purity figure is one number extracted from a picture. How to read the picture — baseline, peak shape, resolution, where the impurities sit, and what an integration boundary can hide.
A purity figure is one number extracted from a picture by a person who made several judgement calls on the way. The picture contains the judgements. This is how to read it.
What the trace is
Reversed-phase HPLC pushes a dissolved sample through a column packed with a non-polar stationary phase — for peptides, usually C18-derivatised silica. A mobile phase gradient, typically water to acetonitrile with a small amount of acid, gradually increases the eluting strength.
Components that interact weakly with the stationary phase leave early. Strongly interacting ones leave later. A UV detector at the column outlet records absorbance over time, and the resulting trace is the chromatogram: time on the x-axis, detector response on the y.
The acid — usually trifluoroacetic acid — is there to keep basic residues protonated and suppress interactions with residual silanols on the silica. It is also the origin of the counter-ion salt that ends up in the final product. See net peptide content versus purity.
The baseline
Start at the bottom of the trace, not the top.
A good baseline is flat and low, with small, even noise. What you are looking for is anything that is not that:
- Drift — a baseline that climbs steadily through the run. Some drift is normal on a gradient because acetonitrile and water have different UV absorbance, but pronounced drift makes integration ambiguous, because where the baseline is drawn determines every peak area on the page.
- Wander — irregular rise and fall, suggesting temperature instability or a poorly mixed mobile phase.
- Noise — a thick, fuzzy baseline. This sets the detection limit: impurities below the noise floor are not being counted, which flatters the purity figure.
A certificate showing a chromatogram with the y-axis scaled so that everything except the main peak is a flat line at zero is not showing you the baseline. It is hiding it.
Peak shape
Symmetry
An ideal peak is symmetric — Gaussian, rising and falling at the same rate.
Tailing — a peak that rises sharply and trails off — is the most common deviation in peptide work. Causes include secondary interactions with residual silanols, column overload, or a column nearing the end of its life. Tailing matters because a tailing main peak overlaps whatever elutes after it, and the purity figure then depends on where the analyst decided the tail ended and the next peak began.
Fronting — a gradual rise and sharp fall — usually indicates column overload or a solvent mismatch between sample and mobile phase.
Width
Peak width is not a quality indicator on its own. Larger peptides produce broader peaks; shallower gradients produce broader peaks. A 44-residue tesamorelin peak will be visibly broader than a 5-residue ipamorelin peak on any comparable method, and neither observation says anything about purity.
Broad and tailing is a different matter, because it degrades the resolution that the whole measurement rests on.
Resolution: the property that determines whether the number means anything
Resolution is how completely two adjacent peaks are separated. It is what determines whether an impurity is counted as an impurity or absorbed into the main peak.
- Baseline resolved — the trace returns to baseline between two peaks. Their areas are unambiguous.
- Partially resolved — a visible valley, but the trace does not return to baseline. Integration requires a decision about where to split them.
- Shoulder — a bump on the side of the main peak. There are two components and the method cannot separate them.
- Co-eluting — no visible indication at all. The impurity is inside the main peak and is being counted as product.
A shoulder on the main peak is the single most informative feature on a peptide chromatogram. In peptide synthesis, the impurities that elute closest to the target are the ones structurally most similar to it — single-residue deletion sequences, diastereomers from partial racemisation, incompletely modified species. These are precisely the impurities you would most want to know about, and precisely the ones a purity percentage can quietly absorb.
Where the impurities sit
The retention time of an impurity relative to the main peak carries information.
Early-eluting (before the main peak) — more polar than the target. Commonly truncated sequences that lost hydrophobic residues, deprotection by-products, or hydrolysis products.
Late-eluting (after the main peak) — more hydrophobic. Commonly incompletely deprotected species still carrying a protecting group, dimers, or aggregated material. On lipidated peptides, over-conjugated species carrying an extra fatty acid chain.
Immediately adjacent — structurally near-identical. The deletion sequences and diastereomers described above.
A chromatogram with a clean baseline and two small, well-separated, distant impurity peaks describes a better material than one with the same purity percentage and an unresolved shoulder — even though both certificates might say 99.1%.
The integration boundaries
Integration is where the numbers are made, and where the judgement lives.
The analyst — or the software, with the analyst's parameters — decides where each peak starts and ends, and where the baseline runs beneath it. Reasonable people make different choices, and those choices propagate directly into the reported percentage.
Practices worth noticing:
- Boundaries placed to absorb a shoulder into the main peak.
- A high detection threshold that drops small peaks from the calculation entirely.
- A truncated time window that ends the run before late-eluting hydrophobic impurities have emerged. If the chromatogram stops two minutes after the main peak, anything slower is not in the total.
- A y-axis scaled to the main peak, which visually flattens everything else to nothing.
None of these are necessarily dishonest. All of them raise the reported number. None of them are visible in the number itself.
Detection wavelength
Usually printed in small type, and it changes what the trace can see.
214 nm detects the peptide bond, which every peptide has. This is the standard for purity work because it responds to all peptide material roughly in proportion to size.
280 nm detects aromatic residues — tryptophan, tyrosine, phenylalanine. Useful for quantitation when a peptide contains them, but as a purity wavelength it is partially blind: any impurity lacking aromatic residues is invisible, and invisible impurities do not count against the purity figure.
A purity figure from a 280 nm run on a peptide with a single tryptophan is not comparable to one from a 214 nm run, and is systematically more flattering.
A reading order
- Baseline — flat, low, quiet?
- Main peak shape — symmetric, or tailing?
- Shoulders — anything unresolved on the sides?
- Impurity peaks — how many, where, how well separated?
- Run length — does the trace continue well past the main peak?
- Axis labels — is the wavelength stated, is the y-axis scaled sensibly?
- Then read the percentage.
The number is the last thing to look at, not the first.
What a chromatogram cannot tell you
It cannot identify the main peak. A beautifully resolved, symmetric, 99.6% peak is a beautifully resolved 99.6% peak of something, and establishing what requires mass spectrometry.
It also says nothing about the non-peptide mass — salt and water are essentially invisible at peptide detection wavelengths, which is why net peptide content is a separate measurement.