Vol. 3, No. 6 — June 2026Independent since 2024

TheCompound Journal

Reporting on incretins, compounding & the peptide supply chain

A monthly journal of record.
30 issues · 32 contributors
Not medical advice. We sell nothing.

Reference intervals

Why your laboratory interval differs from the one in the textbook

Two sources of noise sit under every number: how reproducible the assay is, and how much the analyte varies within the same person on the same day.

Correction

An earlier version reported that serum calcitonin monitoring is recommended for patients taking this drug class. It is not recommended for that purpose, and the sentence has been removed.

The Journal reports laboratory findings from the trials constantly and has come to regard the interpretation of an individual panel as the most consistently mishandled subject in this whole field. The reason is structural rather than educational: the printout gives a number, an interval and a flag, and gives no imprecision estimate, no within-person variation figure and no reference change value. Everything required to interpret the result correctly is omitted from the document that reports it.

What a reference interval actually is

A reference interval is an empirical statement about a population. A laboratory recruits a reference group meeting defined health criteria, measures the analyte, and reports the central ninety-five per cent of the resulting distribution, usually as the 2.5th to 97.5th percentiles. Everything about that construction has consequences. The interval is specific to the assay and platform used to derive it. It is specific to the reference population — its age structure, sex distribution, ethnicity and, for some analytes, its diet and altitude. And it deliberately excludes one in twenty healthy people at each end by design.

Two further points follow. The interval is not a target: for several analytes the optimal value on outcome grounds sits well inside it or below it, and low-density lipoprotein cholesterol is the standard example. And it is not a diagnostic threshold: decision limits, which are what clinical guidelines actually use, are derived from outcome data rather than from a healthy distribution, which is why the diagnostic cut-off for diabetes is not the upper limit of a reference interval.

Laboratories that report both a reference interval and a decision limit are doing the reader a service. Most report one number and one flag, and leave the distinction to be inferred.

The arithmetic of the comprehensive panel

If an analyte’s reference interval excludes five per cent of healthy people, and if analytes were independent, the probability that a healthy person produces at least one flagged result on a panel of n analytes is one minus 0.95 to the power of n. For a twelve-analyte panel that is approximately 46 per cent. For twenty analytes, approximately 64 per cent. For a thirty-analyte comprehensive panel with lipids and thyroid included, approximately 79 per cent.

Analytes are not independent — electrolytes covary, liver enzymes covary, so the true figures are somewhat lower — but the direction and rough magnitude hold. The practical implication is uncomfortable and rarely stated: on a comprehensive panel, the flagged result is the normal outcome, and treating each flag as requiring explanation is a commitment to explaining noise.

This is the strongest single argument for ordering panels against questions rather than by habit. A panel assembled because each analyte answers something the clinician wants to know produces flags that mean something. A panel assembled because it comes as a bundle produces a document in which the interesting result, if there is one, is hidden among four uninteresting ones. The Journal makes this point in a publication whose readers frequently order their own panels privately, and it applies with more force there rather than less.

Everything required to interpret a laboratory result correctly is omitted from the document that reports it.

Perpetua Nwachukwu, Contributing Writer, Laboratory Medicine

The reference change value, worked

Two results in the same person differ for three reasons: the analyte genuinely changed, the assay is imprecise, and the analyte varies within the person from day to day. The last two are quantified in the biological variation literature as the analytical coefficient of variation and the within-subject coefficient of variation, and databases of the latter have been maintained for decades.1

The reference change value combines them: approximately 2.77 times the square root of the sum of their squares, for a two-sided ninety-five per cent probability that a difference is real. The results are instructive. Sodium, with tiny biological variation, has a reference change value of around three per cent. Creatinine is about fourteen per cent. Alanine aminotransferase, with within-subject variation above twenty per cent, requires something like a sixty per cent change. Triglycerides, more variable still, require more.

Apply that to a routine monitoring situation. An ALT moving from 28 to 41 units per litre — a rise of forty-six per cent that crosses no threshold and is unlikely to be flagged — sits inside the reference change value and may be nothing at all. An ALT moving from 28 to 62 has moved. Nothing on the report distinguishes the two cases, and the distinction is the entire question.

Which movements are findings and which are consequences of the weight change
AnalyteDirection during rapid lossPrincipal reasonFinding or artefact?
Serum creatinineFallsReduced muscle massArtefact of composition
eGFR (creatinine-based)RisesFollows creatinineArtefact of composition
Alanine aminotransferaseFallsReduced hepatic fatFinding
TriglyceridesFallImproved insulin sensitivityFinding
LDL cholesterolFalls slightlyWeight lossFinding, small
Lipoprotein(a)Little changeLargely geneticNeither
Free triiodothyronineFallsEnergy restriction adaptationArtefact of deficit
C-reactive proteinFallsReduced adipose inflammationFinding
FerritinFallsBoth inflammation and iron storesAmbiguous
25-hydroxyvitamin DRisesSmaller distribution volumeArtefact of composition
Lipase, amylaseRise modestlyDrug class effectFinding of unclear significance
Directions are typical rather than universal. The classification is the Journal’s own and is offered as an interpretive aid, not as a clinical rule.

The alanine aminotransferase interval is too wide

Most clinical laboratories report an upper limit of normal for alanine aminotransferase somewhere between about 40 and 55 units per litre, with a modest sex difference or none. Those intervals were derived from reference populations that were screened for viral hepatitis and heavy alcohol use but not, in most cases, for hepatic steatosis — which was neither commonly diagnosed nor considered when many of the intervals were established.

Work redefining the healthy range in a large population of prospective blood donors, screened for viral markers, alcohol intake and metabolic risk factors, arrived at substantially lower limits: in the region of 30 units per litre for men and around 19 for women.2 Those figures have been influential in hepatology and have largely not propagated into general laboratory reporting.

The consequence for this population is direct. A person starting treatment with an ALT of 44 has a flagged result by a strict standard and an unflagged one by their laboratory interval; a fall to 31 during treatment represents normalisation by one standard and continued abnormality by the other. Neither reading is wrong. The Journal reports ALT against both where it can, and regards a laboratory report giving only the wider interval as incomplete rather than incorrect.

Against measuring too often

Frequent monitoring in a person doing well is a reliable generator of work. Each comprehensive panel carries a substantial probability of at least one flagged result; the flags are mostly noise; each requires explanation, repetition or investigation; and the cumulative effect over a year of monthly panels is several investigations and no additional information about the person.

There is also a specific problem with monitoring an analyte more frequently than its own window. HbA1c integrates three months. Measuring it monthly produces overlapping windows in which two-thirds of the data is shared between consecutive results, so the apparent trend is smoother than the underlying glycaemia and the independent information per measurement is low. The trials in this class scheduled it quarterly for exactly this reason.

The counter-argument deserves stating fairly: monitoring during escalation, when tolerability problems and their metabolic consequences are most likely, is a different proposition from monitoring during stable maintenance, and the case for closer observation in the first three months is reasonable. What the Journal has not seen is any evidence that a fixed frequent schedule during maintenance detects anything that a symptom-prompted panel would miss. Readers who know of such evidence should write to standards@compoundjournal.com.

How the Journal reports a laboratory figure

Five things accompany a laboratory number in these pages. The units, because international and conventional units differ for several analytes and the same value means different things in each. The reference interval used, with a note where the interval is contested, as it is for alanine aminotransferase. The baseline, because a change of 1.8 percentage points in HbA1c from a starting value of 8.3 is a different claim from the same change from 9.5. The estimand where the figure comes from a trial. And the reference change value where we are discussing an individual delta rather than a group mean.

We also state the assay method where it matters, which is more often than one would like: HbA1c in the presence of a haemoglobin variant, thyroid function in the presence of interfering antibodies, and creatinine measured by enzymatic against Jaffe methods all behave differently, and a comparison across methods is not a comparison.

This is a heavier apparatus than most publications carry and it exists because the alternative, in our experience, is a stream of technically accurate figures that lead readers to conclusions the data does not support. Errors in this apparatus should be reported to standards@compoundjournal.com; the correction log records what came of each one.

What this department is for

The Laboratory Notebook reports what tests measure, how they behave, and what has been found using them. It does not recommend monitoring schedules, interpret readers’ results, or advise on treatment. A laboratory result belongs in a conversation with a clinician who has the rest of the picture, and this publication is emphatically not that conversation.

Two standing notes. Several compounds discussed in these pages are sold for research use only and are not approved for human use in any jurisdiction; the Journal reports on them as commodities and as analytical problems, not as therapies. And where we describe what the pivotal trials monitored, that is reporting on trial protocols and not a template anybody should adopt from a magazine.

Correspondence is welcome at letters@compoundjournal.com. The Journal receives a steady flow of letters containing readers’ own panel results with a request for interpretation, and we do not provide it — not from caution but because a panel without a history, an examination and a reason for ordering it cannot be interpreted by anybody, including us.

The next instalment in this department takes up the analytical side of the same coin: not what a clinical laboratory measures in a person, but what an independent testing service measures in a vial, and why the two documents look more alike than they are. That is a different set of instruments and a different set of failure modes, and it is covered in Analytics.

References

  1. Ricós C, Alvarez V, Cava F, et al. “Current databases on biological variation: pros, cons and progress.” Scandinavian Journal of Clinical and Laboratory Investigation. 1999;59(7):491–500.
  2. Prati D, Taioli E, Zanella A, et al. “Updated Definitions of Healthy Ranges for Serum Alanine Aminotransferase Levels.” Annals of Internal Medicine. 2002;137(1):1–10.

Letters to the Editor

4 printed

Selected from correspondence received on this article. Writers are identified by initial, surname and city, verified before printing. Replies are from the desk that filed the piece or from the standards editor. Write to letters@compoundjournal.com.

Your piece assumes readers have a clinician ordering panels for them. A great many of us are buying our own, from services that supply an interval, a flag and nothing else. The apparatus you describe is not available to us at all.

T. Elorriaga, San Sebastián

The Journal replies

It is available, with effort: published biological variation databases are open, the reference change value is one line of arithmetic, and your own previous result is the comparator that does most of the work. But you are right that the market supplying these panels has no incentive to include any of it, and that is worth stating.

Correction to your HbA1c section. You write that chronic kidney disease lowers HbA1c through shortened erythrocyte survival. It can also raise measured values on some assay platforms through carbamylated haemoglobin interference. The net direction depends on the method.

N. Halvorsen, Trondheim

The Journal replies

Correct, and the omission was ours. The paragraph now says so, and it is a good example of why the assay method belongs alongside the value.

I stopped treatment fifteen weeks ago and my panel is worse than I expected. I had a panel at five weeks that looked fine and I had assumed I had escaped. Your point about the twelve-week timing was the explanation nobody offered me.

H. Baptiste, Fort-de-France

My eGFR has risen from 71 to 84 over fourteen months of treatment and I have lost twenty-six kilograms. My prescriber described this as the drug protecting my kidneys. Having read your creatinine section, I suspect it is mostly that I have less muscle. Which of us is right?

R. Hollenbeck, Spokane, WA

The Journal replies

On the information given, probably you, at least in part. A rise of that size during weight loss of that magnitude is well within what reduced creatinine production can produce. A cystatin C-based estimate alongside the creatinine one would separate the two, and is the measurement worth asking for. It is also possible both things are happening.

Related coverage