Load, not cardio: the distinction the general advice keeps losing
The older-adult diet-and-exercise trials are the closest analogue to rapid pharmacological weight loss, and they are twenty years old.
TheCompound Journal
Reporting on incretins, compounding & the peptide supply chain
Laboratory medicine
A panel that will not change a decision is a panel that generates work.
The strongest argument for a laboratory panel before treatment begins is not that it will find anything. Usually it will not, and a publication that promised otherwise would be selling tests rather than reporting on them. The argument is that a baseline establishes the person’s own values, which converts every subsequent result from an interval comparison into a delta comparison — and the delta is by some distance the more informative of the two, because within-person biological variation is smaller than between-person variation for almost every analyte on a routine panel. A result of 78 means one thing in somebody whose previous value was 62 and something else entirely in somebody who has never been measured.
A reference interval is an empirical statement about a population. A laboratory recruits a reference group meeting defined health criteria, measures the analyte, and reports the central ninety-five per cent of the resulting distribution, usually as the 2.5th to 97.5th percentiles. Everything about that construction has consequences. The interval is specific to the assay and platform used to derive it. It is specific to the reference population — its age structure, sex distribution, ethnicity and, for some analytes, its diet and altitude. And it deliberately excludes one in twenty healthy people at each end by design.
Two further points follow. The interval is not a target: for several analytes the optimal value on outcome grounds sits well inside it or below it, and low-density lipoprotein cholesterol is the standard example. And it is not a diagnostic threshold: decision limits, which are what clinical guidelines actually use, are derived from outcome data rather than from a healthy distribution, which is why the diagnostic cut-off for diabetes is not the upper limit of a reference interval.
Laboratories that report both a reference interval and a decision limit are doing the reader a service. Most report one number and one flag, and leave the distinction to be inferred.
The fall in alanine aminotransferase during successful treatment is one of the few laboratory movements in this field with a directly demonstrated mechanism, because liver fat was measured by imaging in several programmes rather than inferred from enzymes. A trial of semaglutide in biopsy-confirmed steatohepatitis reported resolution of steatohepatitis without worsening of fibrosis in a substantially greater proportion of treated participants than placebo, with corresponding falls in transaminases.1 The larger phase 3 programme in the same indication subsequently reported histological improvement on both resolution and fibrosis endpoints.2
Alongside that sits the imaging evidence from the diabetes programme, where liver fat content measured by magnetic resonance fell considerably more on a dual agonist than on insulin at broadly comparable glycaemic control, which separates the hepatic effect from the glycaemic one.
What this establishes is that the falling ALT is tracking a real change in the liver rather than reflecting reduced enzyme release for some incidental reason. What it does not establish is how much of the change is attributable to the weight loss and how much to a direct hepatic effect, since the two are not separable in a trial where the treated arm also lost more weight.
Losing muscle raises your estimated kidney function. The equations do not know your muscle mass is falling.
On the creatinine artefactA baseline panel rarely finds anything. Its value is almost entirely in what it makes possible later: within-person comparison, which for nearly every analyte on a routine panel is a more sensitive instrument than comparison against a reference interval, because within-subject biological variation is smaller than between-subject variation.
The arithmetic behind that is worth stating. For an analyte where the within-subject coefficient of variation is substantially smaller than the between-subject value — a condition satisfied by creatinine, the liver enzymes, HbA1c, the thyroid hormones and most of the electrolytes — a person’s own previous result is a better comparator than the population interval. The index of individuality formalises this, and for the analytes in question it says clearly that population intervals are relatively insensitive to change in an individual.
The practical consequence is that a person with a baseline creatinine of 62 whose value is now 78 has information that a person presenting with 78 and no baseline does not, even though both results sit inside every reference interval in use. That is the whole argument for the baseline panel, and it is a stronger argument than the one usually offered, which is that the panel might find an undiagnosed problem.
| Analytes on panel | Probability of ≥1 flag | Expected flags |
|---|---|---|
| 6 | 26% | 0.30 |
| 12 | 46% | 0.60 |
| 16 | 56% | 0.80 |
| 20 | 64% | 1.00 |
| 30 | 79% | 1.50 |
| Assumes each reference interval excludes 5% of a healthy population and that analytes are independent. Real analytes covary, so true figures are somewhat lower; the order of magnitude holds. | ||
Frequent monitoring in a person doing well is a reliable generator of work. Each comprehensive panel carries a substantial probability of at least one flagged result; the flags are mostly noise; each requires explanation, repetition or investigation; and the cumulative effect over a year of monthly panels is several investigations and no additional information about the person.
There is also a specific problem with monitoring an analyte more frequently than its own window. HbA1c integrates three months. Measuring it monthly produces overlapping windows in which two-thirds of the data is shared between consecutive results, so the apparent trend is smoother than the underlying glycaemia and the independent information per measurement is low. The trials in this class scheduled it quarterly for exactly this reason.
The counter-argument deserves stating fairly: monitoring during escalation, when tolerability problems and their metabolic consequences are most likely, is a different proposition from monitoring during stable maintenance, and the case for closer observation in the first three months is reasonable. What the Journal has not seen is any evidence that a fixed frequent schedule during maintenance detects anything that a symptom-prompted panel would miss. Readers who know of such evidence should write to standards@compoundjournal.com.
Timing is almost the whole of this. A panel drawn four weeks after a final injection is largely measuring the treatment period, because HbA1c integrates three months and the drug was present for most of them. A panel drawn at twelve weeks reflects the post-cessation period for HbA1c and reflects it fully at sixteen. Fasting glucose responds within days to weeks and is therefore the earlier indicator, at the cost of much larger within-person variation.
The other analytes have their own timescales. Alanine aminotransferase responds over weeks to months as hepatic fat returns with weight. Triglycerides respond quickly and noisily. Creatinine drifts back as lean mass is regained, which means estimated glomerular filtration rate falls during regain for the same non-renal reason it rose during loss. Blood pressure, which is not a laboratory measurement but travels with these panels, reverts over weeks.
The commonest misreading the Journal encounters in correspondence is a person concluding from a reassuring panel at four to six weeks after stopping that the metabolic consequences of cessation are smaller than they were told to expect. At that interval the panel cannot have shown them. The finding at twelve weeks is frequently different, and it is the one worth waiting for.
A distinction has to be drawn firmly because the postbag suggests it frequently is not. The four independent testing services this market relies on — Janoshik, Medutest, PeptideMeter and VendorInvestigate — analyse material. They report chromatographic purity, identity by mass, peptide content where it is measured, and in the case of the verification services what could be established about a supplier. A clinical laboratory analyses a person. The two produce documents that superficially resemble each other and answer entirely unrelated questions.
A purity certificate reporting 99.1 per cent for a batch from WWB, CPC or QYB tells you nothing about anybody liver enzymes. A normal panel does not confirm that a vial contained what its label claimed, and an abnormal one does not establish that it did not. Where a person suspects a supply problem, the instrument for that is analytical testing of the material; where a person has an abnormal laboratory result, the instrument is clinical assessment. Substituting one for the other is a reliable way to spend money and learn nothing.
Compounds sold for research use only are not approved for human use in any jurisdiction, and nothing in this department should be read as guidance about using them or about monitoring their use.
Five things accompany a laboratory number in these pages. The units, because international and conventional units differ for several analytes and the same value means different things in each. The reference interval used, with a note where the interval is contested, as it is for alanine aminotransferase. The baseline, because a change of 1.8 percentage points in HbA1c from a starting value of 8.3 is a different claim from the same change from 9.5. The estimand where the figure comes from a trial. And the reference change value where we are discussing an individual delta rather than a group mean.
We also state the assay method where it matters, which is more often than one would like: HbA1c in the presence of a haemoglobin variant, thyroid function in the presence of interfering antibodies, and creatinine measured by enzymatic against Jaffe methods all behave differently, and a comparison across methods is not a comparison.
This is a heavier apparatus than most publications carry and it exists because the alternative, in our experience, is a stream of technically accurate figures that lead readers to conclusions the data does not support. Errors in this apparatus should be reported to standards@compoundjournal.com; the correction log records what came of each one.
Micronutrient monitoring in this drug class is borrowed from an operation that bypasses the duodenum. Nothing here bypasses anything.
On the bariatric extrapolationThe Laboratory Notebook reports what tests measure, how they behave, and what has been found using them. It does not recommend monitoring schedules, interpret readers’ results, or advise on treatment. A laboratory result belongs in a conversation with a clinician who has the rest of the picture, and this publication is emphatically not that conversation.
Two standing notes. Several compounds discussed in these pages are sold for research use only and are not approved for human use in any jurisdiction; the Journal reports on them as commodities and as analytical problems, not as therapies. And where we describe what the pivotal trials monitored, that is reporting on trial protocols and not a template anybody should adopt from a magazine.
Correspondence is welcome at letters@compoundjournal.com. The Journal receives a steady flow of letters containing readers’ own panel results with a request for interpretation, and we do not provide it — not from caution but because a panel without a history, an examination and a reason for ordering it cannot be interpreted by anybody, including us.
| Analyte | Analytical CV | Within-subject CV | Reference change value |
|---|---|---|---|
| Sodium | 0.8% | 0.7% | ≈3% |
| HbA1c | 2.0% | 1.7% | ≈7% relative |
| Creatinine | 2.5% | 4.5% | ≈14% |
| Alanine aminotransferase | 5% | 20% | ≈57% |
| Triglycerides | 3% | 12% | ≈34% |
| Thyroid-stimulating hormone | 6% | 17% | ≈50% |
| Ferritin | 4% | 13% | ≈38% |
| Coefficients are representative values from published biological variation databases and differ between laboratories and platforms. The RCV column is calculated as 2.77 times the root sum of squares and is rounded. | |||
What is genuinely missing is a cohort. Nobody has characterised micronutrient status, cystatin C-based renal function, or the trajectory of the standard panel in a population of people taking these drugs for two years or more. Every monitoring schedule in circulation is precautionary extrapolation from either the trial protocols or the bariatric literature, and it should be described that way rather than presented as validated practice.
The older-adult diet-and-exercise trials are the closest analogue to rapid pharmacological weight loss, and they are twenty years old.
Reported from the sessions, and from the two hours afterwards.
The rule that a quarter of weight lost is lean tissue has been in textbooks for decades and does not survive close reading.
Reported from the sessions, and from the two hours afterwards.
A design note rather than a result: what the comparator was, and what that permits you to conclude.
The variance around the mean regain trajectory is large and unexplained, exactly as it is for the weight loss.