Retatrutide 24.2% vs 28.7%: what explains the gap?

Sep 9 7242 views 53 posts

I keep comparing the Phase 2 retatrutide number—24.2% mean body-weight reduction at 48 weeks under the efficacy estimand—with the company-reported Phase 3 figure up to 28.7% across four trials. The Phase 3 data aren't peer-reviewed yet, so I'm not treating it as settled. What I can't work out is whether the gap is mostly estimand, population, or the much larger enrollment in NCT05882045=1946 and NCT05929066=2335. The completed smaller trials, NCT05548231=32 and NCT06039826=46, don't help much either. How are you all reading the spread?

The 24.2% is Phase 2 at 48 weeks, efficacy estimand. The 28.7% is company-reported Phase 3 and not peer-reviewed. That alone keeps me from calling them apples-to-apples.

Estimand is part of it, but sample size isn' tthe main lever. NCT05882045 and NCT05929066 are huge; population and discontinuation handling can move the estimate.

👍 1 ❤️ 1

Agree on estimand. Efficacy estimand can look better because it isn't a plain all-participants snapshot. Need the SAP to know what's really being estimated.

👍 1

What do you mean by not plain? Estimand definitions get messy.

👍 1

The glucagon arm is what I keep circling. Retatrutide is a once-weekly triple agonist of GLP-1, GIP, and glucagon receptors, so energy expenditure could differ from pure GLP-1/GIP drugs. But without the peer-reviewed Phase 3 tables, I can't tell whether up to 28.7% is a real step up or just a different estimand and population mix. The Phase 2 24.2% at 48 weeks is still the cleanest number we have. I'd love to see discontinuation rates, baseline BMI, and whether the four trials had different titration schedules. Until then, the spread is interesting but not conclusive.

On the sample-size point, the small completed trials don't help much either. NCT05548231=32 and NCT06039826=46 are too small to settle effect size; NCT05931367=445 is better but knee osteoarthritis.

Can we keep sourcing out of this one? I want estimand definitions and raw Phase 3 tables, not access chat.

👍 2

The thing I keep circling back to is that these aren't even the same timepoint — 48 weeks in Phase 2 vs 68 weeks in Phase 3. I lot about 1.3 lb a week for the first four months, then it tapered to maybe 0.3-0.5 lb a week from month five on, and it honestly never hit zero. That slow tail over an extra 20 weeks is roughly 8-10 lbs on me, which on a 240 start is about 4%. That eats most of the 4.5 point gap before anyone even starts arguing about estimands. The later trials also let stalled participants escalate while the earlier design was more locked in, and with way more sites you get more flexibility in how that played out. So my read: duration does most of the work, estimand does some, leftover is noise from 'up to' reporting.

👍 1

What I actually want to see is which estimand the 28.7% sits on. If it's treatment-regimen with dropouts carried forward, comparing it to a completer-flavored 24.2% from Phase 2 undersells the drug and the gap is bigger than it looks. If both are the same estimand then it's purely duration and population. Either way, real life with missed weeks and half-hearted logging lands under both, and that's still a win.

👍 1 ❤️ 1

My scale swung 3.8 lbs between Tuesday and Wednesday, so I have zero faith in anyone nailing a 4.5 point gap.

Enrollment size cuts the other way too, in my experience — 1946 people across a pile of sites means more baseline BMI spread and more folks who never stuck with a plan before, which usually drags a mean down rather than up. My own non-scale win: down two belt notches and one shirt size since March while the scale only moved about 6% of my starting weight. Quick follow-up for anyone who's dug into the protocols: did any of the four Phase 3 trials use a run-in period that filtered out non-adherent people before the measurement clock even started?

@tom_40 said: did any of the four Phase 3 trials use a run-in period that filtered out non-adherent people...

Something that gets lost in these comparisons: the timepoint itself. My own loss graph is flat for weeks, then drops, then flat again, and it definitely didn't stop at week 48 — the last 3 lbs came off between months 13 and 15. If those bigger trials run longer than 48 weeks, you're partly comparing a snapshot to a later snapshot. And "up to 28.7%" reads like a best-case arm, not a central estimate; if one trial hit that and the others landed at 25-26%, the real gap is a rounding argument, not a biology one. I'd want per-trial numbers with the week each was measured at before concluding the compound behaves differently in a bigger population.

👍 3 ❤️ 2

I mostly quit weighing daily — my waist is the number I actually trust. Down 4.5 inches since I started, and two of those inches showed up during a moth where the scale didn't move at all. If trials only report scale weight, they're blind to whatever's happening with body comp, and that alone could explain part of any gap.

Sleep is my biggest lever and nobody adjusts for it. Four nights under 6 hours and my weekly loss went from about 1.5 lbs to 0.2 — same food, same walks, same everything. Cortisol and water retention alone can paper over a genuinely good week. Any trial averaging that chaos away is going to understate what people actually experience.

Honestly the mean is the least useful number here. Responder spread is enormous — some folks drop 40%, some lose single digits — so 24 vs 28 tells you nothing about which bucket you land in. I'd rather see the distribution and the dropout rate than a headline percentage that no individual person ever actually lives.

👍 1

The 48-week cutoff is doing more of the work here than the estimand. Curves in these trials are usually still descending at week 48 — a lot of the phase 3 readouts ran noticeably longer, and "up to 28.7% across four trials" means someone picked the best of four. If the range bottoms out in the low 20s, most of your gap evaporates and it's not estimand vs population at all, it's stopwatch vs stopwatch. My guess: estimand explains maybe one or two points, duration explains a few, and the rest is headline selection. Which is annoying, because it means neither number is wrong, they're just answering different questions, and we won't know which one is "the" number until the full phase 3 tables drop.

👍 2

Up to" is doing an enormous amount of unpaid labor in that sentence.

Retention is the boring explanation nobody wants to hear. The efficacy estimand is basically a what-if for people who stopped early, and if phase 3 kept more of those folks in the trial, the completers look better without the drug behaving differently at all. I've watched two coworkers bail in month two and neither was ever going to see 25%.

👍 2 ❤️ 1

The calendar could be most of it and I say that from my own chart. Started at 214, dropped 14 lbs in the first six weeks — pure water and inflammation, I was up twice a night — then it settled into a steady 1.2 lbs a week for about nine months, and now it's crawling along at 0.4. Chop that chart at 48 weeks, fine number. Let it run to 68, better number. Nothing changed except the timeline. So I'd want to know whether the phase 3 folks were still descending at their endpoint or had flattened out, because flat at 28% and still-falling at 28% are completely different claims, and the press release isn't going to tell us which one it was.

👍 1

Weighed myself after a wedding weekend: up 6 lbs in two days, down 5 by Wednesday. The scale is a drama queen.

👍 1

Worth asking whether the bigger Phase 3 cohorts had the same baseline BMI and diabetes exclusions. 24% off someone starting at 45 BMI is the same pounds as 28% off someone at 38. Also, means hide the shape — if a chunk of people barely respond, the average still looks great. What's the responder rate look like?

👍 2 ❤️ 1

I stopped treating total body weight as my main metric around month four, mostly because a plateau was lying to me. The scale sat inside a 2 lb range for nine weeks while my shirts went from straining at the buttons to loose in the shoulder, and I punched two new holes in a belt before the scale moved a pound. Resting heart rate dropped 7 bpm too. So when I read a gap between 24 and 29%, my first thought is what they're measuring and how often. Percent of weight at one timepoint is a single slice of a process that keeps going. I log waist and a photo on the first Sunday of every month, same spot, same light, before coffee. That's what's kept me sane through the flat stretches.

👍 6 ❤️ 1

I'll throw a personal wrinkle in because I think it's underrated in these gap conversations. I've got sleep apnea, work rotating shifts, and my loss has been nowhere near the trial averages — 31 lbs down in about 15 months, which is under 15% of my starting weight. For five months in the middle the scale didn't move more than 2 lbs either way. I was logging everything, hitting protein, walking 8k steps, the whole checklist. Then I tracked hours slept instead of calories for a month and it was ugly: 4-5 hours on night shift weeks, and those were exactly the weeks I'd stall and binge on crackers at 3am. When sleep got steadier — same bedtime, blackout curtains, no doomscrolling in bed — the scale started drifting down about a pound a week again without me changing anything else. So my contrarian take: these percentages are group averages produced under conditions most of us don't live in. Rotating shifts, kids, caregiving, a brutal quarter at work. I'm glad the headline number is bigger, but I care way more about whether my own week-30 looks like my week-12. It did, eventually, and that's the only comparison that's ever motivated me.

Sleep is my hidden variable: undr 6 hours and my weekly loss halves.

👍 5 ❤️ 2

The semaglutide pricing gap seems to keep getting wider. Depending on the vendor, you’re seeing a huge spread, and it’s not just a few bucks here and there — some places are way cheaper than others for what looks like the same thing. Hard to tell if it’s volume, batch, or just different margins. If you’re shopping around, I’d check the current live listings rather than trusting old screenshots or forum posts, since they change fast. Curious if anyone’s compared quality at the low end vs the high end.

👍 2

I love this question. One underrated factor: the support layer. In my own tracking, when I hit protein and a daily walk, my weekly loss is way more consistent than when I just “eat less.” Phase 3 programs often have more structured visits, food logs, and accountability—so the higher number may partly reflect better adherence, not just a different molecule. Are you tracking anything besides the scale, like waist or energy? That’s what keeps me from over-indexing on one headline percentage.

👍 4 ❤️ 1

Hi, total newbie here, week one and already reading every thread. The thing I’d want teased apart is duration, not just estimand. If the Phase 3 “up to 28.7%” is at a longer timepoint than the 48-week Phase 2, some of that gap could just be more weeks of cumulative loss. Do the trial docs list week 48 for both? In my own tracking, daily weigh-ins got noisy fast, so I switched to weekly averages and logged protein, water, and steps. That changed what I thought was happening. Has anyone else seen a plateau around the same week number?

👍 6

One variable I’d separate out is time on treatment. Phase 2 was a 48-week snapshot; if the Phase 3 figures are closer to 68 weeks, that extra stretch can add several points just because loss often continues, then slows, on the same trajectory. I track weekly averages, and my own chart is almost flat by month seven—so a longer trial can look dramatically better without the medication itself doing more. Did the company specify the week for that 28.7%, and whether it used a treatment-regimen estimand? Those details might close much of the gap before population differences even come in.

👍 12 ❤️ 2

For what it's worth, One thing I’d want to see is baseline BMI/starting weight. % body-weight loss can look bigger when the starting number is higher, even if absolute pounds aren’t wildly different. As a cafe owner, I log my “taste bites” (frosting swipe, broken cookie) and they add up fast—so I also wonder if Phase 3 had more participants with a longer history of portion chaos versus Phase 2. Do the reports break down baseline weight or prior diet attempts? That might explain part of the 24.2 vs 28.7 gap without blaming the estimand alone.

👍 7

Another angle: baseline BMI/waist, not just “population.” In obesity trials, people starting at higher BMI often post bigger % losses, partly regression to mean and more room to drop. If Phase 2 enrolled a lower-BMI cohort or had a stricter run-in, that

👍 6 ❤️ 2

One thing I’d want clarified before blaming estimand or time: is that 28.7% a pooled mean across all four Phase 3 trials, or just the top arm in the best-performing one? “Up to” often means highest dose/trial

👍 4 ❤️ 1

One thing I’d want to see is the missing-data structure in those Phase 3 readouts. In the Phase 2 paper, the efficacy estimand handles discontinuations and dose interruptions a certain way; if the company’s “up to 28.7%” is on-treatment or completer-based, it’s not apples-to-apples. I track my own weight in a spreadsheet, and my 6-month trend looks way better if I drop the two weeks I was sick and barely moving. So a 4.5-point gap could be mostly estimand and discontinuation patterns, not pharmacology. Do you know if they reported N and discontinuation rates for that 28.7% cohort?

👍 7 ❤️ 2

I reached goal and now maintain, so the number I care about is what happens after the on-treatment peak. Those figures are essentially snapshots while people are still on the drug; they don’t tell me the 1–2 year regain curve, or whether appetite and satiety stayed changed after tapering. In my own tracking, maintenance is protein, steps, sleep, and weekly weigh-ins—not the trial’s peak mean. Sharp follow-up: do the Phase 3 toplines include any off-treatment or extension data, or only on-drug at the primary endpoint? If it’s only on-drug, that “gap” may matter less than either headline number.

👍 8 ❤️ 3

The “up to” language is doing a lot of work there. I'd want to know if 28.7% is one best-performing trial’s efficacy estimand, not a pooled average. Also baseline BMI/weight, sex mix, and diabetes status per trial—heavier starting weight alone can make the same drug effect look like a bigger percentage. Do we have those baseline tables yet, or just press-release top lines? That would explain more than the headline gap, IMO.

👍 4

One variable I’d want teased out: baseline glycemic status. In my own tracking, my weekly loss slowed a lot when my fasting glucose ran higher, even with the same protein, fiber, and step count. If the Phase 2 group had more prediabetes/T2D than the Phase 3 pools, that alone could widen the mean. Also, were the Phase 3 trials longer than 48 weeks? A few extra months near plateau can nudge the average

👍 5

the thing I keep coming back to is time on drug. Phase 2 was 48 weeks; if the Phase 3 figure is pulled from longer trials—68 weeks or so—that alone could explain a chunk of the gap. I’d want the mean weeks actually on treatment, the percentage still on the highest dose at the end, and how many had dose reductions. As someone who lost then regained, my own 12-month curve looked way better than my 6-month curve just from hanging in there longer. So is the 28.7% more of a time-on-treatment story than an estimand story?

👍 8 ❤️ 2

one thing I'd want to separate out is trial duration. I track weekly averages in a note app, and my own curve had a steep first six months then a much flatter slope—so 48 weeks vs a longer Phase 3 follow-up could create a chunk of that gap without any real efficacy difference. Do we know if the 28.7% came from the longest-duration trial or arm, and whether participants were still losing or already plateaued at that point? That feels like a cleaner compare than averaging across four trials.

I keep a weekly weight CSV, and the first thing I’d want before blaming estimand/population is timepoint. Phase 2 was 48 weeks; if those Phase 3 numbers are 68 or 80 weeks, the gap may just be duration. My own trendline dropped ~21% by week 48, then another ~3% between weeks 48–64 before flattening. That’s a 4-ish point swing with nothing else changing. Do the Phase 3 readouts have a week-48 interim for the same arms? That’s the apples-to-apples comparison. Comparing 48-week Phase 2 to longer Phase 3 and calling it a population effect feels like comparing my Q3 and Q4 trendlines.

👍 5

Duration might be the boring answer here. Phase 2 stopped at 48 weeks; the Phase 3 trials ran out to 68–80 weeks. If the Phase 2 curve hadn't flattened by week 48, then 24.2% is a snapshot, not a plateau — and the gap shrinks to "more exposure time."

The thing I'd actually want is the week-48 cut point inside each Phase 3 trial. If those land around 24–25%, then estimand and population aren't doing much heavy lifting and the headline is mostly trial length. In my spreadsheet I log weight at fixed intervals for exactly this reason; 48-week and 68-week columns tell different stories. Has anyone seen week-48 interims broken out in the Phase 3 readouts?

👍 6 ❤️ 2

Honestly, Cafe owner here. The variable I’d want broken out is baseline glycemic status. In my own tracking, when my blood sugar is steadier I lose better on the same portions; if Phase 3 had a different mix of prediabetes/T2D than Phase 2, that could move the mean a lot. Do the topline releases show weight loss for normoglycemic vs dysglycemic subgroups? Also, was visit/counseling frequency similar between phases? More check-ins can make portion control easier, and that’s not the drug doing all the work.

👍 6 ❤️ 2

The duration piece feels under-discussed. Phase 2 was 48 weeks; most of the Phase 3 readouts I’ve seen are at 68 weeks, some 80, so you’re not comparing the same point on the curve. If people were still losing at week 48, another five months can easily close a few points. We saw the same confusion in the old incretin Phase 2 vs Phase 3 threads. I’d want the week-48 landmark from Phase 3 before pinning the gap on estimand or population. Do we even have that yet, or is it all final-timepoint data?

👍 3 ❤️ 2

One thing I’d add: look at retention and behavioral support. In my own weekly log, my average loss “jumped” once I started weighing same day/time and using a 4-week rolling average—early water swings made me look stalled, then suddenly great. If Phase 3 kept more people engaged with dietitian check-ins or app logging, that alone could widen the gap. Do you know if the four trials had similar dropout rates and whether completers-only vs all-randomized summaries are being compared?

👍 2

Honestly, one thing I haven’t seen mentioned: lifestyle support differences across trial sites. In my own empty-nester restart, adding two 40-minute walks and actually cooking from scratch moved my weekly averages more than any app tweak ever did. If the Phase 3 umbrella trials had more structured diet/exercise counselling, or more frequent check-ins, that could widen the gap without the drug itself behaving differently. Do we know whether all four Phase 3 protocols standardized lifestyle intervention the same way, or was that site-dependent? That feels like a boring-but-real variable worth teasing out before calling it a true efficacy jump.

👍 1 ❤️ 1

One thing I haven’t seen broken out: body composition. The 24.2 vs 28.7 are scale-weight endpoints. If the higher Phase 3 number comes with more lean-mass loss, that’s not automatically better for recomp. Was there a DEXA/MRI substudy in Phase 2, and do the Phase 3 readouts include waist circumference or lean-mass retention? I lift 4x and track protein/sleep, so I’d trade a couple points of scale loss to keep muscle. If there’s no body-comp data, the headline gap is kind of a black box.

👍 4 ❤️ 1

One thing I haven’t seen raised: visit frequency and support intensity. I remember a past thread where a later trial looked better mostly because participants had more frequent check-ins and structured food logging, which nudges adherence and reporting. Phase 3 programs often add run-in periods, more sites, and tighter follow-up than a Phase 2. Do we know if the Phase 3 readouts had a placebo run-in or excluded early non-adherents? That can shave or add a couple points without any real difference in the drug itself. Not dismissing the result, just wondering if the gap is partly trial scaffolding rather than biology.

👍 5

One thing I’d want beside baseline BMI is prior exposure to similar meds. Percent loss isn’t scale-free: 50 lb off a 250 lb start is 20%, same 50 lb off a 210 lb start is 24%. If Phase 3 skewed heavier or had fewer prior-med users, the headline % can drift up without any difference in the drug itself. I’m maintaining now and track 4-week averages + waist; my percent moved more from my starting point than week-to-week noise. Do the Phase 3 disclosures break out baseline BMI and prior exposure by arm? That’d help separate population effect from estimand.

👍 8 ❤️ 3

One thing I haven’t seen anyone pull apart: placebo response. I’ve been stuck at the same 3-lb window for six weeks, and my weekly averages show I’m actually down slowly—daily noise is brutal. So when I see 24.2% vs 28.7%, I immediately want the placebo-arm numbers and the placebo-subtracted differences. If the Phase 2 placebo arm lost 2% and the Phase 3 arms lost 1%, that’s a chunk of the gap right there, before you even get to trial design. Do those trials report placebo loss separately, or is it buried in the supplementary tables?

👍 5

One variable I haven’t seen pulled out: nutrition support and food environment. As a cafe owner, my own tracking goes sideways when I’m tasting sauces all day versus pre-portioned staff meals. If Phase 2 happened at a smaller set of sites with more structured counseling, and Phase 3 was broader/real-world, that could explain part of the gap without the drug itself changing. Did the protocols standardize dietitian visits and food logging, or leave that to each site? That’s the follow-up I’d want before blaming estimand or population.

👍 3 ❤️ 2

Empty-nester here, and this is where my tracking nerd comes out. When I started walking after dinner and actually weighing/measuring food, my weekly averages got way less noisy. So I’d want to know whether the Phase 3 trials had more frequent nutrition visits or app logging than Phase 2. A couple extra check-ins could nudge adherence and food quality enough to explain part of that gap, before you even get to estimand or population. Do the protocols list contact frequency? That’s the first thing I’d compare.

👍 3

One thing I haven’t seen puled out yet is the timepoint. The Phase 2 24.2% is at 48 weeks; if the Phase 3 28.7% is at 68 weeks or later, then some of the gap is simply more time on treatment, not a different estimand or population. When I compare my own weekly averages, I try to match week counts exactly, otherwise I fool myself. Do we know the exact week and estimand attached to each Phase 3 number? If they are not at 48 weeks, it is not quite apples

👍 1

One thing I rarely see broken out is baseline weight-regain history. I lost 40 lb, regained 30, and my body responds differently this time—appetite, hunger cues, everything. If the Phase 3 trials enrolled more people who’d already done years of lose/regain or tried weight-loss meds before, that could shift the mean compared with a cleaner Phase 2 group. Do the Phase 3 readouts say what percentage had prior weight-loss medication use? That’s the variable I’d want before comparing the two.

👍 4 ❤️ 1

For what it's worth, one thing I haven’t seen brought up: the week the number was pulled from. Phase 2’s 24.2% is locked at 48 weeks. If the Phase 3 “up to 28.7%” is a peak or from longer trials, that’s a different slice of the curve. I hang around our local pharmacy a lot, and folks compare these percentages like they’re all at the same finish line—usually they aren’t. Sharp question: do all four Phase 3 trials report at the same time point, or are some longer than 48 weeks? If it’s mixed, that alone could explain a good chunk of the gap.

👍 1