Anyone else stuck on retatrutide's 28.7% vs 24.2% gap?

Sep 28 6500 views 54 posts

I've been digging through retatrutide trial data and keep hitting a gap. Phase 2 obesity showed 24.2% mean weight reduction at 48 weeks under the efficacy estimand. Company-reported Phase 3 results from four trials, not peer-reviewed, say up to 28.7%. CTgov lists completed trials with n=1946 and n=2335, plus smaller ones n=32, n=445, n=46 in different populations. Is the Phase 3 number a longer-duration effect, different estimand, or population mix? I'd like to compare protocols without guessing.

Baseline weight matters. Percent loss can look bigger in heavier cohorts.

👍 1

The 'up to' wording is doing a lot of work. Is 28.7% from the primary endpoint or a subgroup? I have not seen that clarified.

Peer review will settle more than forum math.

👍 1

I think the estimand debate is where this gets interesting. Efficacy estimand can look very different from treatment policy if dropout is uneven. If the Phase 3 programs have more discontinuations, the company-reported 'up to' number might reflect a cleaner completers-ish picture. That does not make it fake, but it does mean we cannot line it up against 24.2% at 48 weeks without the protocol and missing-data plan. The small completed trials n=32, n=445, n=46 may never publish weight outcomes either. That is the part I would want to see before trusting any cross-trial comparison.

The big CTgov enrollments, 1946 and 2335, could actually make an 'up to' result less comparable. Four trials pooled is not one trial.

👍 1

Duration might be the simpler answer. 48 weeks in Phase 2; if the Phase 3 trials ran longer, more time could move the mean.

👍 1 ❤️ 1

Fair, but Phase 2 used teh efficacy estimand too. If Phase 3 does the same, the gap is not just press-release spin.

👍 1

The 28.7% is company-reported and not peer-reviewed, so comparing it straight to the 24.2% Phase 2 number is already shaky.

👍 2

The estimand footnote is doing a lot of heavy lifting — efficacy estimand basically tosses out everyone who quit early, and in a 48-week obesity trial that's not a small group. My clinic's real-world numbers look way closer to the lower end, because people pause for surgery, travel, nausea. I track weekly waist and a 7-day rolling scale average now; single weigh-ins lie to me constantly.

👍 1

Duration mismatch is my bet. 48 weeks is right around where the curve tends to flatten hard — steady losses through month 6, then weeks 28-40 often look like bouncing inside a 3 lb window before anything moves again. Another 20 weeks at that slower rate could easily add a few percent, so 24 vs 28 isn't necessarily a contradiction.

Honestly the 4.5% gap is smaller than my week-to-week water swings. I'd take either number.

👍 3

Hot take: the 4.5 point gap is the least interesting number in those datasets. What actually predicts my results is whether I hit protein and steps on the bad weeks, not which trial arm looks better on a slide. I'm 8 months in, down 41 lbs (212 to 171), waist 38 to just under 32, and every plateau I've broken came from boring stuff — 130g protein, 8k steps, lifting 3x a week, and actually sleeping 7 hours instead of scrolling until 1am. Weeks 20 through 36 I basically lost nothing and nearly quit. If you'd shown me a 28.7% headline then I'd have felt like a failure at 19%. Now I just look at the 4-week trend line.

👍 3

I keep coming back to baseline BMI. A cohort averaging around 41 sheds a bigger percentage for the same pounds than one sitting at 34, so percentage points aren't apples to apples across those five trials. I track my own as % of starting weight and it gets flattering fast — 18 lbs off 232 reads way better than 18 off 300. Absolute kg would tell you more.

👍 4 ❤️ 2

Press-release percentages don't meet a journal until someone's looking over their shoulder. I'll beleve 28.7 when it's in print.

👍 4 ❤️ 1

Weekly average or you're just measuring your feelings. Weigh every morning, log it, compare Sunday to Sunday only, because a single day swings a few pounds on nothing — bad sleep, a salty dinner, whatever. My own version of an "efficacy estimand" is the me that skips the two-week vacation where eating goes full chaos. Leave that stretch in and it drags every average down for a month.

👍 2

Those n=46 and n=32 arms are the ones I'd toss first. In a 46-person group, two early dropouts or two super-responders move the mean a couple points by themselves. Not a knock on anything, just arithmetic. I'd rther see per-arm breakdowns from the 1946-person trial than a headline top-line number.

👍 5 ❤️ 1

Baseline BMI is the sneaky one for me—if the Phase 3 pool averaged a higher starting weight, 4.5 points can be partly population math, not a better drug. I’m 5-6 and at 240 vs 190 the same 10% looks totally different on my frame. Also, dropouts: if 15% quit early and aren’t counted the same way, the mean gets rosier. I’d want the denominator before getting excited.

👍 5 ❤️ 1

Chasing trial percentages stops being the point once the week 52 numbers come in. Real-world version people report: scale down 17%, waist down 7 inches, shirts going L to M. The weird part is usually a 9-week stretch around month 5 where the scale barely moves while the belt keeps tightening. And when a Phase 3 reads 28.7%, if it came from a longer run or a group that stayed on treatment, that 4.5-point gap is probably time and adherence more than magic. Then there's the version that goes the other way — 22% down but the waist barely changed, which lines up with muscle coming off, so the tape measure tells on that one. The headline number ends up mattering a lot less than the waist measurement and how the training log looks.

👍 10 ❤️ 1

Honestly the 28.7% won’t mean much if your own week 24 is flat and your jeans fit the same.

👍 3

I’d trade a 28% headline for 5% kept off at year three, no contest.

👍 6

My waist dropped 1.5 inches the same week the scale jumped 2.1 lbs. I’d been doing two 20-minute walks after dinner and finally sleeping more, and my jeans got loose while the scale sulked. Weighing daily made me want to quit; now I log waist every Sunday and weght only Wednesday. That gap between trial percentages? Your own tape measure will tell you more than any headline.

👍 4 ❤️ 3

My Sunday-night scale was up 5.2 lbs and I hadn’t gained an ounce of fat. Friday takeout, salty fries, two glasses of wine, and I’d log Monday and spiral. What fixed it was boring: weigh Thursday and Sunday mornings, same conditions, then only trust the two-week average. I also started drinking 20 oz water before any restaurant meal, ordering the burger without fries, and walking 15 minutes after dinner. The “stall” at week 19 was really a weekend-sodium illusion; my waist was still down 0.5 inch that month. So when I see trial percentages, I think about how much of the gap is just people whose weekends looked like mine. Track the average, not the spike.

👍 4 ❤️ 1

My old jeans fit two weeks befre the scale budged. Bodies are ridiculous.

👍 8 ❤️ 3

Bad sleep week = 3 lb water bump and zero appetite control. Stress is the hidden variable.

👍 3

The boring-but-likely answer is duration. Phase 2’s 24.2% was at 48 weeks and the curves in the paper were still trending down, not flat. The company topline number is probably week 68. I remember a similar freakout in the old sema threads when STEP-1’s 68-week result looked better than earlier shorter data. Twenty extra weeks can add a few points even without any estimand/population wizardry. Did the release actually say 68 weeks? If so, comparing it to the 48-week efficacy number is apples-to-oranges. Still worth watching the peer-reviewed split, but I wouldn’t read 28.7 vs 24.2 as a mystery yet.

👍 7

Not the water-weight stuff—the data nerd in me wonders if you’re comparing the same estimand. Phase 2 “efficacy estimand” often leans on on-treatment data, while company Phase 3 toplines can use different windows, pooled arms, or baseline BMI mixes. There was a thread here last spring where a similar gap shrank once the ITT figure dropped. Do you know if the 28.7% is week 68/72 and from one trial, or pooled across those four? That’s the fork I’d chase before calling it a real efficacy gap.

👍 8 ❤️ 3

One thing I keep returning to is the time point and analysis population. Phase 2’s 24.2% is 48 weeks under an efficacy estimand; the 28.7% headline may be from a longer window or a completer-type analysis. If so, that could explain much of the gap without any new biology. Does the company release specify the week and population for each of the four trials? I track a 4-week rolling average for myself for the same reason—matching the comparison matters more than any single number.

👍 3

Long-timer here. The variable I’d want before blaming estimands is baseline BMI and diabetes status. I remember digging through the old semaglutide threads and the “outlier” numbers usually tracked with whether the cohort was mostly obesity-only vs mixed T2D, plus how much lean mass versus fat folks were carrying. If the Phase 3 registrational pools skew heavier at baseline, a higher ceiling isn’t shocking. Has anyone pulled the baseline characteristics tables for those bigger n=1946/2335 trials yet? I haven’t seen them linked, and that’s the first place I’d look, not the top-line percentage.

👍 4 ❤️ 1

I went down a different rabbit hole: baseline weight. Percent loss isn’t portable across populations. If Phase 2 mean baseline was, say, 105 kg and the Phase 3 obesity cohorts averaged 115 kg, the same absolute kg lost gives a noticeably bigger percent. I track weekly kg and % in a spreadsheet; when my start weight was higher, the same 9 kg drop showed up as ~2 percentage points more. Has anyone pulled baseline BMI/weight and absolute kg lost from the Phase 2 vs those Phase 3 arms? If kg lost are similar, the 28.7% gap may be mostly denominator. If kg lost differ, then it’s real.

👍 9 ❤️ 4

I’ve been thinking about baseline characteristics, not just duration. A 4.5-point gap sounds huge until you convert it: on my 215-lb start, 4.5% is ~10 lb—basically my entire 3-month plateau. If the Phase 3 cohorts skewed heavier or had more men/higher baseline BMI, the sme absolute loss looks bigger in percentage terms. Has anyone actually compared baseline BMI/waist or sex breakdown between the Phase 2 and those Phase 3 arms? I haven’t seen it in the press releases, and that’s the piece I’d want before assuming the drug itself is stronger.

👍 3

One thing I haven’t seen anyone mention: the lifestyle-support backbone. Phase 2 sites often run tighter dietitian check-ins; Phase 3 multinational trials can vary a lot by country in food environment, protein intake, and how often people get weighed. I’ve maintained for a while, and my weekly weigh-ins plus protein/step tracking moved my plateau more than any med tweak. If one dataset had more frequent in-person visits or a stronger counseling protocol, that alone could nudge mean loss a couple points. Does anyone know if the Phase 3 supplements report visit frequency or country-level results? That feels like a missing variable.

👍 4

One thing I’d pin down: is 28.7% a mean, or an “up to” from the best-performing arm/trial? Topline releases love quoting the highest week/

👍 5

Honestly, One thing I haven’t seen unpacked: how much of that 4.5-point gap is lean mass vs fat? Phase 2 often had DEXA substudies; Phase 3 toplines usually don’t. If the higher number comes with more lean loss, then for maintenance it’s not automatically better. I’m at goal now, and I track waist + lifting numbers, not just scale. Does anyone have body-comp data from the larger trials, or are we comparing scale-only endpoints and guessing?

👍 5 ❤️ 1

One thing I haven’t seen anybody dig into is the lifestyle scaffolding in those protocols. Phase 2 often runs at a handful of academic sites with pretty standardized counseling, while the bigger Phase 3 program may have more sites, different dietitian access, and varying retention. That can move scale weight a lot without any real drug difference. I’d love to know

👍 3 ❤️ 1

Something I haven’t seen mentioned: body-comp methodology. If one dataset used DEXA or adjusted for lean mass and the other is just scale weight, a 4.5-point gap could be partly water, glycogen, or lean tissue. I lift 4x and track waist; my scale swings 2–3 lbs from sodium/carbs alone. Do we know if the Phase 3 protocol standardized hydration, training, or protein intake? Also, were prior GLP-1 users excluded? Treatment-naive vs experienced could shift response a lot. Not dismissing the gap, but “mean weight reduction” needs a method and a denominator before it’s comparable.

👍 1

The other number I’d want is retention. If Phase 2 had a chunk of people dropping early for GI stuff, the efficacy math can make the average look better than what a clinic actually sees. Phase 3 press releases usually bury that until the full paper. Also prior GLP-1 exposure: if Phase 3 enrolled more never-treated folks, that can shift the mean independently of baseline weight. Do the CTgov entries list prior med washout or completion rates? If not, I’d treat 28.7 as provisional. I track numbers for fun, and that’s my first missing column.

👍 5

I keep a nerdy spreadsheet of readouts, and the missing column for me is discontinuation/adherence, not just baseline stuff. The Phase 2 48-week number and the “up to 28.7%” Phase 3 number likely aren’t the same timepoint, same trial, or same completers. Even under the same efficacy estimand, a trial that retains more people can show a bigger mean loss without the drug being inherently stronger. Do the company slides break out week-48, per-trial n, and discontinuation rates? I’d also want CIs — if they overlap, the 4.5-point gap may be less dramatic than it looks.

👍 7 ❤️ 2

Tbh Not a stats person, but has anyone looked at lifestyle support intensity? In my own tracking, my loss was way better when I had weekly check-ins and food logging—like 2+ lbs/week vs barely moving when I went solo. If Phase 3 baked in more frequent visits, dietitian calls, or a structured program than Phase 2, that could account for part of the gap without any estimand trickery. Do we know if the protocols were matched on that? I’d want to rule it out before assuming the drug just performs differently at scale.

👍 10 ❤️ 1

One thing I’d want pinned down: the timepoint. Phase 2’s 24.2% is at 48 weeks, but if that 28.7% comes from a later visit—68 or 72 weeks—it’s not really a head-to-head gap. I track my own weekly averages and my slope flatt

👍 3 ❤️ 2

Duration is the missing row in my spreadsheet. Phase 2’s 24.2% was 48 weeks; the company’s 28.7% is likely a later timepoint. My own weekly trendline shows ugly deceleration: slope went from ~0.22%/wk early to ~0.07%/wk by month 10. Extrapolating my first 48 weeks to 68 only adds ~2.5 points, not 4.5. So duration helps but doesn’t close the gap. Does anyone have the actual week-68 timepoint and completer rates for those Phase 3 arms? That denominator would settle more than another BMI argument.

👍 3

Do the company Phase 3 decks show completion/discontinuation rates by trial? That’s the boring variable I’d want before chalking the 4.5-point gap up to baseline stuff. If one trial kept way more people on treatment to the end, an efficacy estimand can look shinier—same drug, different denominator. Sharp follow-up: any rescue or rollover allowed? Also, as a permanent airport person, my hotel-scale weights swing 2–3 lb from sodium and crap sleep, so I’d love a “weeks lived out of a roller bag” subgroup. It wouldn’t settle the gap, but it’d explain why my own tracking graph looks like a heart monitor.

👍 3

Maybe I missed it—what week is the 28.7% actually tied to? Phase 2’s 24.2% is at 48 weeks; if the Phase 3 “up to” figure is a later timepoint, that’s not really a gap, just different finish lines. I live in hotels and airports, and my scale swings 1–2 lb based on sodium and sleep, so timing/window matters even outside trials. Anyone have the exact week and visit window for the Phase 3 number?

👍 4

Quick question: is that 28.7% the pooled mean across the four Phase 3 trials, or just the highest single-trial result? If it’s the latter, comparing it to a single Phase 2 mean is apples-to-oranges. I’ve seen this before—company PR picks the best arm, then everyone quotes it as the new baseline. What’s the range across those four trials?

👍 5 ❤️ 1

Is anyone else wondering if it’s partly a timepoint mismatch? Phase 2 was 48 weeks; if the 28.7% is from a later Phase 3 window (68 weeks? longer?), that gap shrinks a lot. I track my own weekly averages, and my 48-week trend vs 68-week trend looked like totally different results even when my weekly loss was pretty steady. Are the Phase 3 numbers from the same week mark? That’s the first thing I’d want pinned down before comparing.

👍 4

One thing I haven’t seen anyone mention: baseline weight/BMI. If the Phase 3 cohorts started heavier on average than Phase 2, the same absolute kg loss can show up as a bigger percentage. I got burned by this in my own log—my % loss looked amazing after I regained because my start weight was higher, but kg-wise it was basically flat. Do the Phase 3 decks list mean baseline weight by arm/trial? If so, I’d compare absolute kg lost alongside the 28.7 vs

👍 7 ❤️ 4

i’ve been rereading the fine print too. What I keep wondering isn’t the week or pooling — it’s which estimand the 28.7 is under. Efficacy estimand strips out discontinuers; treatment-regimen keeps them in. If the company number is treatment-regimen and still beats phase 2’s efficacy number, that’s genuinely different. If it’s efficacy, the gap may just be analysis population. I track waist and protein now because scale alone lied to me during my regain. Do the posted Phase 3 toplines actually state the estimand for that top-line percentage?

👍 6 ❤️ 1

One thing that never gets unpacked: baseline BMI and diabetes status. Phase 2 obesity trial was non-diabetic with mean BMI around 37; Phase 3 may have a different regional mix or higher starting weights. At those numbers, a 4.5% absolute gap can partly be population/curve-start stuff, not a better drug. I lift 4x and track weekly averages, waist, and strength; my scale weight stalled for a month while my waist dropped, so I care way more about lean-mass sub-studies than the headline %. Did any Phase 3 report body comp/DEXA?

👍 5 ❤️ 1

Are the Phase 3 28.7% figures efficacy estimand or treatment-regimen? The Phase 2 24.2% was efficacy, so if the topline only quotes that, it’s not the same denominator. I track waist-to-height and fasting insulin alongside scale weight—the % can look great while body composition shifts. Do the company slides show both estimands side by side?

👍 2

Didn’t we hash this out in the big Phase 3 thread last spring? The Phase 2 obesity readout was 48 weeks; the Phase 3 topline everyone quotes runs to 68. That’s 20 extra weeks on the same curve. In my own tracking spreadsheet, post-month-10 losses were tiy but not zero, and a few slow responders kept drifting down. Add different baseline weights and press-release math, and the 4.5-point gap gets less spooky. I’d still want the week-48 snapshot from the Phase 3 arms before calling it a true estimand gap.

👍 2

I keep a paper weight log in the kitchen drawer, and the completer-vs-everyone thing bit me when I compared my 6-month and 12-month numbers. My 12-month looked better only because the weeks I fell off weren’t in it. So my sharp question: for tat 28.7%, are they reporting all randomized with an efficacy estimand,

👍 1 ❤️ 1

The thing I’d want pinned down is estimand and discontinuation handling, not just trial size. Lilly’s 28.7% is “up to” across trials—was that efficacy estimand (on-treatment, no rescue?) or treatment-regimen estimand? If the Phase 3 had higher completion or different imputation, that alone could explain a few points. I log weekly weights and a 4-wk average, and missing-data assumptions matter way more to me than a single headline. Do the toplines list completion rates and whether rescue meds were allowed? If not, the 4.5-point gap still feels apples-to-oranges.

👍 1

One piece I have not seen discussed is the estimand and missing-data handling. Phase 2’s 24.2% was under an efficacy estimand, which can exclude data after certain intercurrent events. Company Phase 3 top-line may use a treatment-policy or completers-style analysis, so the 28.7% may not be directly comparable. Does anyone know whether the Phase 3 statistical analysis plans are public yet? On a personal note, I track weekly weights and only look at four-week rolling averages, because single weeks bounce around too much to interpret. I would rather compare like-for-like timepoints and analysis sets than headline percentages.

The thing I keep coming back to is dropout handling. Phase 2’s efficacy estimand can exclude people who stopped, while company Phase 3 headlines sometimes lean on completers or on-treatment numbers. Not saying that’s the whole 4.5-point gap, but that’s exactly the kind of difference you see when 10–15% of a cohort drops out. In my own tracking, I only compare 4-week averages now because daily weights lie. Has anyone found the discontinuation rate buried in those Phase 3 releases? That would tell us a lot.

👍 2