24.2% at 48 weeks vs 28.7% Phase 3 - what explains gap?

Sep 17 4911 views 47 posts

I keep bouncing between the Phase 2 obesity result and the company Phase 3 slides. Phase 2 showed mean body-weight reduction of 24.2% at 48 weeks under the efficacy estimand. Company-reported Phase 3, not peer-reviewed, says up to 28.7% across four trials. That's a ~4.5 point gap. Is it estimand, duration, population, or selected dose arms? Also the completed CTgov trials are small: 32 Chinese participants, 445 with knee OA, 46 postmenopausal women. Do those add anything to the efficacy picture, or are they mostly safety/regional slices?

👍 1

The 28.7% is company-reported, not peer-reviewed. I'd anchor on the 24.2% until full papers land.

The part people forget when they're comparing numbers: longer exposure and a different titration schedule probably explain a lot of that gap.

👍 1

Efficacy estimand versus treatment policy can move obesity numbers a lot. If lots discontinued, estimand choice matters. Do we know dropout rates in those Phase 3 trials?

I looked at the completed CTgov entries after seeing them mentioned here. NCT05548231 enrolled 32 Chinese participants, NCT05931367 enrolled 445 with obesity/overweight and knee OA, and NCT06039826 enrolled 46 postmenopausal women. None list headline weight results in the raw record. So I don't think they resolve the 24.2 vs 28.7 question. The 445-person trial might have useful secondary data, but it's a different population. The 32-person one is tiny. The postmenopausal one is also small. I'd treat them as supporting slices, not the main efficacy story for now, especially without dose-level data.

That 445-person knee OA trial is the only one with enough n to say much, but it's a different population than the big Phase 3.

If 24.2% is efficacy estimand at 48 weeks, I want completers vs ITT side by side.

The thing I keep coming back to is the glucagon component. Retatrutide is a triple agonist, and if that third arm contributes to the weight effect, the difference might widen over longer trials. But I can't verify that from these CTgov records. The 32-person Chinese trial won't answer it. The 46-person postmenopausal trial probably won't either. The 445-person knee OA trial might have more useful data, but that wasn't in the raw record I saw. Until we get the Phase 3 papers, I'm treating 28.7% as a ceiling, not a typical result.

A 4.5-point gap is smaller than my monthly water-weight swings, honestly.

👍 1

Estimand can do a lot: all-randomized vs completers is often worth several points by itself.

👍 1

Those 32-person CTgov arms are basically a pilot; one hyper-responder skews the whole average.

👍 1

The part people forget when comparing programs: the support baked into Phase 3. Weekly check-ins, food logs, a dietitian — that stuff can easily add 4 points over a Phase 2 where folks just got a handout and were sent on their way.

Common surprise from those 48-week logs people keep — food, sleep, training, the whole boring spreadsheet: weeks 1-12 down about 9%, weeks 13-28 down 6%, then weeks 29-48 only 3%, because weekends got loose. Waist kept shrinking the whole time though. Worth knowing before comparing options: scale averages hide a lot.

👍 2 ❤️ 1

My own 48-week number vs my month-3 number differs by 4 points and it's mostly whoosh timing, not estimands.

Waistband check beats the scale for me. Same 12 lbs down at week 16, but jeans went from a 38 to a 34 while the scale sat frozen for nine straight days mid-month. So when I see a 4-point gap between two trials I assume measurement windows and water shifts, not some huge difference in how the thing actually behaves. Nobody's tape measure moves two inches overnight.

👍 3 ❤️ 1

I'm the cautionary tale for weekend math. Ten weeks in I was logging 1,700-1,800 on weekdays, gym three mornings, 8k steps, and I went 199 to 184 and then just parked at 184 for a solid month. I blamed the medication. Then I actually added up Friday night through Sunday: takeout, two beers, brunch, the whole thing was landing near 3,400 a day, which quietly erased about five days of weekday deficit every week. The fix wasn't dramatic. I started eating protein first at every meal, roughly 30g at breakfast, which killed the 10am vending machine habit, and I stopped treating Saturday as a free day and started treating it as a normal day with one actual treat in it. Weeks 14 to 22 went 184 to 171, about 1.6 a week, faster than my first ten weeks. The reversal I didn't expect was sleep: when the late meals stopped I was out by 10:30 instead of midnight, and the morning scale started moving again. So whenever I see 24% vs 28%, I wonder how much of that gap is just how carefully people were counting three days a week.

👍 1 ❤️ 1

I'll push back a bit: the 4.5-point gap might be less about trial design and more about who's still standing at week 48, which is the exact trap I fell into with my own spreadsheet. I logged 48 weeks and my all-weeks average was 41 lbs down; if I'd dropped the six travel weeks and two holiday weeks where I didn't weigh or log, that same stretch reads 46 lbs, a 12% loss instead of 10.6%. Nothing about my body changed, I just picked a friendlier denominator. Retention works the same way, since the people who quit at week 20 because the food noise came roaring back aren't in the week-48 completers number. That number is real, it's just answering a different question than 'what happens if I start.' What moved my own needle more than any of this reading was a boring food tactic: same breakfast every day, Greek yogurt with berries and about 40g of protein, one giant salad at lunch, dinner on a smaller plate. Sixteen weeks of that and I stopped negotiating with myself at 9pm. The number I trust at the end is my 34-inch waist and how I do on the stairs, not a slide.

Estimand debates are above my pay grade, but four points wouldn't change my Tuesday either way.

I slept five hours a night for two weeks during a work crunch and my loss flatlined at 19 lbs down despite the same food. Got back to seven and a half and dropped four more in ten days. Sleep might be a bigger lever than which estimand the company picked. Anyone else see that stall when stress spikes?

👍 1

I've stopped trusting any single morning. Two cups of salty soup on a Wednesday and Thursday's number reads three pounds up, gone by Friday. Weighing daily and only comaring seven-day averages killed the panic for me. The scale is a noisy instrument, not a verdict.

👍 1 ❤️ 1

The number that surprised me was my own chart at 10 weeks: 12 lbs down, then only 1.5 a month for the next four. Draw a line through the whole thing and it looks like failure, but the total kept moving. Trials cut endpoits at fixed weeks, so whoever runs longer tends to show a bigger number, and dose selection stacks on top. That's probably more of the 4.5 points than any estimand fight. Separately, I logged every bite for 90 days and found my 'small' handful of almonds was 400 calories. Fixing that one habit beat anything I did at the gym, and I still ate the almonds, just eleven of them.

Four and a half points is roughly my whole first month, so I get why people are arguing.

👍 2 ❤️ 1

Phase 2 efficacy estimand vs compny slides is apples to oranges, IMO.

Those tiny CTgov arms would not move my personal beief much either.

👍 1

Killed the daily weigh-in around the 12-week mark — those trial averages were making me nuts. Since February the Sunday routine is a cheap tape at the navel plus how a pair of 36 jeans sits. Weeks 16-22 the scale moved 1.5 lbs total, but waist went 38 to 36.5 and the jeans got loose. Then weeks 23-28 the scale dropped 6 lbs without much of anything changing.

Which tells me the Phase 2 vs Phase 3 gap is probably estimand, duration, and which dose arms got stacked into that 28.7%. I'd want completion rates and waist data, not just mean weight. Company slides always look cleaner than a Sunday tape.

👍 3

Belt notches over slide decks, every time. Three notches at week 30 says more than any percentage on a chart.

👍 1

The gap smells like missing data more than biology. Phase 2 efficacy estimand at 48 weeks is a tidy number; the Phase 3 “up to 28.7%” across four trials is likely completers or on-treatment, with longer exposure and a broader population. Ask for completion rates and whether that 28.7% is ITT or completers. A 10–15% differential dropout can absolutely manufacture a 4-point gap. We chased this same headline-vs-paper thing in older threads. Peak loss is the easy part; the maintenance/withdrawal phase is what I actually watch.

👍 4

The “up to” is doing a lot of work. If Phase 3 ran longer, or that 28.7% is the bet arm at some later week, it’s apples-to-oranges. I’d want the week-48, same-arm, similar-baseline-BMI slice before crediting biology or estimand. In maintenance I only compare like weeks—my own chart looked flat around 48 weeks, then kept drifting slowly. Does the deck show a per-arm week-48 table, or just peak numbers? I bet the gap shrinks to 1–2 points once duration and arm selection are matched.

👍 4

One thing I’d dig into: baseline BMI/weight distribution, not just mean percent loss. Percent change isn’t symmetrical—if Phase 3 skewed heavier, 28% could reflect similar absolute pounds as 24% in a lighter Phse 2 cohort. I saw that in my own tracking: my % swings more when I start higher, even when inches lost are steady. Do the slides list baseline BMI or weight ranges? That’d be my first check before assuming biology or estimand.

👍 4

The “up to” is what I’d chase first: is 28.7% one trial’s best arm at a single timepoint, or the max across four trials and multiple timepoints? That’s not apples-to-apples with a pooled 24.2% at 48 weeks. Also, baseline BMI and how much structured lifestyle support each site had can shift absolute percent loss a lot. I’d want baseline BMI by arm. My own tracking moved from daily scale to waist + lifting numbers; those changed less noisily and matched how clothes fit.

👍 8 ❤️ 3

I’d firt check whether the 28.7% is actually at 48 weeks or just the best timepoint across the four Phase 3 trials. Company slides often say “up to” because one arm at week 72 hit that, while the pooled 48-week average is lower. If both are truly 48 weeks, baseline BMI and site-level lifestyle support are my next suspects—more frequent dietitian contact can shift percent loss by a few points. In the trials, the curve also flattens hard after ~9 months, so 4.5 points from extra duration isn’t crazy.

👍 4 ❤️ 1

yo, one thing i haven’t seen mentioned: when was baseline actually measured? i do late-night soda-swap experiments and track daily, and if i cut fizzy stuff for a week before logging my “start” weight, the first month’s % loss looks inflated just from water/glycogen. so if ph2 had a longer run-in/washout or different baseline timing than ph3, 24.2 vs 28.7 isn’t apples-to-apples even at the same week. did both protocols define baseline the same way? that’s the boring footnote i’d hunt for.

👍 3

I’d look at whether the Phase 3 trials had a more standardized lifestyle component—weekly check-ins, protein targets, activity goals—than Phase 2. In my own maintenance, the behavioral stuff moved my six-month average way more than I expected, even with meds staying vague. If the later trials had more hands-on support or better retention, that alone could add a few points. Do the company slides say whether diet/exercise counseling was matched across trials, or was it site-dependent? That’s the first thing I’d overlay before blaming the drug itself.

👍 4

One thing I’d chase: completion rates and what they did about rescue meds or insurance/supply gaps. I work part-time at our little pharmacy counter, and I see folks miss weeks over prior auth or backorder, then weight creeps back. If Phase 2 had tighter check-ins and fewer interruptions while Phase 3 was broader, that can explain a chunk. Do the Phase 3 slides show how many finished all 48 weeks, or just the mean for those still on it? That’s the table I’d want before comparing 24.2% and 28.7%.

👍 3 ❤️ 1

Night shifter here. I’d ask for completion rates and how many were still on it at each timepoint. If Phase 3 had more dropouts, that 28.7% is a survivor number. Efficacy estimand can hide that too. Also check per-trial n, not the pooled headline. Four trials with different sites, food environments, and support isn’t apples-to-apples with one Phase 2. I track protein, steps, sleep. Scale moves way slower than those slides. Ask for the dropouts.

👍 5 ❤️ 2

“Up to 28.7% across four trials” is a max, not a mean. First thing I’d do is rebuild the apples-to-apples table: same week, same estimand, same arm, same baseline BMI. If Phase 3 runs longer or enrolls heavier baseline, percent loss can drift up without any real efficacy difference. My money’s on 28.7 being a selected arm at a later timepoint, not the average at week 48. Do the slides footnote the estimand per trial? That’s the tell.

👍 2

Not a stats person, but the thing I haven’t seen mentioned: baseline BMI/weight and diabetes status. If Phase 3 enrolled more people with T2D or a lower starting BMI, that alone can shave several points off mean % loss. Also “up to 28.7% across four trials” sounds like the best dose arm in the best trial, not a pooled average—do the slides show the range or just the headline max? I’ve been stuck for 9 weeks myself, so I get grasping at these numbers. I track waist and photos because the scale lies. But yeah, I’d want baseline characteristics and whether those Phase 3 numbers are per-arm or pooled.

👍 5 ❤️ 1

Honestly, one thing I’d poke at: behavioral support intensity and visit cadence. In my own tracking, weekly check-ins kept me way more consistent with protein, water, and steps than monthly ones did. If Phase 2 had tighter trial visits or a structured diet/exercise run-in, and Phase 3 was more real-world/global, that alone could shift mean weight by a few points. Do the Phase 3 disclosures show visit frequency, dietitian access, or whether it was open-label? I’d also confirm week 48 is compared with the same week in Phase 3, not just the peak timepoint across trials.

👍 4

Hi, I’m brand new here and still learning the stats, so sorry if this is obvious. Could the gap partly be site/protocol differences? Phase 2 is often a smaller set of motivated sites, while Phase 3 goes global—different food environments, different levels

👍 3 ❤️ 1

One thing I’d dig for: did the Phase 3 trials bake in more structured lifestyle support—dietitian visits, food logs, activity goals—while Phase 2 was closer to “here’s the pen, good luck”? In my n=1 dorm experiment, just tracking cheap protein (eggs, tuna, Greek yogurt) and walking to class moved the scale before anything else. If Phase 3 added regular check-ins or a run-in phase, that could inflate results without any pharmacological difference. Also, were the Phase 2 sites maybe more severe/harder-to-treat? Sharper question: do the Phase 3 protocols list a formal lifestyle intervention, or just “standard advice”?

👍 5 ❤️ 2

Broke dorm-dweller checking in, so take my stats with a grain of instant ramen. The thing I’d chase is the timepoint: is the 28.7% at 48 weeks, or at 72/96 weeks? If it’s later, you’re comparing a 48-week snapshot to a longer movie, and that could eat a big chunk of the gap. I track my own cheap-protein meals (eggs, tuna, lentils) and my trendline looks wildly different at week 12 vs week 30 even when nothing changed. So: what’s the actual week for that 28.7%, and is it the same estimand/timepoint as the Phase 2 24.2%?

👍 4

Daily weigher here. I’d want the average weeks at maintenance on drug, not just completion. Phase 3 programs often have more titration visits and flexible escalation, so a bigger chunk of the average might be spent at a higher tolerated maintenance level. Phase 2 can also run tighter at fewer sites with more standardized behavioral support. So the gap could be exposure time + trial environment, not just estimand. Do the Phase 3 protocols include the same diet/exercise component as Phase 2? If not, that alone can move 4 points. Also, pooled “up to” usually cherry-picks the best arm, so I’d compare like-for-like arms before worrying.

👍 1

Unpopular take: the 4.5 pt gap is probably mostly the word "up to." If 28.7% is one selected active arm at a later week under a secondary estimand, and 24.2% is the Phase 2 efficacy-estimand mean at 48 weeks, you're comparing a highlight reel to a full box score. I’d ask for the ITT mean at

👍 7

Hi, total newbie here (week one, so I’m basically a sponge). One thing I haven’t seen asked: did the Phase 3 trials run longer than 48 weeks? If that 28.7% figure came from a later timepoint, the extra time alone could explain a few points. Also, I’ve been tracking my own daily weigh-ins and the water weight swings are wild—makes me wonder how trial protocols handle that. Sorry if this is a silly question.

👍 2

The four Phase 3 trials probably have different dietary/activity support and visit frequency. If one had weekly coaching and another quarterly, that’s not estimand—that’s adherence to lifestyle. Ask for session attendance and food/activity log completion by arm. My own tracking showed the medication only amplified what my logs did; when I stopped logging, loss flattened. So if Phase 2 had more structured check-ins, that alone can buy a few points. It’s a confounder you can actually measure.

👍 6 ❤️ 2

One thing I haven’t seen raised: trial touchpoints. Phase 2 protocols often have more frequent visits, weigh-ins, and lifestyle check-ins than big Phase 3 programs. That extra accountability can move the scale a couple points on its own. In my own maintenance, weekly check-ins kept me honest; when I went to monthly, my 7-day average crept up until I reset. So I’d ask: were visit frequency and behavioral support identical between the Phase 2 and Phase 3 arms? If not, some of that 4.5-point gap might be study scaffolding, not just estimand or dose.

👍 4 ❤️ 1

Cafe owner here, not a stats person, but my sharp question would be: what exact week is the 28.7% from? If that’s a longer timepoint while Phase 2’s 24.2% is locked at 48 weeks, you’re comparing two different points on the curve, not the same endpoint. In my own tracking, 6 months and a year tell very different stories—habits drift, stress shifts, portions creep. I’d want the 48-week mean from each Phase 3 trial side by side before calling it a real 4.5-point gap. Do the slides show those 48-week numbers?

👍 3

Not a stats person either, but I’d want the actual timepoint behind each number. I keep a simple spreadsheet of my own weekly % change, and my 48→72 week trendline adds only ~1.2–1.8 points—definitely not 4.5. If the 28.7% is a later landmark or a pooled max at different weeks, the gap is partly calendar, not biology. If both are truly 48 weeks, then check whether Phase 2 excluded more early dropouts from the efficacy estimand than Phase 3. Missing-data rules can move means more than people think. Does anyone have the week-by-week curves, not just the headline?

👍 1