GuidesDo you need to train to failure?

Do you need to train to failure?

No. In volume-matched trials, stopping a rep or two short builds about as much muscle and strength as grinding — for meaningfully less fatigue.

21 sources~11 min read
Jump to a section

No, you don't. When the two groups do the same amount of work, training to momentary muscular failure and stopping a couple of reps short produce gains in strength and muscle size that the research cannot tell apart. Grgic's meta-analysis put volume-equated strength at ES 0.01 (95% CI −0.12 to 0.15); Vieira AF and colleagues found that failure's apparent hypertrophy advantage (SMD 0.75, p = 0.005) disappeared entirely once volumes were matched. Both teams conclude the same thing: failure is neither required nor harmful.

The longer answer is that proximity to failure still matters — more for muscle size than for strength, and much more when the load is light.

It also costs a lot of fatigue for the last couple of reps. And the question everyone actually wants answered — exactly how close is close enough — has no tested answer at all. This is a genuinely contested topic, so below is what is settled, what isn't, and which popular arguments don't survive contact with the data.

What "failure" actually means

Momentary muscular failure is the strict version: you cannot complete the concentric portion of another repetition despite maximal effort. It is not the only definition in use. Studies have variously used volitional or self-terminated failure (you decided to stop), velocity-loss thresholds, or self-reported repetitions in reserve — and a good number never defined it at all. Pelland's methods review documents this directly, and it is the main reason the headline findings look muddier than the headlines suggest.

Reps in reserve (RIR) is the alternative currency: 2 RIR means you believe two more repetitions were available. The RPE scale most lifting apps use is anchored to it — RPE 10 is 0 RIR, RPE 9 is 1 RIR, and so on — a convention introduced and validated by Zourdos, whose experienced squatters showed bar velocity correlating with reported RPE at r = −0.88. That validates the scale's internal logic, not the accuracy of any individual rating.

"Stop when your form breaks down" is a coaching convention, not a research definition. It may be good practice; it is not what the studies measured.

The qualifier that settles most arguments

Almost every disagreement here turns on volume-equating. Grgic's team traced it explicitly: pooled across studies where volume was not equated, non-failure looked better for strength (ES −0.32, 95% CI −0.57 to −0.07) — but in most of those studies the non-failure groups simply performed more sets. Equate volume and the effect is gone (ES 0.01, 95% CI −0.12 to 0.15). Vieira AF's meta-analysis shows the mirror image for growth: failure looked substantially better for hypertrophy overall (SMD 0.75), and the advantage vanished once volumes were matched.

Proximity to failure is not simply neutral, though, and it does not act the same on both goals. Robinson's meta-regressions found that across every best-fit hypertrophy model, muscle size increased as sets ended closer to failure, with confidence intervals excluding null; across every best-fit strength model, the intervals contained null. Pareja-Blanco's squat trial shows the same divergence inside a single study — stopping each set at 20% velocity loss matched 40% velocity loss for 1RM gains with 40% fewer repetitions and better jump performance, while the closer-to-failure group gained more quadriceps size. Velocity loss is a proxy for proximity, not failure itself, but the pattern is consistent.

The returns also flatten sharply at the top. Refalo's review found that pushing beyond 25% velocity loss, versus 20–25%, was worth ES 0.08 (95% CI −0.16 to 0.32). The last repetitions before failure buy very little relative to what they cost.

Before you believe any comparison — Ask whether the two groups did the same amount of work. Most of the headline findings on both sides of this debate evaporate when they did.

Does failure help trained lifters? Genuinely unresolved

This is the live dispute, and honest people are on both sides of it.

For: Grgic's resistance-trained subgroup showed a small hypertrophy advantage, ES 0.15 (95% CI 0.03–0.26) — built on 2 of the 7 hypertrophy studies in the analysis. Refalo's pooled set-failure effect was ES 0.19 (95% CI 0.00–0.37, p = 0.045), an interval that touches zero and which, by the authors' own sensitivity analysis, loses significance if the assumed within-subject correlation falls below 0.73 (they used 0.75). Narrowed to momentary muscular failure specifically it is ES 0.12 (−0.13 to 0.37), and that review's stated conclusion is that there is no evidence momentary failure is superior for hypertrophy.

Against: Refalo's later within-participant trial in trained lifters — each person's failure limb against their own 1–2 RIR limb, volume matched, eight weeks — found quadriceps thickness up 0.181 cm versus 0.182 cm. Ruple's group found no condition-by-time interaction for 1RM or vastus lateralis cross-sectional area at 0–1 versus 4–6 RIR.

Both camps are working with 15 to 40 participants over five to eight weeks. That is not enough to resolve an effect this small in either direction, and "no detectable difference" is not the same as "proven identical".

What failure costs

This part is not contested. A separate meta-analysis, Vieira JG and colleagues on acute fatigue — a different paper from the Vieira AF hypertrophy review above — found that sets taken to failure left substantially more acute disruption than stopping short: neuromuscular function SMD −0.96 (95% CI −1.43 to −0.49), creatine kinase at 48 hours SMD 0.86, session RPE SMD 1.93. Training status did not moderate the effect (p = 0.92) — experienced lifters get no discount on the fatigue.

Refalo's crossover puts a dose on it. Six sets of bench press at 75% 1RM: lifting velocity four minutes after the session was down 25% in the failure condition, 13% at 1 RIR and 8% at 3 RIR — for a similar total number of repetitions. The honest counterweight is that the cost was largely transient. By 24 hours the three conditions were within a few per cent of each other, and all had recovered fully by 48 hours. That was one exercise and six sets, not a full session of heavy compounds.

Across eight weeks, failure also felt consistently worse: discomfort 5.1 versus 4.1, session RPE 5.4 versus 4.3, and less positive general feelings. The authors read that as a plausible adherence cost, though adherence itself was never tracked long enough to show that worse feelings actually make people quit.

Where failure is more defensible

Load is the clearest moderator anyone has demonstrated. Lasevicius equated volume across four unilateral knee-extension conditions: at 80% 1RM, failure added nothing (quadriceps CSA +8.1% to failure versus +7.7% stopping short); at 30% 1RM, failure was the difference between growth and none (+7.8% versus +2.8%, with the light non-failure condition not growing significantly). When the load is light, effort is what makes the set count. Those participants were untrained, so applying it to trained lifters is an extrapolation.

From there, the familiar advice — push isolation work close to failure, leave a rep or two on heavy compounds — is a reasonable inference rather than an established finding. It follows from the load interaction, from the fatigue data, and from how much harder a heavy multi-joint set is to judge. The direct evidence is thin and mixed. Davies' compound-exercise subgroup favoured non-failure for strength (ES 0.37–0.38, p = 0.03) while the authors judged the 0.6–1.3% difference unlikely to be meaningful; Grgic's team ran the same kind of subgroup and found no exercise-selection moderation; and no meta-analysis has tested exercise type for hypertrophy at all. The most direct hint is a feasibility note rather than a result: in Robinson's four-arm trial, the all-sets-to-failure squat protocol could not be sustained and produced uninterpretable data.

The related rule — take only the last set of each exercise to failure — is practitioner convention. No trial has tested it as an independent variable.

Judging how close you actually are

There is a gap between the RIR you report and the RIR you have. Halperin's review of 262 effect sizes found people underpredict their remaining repetitions by roughly one on average, with wide individual variation (between-person SD 1.45 reps) — meaning they had more left than they thought, so a self-reported 2 RIR is often nearer 3. Heterogeneity was extreme (I² = 97.9%), so treat that as a direction, not a figure. Accuracy improves when the judgement is made close to failure and on shorter sets of about 12 reps or fewer, and Hermann's group found it improved over eight weeks of practice and was better on the bench press than the squat.

What does not reliably help is general training experience. Halperin found no meaningful moderation by training status (β = −0.006, 95% CI −0.02 to 0.007), and Droguett found none in the back squat, albeit with 16 participants. Zourdos found experienced squatters discriminated better at a true 1RM, so familiarity with a specific lift plausibly matters even where training age does not. Inferring effort from load percentages is worse again: Qin found percentage-based methods overestimated intended RIR by around three repetitions at 60% 1RM, while errors at 80% were small.

Three claims worth dropping

Some of the most repeated arguments in this debate do not hold up.

  • "You need failure to recruit high-threshold motor units, so you need it to grow." That is a mechanistic premise, not an outcome. Ruple's group measured motor-unit firing-rate adaptations that genuinely did differ between 0–1 RIR and 4–6 RIR — with no accompanying difference in strength or muscle size.
  • "Everything before the last few reps is junk volume." Contradicted by every volume-equated comparison, and by Pareja-Blanco, where stopping at 20% velocity loss matched 40% for squat strength with 40% fewer repetitions.
  • "Training to failure is dangerous." We looked for evidence on this specifically and found none in either direction. The long-term injury and joint consequences of habitual failure training have not been studied, and adverse-event reporting across this literature is limited. That is absence of evidence, not a clean bill of health.

Honest limits

No trial has identified an optimal RIR target. Grgic's team state plainly that the literature cannot say whether stopping five reps short differs from stopping two; Refalo's review concludes that the proximity that would maximise hypertrophy is unknown. Any specific prescription you read — 2 RIR, 1–3 RIR, last set only — is convention dressed as a threshold, including the conventions we find reasonable.

The evidence base is also smaller than the confidence of most write-ups suggests. Almost all of it is young adults, over six to fourteen weeks, in samples of fifteen to forty. Studies that size cannot detect the small effects being argued over, so a null result is weak evidence of absence and a barely-significant result is fragile. The one pooled finding most often quoted as proof that closer-to-failure grows more muscle stops being significant if a single analytic assumption is moved by 0.02.

And this is training practice, not medical advice. If you are returning from injury, managing a health condition, pregnant, or new to lifting, the right person to ask is not a website.

How Shojin uses this

Shojin's next-session suggestion turns on a single number: the highest RPE you logged on your heaviest working set. Eight or below and it adds one equipment increment; nine or ten and it holds; a drop in reps at the same top weight on a nine-or-above session triggers a small deload. So the rating does real work, and an honest one is worth more than a heroic one — under-report and the app pushes load onto you, over-report and it freezes you at a weight you could have beaten. Leaving RPE blank counts as "had more in the tank", which is worth knowing if you tend to skip it. The design is defensible rather than precise: it reads the one set where self-reported effort is most reliable, but it is still your estimate, and the literature says that estimate is off by about a repetition.

Common questions

Should I ever train to failure?

There is no evidence you have to. With volume equated, failure and stopping a couple of reps short produce indistinguishable strength gains and a hypertrophy difference too small and too fragile to call settled. Failure is most defensible where the load is light — the one place it has been shown to matter decisively — and least defensible where the fatigue cost is highest and your RIR estimate is least reliable.

How many reps short should I stop?

Nobody knows. Grgic's team state that the literature cannot distinguish stopping five reps short from stopping two; Refalo's review concludes the proximity that would maximise hypertrophy is unclear. Common targets like 1–3 RIR are convention, not tested thresholds — though the evidence on estimation accuracy does suggest your judgement is better when you are close to failure and the set is short.

Does training to failure matter more on light weights?

That is the strongest moderator anyone has demonstrated. In a volume-equated within-participant trial, failure added nothing at 80% 1RM (quadriceps CSA +8.1% versus +7.7%) but was the difference between growth and none at 30% 1RM (+7.8% versus +2.8%). The participants were untrained, so applying it to trained lifters is an extrapolation.

Is training to failure bad for your joints?

We could not find peer-reviewed systematic evidence in either direction. The acute fatigue cost is well quantified — substantially greater loss of neuromuscular function, more muscle damage and higher perceived exertion — but long-term injury outcomes from habitual failure training have not been studied. That is an absence of evidence, not reassurance.

What RPE should I log if I am not sure?

Log what you honestly think, and log it on the heaviest working set where your judgement is most reliable. People underpredict their remaining reps by roughly one on average with wide individual variation, so a reported 8 is often nearer a true 7 — you had more in reserve than it felt. That direction matters here, because Shojin holds the load at RPE 9–10: over-rating a hard set costs you an increase you had actually earned. In Shojin, leaving RPE blank is read as "had more in the tank", so the app will add load.

These are the same answers the page’s FAQ structured data publishes — visible text, no hidden-content mismatch.

Sources

Every number in this article traces to one of these. Where the evidence is contested or thin, the article says so rather than picking a side.

  1. 01

    Grgic J, Schoenfeld BJ, Orazem J, Sabol F (2022). Journal of Sport and Health Science 11(2):202–211

  2. 02

    Vieira AF, Umpierre D, Teodoro JL, et al. (2021). Journal of Strength and Conditioning Research 35(4):1165–1175

  3. 03

    Refalo MC, Helms ER, Hamilton DL, Fyfe JJ (2023). Sports Medicine – Open 9:10

  4. 04

    Robinson ZP, Pelland JC, Remmert JF, et al. (2024). Sports Medicine 54(9):2209–2231

  5. 05

    Pareja-Blanco F, Rodríguez-Rosell D, Sánchez-Medina L, et al. (2017). Scandinavian Journal of Medicine & Science in Sports 27(7):724–735

  6. 06

    Refalo MC, Helms ER, Robinson ZP, Hamilton DL, Fyfe JJ (2024). Journal of Sports Sciences 42(1):85–101

  7. 07

    Ruple BA, Plotkin DL, Smith MA, et al. (2023). Physiological Reports 11(9):e15679

  8. 08

    Vieira JG, Sardeli AV, Dias MR, et al. (2022). Sports Medicine 52(5):1103–1125

  9. 09

    Refalo MC, Helms ER, Hamilton DL, Fyfe JJ (2025). European Journal of Sport Science 25(3):e12266

  10. 10

    Lasevicius T, Schoenfeld BJ, Silva-Batista C, et al. (2022). Journal of Strength and Conditioning Research 36(2):346–351

  11. 11

    Davies T, Orr R, Halaki M, Hackett D (2016). Sports Medicine 46(4):487–502

  12. 12

    Robinson ZP, Macarilla CT, Juber MC, et al. (2025). International Journal of Strength and Conditioning 5(1), Article 393

  13. 13

    Hermann T, Mohan AE, Enes A, et al. (2025). Medicine and Science in Sports and Exercise 57(9):2021–2031

  14. 14

    Halperin I, et al. (2022). Sports Medicine 52(2):377–390 — PMID 34542869

  15. 15

    Droguett FAB, et al. (2025). Journal of Human Kinetics 102:145–158 — PMID 42211811

  16. 16

    Qin X, Liu B, García-Ramos A (2025). BMC Sports Science, Medicine and Rehabilitation 17(1):60

  17. 17

    Zourdos MC, Klemp A, Dolan C, et al. (2016). Journal of Strength and Conditioning Research 30(1):267–275

  18. 18

    Pelland JC, Robinson ZP, Remmert JF, et al. (2022). Sports Medicine 52(7):1461–1472

  19. 19

    Helms ER, Byrnes RK, Cooke DM, et al. (2018). Frontiers in Physiology 9:247

  20. 20

    Bastos V, Machado S, Teixeira DS (2024). Perceptual and Motor Skills 131(3):940–970

  21. 21

    Varela-Olalla D, del Campo-Vecino J, Balsalobre-Fernández C (2025). Journal of Strength and Conditioning Research 39(9):e1129–e1168

Keep reading

Track this properly

Shojin logs sets in seconds, reads your RPE to suggest the next load, and tells you when a lift has actually stalled.

Free to start · No ads · No third-party trackers