Jump to a section
They are the same judgement read on two inverted axes. RIR — reps in reserve — counts the repetitions you believe you had left when you racked the bar. RPE — rating of perceived exertion, in the strength-training version of the scale — counts effort up to a ceiling of 10. In this convention RPE = 10 − RIR, so "RPE 8" and "2 RIR" are identical claims about the same set. There is nothing to choose between them. Log whichever your notebook or app asks for.
The question worth asking is not which scale but how good the rating is.
Briefly: nobody is precise, the error runs in one direction — people believe they are closer to failure than they actually are — and accuracy improves sharply as you approach failure and as the load gets heavier. On a heavy working set taken within a few reps of failure, trained lifters are accurate to well under a rep on average. On light, high-rep work far from failure, the rating is close to noise.
One judgement, two axes
The published mapping comes from Helms and colleagues (2016), and it is the version nearly every lifting app has adopted:
- RPE 10 = 0 RIR — nothing left
- RPE 9 = 1 RIR — one rep left in the tank
- RPE 8 = 2 RIR — two reps left in the tank
- RPE 7 = 3 RIR — three reps left in the tank
- RPE 6 = 4 RIR
- RPE 5 = roughly 4–6 RIR
- RPE 4 and below — light to negligible effort
Two things about that table are worth knowing. It is a stipulated convention, not an empirically derived function — nobody measured their way to "RPE 7 means three reps left"; it was defined that way so the number would refer to repetitions rather than to sensation. And the same paper states plainly that the scale is inappropriate for light, high-velocity power work below about 80% of one-rep max, where the set ends for reasons that have nothing to do with proximity to failure.
The short answer — RPE = 10 − RIR. Use whichever your logbook asks for — the accuracy problem is identical either way.
Careful: four different things are called RPE
Borg's original 6–20 scale, the Borg CR-10, Foster's session-RPE — a whole-session training-load score, a single rating multiplied by session minutes, originally validated against a heart-rate standard (Foster et al. 2001) — and the RIR-anchored 0–10 per-set scale introduced by Zourdos and colleagues (2016). A lifting app means the fourth. If you arrived from running or cycling you have almost certainly met one of the other three, and the numbers do not transfer between them.
That also settles a common misreading. In the RIR convention, RPE 10 is a claim about repetitions — none left — not about how bad the set felt. Anchoring the rating to reps rather than to sensation is the entire point of the scale, and it is what gave the original validation work something to check against: reported RPE tracked bar velocity inversely (r = −0.88 in experienced lifters, −0.77 in novices).
How accurately does anyone rate a set?
The best available synthesis is a scoping review with exploratory meta-analysis covering 12 clusters, 414 participants and 262 effect sizes (Halperin et al. 2022). Pooled, people under-predicted their remaining repetitions by 0.95 reps (95% CI 0.17 to 1.73) — they stopped believing they were nearer to failure than they were. The honest caveat is that heterogeneity was extreme (I² = 97.9%), so that pooled figure summarises a badly scattered literature rather than describing a typical lifter. The authors say outright that whether this inaccuracy is practically acceptable remains undetermined.
There is no single defensible "typical error", and you should distrust anyone who quotes one. The number depends on proximity to failure, load, exercise, and whether the call is made before or during the set:
- 0.65 ± 0.78 reps of absolute error — trained lifters calling RIR mid-set at 1 and 3 RIR on a bench press at 75% 1RM (Refalo et al. 2024, n = 24). The best case, and the closest match to how a lifting app actually uses the number.
- About 1.2 reps of overshoot — velocity-verified back squats at 70% 1RM, in experienced and novice lifters alike (Droguett et al. 2025, n = 16).
- About 2.0 reps (95% CI 0.0 to 4.0) — resistance-trained people who deliberately stopped at their own self-determined limit still had roughly two reps left (Armes et al. 2020; single-joint knee extension, small samples, wide interval).
- 2.6–3.4 reps of measurement error — but this was predicting the count before the set rather than calling it mid-set, a materially harder task (Steele et al. 2017, n = 141).
About a rep of slack is the fair assumption in the conditions the scale was designed for. Two or three reps is fair for everything else.
What actually moves accuracy — and what doesn't
Two moderators replicate across designs. The first is proximity to failure: the closer you are, the better you judge. In the meta-analysis the effect was small but consistent (β = −0.025), and a 26-study systematic review reaches the same conclusion (Russo et al. 2026). The vivid demonstration comes from 259 certified coaches watching video of squats and curls: absolute error of 4.8 reps a third of the way through a set, 2.0 at two-thirds, 1.2 near the end (Emanuel et al. 2022). Note what that implies — the gradient is a property of the judgement itself, not of a lifter's interoception, because an outside observer shows it too. It also disposes of the idea that a coach can simply tell.
The second is load and rep count. Errors grew markedly in sets above 12 reps compared with sets below (β = 0.47 versus 0.06). The cleanest within-subject demonstration: at 85% 1RM bias was negligible and 75–100% of estimates were valid, while at 65% 1RM the same people under-estimated by around a rep, with equipment and sex differences appearing only at the light load (Ruiz-Alias et al. 2024).
Now the claim you will see everywhere, which the evidence does not support as stated: that accuracy improves with training age. In its favour, the largest single dataset reports a tendency in that direction (Steele et al. 2017), and experienced squatters rated a true 1RM closer to 10 than novices did — 9.80 ± 0.18 versus 8.96 ± 0.43 (Zourdos et al. 2016). Against it, and from higher-tier evidence: training status did not moderate accuracy in the meta-analysis (β = −0.006); a velocity-verified squat study found no difference between lifters above and below 18 months of training (d = 0.03 at 3 RIR); intra-set bench-press accuracy showed no relationship with years of experience; and people with a year or more of training still had about two reps left when they stopped at a self-determined limit. Even the coaches' error barely moved with years of coaching.
The honest synthesis is that the direction is plausible and is coaching consensus, but training age is not a reliable moderator, and experienced lifters are not accurate in absolute terms — just less bad in easy conditions. Familiarity with the specific lift and with the rating procedure is the better bet, though we are not aware of a trial that isolates it. Accuracy does improve across successive sets within a session, but only trivially — about 0.07 of a repetition per set (β = −0.07, 95% CI −0.14 to −0.005) — and one trial found ratings statistically equivalent across sets. Real, and far too small to build a routine on.
Why an imprecise number is still worth logging
Because the cost of a one-rep error is small. Training to momentary failure produced no detectable hypertrophy advantage over stopping short (ES 0.12, 95% CI −0.13 to 0.37; Refalo et al. 2023) — the finding already cited on our methodology page. A separate set of meta-regressions suggests proximity matters somewhat for growth and not at all for strength, but its RIR values were retrospectively estimated by the authors from written study protocols rather than measured, model fit was modest, and the authors themselves urge caution (Robinson et al. 2024). Read together: getting the rating a rep wrong is unlikely to cost you an adaptation.
It is also worth being clear about what autoregulation buys. Prescribing load by RIR-based RPE or by bar velocity produced no significant one-rep-max advantage over fixed percentages — a mean difference of 2.07 kg with a confidence interval spanning zero (95% CI −0.32 to 4.46, p = 0.09; Hickmott et al. 2022), from only six studies and 133 participants. The argument for logging RPE is individualisation and day-to-day responsiveness, not a demonstrated size effect. The alternative is no better regulated: applying the same percentage of 1RM across a group produces a wide spread of actual per-set RIR, because people differ in how many reps they get at a given percentage.
Rating more accurately
Most of the practical advice follows directly from the two robust moderators — rate where the judgement is good, and be sceptical where it isn't.
- Rate the sets you take close to failure and heavy. That is where the measurement works. A rating on a light, 15-rep accessory is bookkeeping, not information.
- Rate immediately, before the set fades. Perceived exertion is biased by caffeine, sleep, music, ambient temperature, who is watching and personality — none of which are properties of the set. Consistency of conditions matters more than absolute truth.
- Expect to under-rate. If you wrote 9, there is a fair chance you had two reps left. Knowing the direction of your own bias is worth more than trying to eliminate it.
- Keep it to whole numbers. Given that the honest resolution of the measurement is roughly a repetition, 8.5 is precision the instrument does not have. This is our reading of the accuracy data rather than a tested finding, and Shojin offers integers only.
- Be sceptical of the standard advice to periodically take a set to genuine failure "so you know what 0 RIR feels like". It is practitioner convention. The accuracy studies did familiarise participants with sets to failure before measuring them, which is suggestive, but no trial has tested whether failure exposure improves later RIR accuracy.
How Shojin reads the number
RPE is the single input driving Shojin's progression suggestion, and it is optional. The picker offers six values — <5, 6, 7, 8, 9, 10 — and can be switched off entirely in your profile. Everything below 5 collapses into one bucket, for the reason above: a rating that far from failure carries very little information, and pretending otherwise would be dishonest design.
From your last session's working sets — warm-ups and drop sets never count — the app finds the heaviest weight, takes the highest RPE among the sets at that weight (the maximum, not the average), and applies three rules:
- RPE 8 or below, or nothing logged → add one increment. In kilograms that is 2.5 for barbell, cable and machine work and 2 for dumbbells; in pounds it is a flat 5 for everything, dumbbells included.
- RPE 9 or 10 → hold the weight.
- RPE 9 or 10, and fewer reps than the session before at that same top weight → back off about 5%, rounded to an increment. Only the most recent session needs to be hard; it is the rep count that must have gone backwards.
Timed movements add 10 seconds and unweighted bodyweight movements add a rep on the same RPE ≤ 8 test; timed work takes the highest RPE across all your working sets rather than only the top ones, and neither ever deloads. The full engine, including stall detection, is on the methodology page.
Two consequences are worth naming. Because the app takes the maximum, a single 9 among your top sets holds the load even if a sibling set was rated 7 — deliberately conservative. And because an unlogged RPE is read as "had more in the tank", skipping the rating pushes the suggestion towards adding weight. That is a choice, not a neutral default. Given that the literature says people under-rate, the two biases partly offset: your self-reported 9 was often really an 8, which would have earned an increase anyway.
How Shojin uses this
Shojin asks for one number per set, and only if you want to give it. That number is the whole input to the progression suggestion under each exercise in the logger — the app reads your last session's working sets, takes the highest RPE among the sets at your top weight, and proposes go up, hold, or back off. The thresholds are printed in full on the methodology page, along with the formulas behind every other number in the app, because a suggestion you can't audit isn't worth much. If you would rather not rate sets at all, turn the field off in your profile; the app will read every set as submaximal and keep nudging the load up, which is a defensible default and an honest one to name.
Honest limits
This article cannot tell you your own error. Everything above is a population average, and the scatter around those averages is enormous — the meta-analysis reports I² = 97.9%, which means the studies disagree with each other far more than the pooled number suggests. Nor is the underlying evidence broad. Almost all of it is bench press, back squat or knee extension, performed at 65–85% of one-rep max, by young, mostly male, mostly resistance-trained participants, with samples typically between 14 and 46 people. Nothing here tells you how accurately you rate a lateral raise at 15 reps, and the one systematic review that maps the moderators does so narratively, without pooled effect sizes. Where sex differences appear they are load-dependent, from small samples, and contradicted by a balanced-sample trial — treat them as unsettled. We also do not know whether rating practice makes you better at rating: no trial has tested it, so the familiar advice to take an occasional set to genuine failure as calibration is practitioner convention, and we have labelled it as such rather than dressing it up. Finally, Shojin's own thresholds are conventions, not findings. No trial has compared "add load at RPE ≤ 8" against any alternative, and nothing in the literature supports 5% specifically as a deload, 45 days specifically as a stall window, or 10 seconds specifically as a step for timed work. They are conservative defaults chosen so that being wrong is cheap. This is general fitness guidance, not medical advice.
Common questions
Is RPE or RIR more accurate?
Neither. They are the same judgement expressed on inverted axes — RPE = 10 − RIR — so the accuracy problem is identical whichever you write down. No study has found one framing rated more accurately than the other.
What does RPE 8 actually mean?
Two repetitions in reserve: you believe you could have completed two more before failing. In the RIR convention the number is a claim about repetitions, not about how unpleasant the set felt.
Do experienced lifters rate their sets more accurately?
This is contested, and the higher-tier evidence says no. The meta-analysis covering 414 participants found training status did not moderate accuracy, and two later trials found no difference by years of experience. Some studies report a tendency in that direction, but experienced lifters are still off by roughly a rep in the easiest conditions. Familiarity with the specific lift, and with rating itself, is probably the more useful thing than training age.
Should I use half-points like RPE 8.5?
Not in Shojin, which offers whole numbers only. Given that the honest resolution of the measurement is on the order of one repetition, a decimal place adds precision the instrument does not have. That is our reading of the accuracy data rather than a tested finding.
Do I have to log RPE for Shojin to work?
No. The field is optional and can be hidden entirely from your profile. If you leave it blank, the app treats the set as submaximal and suggests adding load next time — a deliberate choice, and one that leans the opposite way to the population bias, since lifters tend to over-rate how close to failure they were.
These are the same answers the page’s FAQ structured data publishes — visible text, no hidden-content mismatch.
Sources
Every number in this article traces to one of these. Where the evidence is contested or thin, the article says so rather than picking a side.
- 01
Helms ER, Cronin J, Storey A, Zourdos MC (2016). Strength and Conditioning Journal 38(4):42–49
- 02
Foster C, Florhaug JA, Franklin J, et al. (2001). A new approach to monitoring exercise training. Journal of Strength and Conditioning Research 15(1):109–115
- 03
Haddad M, Stylianides G, Djaoui L, Dellal A, Chamari K (2017). Session-RPE Method for Training Load Monitoring. Frontiers in Neuroscience 11:612
- 04
Zourdos MC, Klemp A, Dolan C, et al. (2016). Journal of Strength and Conditioning Research 30(1):267–275
- 05
Halperin I, Malleron T, Har-Nir I, et al. (2022). Sports Medicine 52(2):377–390
- 06
Refalo MC, Remmert JF, Pelland JC, et al. (2024). Journal of Strength and Conditioning Research 38(3):e78–e85
- 07
Droguett FAB, Festa RR, Quintana NAT, et al. (2025). Journal of Human Kinetics 102:145–158
- 08
Armes C, Standish-Hunt H, Androulakis-Korakakis P, et al. (2020). "Just One More Rep!" — Ability to Predict Proximity to Task Failure in Resistance Trained Persons. Frontiers in Psychology 11:565416
- 09
Steele J, Endres A, Fisher J, Gentil P, Giessing J (2017). PeerJ 5:e4105
- 10
Emanuel A, Har-Nir I, Obolski U, Halperin I (2022). Sports Medicine – Open 8:132
- 11
Ruiz-Alias SA, Baena-Raya A, Hernández-Martínez A, et al. (2024). Estimating Repetitions in Reserve During the Bench Press Exercise: Should We Consider Sex and the Exercise Equipment? Sports Health 17(5):1007–1012
- 12
Russo F, Marconcin P, Gomes D, et al. (2026). Physical Therapy Reviews 31(1):46–63; Refalo MC, et al. (2024). JSCR 38(3):e78–e85
- 13
Refalo MC, Helms ER, Trexler ET, Hamilton DL, Fyfe JJ (2023). Influence of Resistance Training Proximity-to-Failure on Skeletal Muscle Hypertrophy: A Systematic Review with Meta-analysis. Sports Medicine 53(3):649–665
- 14
Robinson ZP, Pelland JC, Remmert JF, et al. (2024). Exploring the Dose–Response Relationship Between Estimated Resistance Training Proximity to Failure, Strength Gain, and Muscle Hypertrophy: A Series of Meta-Regressions. Sports Medicine
- 15
Hickmott LM, Chilibeck PD, Shaw KA, Butcher SJ (2022). The Effect of Load and Volume Autoregulation on Muscular Strength and Hypertrophy: A Systematic Review and Meta-Analysis. Sports Medicine – Open 8:9
- 16
Pelland JC, Robinson ZP, Remmert JF, et al. (2022). Methods for Controlling and Reporting Resistance Training Proximity to Failure: Current Issues and Future Directions. Sports Medicine 52(7):1461–1472
- 17
Jukic I, Prnjak K, Helms ER, McGuigan MR (2024). Modeling the repetitions-in-reserve-velocity relationship. Physiological Reports 12(5):e15955
- 18
Larsen S, Kristiansen E, van den Tillaar R (2021). Effects of subjective and objective autoregulation methods for intensity and volume on enhancing maximal strength during resistance-training interventions: a systematic review. PeerJ 9:e10663
- 19
Huang Z, Sun J, Li D, Chen C, Wang D (2025). Autoregulated resistance training for maximal strength enhancement: A systematic review and network meta-analysis. Journal of Exercise Science and Fitness 23(4):360–369