Proven Pathways to EliteLaw Schools and Beyond.

From LSAT mastery to T14 admissions and BigLaw careers,
our students achieve outcomes that transform futures.

|

July 30, 2026

LSAT Plateau: Why Scores Stop Improving and Drop

The most common cause of an LSAT plateau is review failure: doing questions without extracting why you missed them. Feedback roughly doubles the benefit of practice, from g = 0.39 without it to g = 0.73 with it, in the largest meta-analysis of the testing effect. Volume without review buys half of what you paid for.

Lovare Institut is running an IRB-approved randomized controlled trial on test anxiety and LSAT performance with Professor Ian Lyons at Georgetown's Psychology department. That study is ongoing and we will publish its findings when they exist rather than before, so nothing on this page rests on it.

Why is my LSAT score not improving?

Start with the mechanism rather than the mood. Practice improves performance through retrieval, and retrieval works far better when it is followed by feedback that corrects the retrieval.

The pooled testing effect across 61 studies and 159 effect sizes is g = 0.50, and it splits sharply on feedback: g = 0.73 with feedback and g = 0.39 without (Rowland, 2014, Psychological Bulletin, https://pubmed.ncbi.nlm.nih.gov/25150680/). It also grows with delay, at g = 0.69 for retention intervals of a day or more versus g = 0.41 under a day.

Translate that into study behavior. Doing forty questions and checking the answer key is retrieval without feedback; doing twenty questions and writing one sentence per miss explaining why the right answer is right is retrieval with feedback.

The second one is slower, produces less visible output, and is worth roughly twice as much. Most plateaus are the first behavior repeated for eight weeks.

More practice tests is not the answer

This is the intervention students increase when they stall, and it has the weakest evidence of anything they could increase.

The most recent meta-analysis of test preparation for large-scale educational tests, covering 28 studies, 92 effect sizes, and over 16,000 participants, found an overall effect of g = .26. Within that, practice tests alone were the weakest and statistically nonsignificant intervention at g = .19, p = .233, while workbooks reached g = .40 and explicit teaching of test-taking skills g = .35 (Hao, Baird, El Masri and Double, 2025, Review of Educational Research, https://journals.sagepub.com/doi/10.3102/00346543251360775).

InterventionEffectSignificant?Workbooksg = .40YesTeaching test-taking skillsg = .35YesSample itemsg = .23YesPractice tests aloneg = .19No, p = .233

Source: Hao et al. (2025), Review of Educational Research.

The same analysis found no dose-response: neither contact hours (p = .489) nor duration in weeks (p = .314) predicted outcomes. If you have taken twelve practice tests and your score has not moved, a thirteenth is the least likely thing on the menu to change it.

The illusion that keeps people in the wrong method

The reason bad study methods persist is that they feel better than good ones. This is one of the best-replicated findings in learning research and it has a specific shape.

In the classic demonstration, students who studied a passage four times predicted the highest retention, rating themselves 4.8 on a 7-point scale, while students who studied once and tested three times predicted 4.0. At a one-week delay the repeated-study group recalled 40 percent and the tested group 61 percent (Roediger and Karpicke, 2006, Psychological Science, https://colinallen.dnsalias.org/Readings/2006_Roediger_Karpicke_PsychSci.pdf).

The confidence ranking was the exact inverse of the performance ranking. Worse, at a five-minute delay the repeated-study group actually won, 81 percent to 75 percent, which is why students abandon the better method: in the short run it looks worse.

The same inversion shows up for spacing. Across experiments, 78 percent of participants judged massed practice as good as or better than spaced practice while 78 percent actually performed better spaced (Kornell and Bjork, 2008, Psychological Science, https://gwern.net/doc/psychology/spaced-repetition/2008-kornell.pdf).

And in a randomized crossover with 149 Harvard physics students, those taught actively rated their own learning 0.56 standard deviations lower while scoring 0.46 standard deviations higher (Deslauriers, McCarty, Miller, Callaghan and Kestin, 2019, PNAS, https://www.pnas.org/doi/10.1073/pnas.1821936116). If your study sessions feel smooth and productive, that is weak evidence you are learning and moderate evidence you are not.

Is the plateau even real?

Here the honest answer is that psychology has argued about this for over a century and has not settled it. That matters, because how you interpret a flat stretch determines what you do about it.

The classic learning curve, in which performance improves as a power function of practice, turns out to be substantially an artifact of averaging. Across 40 datasets and 7,910 individual learning series, an exponential fit beat a power fit in 82.2 percent of individual series, and the authors concluded there is little empirical evidence from individual learners that a power function describes skill acquisition (Heathcote, Brown and Mewhort, 2000, Psychonomic Bulletin and Review, https://link.springer.com/content/pdf/10.3758/BF03212979.pdf).

The mechanism of that artifact matters for you. Averaging across people manufactures smooth curves that no individual actually follows, and averaging across sessions can both create apparent plateaus and hide real ones.

The current framing in the field is that individual skill acquisition is better described as a series of plateaus, dips, and leaps than as a smooth ascent, and that the century-old debate about whether plateaus are real remains unresolved (Gray, 2017, Current Directions in Psychological Science, https://journals.sagepub.com/doi/abs/10.1177/0963721416672904). Practically: a three-week flat stretch inside noisy practice-test data may be nothing at all.

What actually breaks a plateau

The framework with the best support is that plateaus break on new techniques rather than on more repetition. In the Tetris research program, the limit on expertise was not raw speed but the techniques players had acquired, and one experienced player spent six months unlearning a habitual rotation strategy to acquire a better one (Gray and Banerjee, 2021, Topics in Cognitive Science, https://openreview.net/pdf?id=KvSo086RzY).

The strongest hard evidence for changing practice structure comes from interleaving. In a cluster randomized trial of 787 seventh-grade students across 54 classrooms, interleaved practice produced 61 percent correct against 38 percent for blocked practice at roughly a 33-day delay, d = 0.83 (Rohrer, Dedrick, Hartwig and Cheung, 2019, Journal of Educational Psychology, https://gwern.net/doc/psychology/spaced-repetition/2019-rohrer.pdf).

So the prescription for a stalled score is structural, not motivational. Stop full tests for three weeks, interleave question types instead of blocking them, and convert every miss into a written explanation rather than a tally mark.

One caution about a related claim. The intuitive idea that you should target your specific limiting subskill is theoretically motivated and has indirect support, but we could not locate a meta-analytic effect size isolating it, so we present it as a working hypothesis rather than an established result.

Why did my LSAT score drop on test day?

Three candidate explanations, in descending order of how often they are the real one.

First, regression to the mean. Your practice-test average is a distribution, and your best practice score is by definition its upper tail, so a test-day score near your average is not a drop, it is the expected value you were comparing against the wrong number.

Second, execution under novel conditions. The scored test is two Logical Reasoning sections and one Reading Comprehension section, each 35 minutes, with one unscored variable section and a 10-minute intermission after section two (https://www.lsac.org/lsat/register-lsat/accommodations/specifications-lsat-and-lsat-argumentative-writing). Beginning with the August 2026 administration, almost all test takers now sit the exam in person at Prometric centers rather than at home (https://www.lsac.org/lsat), which makes home-desk practice a less faithful rehearsal than it was.

Third, and much less often than people assume, anxiety. We take that up below, with the evidence and its ceiling.

For calibration, LSAC's repeater data shows mean second-attempt gains between 2.18 and 2.69 points across recent testing years, at 2.39 in 2024-25, with the smoothed distributions showing most score changes falling between -10 and +15 (LSAC TR 26-01, https://www.lsac.org/sites/default/files/research/TR-26-01.pdf). Scores moving down on a retake is a normal part of that distribution, not an anomaly requiring explanation.

Sleep: what the evidence supports and what it does not

The popular advice is directionally right and specifically wrong in a way that matters.

In the best-specified meta-analysis of sleep deprivation and cognition, covering 70 articles and 209 aggregated effect sizes, simple attention was hit hardest: lapses at g = 0.762 and reaction time at g = 0.732. Reasoning was the least affected domain, with accuracy at g = 0.125 and not statistically significant (Lim and Dinges, 2010, Psychological Bulletin, https://www.med.upenn.edu/uep/assets/user-content/documents/LimDinges2010MetaAnalysis.pdf).

So a sleep-deprived test taker most likely loses points to attentional lapses across a long timed test, not to a collapse in reasoning ability. That is an argument for stamina protection rather than for panic about cognitive damage.

The chronic restriction finding is the one that should change your schedule. After 14 days at six hours a night, attention lapses reached the equivalent of one full night of total sleep deprivation, and at four hours the equivalent of two nights, while subjects rated themselves only slightly sleepy (Van Dongen, Maislin, Mullington and Dinges, 2003, Sleep, https://academic.oup.com/sleep/article-abstract/26/2/117/2709164).

Your sense of how tired you are does not track your impairment. That is the whole finding, and it is why a study schedule that runs on six hours for a month feels sustainable while quietly degrading the exact faculty the test measures.

One nuance that cuts against folk wisdom: in a semester-long study of 88 students, sleep on the single night before each exam showed no significant correlation with performance, while sleep across the month before correlated at r = .25 to .34 (Okano, Kaczmarzyk, Dave, Gabrieli and Grossman, 2019, npj Science of Learning, https://www.nature.com/articles/s41539-019-0055-z). Correlational, single course, small sample, and still a useful corrective: the month matters more than the night.

Does anxiety affect LSAT scores?

Yes, and less than the internet says. Both halves of that sentence are supported.

The best modern pooled estimate of the relationship between test anxiety and achievement is r = -.23, 95 percent CI [-.26, -.19], across 177 studies and 906,311 participants (Caviola, Toffalini, Giofre, Ruiz, Szucs and Mammarella, 2022, Educational Psychology Review, https://link.springer.com/article/10.1007/s10648-021-09618-5). A correlation of that size is real and modest.

The ceiling is the number to hold onto. In a study of 730 university students, test anxiety explained only 2 to 5 percent of variance across standardized tests, and the components that carried unique variance were interference and lack of confidence rather than worry or arousal (Schillinger, Mosbacher, Brunner, Vogel and Grabner, 2021, Educational Psychology Review, https://d-nb.info/1231607017/34).

That arithmetically caps how many points eliminating anxiety could recover. It is not zero and it is not the twelve points people attribute to it.

What anxiety interventions actually deliver

The asymmetry in the intervention literature is the most useful finding for anyone deciding where to spend effort.

Across 44 randomized controlled trials with 2,209 university students, interventions reduced test anxiety at g = -0.76 but improved academic performance at only g = 0.37 (Huntley, Young, Temple, Longworth, Smith, Jha and Fisher, 2019, Journal of Anxiety Disorders, https://pubmed.ncbi.nlm.nih.gov/30826687/). Reducing distress is roughly twice as large an effect as raising scores.

The authors also caution that evidence of publication bias was found and that poor reporting quality means confidence in the results should be moderated. We repeat that caveat rather than burying it.

One widely repeated intervention deserves a direct correction. The expressive-writing finding, in which students write about their worries for ten minutes before an exam, was one of 21 Nature and Science social science experiments subjected to high-powered replication, and it is among the eight that did not produce a significant effect in the same direction (Camerer et al., 2018, Nature Human Behaviour, https://www.nature.com/articles/s41562-018-0399-z).

It may still help you personally, it costs ten minutes, and it is not the established result it is usually presented as. Anxiety reduction is worth pursuing because feeling less awful during a hard process is a good in itself, and it is a weaker score lever than the review protocol.

The diagnostic sequence: find your actual bottleneck

Run these four checks in order before changing anything about your study plan.

One, split your last three timed sections into blind review piles: questions you got right on untimed reattempt are timing problems, questions you got wrong are knowledge problems. Two, count your error log entries by question type and look for a concentration.

Three, check whether your review time exceeds your practice time. If it does not, you have found the plateau and no further diagnosis is needed.

Four, check the boring variables: sleep across the last month, whether you are practicing in conditions resembling a test center, and whether you have taken more than four full tests in the last month. The full protocol sits inside our three-month study plan.

Deliberate practice, and why the popular version overpromises

The idea that ten thousand hours of focused practice produces expertise is the folk theory behind most LSAT study plans. The evidence for it is considerably weaker than its cultural status suggests, and the weakness matters most in exactly the domains closest to test taking.

Across 157 studies and 11,135 participants, deliberate practice explained 12 percent of performance variance overall. Broken out by domain: games 26 percent, music 21 percent, sports 18 percent, education 4 percent, and professions under 1 percent (Macnamara, Hambrick and Oswald, 2014, Psychological Science, https://gwern.net/doc/psychology/2014-macnamara.pdf).

Education and professions, the two domains most like preparing for a professional entrance exam, show the weakest effects in the entire analysis. The estimate also shrinks as measurement improves, from r = .45 with retrospective interviews to r = .22 with contemporaneous logs.

A direct replication of the original violinist study found the good group had out-practiced the best group, with the difference between them nonsignificant (Macnamara and Maitra, 2019, Royal Society Open Science, https://royalsocietypublishing.org/doi/full/10.1098/rsos.190327). The practical reading is not that practice fails, it is that hours logged is a poor proxy for the thing that works.

Feedback is not automatically good either

Since this page rests heavily on feedback, the honest caveat belongs here rather than in a footnote. Feedback interventions average a positive effect of d = 0.41 across 607 effect sizes and 12,652 participants, and in 38 percent of cases they decreased performance (Kluger and DeNisi, 1996, Psychological Bulletin, https://mrbartonmaths.com/resourcesnew/8.%20Research/Marking%20and%20Feedback/The%20effects%20of%20feedback%20interventions.pdf).

The pattern in that literature is that feedback directed at the task helps and feedback directed at the self does not. Applied to LSAT review: writing why the correct answer is correct is task feedback, and concluding that you are bad at inference questions is self feedback.

Keep your error log free of self-assessment. Every entry should describe the question and your reasoning, not your aptitude.

Stereotype threat: what the current evidence shows

This comes up in test anxiety conversations and deserves an accurate account rather than either the popular version or a dismissal.

The largest meta-analysis covers 212 adult samples and reports an overall d = -.31 with a 90 percent credibility interval from -.93 to .32, which spans zero (Shewach, Sackett and Quint, 2019, Journal of Applied Psychology, https://gwern.net/doc/psychology/cognitive-bias/stereotype-threat/2019-shewach.pdf). Restricted to operationally relevant conditions, the effect falls to d = -.14.

In actual high-stakes settings, real college placement exams, the effect was essentially zero at d = -.01 across four studies, significantly different from the lab estimate of -.36. With financial incentives present, d = .00.

After correction for publication bias, the operational estimate moves to d = -.09 with a confidence interval crossing zero. The same analysis identified an analytic error inflating roughly 15 percent of studies by about 67 percent.

We report this because the honest version is more useful than the reassuring one. On a high-stakes test you are paying to take, the best available evidence does not support treating stereotype threat as a major driver of your score.

The three plateaus that are not plateaus

Before restructuring your entire method, rule out the three cases where nothing is actually wrong.

First, noise. Practice test scores carry measurement error, and a three-point band across four tests is consistent with no change and with slow improvement equally well. Four data points cannot distinguish them.

Second, the wrong comparison. If you are measuring against your best score rather than your average, you will perceive a decline every time you produce a normal result, which is most of the time.

Third, an underlying gain hidden by rising difficulty or by fatigue accumulated across a heavy week. Testing on Saturday after a fifty-hour work week measures something other than your ability.

Fix the measurement before you fix the method. Take the score seriously only when it is a rolling average across at least four tests taken in comparable conditions.

A four-week plateau protocol

If the measurement is clean and the score is genuinely flat, run this.

Week one, no timed work at all. Re-review your last three tests using blind review from scratch, sorting every miss into timing versus knowledge, and write the one-sentence explanation for each.

Week two, drill only the two question types that dominate your knowledge pile, interleaved rather than blocked, untimed. The interleaving evidence from classroom mathematics is the strongest available support for changing practice structure, at d = 0.83 in a cluster randomized trial of 787 students.

Week three, reintroduce single timed sections with the same two types heavily represented. Week four, one full test, then compare its error concentration to your log.

If the concentration has moved, the protocol worked and you repeat it on the next two types. If it has not, the problem is more likely execution under time than knowledge, which is a different fix and usually a pacing one.

A note on how to read the numbers on this page

Effect sizes like g and d express differences in standard deviations, and they are easy to over-read. A g of 0.50 means the average person in the better condition scored about half a standard deviation above the average person in the worse one, with substantial overlap between the groups.

Two implications for you. First, none of these findings guarantee an individual outcome, since they describe distributions rather than people. Second, a nonsignificant result like the g = .19 for practice tests does not prove practice tests do nothing, it means the evidence did not distinguish their effect from zero.

We report confidence intervals and caveats throughout because prep marketing generally does not. If a claim on this page looks weaker than what a course sold you, that is the point of the page.

The evidence, summarized

ClaimBest estimateSourceTesting beats restudyg = 0.50Rowland (2014), Psych BulletinWith feedback vs withoutg = 0.73 vs 0.39Rowland (2014)Test prep overallg = .26Hao et al. (2025), Rev Educ ResearchPractice tests aloneg = .19, nonsignificantHao et al. (2025)Interleaved vs blocked practiced = 0.83Rohrer et al. (2019), J Educ PsychTest anxiety and achievementr = -.23Caviola et al. (2022), Educ Psych RevAnxiety interventions: anxiety vs performanceg = -0.76 vs 0.37Huntley et al. (2019), J Anxiety DisordersSleep deprivation: reasoning accuracyg = 0.125, nonsignificantLim and Dinges (2010), Psych BulletinSleep deprivation: attention lapsesg = 0.762Lim and Dinges (2010)Mean LSAT second-attempt gain, 2024-252.39 pointsLSAC TR 26-01

Every row links to a source cited in full above. Where a result is contested or has failed replication, that is stated in the section rather than in this summary.

FAQ

Why is my LSAT score not improving?

Most often because practice is not being converted into feedback. The testing effect roughly doubles when practice is followed by corrective review, g = 0.73 with feedback versus g = 0.39 without, so drilling without written review captures about half the available gain.

Will taking more practice tests break my plateau?

Probably not. In the most recent meta-analysis of test preparation, practice tests alone were the weakest intervention at a nonsignificant g = .19, and neither contact hours nor program duration predicted outcomes.

Why did my LSAT score drop on test day?

Usually regression to the mean: your best practice score is the top of a distribution, not your true level. Execution under unfamiliar conditions is the second cause, which matters more now that almost all test takers sit the exam in person at Prometric centers rather than at home.

Does anxiety affect LSAT scores?

Modestly. The best pooled estimate of the test anxiety and achievement relationship is r = -.23, and one careful study found anxiety explained only 2 to 5 percent of variance in standardized test performance. Interventions reduce anxiety about twice as effectively as they raise scores.

Does sleep matter more than studying?

Sleep across the month before matters more than the night before. Chronic restriction to six hours produced attention deficits equivalent to a full night of total sleep loss within two weeks, while subjects reported feeling only slightly sleepy.

Written by Ali, Georgetown Law, founder of Lovare Institut.

Book a call

Read more guides

July 30, 2026

How Much Does Bar Prep Cost Beyond the Sticker Price
How much does bar prep cost is the wrong question, because the sticker is the smallest part. Courses run $1,199 to $3,099 street. Then come $1,320 in...
Read More

July 30, 2026

NextGen Bar Prep Courses: Who Has Shipped and Who Has Not
NextGen bar prep courses have shipped at seven providers and not at four, with SmartBarPrep a full exam cycle behind. The finding that matters: no...
Read More