Sightline PrepGet started

Why a third-party practice test score is close to meaningless

A student came to me last year convinced they had dropped 120 points in a fortnight. They had taken one practice test from a well-known prep company, then another from a different one. Same student, same fortnight, same amount of studying.

Neither number meant anything. That is not a criticism of how hard the student worked — it is a statement about how those tests are scored.

What a curve actually is

Turning a number of correct answers into a score out of 800 is not arithmetic. It requires knowing how a large, representative population performed on those exact questions. College Board builds that from real administrations with hundreds of thousands of test takers, and it is the expensive part of writing a test — far more expensive than writing the questions.

A third-party publisher has none of that. They write questions that look like SAT questions, then estimate a conversion. Sometimes the estimate comes from a small internal sample. Often it is a copy of a published official curve stapled onto a different set of questions.

The result is that the same raw performance can come out 50 to 200 points apart depending on whose test you sat. I have seen a student score 1340 and 1490 in the same week without their ability changing at all.

Why the questions are wrong even when they look right

The curve is the bigger problem, but not the only one.

Unofficial questions tend to be harder in the wrong way. Writing a genuinely hard SAT question — one where the wrong answers are tempting for specific, diagnosable reasons — is difficult. Writing a question that merely takes longer is easy. So third-party math sections drift towards heavy computation, and third-party reading drifts towards obscure vocabulary, neither of which is what the real test does.

They also miss the traps. A real SAT question usually has one wrong answer built for the student who did the right work and stopped one step early, and another for the student who misread a single word. That construction is the whole point of the test. Practice questions without it train you to arrive at an answer, but not to check whether it is the answer to the question that was asked.

And on the digital SAT there is a structural problem: an unofficial test cannot reproduce the adaptive routing. A fixed-form practice test does not have a second module that changes based on your first. So the number at the end is answering a question the real test does not ask.

What they are genuinely good for

I use third-party material constantly, just never for scores.

For drilling a single topic, unofficial questions are excellent, and there are far more of them than the official supply. If a student needs thirty questions on circle equations this week, no official source has thirty spare.

For rebuilding a broken rule, they are ideal, because I want volume on one narrow thing rather than a realistic mix.

For untimed practice, they are fine. The distortions that matter are distortions of difficulty and pacing, and both stop mattering when the clock is off.

What I will not do is let a student read a score off one and draw a conclusion.

What to use instead

The official Bluebook app has full practice tests, and they are free. They adapt like the real thing, they are scored on real curves, and there are enough of them to structure several months around if you do not waste them.

Which is the constraint worth planning for: the official supply is finite. A student who burns through them in six weeks has nothing left to measure with in October. I space them out and use unofficial material for everything in between.

The number to watch instead of the score

Even on an official test, the total is the least informative thing on the report. Two students with the same 1300 can need completely different months of work — one is losing points to pacing on questions they know, the other is confidently applying a rule that is wrong.

That distinction does not appear in any score, official or not. It only appears when you look at how each question was answered: how long it took, and how sure the student felt. Our free diagnostic shows that split on eight questions in a few minutes, and it is the same thing I look at on a full official form for every student I work with.