Answer
Why is accuracy a poor way to judge a classification model?
The short answer
Accuracy counts how often a model's yes/no labels are right, and that throws away most of what matters. With a rare outcome, a model that always says "no" looks excellent. Accuracy also ignores how confident the predicted probabilities are and treats a missed case and a false alarm as equally bad. Judge the predicted probabilities with a proper scoring rule such as the Brier score or log loss, and choose a cut-off only when you know the costs of each kind of error.
The short answer
Accuracy is the proportion of cases a classifier labels correctly. It feels like the obvious score, but it measures a thin slice of model quality, and in three common situations it points you the wrong way.
- Rare outcomes. If 5% of people have a condition, a rule that says "nobody has it" is right 95% of the time while finding no one. A useful model can struggle to beat that number.
- It ignores the probabilities. Most models (logistic regression, random forests, neural networks) produce a probability, and accuracy only asks which side of 0.5 it lands on. A forecast of 0.51 and one of 0.99 count the same, so a model with badly overconfident probabilities can score exactly as well as a well-calibrated one.
- It hides the costs. Accuracy weighs a missed case and a false alarm equally. In screening, fraud detection or safety checks they rarely cost the same, and the cut-off that maximises accuracy is then the wrong cut-off.
Our intuition misleads us because we picture a balanced problem with a fixed cut-off and equal costs. Real problems seldom look like that.
See it in R and Python
The code simulates an outcome that about 1 person in 20 has, driven by one predictor. It fits a logistic regression on 5,000 training cases and scores three sets of predictions on 5,000 new cases: a "nobody has it" rule that always predicts the training prevalence, the fitted model, and an overconfident copy of the model that keeps its labels but replaces every probability with 0.99 or 0.01. Lower is better for the Brier score (mean squared error of the probabilities) and log loss. The last part assumes a missed case costs ten times as much as a false alarm.
R
set.seed(312780)
# Rare outcome (about 5%): one predictor, true logistic model
sim <- function(n) {
x <- rnorm(n)
data.frame(x = x, y = rbinom(n, 1, plogis(-4 + 1.5 * x)))
}
train <- sim(5000); test <- sim(5000)
round(c(prevalence = mean(test$y)), 3)
fit <- glm(y ~ x, family = binomial, data = train)
p_model <- predict(fit, newdata = test, type = "response")
p_none <- rep(mean(train$y), nrow(test)) # "nobody has it"
p_over <- ifelse(p_model > 0.5, 0.99, 0.01) # same labels, overconfident
scores <- function(p, y, cut = 0.5) {
pred <- as.numeric(p > cut)
pc <- pmin(pmax(p, 1e-15), 1 - 1e-15)
c(accuracy = mean(pred == y),
sensitivity = mean(pred[y == 1] == 1),
brier = mean((p - y)^2),
log_loss = -mean(y * log(pc) + (1 - y) * log(1 - pc)))
}
round(rbind(always_no = scores(p_none, test$y),
logistic = scores(p_model, test$y),
overconfident = scores(p_over, test$y)), 3)
# If a missed case costs 10 times a false alarm, the best cut-off is 1/11
cost <- function(p, y, cut) {
pred <- as.numeric(p > cut)
mean(10 * (pred == 0 & y == 1) + 1 * (pred == 1 & y == 0))
}
cut2 <- 1 / 11
round(rbind(cut_0.50 = c(accuracy = mean((p_model > 0.5) == test$y),
cost = cost(p_model, test$y, 0.5)),
cut_0.09 = c(accuracy = mean((p_model > cut2) == test$y),
cost = cost(p_model, test$y, cut2))), 3)Python
import numpy as np
from scipy.special import expit
from sklearn.linear_model import LogisticRegression
rng = np.random.default_rng(312780)
# Rare outcome (about 5%): one predictor, true logistic model
def sim(n):
x = rng.normal(size=n)
return x, rng.binomial(1, expit(-4 + 1.5 * x))
x_tr, y_tr = sim(5000)
x_te, y_te = sim(5000)
print("prevalence", round(y_te.mean(), 3))
fit = LogisticRegression(penalty=None).fit(x_tr[:, None], y_tr)
p_model = fit.predict_proba(x_te[:, None])[:, 1]
p_none = np.full(len(y_te), y_tr.mean()) # "nobody has it"
p_over = np.where(p_model > 0.5, 0.99, 0.01) # same labels, overconfident
def scores(p, y, cut=0.5):
pred = (p > cut).astype(int)
pc = np.clip(p, 1e-15, 1 - 1e-15)
return {"accuracy": np.mean(pred == y),
"sensitivity": np.mean(pred[y == 1] == 1),
"brier": np.mean((p - y) ** 2),
"log_loss": -np.mean(y * np.log(pc) + (1 - y) * np.log(1 - pc))}
for name, p in [("always_no", p_none), ("logistic", p_model), ("overconfident", p_over)]:
print(name, {k: round(float(v), 3) for k, v in scores(p, y_te).items()})
# If a missed case costs 10 times a false alarm, the best cut-off is 1/11
def cost(p, y, cut):
pred = (p > cut).astype(int)
return np.mean(10 * ((pred == 0) & (y == 1)) + 1 * ((pred == 1) & (y == 0)))
for label, cut in [("cut_0.50", 0.5), ("cut_0.09", 1 / 11)]:
print(label, "accuracy", round(float(np.mean((p_model > cut) == y_te)), 3),
"cost", round(float(cost(p_model, y_te, cut)), 3))prevalence
0.044
accuracy sensitivity brier log_loss
always_no 0.956 0.000 0.042 0.182
logistic 0.956 0.032 0.037 0.145
overconfident 0.956 0.032 0.044 0.214
accuracy cost
cut_0.50 0.956 0.431
cut_0.09 0.873 0.303
The output is from one seeded run of the R code. Python's random numbers differ from R's, so its figures differ slightly (prevalence 4.8%, and the model's accuracy edges the "nobody" rule by 0.954 to 0.952), but they show the same pattern: near-identical accuracy, clearly better Brier score and log loss for the model, the worst log loss for the overconfident copy, and a lower cost at the 1/11 cut-off.
What the numbers show
- Accuracy cannot tell the three apart. All three score 0.956. The "nobody" rule finds none of the cases, and at the 0.5 cut-off the model catches only about 3% of them, because a rare outcome seldom pushes anyone's probability above 0.5.
- Proper scoring rules can. The model's Brier score (0.037 vs 0.042) and log loss (0.145 vs 0.182) are clearly better than the "nobody" rule, so the predictor carries real information that accuracy missed.
- They also punish bad probabilities. The overconfident copy makes the same yes/no calls as the model, yet it has the worst Brier score (0.044) and log loss (0.214). Saying 0.99 or 0.01 when the truth is less certain is costly, and accuracy never notices.
- The best cut-off depends on costs. With a missed case ten times as costly as a false alarm, flagging everyone whose probability is above 1/11 lowers accuracy from 0.956 to 0.873 but cuts the average cost per case from 0.431 to 0.303, about 30% less. Maximising accuracy would choose the more expensive rule.
What to use instead
A scoring rule is proper when you get the best expected score by reporting the probabilities you truly believe. The Brier score and log loss are proper; accuracy is not, because it rewards pushing every probability past the cut-off and ignores everything else. So evaluate the probabilities first and treat classification as a separate decision:
- Brier score or log loss for overall quality of the probabilities. Log loss punishes confident mistakes very heavily; the Brier score is easier to explain.
- AUC (the c-statistic) for how well the model ranks cases above non-cases, regardless of cut-off. It says nothing about calibration, so pair it with a calibration plot. See our AUC explainer.
- A cut-off chosen from the costs. If a missed case costs c times a false alarm and the probabilities are calibrated, flag cases whose probability exceeds 1/(1 + c). Then report sensitivity, specificity and the predictive values at that cut-off.
- Do not rebalance the data just to rescue accuracy. Oversampling the rare class changes the predicted probabilities, so they no longer match the real prevalence unless you correct them afterwards.
Accuracy is still fine as a plain description when classes are roughly balanced and both errors cost about the same, but even then it is a blunt tool for comparing models. For choosing a statistical method more generally, try our test chooser or the DASS blog.
How to report classification accuracy in APA style (7th edition)
Never report accuracy alone. Give the outcome's prevalence, a proper score and the cut-off, so readers can see what the model adds over guessing the majority class. Using the R example:
"In the test sample (n = 5,000; prevalence 4.4%), the logistic regression model had a Brier score of .037, compared with .042 for a model that always predicted the base rate. At a cut-off of .50, accuracy was 95.6%, identical to the base-rate model, and sensitivity was 3.2%."
In the method section, say how the cut-off was chosen (for example, from the relative cost of missed cases) and whether performance was measured on new data. A table can list accuracy, sensitivity, specificity, Brier score and AUC side by side for each model.
Related tools and guides
- What is AUC? The area under the ROC curve, explained
- Logit vs probit: what is the difference, and which should you use?
- Is R-squared useful or misleading?
- Brier score (Wikipedia)
More answered questions
- What is AUC? The area under the ROC curve, explained
- Logit vs probit: what is the difference, and which should you use?
- AIC vs BIC: which model selection criterion should you use?
Working with your own data?
General answers only go so far. Send us your situation and we'll reply by email within two business days.
Ask your questionWritten by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.