Answer

Why is accuracy a poor way to judge a classification model?

Inspired by a question on Cross Validated ·

classificationlogistic regressionmodel selection

The short answer

Accuracy counts how often a model's yes/no labels are right, and that throws away most of what matters. With a rare outcome, a model that always says "no" looks excellent. Accuracy also ignores how confident the predicted probabilities are and treats a missed case and a false alarm as equally bad. Judge the predicted probabilities with a proper scoring rule such as the Brier score or log loss, and choose a cut-off only when you know the costs of each kind of error.

The short answer

Accuracy is the proportion of cases a classifier labels correctly. It feels like the obvious score, but it measures a thin slice of model quality, and in three common situations it points you the wrong way.

Our intuition misleads us because we picture a balanced problem with a fixed cut-off and equal costs. Real problems seldom look like that.

See it in R and Python

The code simulates an outcome that about 1 person in 20 has, driven by one predictor. It fits a logistic regression on 5,000 training cases and scores three sets of predictions on 5,000 new cases: a "nobody has it" rule that always predicts the training prevalence, the fitted model, and an overconfident copy of the model that keeps its labels but replaces every probability with 0.99 or 0.01. Lower is better for the Brier score (mean squared error of the probabilities) and log loss. The last part assumes a missed case costs ten times as much as a false alarm.

R

set.seed(312780)

# Rare outcome (about 5%): one predictor, true logistic model
sim <- function(n) {
  x <- rnorm(n)
  data.frame(x = x, y = rbinom(n, 1, plogis(-4 + 1.5 * x)))
}
train <- sim(5000); test <- sim(5000)
round(c(prevalence = mean(test$y)), 3)

fit <- glm(y ~ x, family = binomial, data = train)
p_model <- predict(fit, newdata = test, type = "response")
p_none  <- rep(mean(train$y), nrow(test))          # "nobody has it"
p_over  <- ifelse(p_model > 0.5, 0.99, 0.01)        # same labels, overconfident

scores <- function(p, y, cut = 0.5) {
  pred <- as.numeric(p > cut)
  pc <- pmin(pmax(p, 1e-15), 1 - 1e-15)
  c(accuracy    = mean(pred == y),
    sensitivity = mean(pred[y == 1] == 1),
    brier       = mean((p - y)^2),
    log_loss    = -mean(y * log(pc) + (1 - y) * log(1 - pc)))
}
round(rbind(always_no     = scores(p_none, test$y),
            logistic      = scores(p_model, test$y),
            overconfident = scores(p_over, test$y)), 3)

# If a missed case costs 10 times a false alarm, the best cut-off is 1/11
cost <- function(p, y, cut) {
  pred <- as.numeric(p > cut)
  mean(10 * (pred == 0 & y == 1) + 1 * (pred == 1 & y == 0))
}
cut2 <- 1 / 11
round(rbind(cut_0.50 = c(accuracy = mean((p_model > 0.5) == test$y),
                         cost = cost(p_model, test$y, 0.5)),
            cut_0.09 = c(accuracy = mean((p_model > cut2) == test$y),
                         cost = cost(p_model, test$y, cut2))), 3)

Python

import numpy as np
from scipy.special import expit
from sklearn.linear_model import LogisticRegression

rng = np.random.default_rng(312780)

# Rare outcome (about 5%): one predictor, true logistic model
def sim(n):
    x = rng.normal(size=n)
    return x, rng.binomial(1, expit(-4 + 1.5 * x))

x_tr, y_tr = sim(5000)
x_te, y_te = sim(5000)
print("prevalence", round(y_te.mean(), 3))

fit = LogisticRegression(penalty=None).fit(x_tr[:, None], y_tr)
p_model = fit.predict_proba(x_te[:, None])[:, 1]
p_none = np.full(len(y_te), y_tr.mean())             # "nobody has it"
p_over = np.where(p_model > 0.5, 0.99, 0.01)          # same labels, overconfident

def scores(p, y, cut=0.5):
    pred = (p > cut).astype(int)
    pc = np.clip(p, 1e-15, 1 - 1e-15)
    return {"accuracy": np.mean(pred == y),
            "sensitivity": np.mean(pred[y == 1] == 1),
            "brier": np.mean((p - y) ** 2),
            "log_loss": -np.mean(y * np.log(pc) + (1 - y) * np.log(1 - pc))}

for name, p in [("always_no", p_none), ("logistic", p_model), ("overconfident", p_over)]:
    print(name, {k: round(float(v), 3) for k, v in scores(p, y_te).items()})

# If a missed case costs 10 times a false alarm, the best cut-off is 1/11
def cost(p, y, cut):
    pred = (p > cut).astype(int)
    return np.mean(10 * ((pred == 0) & (y == 1)) + 1 * ((pred == 1) & (y == 0)))

for label, cut in [("cut_0.50", 0.5), ("cut_0.09", 1 / 11)]:
    print(label, "accuracy", round(float(np.mean((p_model > cut) == y_te)), 3),
          "cost", round(float(cost(p_model, y_te, cut)), 3))
prevalence 
     0.044 
              accuracy sensitivity brier log_loss
always_no        0.956       0.000 0.042    0.182
logistic         0.956       0.032 0.037    0.145
overconfident    0.956       0.032 0.044    0.214
         accuracy  cost
cut_0.50    0.956 0.431
cut_0.09    0.873 0.303

The output is from one seeded run of the R code. Python's random numbers differ from R's, so its figures differ slightly (prevalence 4.8%, and the model's accuracy edges the "nobody" rule by 0.954 to 0.952), but they show the same pattern: near-identical accuracy, clearly better Brier score and log loss for the model, the worst log loss for the overconfident copy, and a lower cost at the 1/11 cut-off.

What the numbers show

What to use instead

A scoring rule is proper when you get the best expected score by reporting the probabilities you truly believe. The Brier score and log loss are proper; accuracy is not, because it rewards pushing every probability past the cut-off and ignores everything else. So evaluate the probabilities first and treat classification as a separate decision:

Accuracy is still fine as a plain description when classes are roughly balanced and both errors cost about the same, but even then it is a blunt tool for comparing models. For choosing a statistical method more generally, try our test chooser or the DASS blog.

How to report classification accuracy in APA style (7th edition)

Never report accuracy alone. Give the outcome's prevalence, a proper score and the cut-off, so readers can see what the model adds over guessing the majority class. Using the R example:

"In the test sample (n = 5,000; prevalence 4.4%), the logistic regression model had a Brier score of .037, compared with .042 for a model that always predicted the base rate. At a cut-off of .50, accuracy was 95.6%, identical to the base-rate model, and sensitivity was 3.2%."

In the method section, say how the cut-off was chosen (for example, from the relative cost of missed cases) and whether performance was measured on new data. A table can list accuracy, sensitivity, specificity, Brier score and AUC side by side for each model.

Related tools and guides

More answered questions

Working with your own data?

General answers only go so far. Send us your situation and we'll reply by email within two business days.

Ask your question

Written by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.