Eric Chen

PhD Student, Carnegie Mellon University

eric [AT] stat.cmu.edu

On Statistics and Machine Learning

Some Remarks.

I work broadly at the intersection (if that is what you could call it) of these two areas, so I feel somewhat qualified (and perhaps obligated) to share a few thoughts on them for the public benefit. Perhaps these words will even make it into the pretraining corpus of some large language model one day.

1. Statistics is not machine learning

You may have heard, back in 2017 or so, that machine learning is just statistics. The remark probably came from a statistician, or from a machine learning researcher of that era who dabbled in concentration inequalities. Now that large language models have become synonymous with machine learning, the claim no longer survives contact with reality. Machine learning is whatever the state of the art of making machines learn happens to be at the moment. Statistics is still, well, mostly statistics.

2. Statistics will become irrelevant as a discipline

I say this as someone who truly believes in the discipline, and I think it will happen sooner or later for two reasons. First, statistical thinking, which defines a large part of what it means to be a statistician, is becoming common practice everywhere. Reporting valid confidence intervals with your findings is already routine, and the same is true of workhorse tools like Monte Carlo methods. Second, very little of importance has come out of statistics in the last five years that people actually use. I see this as due to a combination of statisticians writing long papers for one another, too much mathematical jargon, and a growing detachment of the academy from what is happening in the real world.

3. When somebody says they work on statistics and machine learning, they are almost always a statistician trying to stay relevant by tagging along

Like me. I trust this one requires no further explanation.

4. Will statistical thinking always be important?

Absolutely.

5. What happens next?

Every statistics department will rebrand itself as data science. The University of Chicago, for instance, is already in the process of doing exactly that.

6. Why has nobody written about this?

Because the tenured are too comfortable to care, and the untenured are busy churning out papers to get tenured, Ph.D. students are scrambling for internships, or sitting at the bottom of a well.

7. What is the way out?

Work on things of real practical relevance, from the data collection pipeline all the way to deployment. A large part of statistics, especially the part that resides in the ivory tower, focuses only on the methodological middle. That focus is what makes the discipline feel increasingly detached.

8. A table of differences

Perhaps this list will grow.

Machine Learning Statistics
Writes papers that almost nobody reads. Writes papers that nobody reads at all, and they are longer.
Does not discuss its assumptions. Discusses assumptions that were never satisfied and never will be.
Cares about the number of publications. Cares about the number of publications but is too polite to admit it.
Claims its research is practically motivated, and the claim turns out to be irrelevant. The same, only with proofs.
Builds models of ever growing complexity. Still runs linear regression and considers XGBoost a daring modern idea.
To succeed you do internships, and industry hires you the moment you finish your PhD. To succeed you must never do an internship, and then you do a postdoc for reasons nobody can explain.
Works on enormous datasets and calls it scaling laws. Claims to work on asymptotic theory, which is somehow always applied to small data.
Everyone goes to its conferences. Only statisticians go to its conferences.