LLMs are estimators

In response to a post in which it is asked what topics an emerging statistician should be familiar with, and whether they should know some basic theory about large language models (LLMs).

LLMs are estimators, completely in the statistician’s wheelhouse. I would say that the average statistician understands LLMs very well, because the average statistician understands estimators (not just how to estimate them, but also how they behave (inference)). However, the current LLM terminology makes them seem like some new species. We see the standard things like estimation, extrapolation, prediction, etc anthropomorphized as learning, hallucinating, reasoning, etc.  It is though important (and in my opinion interesting) to study text modeling.

Further, decision theory interacts very closely with the reinforcement learning used to then take an LLM and build a human-in-the-loop chatbot (in fact it is an application of decision theory).

I talk about all this in the paper “Treatment, evidence, imitation, chat.”

I hope that statisticians engage fully here, for the sake of the healthcare system, at least, and probably many other parts of society, too.  

Leave a comment