Glossary

Inference

Inference means using a fully trained model. It receives an input and computes an output with its fixed parameters. For a language model, every answer to a prompt is inference.

Unlike in training, nothing is compared and no parameter is adjusted. That is why inference needs far less computing time and memory per request than training does.

Mental image: the concert after the sound check — the faders stay put, and the desk processes whatever comes in.

An example: A trained spam filter receives a new email and computes a verdict with its weights: spam or inbox. The weights stay unchanged. The same goes for a chatbot: it produces its answer to your question one text piece at a time, and for every piece it computes with the same parameters as for every other request.

Not to be confused with learning during a conversation: A chatbot does not learn while you chat with it. During inference, no parameter moves, whatever you type. Whatever it “remembers” within a conversation is sent along as input with every message. And “inference” here does not mean drawing a logical conclusion; it simply means using the model.

Where you’ll come across it: With providers of AI services, who often bill the use of their models by tokens. In news about data centers and AI chips that distinguishes hardware for training from hardware for inference. A single request is cheap, but because millions of people ask questions, inference as a whole also needs large data centers.

Introduced in Parameters, Training vs. Inference, Hardware: How a Model Runs.

Last changed on · commit e7392c5