How AI works
What is a model weight, and what is a transformer?
Short answerA weight is a number that survived training. A model is billions of them, arranged in layers, and together they decide how an input becomes an output. Nobody, including the people who trained the model, can point at one weight and say what it is for.
Start with a straight line
In y = mx + b, the slope m is the number you adjust until the line fits the data. In a neural network the same quantity is called a weight. Training is the process of adjusting billions of them at once.
One neuron is a weighted vote
An artificial neuron multiplies each input by a weight, adds the results together, adds an offset (the b in the line), and passes the total through a simple nonlinear function.
That last step is what lets layers add capability. Without it, any stack of layers collapses back into a single straight line, however many layers there are.
Billions of them
A model is billions of those weights, in layers. They are also called parameters: the numbers in a model that determine how an input becomes an output.
The weights are the only thing training produces. Everything a model learned is in those numbers, and no one can read a reason out of any single one. That is why a learned system cannot have its logic audited the way a rule system can, and why publishing a model's weights reveals so little about what went into them.
What a transformer is
The transformer is the architecture behind current language models. Google researchers published it in 2017, and most of the improvement in generative systems since then rests on it.
A transformer processes whole sequences at once. It uses a technique called attention to weigh how much each part of an input bears on every other part. For almost every purpose inside an organization, that is as deep as the architecture needs to go.
Where this shows up in practice
- Explainability. Asking why a learned model decided something is asking someone to read its weights. Ask for test results and a record of behavior, which can be produced.
- Size. A parameter count is a count of weights. It says how much there is to store and move, and it no longer predicts running cost. See Does all AI need a hyperscale data center?
- Control. Whoever holds the weights holds the model. An open-weight release hands them over; a hosted model keeps them with its provider.
Sources
- Congressional Research Service, Generative Artificial Intelligence: Overview, Issues, and Considerations for Congress (IF12426). Parameters, transformers, and attention.
Reviewed