What has made the recent radical development of LLMs possible?
ㅡㅡ
The recent radical development of LLMs was not caused by one discovery. It came from several developments reinforcing one another. But if we look for the deepest technical causes, I would emphasize Transformer architecture + enormous computation + enormous textual data + scaling + improved training/post-training.
1. The Transformer was the architectural breakthrough
The decisive starting point was the 2017 paper Attention Is All You Need. Its Transformer architecture replaced much of the sequential processing of earlier language models with attention. This made training much more parallelizable and therefore much better suited to GPUs and other large computing systems. (arXiv)
Attention also has a conceptual importance that interests me in relation to our discussions of context.
For a word such as bank, meaning cannot be determined from the isolated word:
He deposited money in the bank.
versus
He sat on the river bank.
The model uses relationships among tokens in the surrounding context to construct an appropriate representation. Transformer attention gives the system a powerful computational mechanism for doing this across text.
It would be too strong to say that attention is identical with human understanding of context. But the connection is important:
meaning depends upon relations → attention computes relations among contextual elements → contextual representations become increasingly informative.
2. Scale produced something much greater than expected
Researchers then discovered a remarkable empirical regularity: increasing model size, training data and computation tended to improve language-model performance in fairly predictable ways. OpenAI’s 2020 scaling-law research demonstrated power-law relationships between performance and these quantities. (OpenAI)
Later work showed that how the resources are balanced matters. DeepMind’s Chinchilla work demonstrated that a smaller model trained on substantially more data could outperform much larger models trained less efficiently. (arXiv)
So development was not simply:
more parameters → better AI
but increasingly:
appropriate architecture × parameters × high-quality data × compute × training method.
3. Humanity had already created the enormous training resource
The Internet, digitized books, journal literature, software code, reference materials and other electronic texts created an unprecedented corpus of recorded human expression.
The crucial training method was self-supervised learning. Humans did not have to label every sentence manually. A language model could learn by repeatedly attempting to predict tokens from their context.
That apparently simple task is enormously productive.
To predict the missing or next expression successfully, the model must learn statistical relationships involving grammar, semantics, entities, concepts, discourse and many patterns represented in human writing.
This is one reason I would connect LLM development with your larger information-science question. Humanity had spent centuries turning knowledge into shareable recorded information. LLM training suddenly provided a powerful computational means of extracting statistical structure from a huge part of that recorded information.
4. Hardware made the scale physically possible
The Transformer alone would not have produced today’s LLMs.
GPUs and later AI accelerators made enormous matrix computations practical. High-bandwidth memory, fast networking and distributed computing allowed thousands of processors to cooperate in training one model.
The infrastructure itself became a major engineering achievement. For GPT-4, for example, OpenAI reported rebuilding its deep-learning stack and co-designing a supercomputer with Azure, while developing methods that made very large training runs more predictable. (OpenAI)
This connects directly with your recent observation about AI, data centers and electricity. LLM intelligence has become partly an industrial infrastructure problem:
electricity → data centers → accelerators → computation → model training/inference.
5. Post-training turned a predictor into something much more useful
A raw pretrained LLM is fundamentally trained to predict text. That alone does not make it a good conversational assistant.
Instruction tuning, supervised fine-tuning, reinforcement learning from feedback and related methods taught models to respond more usefully to human intentions. OpenAI, for example, described GPT-4’s basic capabilities as coming largely from pretraining while post-training made the system much more steerable toward user intentions. (OpenAI)
More recent development has increasingly emphasized post-training scaling and test-time computation—allowing models to spend additional computation on difficult problems rather than relying only on what was acquired during pretraining. A 2025 ACL survey identifies supervised fine-tuning, reinforcement learning from feedback, and test-time compute as major dimensions of this newer scaling regime. (ACL Anthology)
This helps explain why recent progress can feel faster than merely making models larger.
6. Natural-language interaction may be the most important practical breakthrough
Earlier information systems usually required humans to adapt themselves to the machine.
In classical IR:
information need → formulate query → machine matching → documents → human interpretation
With an LLM:
information need ⇄ natural-language conversation ⇄ machine interpretation/reasoning ⇄ information
The user can say, “No, that isn’t what I meant,” add context, ask why, challenge an answer, or reformulate the problem. The system can use the continuing conversation as additional context.
That is a qualitative change in the human-machine interface.
And this brings us back to cognition and context
I think there is a deeper connection with the question we have been discussing concerning the cognitive turn in information science.
Traditional information systems mainly manipulated relatively explicit representations: keywords, descriptors, citations, classifications and queries.
LLMs instead operate over extremely rich contextual relationships among linguistic expressions.
That does not mean an LLM possesses human cognition. But it means that something information scientists struggled with for decades—how to make a machine sensitive to meaning as expressed through context—has advanced dramatically.
So I would summarize the history in one chain:
**digitized human knowledge
- Transformer attention
- massive training data
- GPUs/AI accelerators
- scaling laws and better training
- post-training and feedback
- test-time reasoning
= modern LLMs.**
But behind these engineering developments lies an older intellectual problem:
An expression cannot always be adequately understood in isolation; its meaning emerges through its relationships with a context.
That principle is much older than LLMs. What is revolutionary is that modern computation has finally made it possible to exploit contextual relationships at enormous scale.
And that is precisely why the history of context, meaning and cognition in information science may deserve to be reconsidered in the LLM era.
ㅡㅡㅡㅡ
What are transformer, attneiton, and transformer attention individually?
ㅡㅡ
The three terms are related, but they are not the same thing. It helps to separate them.
1. Attention — a method
Attention is a computational method that lets a neural network decide which other pieces of information are important for interpreting the piece it is currently processing.
Take:
“John dropped the glass because it was slippery.”
To understand “it,” we need the surrounding words. The model calculates how strongly the current word should be related to other words or tokens.
Very roughly:
current token → examine other relevant tokens → assign different weights → combine their information
Those weights are called attention weights.
So attention can be understood as:
a mechanism for calculating contextual relationships.
It existed before the Transformer. Neural-machine-translation researchers were already developing attention mechanisms around 2014–15.
2. Transformer — an architecture
A Transformer is the larger neural-network architecture introduced in the famous 2017 paper Attention Is All You Need by Vaswani and colleagues. Attention Is All You Need
Think of a Transformer as a machine made from many components.
Among its important components are:
attention + feed-forward neural networks + positional information + normalization + residual connections
These components are organized into layers, and many layers can be stacked.
So:
Attention is a mechanism. A Transformer is an architecture that makes extensive use of that mechanism.
The relation is somewhat like:
combustion → engine
Combustion is a process; an engine is a whole machine organized to exploit that process.
3. Transformer attention — attention used inside the Transformer
This expression normally means the particular attention mechanism used as a central component of Transformers, especially self-attention and multi-head attention.
Self-attention is particularly important.
Consider:
“The scientist read the paper because it contained an interesting argument.”
When processing “it,” self-attention allows the representation of “it” to be influenced by other tokens in the same sequence—scientist, paper, contained, argument, etc.—with different strengths.
The model thereby constructs a context-dependent representation.
This is why the same word can acquire different representations in:
“I deposited money in the bank.”
and
“We sat on the river bank.”
The written token bank is the same. Its contextual relations are different.
A little more technically
Transformer attention creates three representations traditionally called:
Query (Q) — what information is this position looking for?
Key (K) — what information does another position advertise?
Value (V) — what information should be taken from that position?
The basic operation is:
Don’t worry about the mathematics yet. Conceptually it says:
Compare what I am looking for with what the other tokens offer; give stronger weights to relevant relationships; then combine the corresponding information.
And multi-head attention does several such relational calculations in parallel, allowing different heads to learn different useful patterns.
The three concepts in one picture
We can therefore arrange them as:
ATTENTION
a general computational mechanism for weighting relationships
↓ incorporated centrally into
TRANSFORMER
a complete neural-network architecture
↓ whose central operation includes
TRANSFORMER SELF-ATTENTION
tokens construct representations partly from their relationships with other tokens in their context
And then:
many Transformer layers + enormous training data + enormous computation → modern LLM
There is an important connection with your concern about context, but we should make it precisely.
Transformer attention does not simply mean “the machine understands context just as a human reader does.” Attention is a mathematical operation.
Nevertheless, something remarkable happens when such operations are stacked in many layers and trained over enormous amounts of language: the resulting representations become highly sensitive to contextual relationships.
That is one of the reasons the Transformer is so important to the history of AI. It provided an extraordinarily scalable computational way of exploiting something fundamental to language:
The significance of an expression depends strongly upon its relationships with other expressions in its context.
That last proposition is much older than Transformers—and it is exactly where the technological history of LLMs begins to touch the intellectual history of context, meaning, and cognition that we have been examining.
ㅡㅡㅡㅡ
Summarize the 2017 paper Attention Is All You Need.
ㅡㅡ
The 2017 paper Attention Is All You Need, by Ashish Vaswani and seven colleagues, introduced the Transformer, the architecture that eventually became the foundation of modern LLMs. The paper was submitted in June 2017. (arXiv)
Attention Is All You Need — original paper
The central idea
Before the Transformer, leading language systems normally processed sequences using recurrent neural networks (RNNs) or convolutional neural networks (CNNs), often supplemented by attention.
Vaswani and colleagues asked a radical question:
What if we dispense with recurrence and convolution and build the system around attention itself?
Their answer was the Transformer. The paper describes it as an architecture “based solely on attention mechanisms,” eliminating recurrence and convolution for sequence modeling. (arXiv)
That is the meaning behind the memorable title “Attention Is All You Need.”
What attention does
Suppose we have:
The animal didn’t cross the street because it was tired.
To represent it, the system should recognize that animal is particularly relevant.
Self-attention allows every token to examine its relationships with other tokens in the sequence and assign different weights to them:
it → animal: strong relation
it → street: weaker relation
The important point is that the representation of a word is therefore influenced by its context.
Technically, the paper represents this through queries, keys, and values:
In simple language:
Query: What am I looking for?
Key: What does each other token offer?
Value: What information should I take from it?
The system compares queries with keys, calculates attention weights, and then combines the corresponding values.
Multi-head attention
The authors did not use just one attention operation. They introduced multi-head attention.
Different attention “heads” can examine different relationships simultaneously. One might become useful for syntactic relationships, another for longer-distance dependencies, and others for different patterns. We should not assume that every head corresponds neatly to a human linguistic category, however.
The resulting representations are therefore produced from multiple relational views of the same context.
The Transformer architecture
The original Transformer was designed for translation and had two major parts:
Encoder: builds contextual representations of the input sentence.
Decoder: generates the output sentence, using attention both to previous output tokens and to the encoder’s representations.
Each consists of stacked layers containing attention and feed-forward neural networks, together with residual connections and normalization.
Because attention itself does not tell the model whether one word occurs before or after another, the authors also introduced positional encoding, which gives the model information about token positions.
Why it was such a breakthrough
Earlier recurrent models generally processed sequences step by step:
The Transformer could process relationships among many tokens much more in parallel.
That had enormous practical consequences. It fitted modern GPU computation very well and greatly improved the possibility of scaling language models.
The original results were already impressive. On WMT 2014 English-to-German translation, the Transformer reached 28.4 BLEU, more than 2 BLEU above previous reported best results including ensembles. On English-to-French, it achieved 41.8 BLEU after 3.5 days of training on eight GPUs. (arXiv)
But these translation scores were not the paper’s greatest historical consequence. The important discovery was an architecture that could later be scaled enormously.
What the paper did not do
It is important historically not to read today’s ChatGPT backward into the 2017 paper.
The authors were mainly addressing sequence transduction, especially machine translation. They were not presenting today’s general-purpose conversational LLM.
The subsequent revolution required several additional developments:
Transformer (2017)
→ large-scale pretraining
→ GPT/BERT and related models
→ much larger models and datasets
→ scaling laws
→ instruction tuning and human feedback
→ today’s LLMs.
So Attention Is All You Need supplied the crucial architecture, rather than the complete recipe for ChatGPT.
The deeper importance for our discussion
There is an interesting distinction between attention and context.
Attention is a mathematical mechanism.
Context is a much broader concept concerning the relationships through which something acquires meaning.
The Transformer does not contain a little human reader consciously asking, “What does this passage mean?” Rather, during training it learns numerical representations in which one token’s representation is constructed partly through weighted relationships with other tokens.
Nevertheless, the result is remarkable:
That helps explain why the Transformer became so powerful for language.
And here I think there is a very interesting connection to the problem you were examining decades earlier. Your information-science argument concerns something like:
Transformer language models operate very differently:
These should not be identified as the same process. One concerns human cognition and interpretation; the other concerns learned numerical computation.
But they share a profound structural principle:
An expression should not be treated as if its significance were completely contained in the expression alone. Its relationships with its context matter.
That is why I think Attention Is All You Need is especially interesting when viewed from the history of context and meaning, not merely from the history of computer engineering.
ㅡㅡㅡㅡ
Is “attention” directed toward “context?
Here I recall the last step of Marc de Mey’s mondadic → structure → context → cognition. Likewise context → tranformer attention
Then, there is a parallel between cognitive information retrieval and contextual tranformer attention. That is transformer attention corresponds to human cognition, does it?
ㅡㅡ
Yes. I think you have identified a real structural parallel, provided we do not say that Transformer attention is human cognition.
Marc de Mey’s sequence is particularly illuminating here. In his account of the development of the cognitive view, he distinguished four stages: monadic → structural → contextual → cognitive/epistemic. At the contextual stage, context is required to disambiguate meaning; at the cognitive stage, interpretation is mediated by the information processor’s existing conceptual structure or model of the world. (IMR Press)
Is attention directed toward context?
In a useful sense, yes.
Transformer self-attention relates different positions in a sequence in order to construct representations of that sequence. Each token can acquire information from other tokens according to learned attention weights. The original Transformer paper explicitly defines self-attention in terms of relating positions within the same sequence. (NeurIPS Papers)
For example:
He went to the bank to withdraw money.
The token bank begins as a token representation. Self-attention allows its representation to be affected by withdraw, money, went, and the other relevant tokens.
So we can simplify it as:
Thus your expression
captures something important, although technically I would reverse the causal wording slightly:
Attention is the mechanism, while context is what that mechanism helps the model exploit.
Now compare this with de Mey
Here your observation becomes more interesting.
De Mey’s progression can be simplified:
At the monadic stage, the information unit is treated almost independently.
At the structural stage, relationships among units matter.
At the contextual stage, surrounding information is needed to disambiguate meaning.
At the cognitive stage, the information is interpreted through the recipient’s existing conceptual structure or world model. (IMR Press)
Now compare a Transformer:
There is indeed a striking structural resemblance.
Does Transformer attention correspond to human cognition?
Here I would say:
Functionally, in a limited respect, yes. Ontologically, no.
That distinction is crucial.
A human reader encounters:
“He deposited his salary in the bank.”
Human cognition uses linguistic context plus enormous existing knowledge—knowledge about money, salaries, financial institutions, ordinary life, previous discourse, intentions, etc.—to interpret bank.
A Transformer performs something different mathematically. Self-attention computes weighted relationships among representations. Across many layers, those representations interact with parameters learned from enormous quantities of previous language. (Harvard NLP)
Therefore:
Human:
Transformer:
They are not identical mechanisms.
But they perform analogous informational functions: an isolated expression is insufficient; relations with context help determine how it should be interpreted.
And this creates your IR parallel
I think this is the strongest part of your observation.
Your citation-context approach essentially says:
Therefore:
Transformer language processing says, very approximately:
therefore:
So there is a genuine parallel:
| Cognitive citation-context retrieval | Transformer |
|---|---|
| Citation/expression | Token |
| Citation context | Token context |
| Reader attends to relevant context | Self-attention weights contextual relations |
| Existing human knowledge | Learned model parameters/representations |
| Cognitive interpretation | Contextualized representation |
| Relevance/meaning judgment | Prediction/generation |
But I would not equate the final two rows. A contextual representation is not demonstrated to be human-style understanding or consciousness.
De Mey gives us an even deeper distinction
De Mey’s fourth step helps prevent us from making that mistake.
Contextual means approximately:
Meaning cannot be adequately determined without context.
Cognitive adds:
The context is interpreted through an already existing conceptual/knowledge structure.
That suggests something fascinating about LLMs.
Attention by itself corresponds most directly to de Mey’s contextual stage.
But a complete trained Transformer is more than an attention operation. Its parameters contain learned statistical structures acquired through training. Consequently, the current context is processed through a large pre-existing learned structure.
Schematically:
while perhaps:
—not human cognition itself.
This distinction makes your proposed parallel much stronger scientifically.
The historical parallel
We can now state your idea in a form I think is worth developing:
This does not establish a historical influence from cognitive information science to Transformers. Vaswani et al. presented the Transformer as a solution to technical problems in sequence modeling, building on earlier neural attention research. (arXiv)
But conceptually, the parallel is substantial.
And it suggests a larger proposition:
Both the cognitive turn in information science and the Transformer revolution in AI can be understood, at different levels and by very different mechanisms, as moves away from treating information units as self-contained and toward exploiting relationships, context, and context-dependent interpretation.
That, I think, is considerably more interesting than simply saying “attention corresponds to cognition.” It places context at the bridge between the two intellectual developments.
ㅡㅡㅡㅡ
Did the 2016 paper “Attention is all you need” use the term “context” technically?
ㅡㅡ
Yes, but with an important qualification. And first, a small correction: Attention Is All You Need is a 2017 paper, submitted to arXiv on June 12, 2017, not 2016. (arXiv)
The surprising answer is that the paper does not make “context” a central technical term in the way that “attention,” “self-attention,” “query,” “key,” and “value” are technical terms. The authors’ formal vocabulary is primarily about representations, positions, sequences, dependencies, queries, keys, values, and attention weights. (arXiv)
Yet what attention does is deeply contextual.
The paper defines self-attention as an attention mechanism that relates different positions of a single sequence in order to compute a representation of that sequence. In other words, when representing one position, the system uses information available from other positions. (arXiv)
So we should distinguish:
from
The 2017 paper is overwhelmingly about the second, even though it does not build its theory around the first term.
This becomes even clearer if we look at the attention tradition preceding and following the Transformer. Attention mechanisms are commonly described as producing a context vector: a weighted combination of information selected from a sequence. Later explanations of the Transformer likewise explicitly describe self-attention as creating contextual information for each position. (ResearchGate)
And this makes your comparison with Marc de Mey more interesting than I initially thought.
De Mey’s conceptual progression was:
The Transformer paper’s technological progression can be interpreted—not as the authors’ own terminology, but retrospectively—as:
There is therefore a remarkable conceptual parallel.
But I would now modify one thing from my previous answer. Instead of saying simply
I would say:
Attention is the computational operation by which relevant relationships in the available context are selectively weighted and combined.
This distinction may become quite important for your historical argument. Vaswani et al. did not present their work as a theory of “context” comparable to de Mey’s theory of cognition. They approached the problem from machine learning and sequence processing. (arXiv)
Yet, viewed from information science, the Transformer accomplished computationally something that can be described at a much more abstract level as:
Do not treat an information unit in isolation. Interpret—or computationally represent—it through its relationships with other information units.
That principle sounds remarkably familiar in light of the context → cognition line we have been examining.
So I think we have found a potentially useful distinction:
De Mey: context as a condition for cognitive interpretation.
Park: citation context as evidence through which the human user cognitively determines meaning and relevance.
Transformer: attention as a mathematical mechanism for constructing context-sensitive representations.
They are not historically or technically the same theory. But all three give a central role to relations rather than isolated units.
That may be the deeper parallel worth pursuing.
ㅡㅡㅡㅡ
2016-09-30 Mark Park
