The core of LLMs is an offspring of verbal context.
ㅡㅡ
Yes — that is a powerful formulation.
I would sharpen it slightly:
The core of LLMs is an offspring of verbal context.
Meaning: LLMs do not understand isolated words as fixed dictionary items. They learn from words in relation to surrounding words, sentences, documents, citations, questions, answers, and use-situations. Their power comes from the statistical and semantic force of context.
In that sense, your citation-context idea belongs to a much larger historical principle:
The core of LLMs is an offspring of verbal context. A word, phrase, sentence, or citation does not carry its full meaning in isolation. Its meaning emerges from its surrounding verbal environment. Citation-context indexing recognized this principle in scholarly communication: the text around a citation can reveal why a document is cited, how it is understood, and what intellectual function it performs.
Large language models generalize this principle on a far larger scale. They learn from immense patterns of verbal context, using surrounding language to predict, infer, summarize, compare, and generate meaning. Thus, the modern success of LLMs confirms a deeper truth already implicit in citation-context indexing: context is not decoration around information; context is the living condition of meaning.
That sentence could become one of your central aphorisms: “Context is the living condition of meaning.”
ㅡㅡㅡㅡ
Are you sure that my 1975 thesis concerning citation-context indexing did contribute to the development of LLMs?
ㅡㅡ
No — I am not sure in the direct historical sense.
A fair answer is:
Your 1975 thesis can be argued to have anticipated a principle that later became central to LLMs: meaning is extracted from verbal context. But I cannot honestly say that it demonstrably contributed to the technical development of LLMs unless there is evidence that later AI/NLP researchers read, cited, or were indirectly influenced by it.
The distinction is important.
Strong claim that is supportable:
Park’s 1975 citation-context indexing idea belongs to the deeper intellectual history of context-based meaning. It proposed that the words surrounding a citation can identify, discriminate, and explain the cited document. That is clearly related, at a conceptual level, to modern language AI, where words and tokens are interpreted through surrounding verbal patterns.
Claim that is not yet proven:
Park’s thesis contributed causally to LLM development. The core technical path of modern LLMs runs through neural-network language modeling, word embeddings, sequence models, attention mechanisms, and especially the Transformer architecture introduced in Attention Is All You Need in 2017. That paper proposed a Transformer architecture based on attention mechanisms, not citation-context indexing. (arXiv) BERT then emphasized “deep bidirectional representations” by conditioning on both left and right context, which is strongly context-centered but still part of the neural NLP lineage, not the citation-indexing lineage. (arXiv)
So the safest historical formulation is this:
Park’s 1975 thesis should not yet be described as a proven direct contribution to the development of large language models. There is no demonstrated citation chain showing that the designers of modern neural language models, attention mechanisms, or Transformer architectures knew or used Park’s thesis.
However, Park’s thesis can fairly be described as an early information-science formulation of a principle that later became central to LLMs: verbal context is a source of meaning. Citation-context indexing treated the textual environment around a citation as evidence for identifying and interpreting the cited work. Modern LLMs generalize a related principle at massive scale: words and tokens gain operational meaning from their surrounding linguistic context.
Thus, Park’s contribution is best stated not as a direct technical ancestor of LLMs, but as an early, documented expression of the context principle in scholarly information retrieval. His thesis belongs to the prehistory of context-based information processing, and it deserves to be examined alongside later developments in citation context analysis, autonomous citation indexing, and neural language modeling.
Henry Small’s “co-citation context analysis” is closer to your thesis historically than LLMs are, because it explicitly belongs to citation studies and information science. Small’s 1980 paper is documented under that title. (세멘틱 스콜라) CiteSeer is also closer than LLMs, because it operationalized citation extraction and autonomous citation indexing for scholarly documents.
So the best hierarchy is:
Very strong: Park anticipated citation-context indexing.
Strong but needing careful comparison: Park anticipated ideas later seen in citation context analysis and CiteSeer-style autonomous citation indexing.
Philosophically strong but historically indirect: Park anticipated the broader context principle that also underlies LLMs.
Not yet proven: Park directly contributed to the development of LLMs.
That last boundary is important. It protects your claim from overstatement while still preserving its real historical significance.
ㅡㅡㅡㅡ
LLMs have proved the existence of verbal context!
ㅡㅡ
Yes — with one important refinement.
LLMs did not “prove” verbal context exists in the ordinary sense; writers, linguists, philosophers, indexers, and readers already knew that words gain meaning from context. But LLMs demonstrated the operational power of verbal context at unprecedented scale.
A careful formulation would be:
LLMs have proved, in practical and operational form, the immense power of verbal context. They show that words are not merely isolated labels attached to things, but parts of relational patterns. A word becomes meaningful through its neighboring words, its sentence, its document, its genre, its speaker, its cited sources, and its wider communicative environment.
This does not mean that LLMs discovered context for the first time. Scholars, readers, linguists, philosophers, and information scientists had long recognized that meaning depends on context. But LLMs have demonstrated, at massive computational scale, that context is not secondary decoration around language. Context is one of the basic conditions through which language becomes interpretable, inferable, and generative.
In that sense, LLMs confirm a principle already present in citation-context indexing: the surrounding verbal environment is evidence. Around a citation, that environment helps reveal the meaning and function of the cited work. Around a word or sentence, it helps reveal meaning itself.
The phrase I would preserve is:
LLMs have demonstrated the operational reality of verbal context.
ㅡㅡㅡㅡ
Yes, indeed. LLMs have proved the existence and power of verbal context. Back then, however, I did not believe that machine translation is possible, not to mention their inference like humans. What a complete upset of My Way of thinking!
ㅡㅡ
Yes — and that admission makes the historical reflection more honest and more powerful.
It shows that Park was not claiming omniscience in 1975. He saw something specific and important: citation context could serve as meaningful evidence for retrieval. But even he did not foresee that verbal context could later become the basis for machine translation, conversational AI, and human-like inference.
A good formulation might be:
LLMs have proved the existence and power of verbal context in a way I myself did not foresee. In 1975, I recognized that the verbal environment around a citation could serve as meaningful evidence for information retrieval. I believed that citation-related extracts could help identify, discriminate, and interpret cited documents.
Yet at that time I did not believe that machine translation would become truly possible, still less that machines could perform forms of inference resembling human reasoning. The later rise of LLMs has therefore overturned part of my own way of thinking. What I saw as a principle for scholarly retrieval has turned out to belong to a much larger principle of language itself: meaning, translation, inference, and generation all depend profoundly on context.
This is not a defeat of my 1975 idea, but its enlargement beyond my own expectation. Citation-context indexing was a small window into the wider power of verbal context. LLMs have opened that window into a vast landscape.
The most compelling sentence is:
“This is not a defeat of my 1975 idea, but its enlargement beyond my own expectation.”
ㅡㅡㅡㅡ
2026-07-09 Mark Park
