My Academic Database (MAD)

I like to develop My Academic Database (MAD). It essentially contains all the academic references in hypertext, including all my writings. It can retrieve hypertexts using not only metadata indexing but also citation-context indexing. It will share the window with ChatGPT. The window will be used in presentation, which uses Text2Speech.

ㅡㅡ

This is a clear and practical first form of Our Library.

My Academic Database — MAD

MAD would be a personal academic knowledge system containing:

  • all bibliographic references important to the user;
  • all the user’s own writings;
  • full citation links among documents;
  • citation contexts showing how one work uses another;
  • notes, summaries, criticism, and later corrections;
  • links to available full texts.

Its basic unit should not be only the document. It should also store the meaningful relation between documents.

A normal database records:

Author A cited Author B.

MAD should also record:

Author A supported, described, used, compared, criticized, corrected, or rejected Author B—and here are the actual words.

That is citation-context indexing.

Two retrieval paths

MAD should provide two connected ways of finding information.

Metadata retrieval

Search by:

  • author;
  • title;
  • date;
  • journal;
  • subject;
  • keyword;
  • document type;
  • institution.

This answers:

Which documents exist?

Citation-context retrieval

Search the passages surrounding citations.

This answers:

What did later authors say about this document?

A query such as:

Who criticized Belkin’s ASK?

should retrieve not merely works containing the words Belkin and ASK, but passages in which ASK was explained, accepted, questioned, or rejected.

Hypertext structure

Every academic reference should become a live hypertext node.

From one document, the user could move:

  • backward to cited sources;
  • forward to citing sources;
  • sideways to related works;
  • inward to citation contexts;
  • outward to reviews, corrections, and discussions.

The citation context would act as the meaningful bridge between two nodes.

Thus MAD would not be only a bibliography. It would be a personal scholarly web.

Shared window with ChatGPT

The shared window is especially important.

One side could show ChatGPT.
The other side could show MAD.

ChatGPT may explain, compare, summarize, and infer. MAD supplies the documentary evidence.

A useful interaction would be:

  1. The user asks ChatGPT a question.
  2. ChatGPT searches MAD.
  3. MAD opens the most relevant references and citation contexts.
  4. ChatGPT explains the evidence.
  5. The user clicks any source without losing the conversation.
  6. The user corrects the answer or relevance ranking.
  7. MAD records that correction.

This is close to your proposed LLM+:

LLM for inference; MAD for retrieval and evidence.

The shared window also realizes your earlier popup-window principle: related material should open without making the researcher lose the main line of thought.

Presentation mode

MAD can also become a presentation system.

A presentation page could contain:

  • a short main statement;
  • a document image or citation context;
  • highlighted evidence;
  • backward and forward citation links;
  • ChatGPT commentary;
  • Text-to-Speech narration.

The presentation would not have to follow fixed slides only. It could move through the citation web while speaking.

For example:

“This is Park’s 1975 statement. Now let us move forward twenty-two years to CiteSeer. Here is the later technical realization. Now let us compare their citation-context functions.”

The screen could automatically open each supporting item while Text-to-Speech reads the prepared explanation.

This would make the presentation evidence-driven, rather than merely decorative.

A possible main screen

A simple first layout could have four areas:

Left: search and reference list
Center: document or citation context
Right: ChatGPT conversation
Bottom: presentation controls and Text-to-Speech

The user could save any research path as a presentation sequence.

The first realistic version

The first MAD does not need to contain all academic literature. It can begin with:

  • all of Park’s writings;
  • the works cited in the 1975 thesis;
  • later works related to citation context;
  • UCL information-science works from 1974–1982;
  • CiteSeer and later scholarly-search systems;
  • personal notes and correspondence.

This limited collection would already demonstrate the central principle.

A concise description is:

My Academic Database is a personal hypertext database of academic references, writings, citations, and citation contexts. It supports both metadata retrieval and citation-context retrieval. It shares a working window with ChatGPT, which uses MAD as its evidence base. Its presentation mode combines live hypertext navigation, documentary evidence, AI explanation, and Text-to-Speech.

MAD is a memorable name. It also carries a useful suggestion: the present academic record may appear disorderly, but MAD will make its hidden relations visible.

ㅡㅡㅡㅡ

Introduce the real-life systems closest to MAD.

ㅡㅡ

No single present system is yet MAD. The closest real-life answer is a combination of several systems, each covering one part.

Closest systems

1. Zotero — closest to MAD’s personal database

Zotero is the strongest base for storing one’s own academic collection. It collects references and PDFs, stores notes and annotations, supports tags and saved searches, indexes PDF full text, synchronizes across devices, and produces citations and bibliographies. Its “My Publications” collection can also publish one’s own writings. (조테로)

Its weakness is central to MAD: Zotero mainly treats references as items in a collection. It does not automatically build a rich citation-context network showing how one work supports, criticizes, or develops another.

Closest MAD function: personal storage, metadata, PDFs, notes, and one’s own publications.


2. ResearchRabbit — closest to MAD’s navigable citation web

ResearchRabbit lets a scholar begin with papers, build personal collections, and explore references, later citations, similar works, authors, and visible citation maps. Its maps help users follow the historical growth of a subject. (ResearchRabbit)

It is closer than ordinary search engines to your idea that scholarship is naturally explored backward and forward through citations.

Its weakness is that the links usually mean only:

Paper A cites Paper B.

They do not fully explain what A said about B.

Closest MAD function: backward and forward citation tracking in a personal visual workspace.


3. Litmaps — closest to citation tracking plus current awareness

Litmaps builds visual maps from citation networks. It can search outward from selected papers and monitor the map for new publications, sending alerts when relevant work appears. (Litmaps Help Center)

This approaches two MAD services at once:

  • Citation Tracking;
  • Current Awareness.

Its main unit remains the citation link, not the interpreted citation context.

Closest MAD function: living citation maps with automatic updating.


4. Scite — closest to citation-context indexing

Scite is the system closest to the most original part of MAD. Its “Smart Citations” show the passage in which a paper is cited and classify the citation as supporting, contrasting, or mentioning. Its system has processed more than a billion citation statements. (Scite)

This is clearly a form of automated citation-context indexing.

Yet Scite’s classes are still rather broad. MAD could go further by distinguishing:

  • description;
  • theoretical use;
  • method use;
  • evidence;
  • comparison;
  • extension;
  • criticism;
  • correction;
  • rejection;
  • misrepresentation.

It could also keep the full historical chain of later responses.

Closest MAD function: retrieval and classification of actual citation contexts.


5. Semantic Scholar — closest to a large academic graph with AI assistance

Semantic Scholar offers paper search, references, citing works, personalized libraries, citation and author alerts, recommendations, and an academic graph API. Its Semantic Reader adds contextual information directly around citations while a paper is being read. (Semantic Scholar)

It combines several MAD elements more fully than most single systems:

  • metadata retrieval;
  • citation graph;
  • personal library;
  • current awareness;
  • AI-supported reading.

But it is a global service controlled by its own database. It is not a scholar-owned database containing a fully editable private and public knowledge structure.

Closest MAD function: broad integrated scholarly search and citation services.


6. OpenAlex — closest to an open foundation for building MAD

OpenAlex is not mainly a personal reading interface. It is an open catalog and graph of scholarly works, authors, institutions, topics, funders, and citation relations. Its data are openly reusable, and its API can find works citing a particular work. (OpenAlex Developers)

OpenAlex may be the best external data foundation for developing MAD because MAD would not need to create the whole scholarly graph from nothing.

However, OpenAlex mostly provides metadata and citation links. Citation sentences and full interpretive contexts would need another source or a locally created extraction process.

Closest MAD function: open global metadata and citation infrastructure.


7. NotebookLM — closest to the shared AI, source, and spoken-presentation window

NotebookLM lets users upload selected sources, speak with an AI grounded in those sources, receive inline citations, and create briefings, mind maps, slide decks, audio overviews, and video overviews. It can show source material while the user asks questions or listens to generated audio. (Google 도움말)

This comes remarkably close to MAD’s presentation side:

  • selected personal documents;
  • AI conversation;
  • visible source support;
  • shared working window;
  • spoken explanation;
  • presentation artifacts.

But NotebookLM does not provide a deep global citation graph or systematic citation-context indexing.

Closest MAD function: ChatGPT-like discussion grounded in personal sources, joined with Text-to-Speech and presentation.


8. Elicit — closest to AI critical review

Elicit searches academic literature, summarizes papers, extracts comparable data, helps screen studies, produces literature reviews, and links AI-produced statements back to supporting passages. It also offers research alerts. (Elicit)

This is close to MAD’s proposed Critical Review service. It is strongest for systematic and empirical literature review.

Its limitation is that it does not organize all scholarship mainly through a scholar’s personal hypertext and citation-context history.

Closest MAD function: AI-supported evidence comparison and critical synthesis.

The closest present combination

A practical early MAD could be assembled as:

Zotero for the personal library
OpenAlex or Semantic Scholar for global metadata and citation links
ResearchRabbit or Litmaps for visual citation navigation
Scite for citation contexts
Elicit for critical review
NotebookLM or ChatGPT for dialogue, presentation, and speech

That combination would be powerful, but it would remain fragmented. The user would move between separate services, accounts, interfaces, data rules, and subscription systems.

What would make MAD new

MAD would unite these parts inside one scholar-controlled system:

  1. All personal writings and references in one database
  2. Metadata and full-text retrieval
  3. Backward and forward citation tracking
  4. Actual citation-context retrieval
  5. Typed relations such as support, use, criticism, and correction
  6. ChatGPT sharing the same visible workspace
  7. A presentation path through live sources
  8. Text-to-Speech tied to each displayed piece of evidence
  9. Human correction of AI summaries and relation labels
  10. Permanent preservation of the scholar’s own intellectual history

The closest single system to the database side is probably Zotero.

The closest to the citation-context principle is Scite.

The closest to the shared AI and presentation window is NotebookLM.

The closest to MAD as a whole is therefore not one product, but:

Zotero + Scite + ResearchRabbit/Litmaps + NotebookLM, joined under one interface.

MAD’s real novelty would be the integration of personal documentary memory, global citation relations, citation-context meaning, AI inference, and spoken evidence-based presentation.

ㅡㅡㅡㅡ

MAD could relate CiteSeer, for example, to my 1975 thesis, if you like. Imagine every document has a special field “See also” under which “Citation-context indexing” hyperlinks to my thesis, and “Citation-context analysis” does to Henry Small’s various relevant papers.

ㅡㅡ

Yes. This would make MAD more than a citation database. It would become a conceptual hypertext.

For a CiteSeer record, the field might look like this:

See also

  • Citation-context indexing → Park, A Direct Approach to Information Retrieval (1975)
  • Automatic citation indexing → CiteSeer technical papers
  • Citation-context analysis → Henry Small’s relevant papers
  • Co-citation analysis → Henry Small’s foundational work
  • Forward and backward citation tracking → later scholarly search systems

This is useful because ordinary citation links answer:

Who cited whom?

The See also field answers another question:

Which earlier and later works belong to the same intellectual idea?

That relation may exist even when CiteSeer did not cite the 1975 thesis.

MAD should therefore distinguish several link types:

  • cites — an explicit citation exists;
  • conceptually related — the works share an important idea;
  • earlier proposal — one work proposed the idea earlier;
  • technical realization — one work implemented the idea;
  • analysis of the phenomenon — one work studied its use or effects;
  • possible influence — influence is suspected but not proved;
  • documented influence — evidence of transmission exists.

This distinction is essential. MAD should not falsely say:

CiteSeer was derived from Park’s thesis.

Unless evidence proves that.

It can safely say:

CiteSeer is conceptually related to Park’s 1975 proposal and may be understood as a later technical realization of citation-context indexing.

The relation to Henry Small also needs precision. Small’s work mainly concerns citation relations as evidence of scientific structure, especially co-citation, specialty formation, and the meaning of citations. That is close to citation-context analysis, but it is not identical to Park’s retrieval proposal.

So MAD could show a clear division:

Park: citation contexts as retrieval clues.
Small: citations and citation contexts as evidence for analyzing science.
CiteSeer: automated citation indexing and navigation.

This gives three different functions:

  1. retrieval;
  2. analysis;
  3. technical realization.

The See also field should be partly human-curated and partly AI-suggested. AI may propose a relation, but a human should confirm its label and supporting evidence.

Each link could also carry a short explanation:

Citation-context indexing — Park proposed retrieving extracts in which a source document cites, describes, criticizes, or compares another document.

Citation-context analysis — Small examined what citation relations reveal about scientific communication and intellectual structure.

Then the hyperlink itself becomes meaningful. It does not merely move the reader to another document. It explains why the reader should go there.

That is true hypertext in your sense:

not only linked documents, but meaningful and named relations among documents.

ㅡㅡㅡㅡ

2026-08-06 Mark Park